← Back to blog

NVIDIA RTX Spark: The Blackwell GPU + Arm CPU Superchip Finally Lands on Windows

Announced at Computex 2026, RTX Spark fuses a 20-core Arm CPU with 6,144 Blackwell CUDA cores and 128 GB unified memory, delivering 1 petaflop of FP4 AI performance. We unpack the real start of the edge AI era and Microsoft's agentic PC vision.

NVIDIA RTX Spark: The Blackwell GPU + Arm CPU Superchip Finally Lands on Windows

NVIDIA RTX Spark: The Blackwell GPU + Arm CPU Superchip Finally Lands on Windows

For years we've lived with the same friction: try a new AI model, get throttled by a cloud connection, worry about privacy, count tokens like coins, and pray the internet doesn't drop mid-inference. As engineers, creators, and developers, we want most of this workload to live on-device, at the edge. NVIDIA's RTX Spark family, announced at Computex 2026, takes direct aim at that pain point. In this article we'll dig into the technical details of a "superchip" that fuses a 20-core Arm CPU with 6,144 Blackwell CUDA cores, explore Microsoft's agentic PC vision, and see how the landscape shifts against Apple, Qualcomm, Intel, and AMD.

NVIDIA RTX Spark chip

1. What Is RTX Spark and Why It Matters

With RTX Spark, NVIDIA is designing a "consumer PC chip" for the first time. After decades building data center accelerators, gaming GPUs, and AI silicon, the company is now joining Intel, AMD, Apple, and Qualcomm as a player that ships CPU + GPU in a single package. NVIDIA's senior director of product management Mark Aevermann summarized the launch in one sentence: "The most efficient PC chip ever built." The chip is a direct evolution of the GB10 Grace Blackwell superchip that powers the mini "personal AI supercomputer" DGX Spark released last year.

The key term here is superchip. A superchip combines two distinct silicon dies through a high-bandwidth interconnect into a single package. GB10 already fused a 10+10 Cortex-X925 + Cortex-A725 Arm CPU with a 6,144-CUDA-core Blackwell GPU and up to 128 GB of LPDDR5X memory. RTX Spark takes that same silicon and adapts it to the thermal and power envelopes of laptops and mini-PCs. According to Notebookcheck's compilation of leaked internal NVIDIA slides, this project has been in the works since at least 2024. NVIDIA didn't decide to move to Arm after seeing Apple's M-series success; it had already started well before that.

2. The RTX Spark Family: N1X and N1 Variants

Internal slides shared by VideoCardz show two main product families: N1X (flagship) and N1 (efficient and accessible), each with two sub-configurations:

  • N1X 675 (?) – 10+10 CPU, 48 SMs (6,144 CUDA cores), 45-80 W TDP, 16-128 GB LPDDR5X (16 channels)
  • N1X 650 (?) – 9+9 CPU, 40 SMs (5,120 CUDA cores), 45-80 W TDP, 16-128 GB LPDDR5X (16 channels)
  • N1 #1 – 8+4 CPU, 20 SMs (2,560 CUDA cores), 18-45 W TDP, 8-64 GB LPDDR5X (8 channels)
  • N1 #2 – 7+3 CPU, 16 SMs (2,048 CUDA cores), 18-45 W TDP, 8-64 GB LPDDR5X (8 channels)

These numbers tell us two things. First, NVIDIA wants a full scenario spectrum. The flagship variant at 80 W package power already carries CUDA core counts matching the RTX 5070, but inside a 14 mm thick ultrabook. Second, the N1 variant scaling down to 18 W directly targets the thin-and-light segment where Qualcomm Snapdragon X Elite currently dominates. As OC3D points out, the 48-SM Blackwell GPU matches the CUDA core count of NVIDIA's own desktop RTX 5070, though the much lower power budget means clock speeds will scale down accordingly. Still, it makes ray tracing and DLSS 4 viable on a laptop in ways we haven't seen before.

RTX Spark agentic PC vision

2.1. One Petaflop of Local AI Performance

The most-quoted spec is 1 petaflop of FP4 AI performance. To put that number in context, the 2023 Apple M3 Max delivered around 18 TFLOPS at FP16. FP4 is a quantized format, meaning lower precision, but it also means dramatically higher throughput at the same power budget. NVIDIA's secret sauce is that the new Blackwell Tensor Cores accelerate this format in hardware. The practical outcome: a 120-billion-parameter LLM agent can run locally thanks to 128 GB of unified memory. This is the threshold we predicted in last year's SLM trend piece being officially crossed: "running 100B+ models on edge devices."

3. The Agentic PC Vision

Beyond the silicon, the equally important story is software and OS integration. NVIDIA and Microsoft are positioning this chip not as a "fast laptop" but as an agentic PC: a machine that doesn't just run tools, but actively takes over routine tasks as a teammate.

Microsoft's "Windows security and containment primitives" unveiled at Build 2026 combine with NVIDIA's OpenShell runtime to enable personal agents that run safely inside a sandbox, under full user control. Examples on NVIDIA's official product page are bold: an esports streamer tells the PC to dim lights, mute the mic, and switch broadcast mode when stepping away. A designer asks Adobe to turn a sketch into a full image, render a 3D model of it, then create an AI video, all by voice. A developer hooks their GitHub repo to an agent that monitors QA issues and takes over the keyboard and mouse to handle "repetitive and boring" tasks autonomously.

This vision aligns with the philosophy behind Karpathy's Autoresearch: the human becomes the director, the agent becomes the executor. As NVIDIA's leadership put it, this is a "new personal computing paradigm where AI is the UX." You won't need to master complicated app UIs anymore. Just state your intent and the machine handles the rest.

4. The Windows-on-Arm Reality: Prism Emulation and Developer Support

Arm-based Windows machines have been treated as second-class citizens for years, mainly because x86 software had to run through an emulation layer with noticeable performance hits. Microsoft's Prism emulator has been closing that gap, and now NVIDIA brings Blackwell-class graphics into the fold.

The software ecosystem is improving fast. According to NVIDIA's official site, professional tools like Blender, DaVinci Resolve, Maxon Cinema 4D, Maxon Redshift, Topaz Photo, CapCut, Cubase, Bitwig Studio, and Affinity all run natively on Arm. Adobe prepared special optimizations for Premiere and Photoshop. Easy Anti-Cheat, BattlEye, and Denuvo are now Windows-on-Arm compatible, which opens the door for Valorant, League of Legends, and PUBG. Epic's Fortnite arrived last year.

For developers, the question is whether the toolchain is ready. CUDA, naturally, runs natively on RTX Spark, and CUDA remains the de facto standard for AI development. As we noted in our Mojo 1.0 Beta coverage, anyone trying to supercharge Python is already navigating the Arm ecosystem's nuances. RTX Spark extends that compatibility to the GPU side. Techniques like Multi-Token Prediction benefit heavily from low-precision formats (FP4), which favors NVIDIA's hardware-software stack.

5. Practical Implications for Developers and Creators

This chip could reshape the hidden costs of desktop AI development. 128 GB of unified memory means running 70B+ parameter open models in quantized form, or even 120B parameter agents at full precision, on a single machine. More importantly, prototyping, fine-tuning, and inference can all happen on the same device, with no data leaving the corporate network. This unlocks on-prem AI development for healthcare, legal, and finance, where data residency is non-negotiable.

On the creative side, NVIDIA highlights workloads like:

  • 90 GB 3D scene rendering – a scale that used to require a data center
  • 12K video editing – 4:2:2 hardware encode/decode
  • AV1 encoder and NVIDIA Broadcast for sharper streams
  • DLSS 4 and Ray Tracing for 100 FPS Indiana Jones and the Great Circle at 1440p

So this is not just an "AI chip". It's an all-in-one package designed to accelerate a content creator's entire daily workflow. The claim that you can do all of this in a 14 mm thick ultrabook without a power cord, if it holds up in real reviews, will seriously disrupt the mobile workstation market currently dominated by Apple's MacBook Pro M-series.

6. Market Impact: NVIDIA vs. Apple, Qualcomm, Intel, AMD

RTX Spark is not without competition. Apple M4 Max/Ultra remains the gold standard for unified memory architecture and performance-per-watt. Qualcomm Snapdragon X Elite 2 is a fierce rival in the thin-and-light category. Intel and AMD are pushing Lunar Lake and Strix Halo to claim a slice of the AI PC pie. Where does NVIDIA differentiate?

  1. AI acceleration leadership – Blackwell Tensor Cores are among the rare mobile solutions with native FP4 hardware support. Most rivals cap out at INT8 or FP8.
  2. CUDA ecosystem – The entire AI toolchain (PyTorch, TensorRT, vLLM, Ollama) is optimized for CUDA. By shipping hardware and software together, NVIDIA minimizes friction.
  3. 128 GB unified memory option – Apple's top tier reaches 192 GB, but NVIDIA's 128 GB will likely land in more aggressively priced machines. AMD's Strix Halo also offered 128 GB, but its GPU performance trails Blackwell.
  4. Deep partnership with Microsoft – First-party devices like Surface Laptop Ultra show how deep the optimization goes. Embedding "Personal AI" into Windows' UX fabric mirrors Apple's strategy with Apple Intelligence on macOS.

NVIDIA's short-term disadvantage is compatibility. Windows-on-Arm still emulates decades of x86 software. Unlike Apple's clean break, the existing game library is enormous, and porting each title will take years. NVIDIA is working with anti-cheat vendors to close that gap fast, but at launch some titles will likely still feel sluggish.

7. Conclusion: The Real Beginning of the Edge AI Era

The biggest story out of Computex 2026 isn't just a new chip. With RTX Spark, NVIDIA aims to eliminate cloud dependency, standardize local AI agents, and bring the "personal supercomputer" concept to mainstream consumers. Eight OEMs (Asus, Dell, HP, Lenovo, Microsoft, MSI, Acer, Gigabyte) have already announced over 30 laptops and more than 10 desktops. Microsoft's Surface Laptop Ultra is being marketed as "the most powerful thing we've ever made." This might be the most tangible expression of Microsoft's decade-long "Windows everywhere" strategy.

For us developers, the takeaway is clear: in the next 12-18 months, open-source models fine-tuned on your own laptop, autonomous agents running without a cloud API, and privacy-preserving local inference will all become commonplace. The lineer attention kernels we discussed in FlashQLA will make 100B+ models run at conversational speed on devices like RTX Spark. Newer open-source coding models like Kimi K2.6 will become dramatically more accessible with this kind of hardware.

RTX Spark has yet to prove itself in the field, but on paper it carries the biggest leap in personal computing we've seen in a decade. If NVIDIA's performance claims hold up, the sentence "I'll just hit the cloud for that AI task" will soon feel as outdated as "I'll just reboot the server."


This article was prepared with the assistance of the minimax-m3 model itself. The model's technical spec comparisons and architectural commentary were compiled from the official NVIDIA product page and The Verge's coverage. For deeper technical details, see VideoCardz's leaked slide analysis and Notebookcheck's summary.

Efe Hüseyin Özkan

Software Engineer & AI Developer

Working on AI systems, full-stack development, and scalable product architecture. Follow the blog for more technical articles.