NVIDIA RTX Spark: NVIDIA Redefines the AI PC, 128G Unified Memory Runs a 120B Model Locally

·Toolin Editorial Team

NVIDIA has released the RTX Spark consumer AI chip: 128GB of unified memory and 1 PFLOP of compute, able to run a 120B large model locally on a 14mm laptop — the Windows ecosystem enters the AI PC era

NVIDIA RTX Spark: NVIDIA Redefines the AI PC, 128G Unified Memory Runs a 120B Model Locally

At NVIDIA GTC Taipei 2026, NVIDIA unveiled RTX Spark -- a consumer-grade AI chip that gives a Windows PC a unified memory architecture for the first time. It means you can run a 120B-parameter large model locally on a 14mm-thick laptop.

NVIDIA calls this "a redefinition 40 years in the making, since the birth of the personal computer."

RTX Spark chip

The RTX Spark chip in the flesh

What Is RTX Spark

RTX Spark grew out of last year's developer-focused DGX Spark, and has now been officially upgraded into NVIDIA's brand-new consumer product line. Under the hood it uses the same GB10 chip as DGX Spark, with these key specs:

  • AI compute: Up to 1 PFLOP (FP4 precision)
  • CPU cores: 20
  • GPU cores: 6144
  • Unified memory: 128GB LPDDR5X

RTX Spark specs

Flagship specs: 1 PFLOP AI performance, 128GB unified memory

Why Unified Memory Is Critical for Running Large Models

The Memory Dilemma of Traditional PCs

In a traditional PC, the CPU has its own system memory (RAM) and the GPU has its own video memory (VRAM), connected through a PCIe lane. The problem:

  • The GPU reading its own video memory gets roughly 1 TB/s of bandwidth
  • The PCIe lane's bandwidth is only about 32 GB/s — a 30x gap

When you want to run a quantized 70B model locally (which needs tens of GB of memory), even if your system has 64GB of RAM, the GPU can actually use at high speed only its 16GB of VRAM. If the model is too large, data has to be shuttled back and forth over PCIe, and speed is severely limited.

Unified memory: CPU and GPU share one memory pool, eliminating the data-transfer bottleneck

The Unified Memory Solution

RTX Spark turns the CPU's and GPU's memory into a single 128GB shared pool. The GPU can directly use the vast majority of this large pool, no longer constrained by the 16GB, 24GB, or 32GB VRAM of traditional graphics cards.

This means you can run a 120B-parameter large model directly on your local machine — no cloud inference needed, extremely low latency, and data kept entirely local.

Why Not Just Buy a Mac

Macs do come in 128GB unified-memory configurations, but RTX Spark has one killer advantage a Mac cannot replace: the CUDA ecosystem.

CUDA is not just a graphics driver — it is an entire GPU computing ecosystem built up over nearly 20 years. The vast majority of AI frameworks, inference engines, and training tools are built on CUDA. Running these tools on a Mac means compromises in both compatibility and performance.

RTX Spark = unified memory + the full CUDA ecosystem, a first on consumer devices.

Partner Hardware Form Factors

NVIDIA showed off devices from multiple partners built on RTX Spark:

Ultrathin laptops

  • Just 14mm thick
  • Can render a 90GB 3D scene while unplugged
  • Can edit 12K-resolution video

RTX Spark partner devices

The RTX Spark ultrathin laptop and mini desktop

Mini desktops

  • A small box similar to a Mac Mini
  • Low power draw, well suited as a home AI server

Native AI Support in Windows

Microsoft will partner with NVIDIA to comprehensively rearchitect Windows so that RTX Spark machines natively support local Agent execution. This means:

  • Agents can directly tap local GPU compute
  • Complex AI tasks can run without an internet connection
  • Data and privacy stay protected locally

Application Scenarios

  • Local large-model inference: Run a 120B model directly on a laptop, no cloud API needed
  • Local model fine-tuning: Use the 128GB unified memory to personalize fine-tune models
  • 3D rendering and video editing: 90GB 3D scene rendering, 12K video editing
  • Privacy-sensitive scenarios: Data in healthcare, finance, legal, and similar fields never leaves the machine
  • Development and testing: Iterate on Agent apps locally at speed, without depending on cloud services

Expected Impact

RTX Spark's launch will most likely bring several changes:

  1. A Windows replacement wave: RTX Spark devices are expected to go on sale in 2027 and could trigger large-scale upgrades
  2. Local AI goes mainstream: Large models will no longer be cloud-only; ordinary developers can deploy and debug locally too
  3. A new Mac vs Windows rivalry: For the first time, the Windows camp has a hardware foundation that can stand up to the Mac in local AI inference
  4. The CUDA ecosystem gets further entrenched: Unified memory + CUDA deepens NVIDIA's moat in consumer AI hardware

Frequently Asked Questions

Q: When can I buy one? A: Partner devices are expected to roll out over the course of 2027; exact timing awaits vendor announcements.

Q: Roughly how much will it cost? A: NVIDIA has not yet announced standalone pricing for the RTX Spark chip, but going by the DGX Spark's positioning, equipped devices are expected to land in the premium thin-and-light price band.

Q: Which models can it run? A: 128GB of unified memory can run 120B-parameter models directly, and quantization allows even larger ones. Mainstream open-source model families such as Llama, Qwen, and DeepSeek can all run locally.

Q: Does it need an internet connection? A: Local inference and fine-tuning need no internet. You only need a network when downloading models and syncing data.

Related articles

Google Workspace CLI: One Command Lets AI Agents Take Over Your Mail, Drive, and Calendar
AI Products

Google Workspace CLI: One Command Lets AI Agents Take Over Your Mail, Drive, and Calendar

An open-source CLI with nearly 30K stars wraps Gmail, Drive, Calendar, and Sheets into one unified interface, ships 100+ Agent Skills built in, and works for humans and AI alike.

Toolin Editorial Team
Deemos Hyper3D Rodin Gen-2.5: 3D Generation Enters Its Thinking Era
AI Products

Deemos Hyper3D Rodin Gen-2.5: 3D Generation Enters Its Thinking Era

The world's first 10M-polygon-class 3D generation model: 1M polygons in 4 seconds, 12K high-res textures, five switchable thinking depths, with B2B revenue exceeding the rest of its niche combined.

Toolin Editorial Team
Xiaomi MiMo-V2.5-Pro-UltraSpeed: A 1T-Parameter Model Running at 1000 tokens/s
AI Products

Xiaomi MiMo-V2.5-Pro-UltraSpeed: A 1T-Parameter Model Running at 1000 tokens/s

Xiaomi's flagship ultra-speed inference model breaks 1000 tokens/s of output on a standard 8-GPU commodity node, with output priced at 18 RMB per million tokens.

Toolin Editorial Team
Google Ships Two Lightweight Creative Models: Images in 4 Seconds, Video in 10
AI Products

Google Ships Two Lightweight Creative Models: Images in 4 Seconds, Video in 10

Google releases the Nano Banana 2 Lite image model and the Gemini Omni Flash video model, built for speed and savings — $0.034 per image and $0.10 per second of video, already embedded in the Gemini App and YouTube Shorts for free.

Toolin Editorial Team
Claude Tag Turns Claude Into a Team AI Colleague That Lives in Slack
AI Products

Claude Tag Turns Claude Into a Team AI Colleague That Lives in Slack

Anthropic releases Claude Tag, evolving Claude Code into a team collaboration agent with shared context, persistent memory, and proactive intervention, now in Beta for Enterprise and Team users.

Toolin Editorial Team
Doubao Seed 2.1 Pro Hands-On: Coding Crosses the Usability Line, and Multimodal Recognition Holds a Surprise
AI Products

Doubao Seed 2.1 Pro Hands-On: Coding Crosses the Usability Line, and Multimodal Recognition Holds a Surprise

Volcengine launches the Doubao Seed-2.1 series: the Pro tier reaches production-grade usability for coding and agent calls, photo-based fish recognition outperforms Gemini 3.1 Flash, with 6 tested prompts included.

Toolin Editorial Team