Surface Laptop Ultra and RTX Spark: Can It Run Local AI

The first Windows laptop on Nvidia's RTX Spark has a price and ship date. Here is what fits in 128GB of unified memory - and what it costs you.

Microsoft has priced and dated the Surface Laptop Ultra - the first shipping Windows laptop built on Nvidia’s RTX Spark Arm chip - and the configuration that matters for local AI starts at $2,599. Preorders opened on October 7 and units ship October 16, so anyone deciding whether to wait for this machine instead of buying a Mac Mini or a used RTX 4090 has about eight days to make up their mind (The Verge; Ars Technica).

The Hardware

The Surface Laptop Ultra is built around Nvidia’s RTX Spark, a Blackwell-architecture system-on-chip that Nvidia describes as “the world’s first Windows PCs purpose-built for personal agents” (NVIDIA Newsroom). The top-end RTX Spark pairs a Blackwell RTX GPU with 6,144 CUDA cores and fifth-generation Tensor Cores with FP4 precision, a 20-core Grace CPU, and up to 128GB of unified LPDDR5x memory, fused by NVLink (NVIDIA Newsroom; NVIDIA Blog).

Surface Laptop Ultra buyers do not get the full 128GB SKU on the cheap. The $2,599 entry model comes with an 8-core CPU, 24GB of RAM, and 512GB of storage (The Verge). Ars reports that higher configurations include either a 5,120-core or 6,144-core Blackwell GPU and top out at 128GB of unified memory - the variable that decides what fits on-device (Ars Technica). The Verge’s event roundup, listing a $2,599 base and an October 16 ship date, cross-checks the higher-end numbers reported by Ars.

Microsoft also revealed a second machine: the Surface RTX Spark Dev Box, a $5,999 mini-PC “designed to meet the needs of frontier developers.” It carries the 128GB unified-memory configuration and “claims one petaflop (that’s 1,000 teraflops) of AI compute performance,” per Ars Technica (Ars Technica). The Verge lists the same $5,999 starting price and says the Dev Box is available for preorder ahead of a November ship date (The Verge). TechCrunch rounds the figure slightly differently, at “starts priced at $6,000” with the same November shipping window (TechCrunch).

What It Actually Runs

The RTX Spark exists so that the entire local-AI software stack - llama.cpp, Ollama, LM Studio, vLLM, the DeepSeek model family, Nvidia’s own Nemotron family - can hit the same accelerator silicon that Nvidia already sells for data-center inference. Microsoft is shipping Surface RTX Spark machines with “native support for local frameworks like llama.cpp and models from the Deepseek and Nvidia Nemotron families,” per Ars (Ars Technica). That is the most practical answer for readers wondering whether ChatGPT-style agents can finally run offline: yes, with the right model size and the right RTX Spark SKU.

Because the GPU and CPU share the same LPDDR5x pool, the strict VRAM limits of a discrete Nvidia card do not apply. The Dev Box’s 128GB lets a quantized 70B model sit entirely in memory with headroom for context; smaller 7B and 13B models become essentially free to run. Ars notes that the unified memory design is what let Microsoft show Gears of War: E-Day running on the Surface Laptop Ultra - “even without a traditional gaming GPU, AAA titles are on the table” (Ars Technica). The same property makes a 30B-class reasoning model realistic on the $2,599 base config - a noticeable jump over the 8GB VRAM floor that bottlenecks most consumer Nvidia cards.

The pitch is broader than inference. Microsoft is repositioning Windows 11 around what it calls “Hybrid Intelligence,” a term that includes a rebuilt Copilot split into Home, Code, and Autopilot tabs. Autopilot is the agent-oversight surface; Code is meant to let non-developers prompt their way to widgets and Windows apps; Home is the conversational front door (Ars Technica). Satya Nadella made the framing explicit at the San Francisco event. Per TechCrunch’s Julie Bort, he told the audience: “One of the things we realized in the last 3-4 years is that just having a model doesn’t do much for anything. You really do need to orchestrate, and you need to have memory outside of the model. You need to have this harnessed layer that is able to take multiple models plus context, plus memory, and the action space” (TechCrunch).

The Trade-Offs

Two things to weigh before pre-ordering. First, the price ladder versus alternatives. A $2,599 Surface Laptop Ultra with 24GB is not a workstation; it is a portable RTX 5060-class machine in a chassis, and the Dev Box at $5,999 is what you actually want for sustained 70B inference. For the same money you can buy a desktop RTX 4090 plus a used workstation, or a Mac Studio with more raw memory bandwidth for LLM throughput. The Surface value is portability plus a single-vendor guarantee that llama.cpp and DeepSeek checkpoints work on day one.

Second, the OS cost. Microsoft paired the hardware reveal with “Execution Containers” (MXC), a policy layer for sandboxing AI agents so they cannot reach files, network, or capabilities outside a declared scope. Ars reports it can route a workload to a process container, a session container, a WSL container, or an experimental hardware-backed MicroVM (Ars Technica). That is a real privacy improvement for an agent beat where recent weeks have produced jailbreaks, rogue Wikipedia editors, and abliterated-model backdoors. The catch: the sandbox only helps if the agent you want to run actually opts in to MXC. The Verge and TechCrunch confirm Execution Containers “will be available for all Windows 11 users,” not just Surface buyers (The Verge; TechCrunch).

What This Means

For local-AI readers, RTX Spark is the first Windows-native answer to the Apple Silicon question. Strix Halo and the Mac Studio are faster for many model loads because of memory bandwidth, but they live in Apple’s walled garden. RTX Spark is a Blackwell-class accelerator with the full Nvidia software stack - CUDA, RTX, TensorRT, FP4, NVLink-C2C - inside a Windows laptop you can carry to a coffee shop. Jensen Huang’s framing in the Nvidia Newsroom release is that RTX Spark brings “everything NVIDIA has built - CUDA, RTX, our AI platform - into a single superchip” (NVIDIA Newsroom). The llama.cpp team, Adobe, ComfyUI, Blackmagic Design, OTOY and others are quoted in the same release as already building against the platform; if their pre-launch commitments hold, llama.cpp, Ollama, and Hugging Face’s transformers ecosystem should follow the same drop-in path the DGX Spark already proved out.

For privacy readers, the bigger story is Execution Containers. The site has tracked a steady drumbeat of agent jailbreaks and supply-chain compromises over the last year; an OS-level containment primitive that ship with every Windows 11 PC is more consequential than any one laptop, because it changes the threat model for the millions of existing users who will never buy a $2,599 Surface.

For hardware buyers: do not pre-order the $2,599 model and expect it to replace a desktop. Wait for independent 8-core vs. 20-core benchmarks once the 16th passes, especially for quantized 70B inference, where memory bandwidth and CPU offload matter as much as CUDA cores.

The Bottom Line

The Surface Laptop Ultra is the first widely shipped Windows laptop with the silicon headroom to run modern open-weight models locally, and the Dev Box is a fair-priced developer kit if the 128GB SKU is what you need. Hold off on the base model until RTX Spark memory-bandwidth numbers and real 70B-inference benchmarks land.

  • Local AI - more coverage of self-hosted models, hardware picks, and licensing.