memujo
AI7 min read

FreeToken: Running 753B Models on Consumer GPUs

FreeToken from UC Berkeley and MIT runs 753B MoE models on one consumer GPU by co-designing CPU-GPU execution around the real limits of local hardware.

By Alice

In this article
  1. 01What FreeToken actually does
  2. 02The memory math under the hood
  3. 03Benchmarks and hardware reach
  4. 04The roadmap beyond NVIDIA
  5. 05Our Read
  6. 06References

FreeToken is a new open-source inference engine that lets you run massive Mixture-of-Experts models on hardware most people already own. Developed by researchers from UC Berkeley and MIT, including Databricks co-founders Matei Zaharia and Ion Stoica alongside Song Han and Kurt Keutzer, the project treats your laptop or desktop not as a small datacenter node but as a unified elastic platform combining GPU, CPU, RAM and the PCIe interconnect between them.

The team released the project under Apache 2.0 on GitHub under the FlashML-org organization, and it has already attracted 12,600 stars and 1,200 forks since its initial release. A peer-reviewed paper describing the system appeared on arXiv in August 2026, and it is the primary source for all the technical claims below.

What FreeToken actually does

FreeToken targets one specific problem: serving frontier-scale open-weight MoE models on consumer hardware where GPU memory is limited and PCIe bandwidth is nowhere near NVLink speeds.

Sparse MoE architectures compute only a fraction of total parameters per token. DeepSeek-V4-Flash, for example, activates just 6 of 256 routed experts in each of its 43 layers, meaning only 13B of its 284B parameters participate in any single token. That active footprint fits inside the 32GB of an RTX 5090. But the full expert pool still exceeds GPU memory, forcing inactive experts to live in host RAM and stream over PCIe when needed. This is the same fundamental tension explored in the full DeepSeek-V4-Flash analysis, except that work assumed a datacenter cluster while FreeToken assumes a single PCIe slot. The full peer-reviewed paper behind FreeToken is available at arXiv:2608.16157.

In a datacenter, high-bandwidth interconnects hide this transfer cost. On consumer hardware, PCIe throughput of 16 to 64 GB/s plus host RAM latency creates severe decode bottlenecks. Existing edge runtimes like Ollama and llama.cpp rely on static expert offloading, where inactive weights sit in system RAM and stream synchronously to the GPU, stalling execution on every cache miss.

FreeToken replaces this rigid approach with a dynamic scheduling policy the paper terms the q* policy. Instead of halting the GPU during cache misses, it splits token computation between CPU cores and GPU tensor cores based on real-time interconnect throughput. The system uses a fast weight format alongside full-layer double buffering so weight streaming over PCIe overlaps entirely with active computation layers. An elastic memory manager also dynamically reallocates VRAM between KV cache entries and resident expert slots during runtime without triggering model reloads.

The memory math under the hood

What makes FreeToken structurally different from a static offloader is how it treats VRAM as a shared pool rather than a fixed partition. Most edge runtimes carve VRAM at startup: a slice for weights, a slice for KV cache, and the rest unreachable. FreeToken instead runs a runtime reallocator that shifts bytes between the KV cache and resident expert slots as the model moves through prefill and decode. For a coding agent, prefill is memory-hungry because the growing context inflates the KV cache, while decode is bandwidth-hungry because each step streams experts over PCIe. The same 32GB of VRAM can therefore serve both phases without the agent hitting a hard out-of-memory error.

That tradeoff has a real latency cost if managed poorly. On a Gen5 x16 slot running at roughly 128 GB/s, streaming a single 4.7GB expert tensor from host RAM takes on the order of 37 milliseconds round trip. A 35B model at 39 tok/s on an 8GB laptop GPU means roughly 25 milliseconds are already consumed by the compute itself, leaving a tight margin before PCIe transfers dominate the token budget. The q* policy exists precisely to keep those transfers below the compute floor by predicting the next cache miss and issuing the read in parallel with the current layer. The same dynamic VRAM tension appears elsewhere in the local-AI stack, as the HBM4 cost-curve piece showed for datacenter memory. FreeToken simply accepts the constraint instead of trying to defeat it.

Benchmarks and hardware reach

The paper reports concrete throughput numbers across three tiers of consumer hardware:

Model Parameters Hardware Speed
Qwen3.6-35B 35B (active) 8GB RTX 4060 laptop ~39 tok/s
DeepSeek-V4-Flash 284B (13B active) RTX 5090 desktop Interactive
GLM-5.2 753B Single workstation GPU Serving

The system currently supports 20+ MoE models including Qwen3.5, Qwen3.6, Qwen3.8, DeepSeek-V4-Flash, GLM-5.2 and Gemma-4 families. Installation is straightforward via uv or pip on Linux, with a desktop app available at flashml.ai for both Windows and Linux. Official support currently requires x86_64 hardware with an NVIDIA GPU from the Ampere generation onward, CUDA 13 and driver r580+.

FreeToken also provides Anthropic and OpenAI-compatible APIs, making it compatible with real-world coding and tool-using agents like Codex, Claude Code, OpenCode and OpenClaw. That integration detail matters more than many benchmarks: a model that runs locally but cannot plug into your agent harness has limited practical value. The same serving-interface problem is exactly what datacenter platforms like vLLM solved for servers, and FreeToken is attempting the equivalent for the edge.

The roadmap beyond NVIDIA

The project has published a 2026 roadmap that extends beyond its current NVIDIA-only launch. The planned additions include:

  • Apple Silicon through a native Metal backend, which would open FreeToken to machines with large unified-memory configurations
  • AMD GPUs through ROCm, with community contributors already reporting successful RDNA4 bring-up on Windows 11 using an RX 9070 XT
  • DGX Spark support through aarch64 wheels and sm_121 kernels
  • Multi-GPU tensor parallelism for spanning a single model across multiple GPUs in one workstation
  • Multimodal image input for vision-language models
  • Speculative decoding techniques including MTP, DFlash and DSpark to reduce decode latency

Those community AMD ports are not official upstream support yet. One report noted that MoE expert offload had not been fully tested on the ROCm fork, and another identified low-level host-memory and device-pointer problems on RDNA4. The roadmap is work in progress, and treating it as a shipping announcement would overstate the current state of the project.

Our Read

FreeToken matters because it solves a systems problem that has been quietly bottlenecking the open-weight movement: the gap between releasing model weights and actually being able to serve them at interactive speed on consumer hardware.

From a data science perspective, the q* policy is the most interesting contribution. The closed-form optimal split calculation per layer is mathematically clean, but the real question is whether those calculations accurately reflect real-world CPU dispatch latency, memory contention and varying expert residency under concurrent agent workloads. The Hacker News and LocalLLaMA discussions have been rightly skeptical about this point. The q* policy optimizes for PCIe bandwidth in isolation, but it does not yet model the second-order effects: CPU cache pressure when a parallel agent process competes for the same cores, page-fault stalls when host RAM is already near capacity, and the way expert routing itself can shift between requests and invalidate a previously computed schedule.

From an engineering standpoint, FreeToken's semantic anchor checkpointing addresses a real pain point for agentic workloads. Coding assistants constantly modify their context through tool calls, thinking blocks and external execution output. Standard engines discard linear KV caches when prefixes mutate, triggering costly full-sequence recomputations. FreeToken caches intermediate attention states at logical task boundaries so that editing intermediate tool arguments only requires recomputing the new suffix. For agents that make dozens of tool calls per session, this is a meaningful optimization. The design echoes the same tradeoff discussed in the DeepSeek-V4-Flash work, where context-length management was a first-class concern, but here the target is not a single long context but a mutating one.

The broader implication is worth noting. If FreeToken can maintain the same serving interface across CUDA, ROCm and Metal, it could become a portability layer for large local models across heterogeneous personal machines. That would be a significant shift from a specialized MoE runtime for RTX hardware to something closer to what datacenter platforms like vLLM and SGLang are for server deployments, but for the edge. The barrier is not the algorithm. It is the driver maturity and the fragmenting GPU market, where each vendor still speaks its own memory and scheduling language.

The current NVIDIA-only limitation and the unanswered question of whether the q* policy holds up under real-world agent workloads are the two biggest caveats. The community is already stress-testing both, and the project's open development model means those findings will be public. That is a feature, not a bug.

For anyone running frontier models locally, FreeToken is worth trying. It turns the machines you already own into a practical platform for frontier-scale intelligence, and the roadmap suggests it will become even more useful over the next few months.

References

  1. FreeToken paper: "FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution" by Shuo Yang, Xiaoze Fan, Melissa Pan, Haocheng Xi, Zhe Wang, Shanlin Sun, Kurt Keutzer, Song Han, Matei Zaharia, Chenfeng Xu, Ion Stoica. arXiv:2608.16157, August 2026. View paper
  2. GitHub repository: FlashML-org/FreeToken. View source
  3. AI Polix coverage of FreeToken roadmap. Roadmap analysis
  • #moe
  • #inference
  • #local-ai
  • #edge-computing
  • #open-source

Sources

Share this story