DocsQuick StartAI News
AI NewsRun a 125B model on 12GB of VRAM: Strata takes off.
Industry News

Run a 125B model on 12GB of VRAM: Strata takes off.

2026-10-06T11:06:35.399Z
Run a 125B model on 12GB of VRAM: Strata takes off.

The open-source inference engine Strata enables consumer-grade GPUs with 12 GB of VRAM to run the quantized Qwen3.8-Flash-Next using expert caching, heterogeneous memory coordination, and speculative decoding. In testing, the RTX 5070 reached up to 94 tokens per second, but this speed advantage relies on low-bit quantization and model-specific optimizations.

Running a 125B Model on 12GB of VRAM: Strata Takes Off

A model with 125 billion parameters has begun appearing on gaming GPUs with just 12GB of VRAM.

On October 6, the open-source inference engine Strata attracted attention from the developer community. Open-sourced by developer Niko1221, it has a clear objective: to run the quantized version of Qwen3.8-Flash-Next on ordinary consumer hardware instead of restricting model inference to data-center GPUs.

According to the project's disclosed test results, on a system equipped with an RTX 5070, Ryzen 5 7600, and 64GB of memory, the Q2_0 quantized version reaches a generation speed of up to 94 tokens per second, while the IQ3_S version reaches 53 tokens per second. In other words, a graphics card with only 12GB of VRAM can now handle inference tasks that in the past typically required server-grade memory capacity.

But what really deserves attention here is not the headline "running a 125B model on 12GB of VRAM" itself. It is that Strata combines the sparsity of MoE models with CPU memory and SSD storage, rearranging the role of each type of hardware during inference.

Architecture diagram showing an RTX 5070 consumer graphics card, system memory, and SSD working together to run a 125B MoE model

It Does Not Force the 125B Model Into VRAM

Qwen3.8-Flash-Next is a compressed version of the multimodal mixture-of-experts model launched by the Qwen team in August this year. It has a total parameter count of 125B, includes an n-gram embedding table of approximately 51B parameters, and natively supports a context length of 262K tokens.

Here, "125B" does not mean that all 125 billion parameters must be computed every time a token is generated. MoE models typically divide the network into a large number of expert modules, with each token activating only a small subset of them. Community analyses of Qwen3.8-Flash-Next indicate that the model has approximately 24,576 experts, but each token actually invokes only around 10 of them, resulting in an active parameter count of roughly 6B.

This is the prerequisite that makes Strata possible.

When traditional inference frameworks perform CPU offloading, they often move the model layer by layer: the current layer is transferred from system memory to VRAM, swapped out after computation completes, and then the next layer is processed. The model can run, but the PCIe bus and memory bandwidth quickly become bottlenecks. For MoE models with enormous parameter counts, this approach wastes a great deal of time moving experts that are not actually activated.

Strata takes a different approach:

  • Keep the complete set of experts in system memory;
  • Cache the most frequently used experts in GPU VRAM according to usage frequency;
  • Have the CPU participate in computation for experts that are not found in the VRAM cache;
  • Store the relatively large n-gram embedding table on the SSD and read the required data when needed;
  • Use a lightweight model to predict subsequent tokens, then have the main model verify the predictions.

This can be understood as a multilevel caching system designed for MoE models. VRAM acts as the high-speed cache, system memory as main memory, and the SSD as a larger but slower storage layer. It does not eliminate the model's parameter count; it simply avoids having all parameters occupy the most expensive VRAM resource at the same time.

What Conditions Enable 94 Tokens per Second?

The figure most widely circulated so far is 94 tokens per second on an RTX 5070 with 12GB of VRAM. However, this figure corresponds to the Q2_0 quantized version and cannot be directly compared with the inference speed or capabilities of a high-precision model.

On the same NVIDIA hardware, Strata reports roughly the following results for different quantization versions:

| Quantization version | Generation speed | Prompt processing speed | | --- | ---: | ---: | | Q2_0 | 94 tokens/s | 2650 tokens/s | | IQ2_XS | 79 tokens/s | 2090 tokens/s | | IQ3_XXS | 62 tokens/s | 1750 tokens/s | | IQ3_S | 53 tokens/s | 1620 tokens/s | | Coder | 55 tokens/s | 2180 tokens/s |

Lower quantization precision generally means smaller weight footprints and lower memory bandwidth pressure, but it also means more noticeable information loss. Q2_0's speed is impressive, but it should not be treated as a direct representation of the model's original capabilities. For tasks such as code generation, complex reasoning, and long-text consistency, IQ3_S or other higher-quality quantized versions may be more reliable, although their execution speed will be lower.

AMD platforms are not entirely absent either. In a test environment using a Radeon RX 9070 XT with 16GB of VRAM, the Q2_0 version achieved a generation speed of approximately 60 tokens per second. Community tests have also shown that some AMD ROCm environments can produce higher or different results, but these figures are affected by driver versions, context length, quantization format, and compilation parameters, so they cannot be compared directly.

More importantly, Strata uses speculative decoding. It first uses a lightweight model to predict several subsequent tokens, then has the main model verify them all at once. When the predictions are correct, the main model does not need to perform the complete computation for each token individually, producing an additional speedup. Estimates disclosed by the community suggest that speculative decoding can provide an improvement of approximately 1.6 to 1.8 times.

This is not the same as simply saying that "the GPU computes faster." Strata's speed comes from several conditions working together: MoE's sparse activation, expert cache hit rates, the lower bandwidth requirements of quantization, and the prediction accuracy of speculative decoding. Changing the graphics card, using a different quantization method, or extending the context to its limit can all produce significantly different results.

The System Requirement Is Not the GPU, but Memory

Strata's minimum hardware requirements are approximately:

  • An NVIDIA or AMD graphics card with at least 12GB of VRAM;
  • At least 32GB of system memory, with 64GB recommended for the full version;
  • At least 80GB of storage space, preferably on an SSD;
  • Windows 10/11 or Linux.

This means that it is not a matter of being able to run the model unconditionally as long as you have a 12GB graphics card. The graphics card is merely the entry point; the actual experience is often determined by system memory and storage.

During initial loading, the project may need to map or lock tens of gigabytes of data in memory. When memory is insufficient, the operating system will frequently swap pages, and inference speed can quickly drop from "interactive" to "practically unusable." With only 32GB of memory, users will typically need to choose the more streamlined Coder version or accept greater loading and runtime pressure.

The SSD is not an optional component either. Qwen3.8-Flash-Next's n-gram embedding table is very large, and Strata places some of the data on the SSD for on-demand access. A mechanical hard drive can provide the capacity, but it is unlikely to deliver a stable interactive experience. An NVMe SSD is better suited to serving as the slower cache layer in this type of inference pipeline.

Therefore, Strata's real threshold is closer to "a modern PC with 12GB of VRAM and 64GB of memory" than "12GB of VRAM is enough." The two claims differ little in terms of how they spread, but they differ substantially when it comes to building a system and using it in practice.

It Is More Like a Specialized Runtime Than a General-Purpose Inference Framework

Strata's strengths are also its boundaries.

At present, it is primarily optimized around Qwen3.8-Flash-Next. It provides installation methods for Windows and Linux, and can launch OpenAI-compatible and Anthropic-compatible APIs locally, making it easier to integrate with local coding agents, editors, and other tools. For developers who want to integrate the model into workflows involving tools such as Claude Code or Cursor, this is considerably more practical than opening a standalone chat interface.

However, it should not simply be understood as the next general-purpose version of llama.cpp or vLLM. The latter two need to cover a large number of model architectures, hardware backends, and deployment scenarios. Strata is more like a specialized runtime for a specific model: it knows which experts are used more frequently, which weights are suitable for remaining resident in VRAM, and how to schedule around specific quantization formats and speculative decoding paths.

This creates three practical limitations:

  1. Limited model compatibility. Switching to another architecture may involve different expert layouts, routing methods, and weight formats, meaning the existing caching strategy may no longer apply.
  2. Performance is highly workload-dependent. Short conversations, repetitive tasks, and workloads that hit frequently used experts are more likely to achieve high throughput. Extremely long contexts, complex multi-turn reasoning, or tasks with significant changes in expert distribution may run more slowly.
  3. The quality loss from low-bit quantization cannot be ignored. 94 tokens per second is an engineering optimization result, not a measure of model capability. When evaluating output quality, the specific quantization version and test task must be stated.

Community tests show that on an RTX 5070, short-conversation generation speeds are approximately 52 to 90 tokens per second. When the context reaches 128K, speeds may still remain between 41 and 67 tokens per second. For comparison, tests of the same model using native llama.cpp show a speed of approximately 20 tokens per second. However, these results currently come mainly from project maintainers and community tests. They are not yet sufficient to replace rigorous evaluations using a unified test set, standardized prompts, and consistent context lengths.

The Value of Local Deployment Is Not Just Saving on API Costs

If viewed solely in terms of price, local deployment may not be worthwhile. Model API prices continue to decline, and the number of tokens used by an ordinary developer each month is often insufficient to offset the cost of buying a graphics card, memory, and an SSD. Upgrading an entire system to save a few dollars in API fees is generally not a rational calculation.

Local inference is genuinely attractive mainly in four scenarios:

  • Privacy and data boundaries. Company code, internal documents, medical records, and personal notes do not need to leave the local machine.
  • Offline and intranet environments. When there is no public internet, external APIs cannot be accessed, or data-export requirements apply, local models are more reliable.
  • Long-running agents. A local OpenAI-compatible API can serve as the backend for a coding agent, avoiding frequent dependence on cloud quotas, rate limits, and network instability.
  • Long-context experiments. Qwen3.8-Flash-Next natively supports a 262K-token context, allowing users with sufficient memory to explore local long-context workflows.

If the requirement is stable production service, high-concurrency requests, or strict quality consistency, Strata is clearly not a replacement for cloud inference infrastructure at present. It is better suited to personal development environments, laboratories, intranet tools, and engineers interested in hardware.

For developers in China, local execution and cloud APIs are not mutually exclusive. When rapid access to different models such as GPT, Claude, Gemini, and DeepSeek is needed, an aggregation service compatible with the OpenAI format, such as OpenAI Hub, can be used. For sensitive code, offline tasks, or long-running local agents, part of the workload can then be moved to Strata. The former addresses model and network access, while the latter keeps data on the local machine. They solve different problems.

The Next Step for Open-Source Inference May Be "Assembling the Hardware"

Strata's significance is not that it proves consumer graphics cards are already sufficient to replace high-end GPUs. Rather, it demonstrates an increasingly realistic path: inference is no longer determined solely by VRAM capacity, but by whether the GPU, CPU, system memory, SSD, and software scheduler can work together.

In the past, model deployment was often reduced to a single question: can the model weights fit entirely into VRAM? The emergence of MoE models has changed that question. Since each token activates only a small subset of experts, a more sensible approach is not to put all experts on the GPU, but to tier them according to access frequency and have different hardware handle different types of work.

This approach is not new, but Strata has advanced it to a level that is tangible on consumer hardware. Running a 125B model on 12GB of VRAM sounds like an extreme demonstration. Its real value is that it prompts developers to rethink the hardware composition of "local large models": the graphics card handles hot computation, memory handles capacity, the SSD handles cold data, the CPU provides fallback computation, and the lightweight model reduces the number of times the main model must be invoked.

Of course, Strata is still evolving rapidly, and project versions, model weights, drivers, and quantization implementations can all affect the results. Developers who want to try it should treat it as an experimental runtime for a specific model rather than as a mature general-purpose production solution.

But at least as of October 2026, Strata has put an important fact on the table: the bottleneck for local large models may not always be insufficient VRAM. It may simply be that we have not yet organized the hardware we already have.

Sources

Related Articles

View All

Contact Us

We usually reply quickly during business hours

Scan WeChat

Support: Hub Assistant

WeChat ID: