For most of the past year, running a large mixture-of-experts model at home meant one of two things: renting cloud time, or buying a rack of data-centre cards that cost more than a family car. A new open-source project called Strata is trying to collapse that gap. It runs Qwen3.8-Flash-Next, a 125-billion-parameter mixture-of-experts model, on a single consumer gaming GPU with as little as 12GB of video memory, paired with 64GB of ordinary system RAM.

The trick is in how it treats the model's weights. A mixture-of-experts model does not fire all of its parameters on every token. Only a handful of the model's many expert sub-networks are active at any moment. Strata exploits that by splitting one checkpoint across three tiers of storage: the GPU holds the experts that come up most often, system RAM holds the rest, and read-only NVMe storage backs the whole thing when the model is too big to sit in memory at once. An adaptive cache keeps the hot experts on the card while the processor computes the colder ones in parallel, using the AVX-512 and AVX2 instruction sets built into modern CPUs.

Fast enough to use

The numbers people are reporting are the surprising part. On an RTX 3090, a card that is now several years old, users describe roughly 70 tokens per second with a 128,000-token context using an aggressive IQ2_XS quantisation. That is comfortably past reading speed and well into territory where the model is pleasant to work with rather than something you set running and walk away from. Written in C++20 and CUDA, Strata serves an API on localhost that mimics both OpenAI's and Anthropic's formats, so it drops straight into tools and agents already built for those services.

The democratisation question, again

Strata lands in a run of releases pushing frontier-scale capability onto hardware people already own. It follows efforts like DeepSeek's open toolkit for non-Nvidia chips and the steady stream of large open-weight models such as NaiveAI's 309-billion-parameter release. Each chips away at the assumption that serious models can only live in a data centre.

There are trade-offs, and they are real. Heavy quantisation shrinks a model's precision, and a two-bit quant of a 125B model is not the same thing as the full-weight version running on proper hardware. Splitting a checkpoint across disk and memory adds latency that a single fast card would not. But for researchers, tinkerers and small shops that cannot justify cloud bills or data-centre GPUs, the pitch is hard to ignore: a model this size, running on the machine already under the desk, with the weights and the code open for anyone to inspect. That combination keeps moving the floor of what local AI can do.

Sources

  1. i. alphasignal.ai
  2. ii. github.com

Commentarii · 0

Add · a · Comment