Strata: A 125B Model on a 12GB Gaming Card, at Reading Speed
A 125-billion-parameter model on a gaming PC used to be a party trick that ran at a few tokens per second. Strata makes it usable. The open source engine runs Qwen3.8-Flash-Next on any NVIDIA or AMD card with 12GB of VRAM or more, Windows or Linux, and hit 508 points on Hacker News on Sunday. The repo is at 10.7k stars after ten days.
The measured numbers are on ordinary hardware. An RTX 5070 with 12GB, a Ryzen 5 7600 and 64GB of RAM writes 94 tokens per second at the Q2_0 quant and reads prompts at 2,650 tokens per second on a 32K document. At IQ3_S it's 53 tokens per second. An AMD RX 9070 XT gets 60. The project estimates 100 to 140 tokens per second on a 24GB RTX 3090. You need 32GB of system RAM at minimum, 64GB runs every size, and about 80GB of disk.
The agent angle is in the install section. Strata serves both OpenAI- and Anthropic-compatible APIs on localhost, so Claude Code, Cursor or Codex can point at it directly. The recommended install is to paste one line into your coding agent and let it follow an AI_SETUP.md doc that checks your GPU, RAM and disk and picks the quant. There's also an MCP server so agents can start and stop the model themselves.
The trade-off is plain. A 2-bit or 3-bit quant of a 125B model is not the model Qwen benchmarked, and the README doesn't claim it is. But "frontier-adjacent coding model, private, on the card you already own, fast enough that the agent loop doesn't stall" is a new price point for local agents. The installer written for agents to run, not humans, is the detail to copy.
Link: github.com/Niko1221/Strata
← Back to all articles
The measured numbers are on ordinary hardware. An RTX 5070 with 12GB, a Ryzen 5 7600 and 64GB of RAM writes 94 tokens per second at the Q2_0 quant and reads prompts at 2,650 tokens per second on a 32K document. At IQ3_S it's 53 tokens per second. An AMD RX 9070 XT gets 60. The project estimates 100 to 140 tokens per second on a 24GB RTX 3090. You need 32GB of system RAM at minimum, 64GB runs every size, and about 80GB of disk.
The agent angle is in the install section. Strata serves both OpenAI- and Anthropic-compatible APIs on localhost, so Claude Code, Cursor or Codex can point at it directly. The recommended install is to paste one line into your coding agent and let it follow an AI_SETUP.md doc that checks your GPU, RAM and disk and picks the quant. There's also an MCP server so agents can start and stop the model themselves.
The trade-off is plain. A 2-bit or 3-bit quant of a 125B model is not the model Qwen benchmarked, and the README doesn't claim it is. But "frontier-adjacent coding model, private, on the card you already own, fast enough that the agent loop doesn't stall" is a new price point for local agents. The installer written for agents to run, not humans, is the detail to copy.
Link: github.com/Niko1221/Strata
Comments