⬢github C++ · 932 ★ +216 since we first saw it · pushed 27 min ago · MIT
Niko1221/Strata
Qwen3.8-Flash-Next (125B MoE) on a 8GB+ NVIDIA GPU: one-click install for Windows / Linux. Strata inference engine, OpenAI/Anthropic API on localhost, optional image input.
Strata is a one-click installer and C++ inference engine that runs the 125B-parameter MoE model Qwen3.8-Flash-Next on a single NVIDIA GPU with 12GB+ VRAM plus 64GB of system RAM, on Windows or Linux. It serves an OpenAI/Anthropic-compatible API on localhost, supports optional image input, and quantized variants trade quality for speed, reaching roughly 40-95 tokens per second.
Why now: It is trending on GitHub's most-starred new repos because it makes a huge model that normally needs a server runnable on a normal gaming PC with surprisingly fast speeds.
Who it is for: PC gamers and local-AI enthusiasts with an RTX 30/40/50 card and enough RAM who want to run large LLMs locally without a server.
Stars over our 22 snapshots: 716 to 932, since 5 h ago.
Where people talked about it
- ⬢github new repos, most starred 5 min ago
API: https://socialmediatrends-api.osmike.com/v1/repos/Niko1221/Strata