MikeTrendsTrends right now

⬢github C++ · 932 ★ +216 since we first saw it · pushed 27 min ago · MIT

Niko1221/Strata

Qwen3.8-Flash-Next (125B MoE) on a 8GB+ NVIDIA GPU: one-click install for Windows / Linux. Strata inference engine, OpenAI/Anthropic API on localhost, optional image input.

Strata is a one-click installer and C++ inference engine that runs the 125B-parameter MoE model Qwen3.8-Flash-Next on a single NVIDIA GPU with 12GB+ VRAM plus 64GB of system RAM, on Windows or Linux. It serves an OpenAI/Anthropic-compatible API on localhost, supports optional image input, and quantized variants trade quality for speed, reaching roughly 40-95 tokens per second.

Why now: It is trending on GitHub's most-starred new repos because it makes a huge model that normally needs a server runnable on a normal gaming PC with surprisingly fast speeds.

Who it is for: PC gamers and local-AI enthusiasts with an RTX 30/40/50 card and enough RAM who want to run large LLMs locally without a server.

llminferencelocalgpuquantizationc++

Open on GitHub →

Stars over our 22 snapshots: 716 to 932, since 5 h ago.

Where people talked about it

API: https://socialmediatrends-api.osmike.com/v1/repos/Niko1221/Strata