⬢github C++ · 847 ★ +1 since we first saw it · pushed 5 h ago · Apache-2.0
incoai/splash
A local inference engine for Apple silicon, built around the model.
Splash is a local LLM inference engine for Apple silicon Macs, written in C++ with Metal kernels and DFlash 2 speculative decoding. It runs coding agents and OpenAI/Anthropic-compatible apps on one machine, with vision, tool calling, a built-in chat page, prefix caching, automatic memory planning, and optional KV cache offload to SSD.
Why now: It's trending on GitHub's most-starred new repos shortly after release, and its README references a fresh DFlash 2 blog post and an LM Studio integration guide, suggesting recent ecosystem buzz.
Who it is for: Mac owners with M3 or newer chips (ideally 24–48 GB+ of unified memory) who want to run coding agents or LLM apps locally without the cloud.
apple-siliconcoding-agentsllm-inferencemacosmetalspeculative-decoding
Stars over our 9 snapshots: 846 to 847, since 1 h ago.
Where people talked about it
- ⬢github new repos, most starred 1 min ago
API: https://socialmediatrends-api.osmike.com/v1/repos/incoai/splash