MikeTrendsTrends right now

⬢github C++ · 847 ★ +1 since we first saw it · pushed 5 h ago · Apache-2.0

incoai/splash

A local inference engine for Apple silicon, built around the model.

Splash is a local LLM inference engine for Apple silicon Macs, written in C++ with Metal kernels and DFlash 2 speculative decoding. It runs coding agents and OpenAI/Anthropic-compatible apps on one machine, with vision, tool calling, a built-in chat page, prefix caching, automatic memory planning, and optional KV cache offload to SSD.

Why now: It's trending on GitHub's most-starred new repos shortly after release, and its README references a fresh DFlash 2 blog post and an LM Studio integration guide, suggesting recent ecosystem buzz.

Who it is for: Mac owners with M3 or newer chips (ideally 24–48 GB+ of unified memory) who want to run coding agents or LLM apps locally without the cloud.

llminferenceapple-siliconmetalcoding-agentsspeculative-decoding

apple-siliconcoding-agentsllm-inferencemacosmetalspeculative-decoding

Open on GitHub →

Stars over our 9 snapshots: 846 to 847, since 1 h ago.

Where people talked about it

API: https://socialmediatrends-api.osmike.com/v1/repos/incoai/splash