
Jeff's 0.8B decision models claim ~30 ms, trained at home
Jeff, a GitHub repo of 0.8B decision models trained at home, claims ~30 ms. Here is the trend it belongs to and what to test before trusting it.
~30 ms. That is the latency in the title of Jeff, a GitHub repo of 0.8B “decision models” trained at home. Small models are no longer just shrunk. They are being built for one job.
A caveat first. The source here is the repo title and link, not a benchmark table. The title says the models are Jev-compatible, and it gives no hardware or test conditions. Treat ~30 ms as the author’s claim about the author’s setup until you run it yourself.
The tempo is the story. Here is how it got here:
September 2024: Meta released the smallest Llama 3.2 text models, aimed at phones and edge devices. They were scaled-down generalists.
August 2025: Google released the smallest Gemma 3 model yet, pitched for fine-tuning on a specific task. The job description narrowed.
Today: Jeff lands on Hacker News. It is a 0.8B family trained at home, named for decisions, with a latency figure in the title.
The through-line is who is doing the narrowing. First the big labs shipped small versions of their chatbots. Then they told you to specialize them. Now an individual trains a specialist on home hardware and leads with latency, not with a leaderboard score.
The mechanism explains why that matters. A decision is usually a pick from a short list: route here, approve, retry, escalate. That takes very few output tokens. If Jeff works as its name suggests, the cost is mostly reading the prompt and emitting a handful of tokens. That is why a model this small can be fast enough to sit in a hot path.
Here is the practitioner read. If your agent calls a frontier model just to choose the next tool, judge whether output passed, or decide whether to retry, you are paying for a network round trip and a big model to do a small model’s job. The same logic sits behind Forge, which wraps a small model in guardrails instead of reaching for a bigger one. In a full agentic pipeline, those small decisions run constantly. They are where latency and cost pile up.
So do the boring test this week. Clone the repo. Replay a labeled sample of your own logged decisions through it. Measure accuracy and latency on your hardware, not the author’s. If it holds, the gate step in your agent loop is the first thing to move off the API. If it doesn’t, you’ve spent an afternoon.
Prediction: by December 2026, at least one widely used open-source agent framework will document a sub-billion-parameter local model as a built-in router or gate. If none does, this read was wrong, and decisions stayed with the big models.
I’ll flag the next specialist small model that ships, and whether it survives independent testing. To get that, subscribe and it lands in your inbox.