
Build your own decision model: replace an LLM call this afternoon
A Hacker News post titled 'Build your own decision model' is a prompt to stop paying an LLM for routing calls. Here is a recipe and the code to start.
In one afternoon you will know how to replace a routing call to an LLM with a small decision model you own, and fall back to the LLM only when it is unsure.
Why now: a Hacker News post titled “Build your own decision model” is circulating, and it points at a habit worth breaking. Many builders send every yes/no or which-tool question to a frontier model. The recipe below is my general version of the idea in that title, not a walkthrough of that post.
The mechanism is simple. An LLM call that only picks a label (route, escalate, skip, approve) is a classification problem. Classification does not need a model that can write sonnets. It needs labeled examples of your own traffic. You already have them if you log your calls.
- Log every input and the decision the LLM returned into a
calls.jsonlfile. - Review a sample of the labels by hand and fix the wrong ones.
- Train a plain text classifier on input to decision.
- Set a confidence floor and keep it in config, not in code.
- Route confident predictions to the small model and everything else to the LLM.
- Keep logging the fallback cases and retrain on them regularly.
Here is the whole thing. Paste it, point it at your log, and set CONFIDENCE_FLOOR in your environment.
import json, os, joblib
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline
rows = [json.loads(line) for line in open("calls.jsonl")]
X = [r["input"] for r in rows]
y = [r["decision"] for r in rows]
model = make_pipeline(TfidfVectorizer(), LogisticRegression())
model.fit(X, y)
joblib.dump(model, "decision.joblib")
FLOOR = float(os.environ["CONFIDENCE_FLOOR"])
def decide(text, ask_llm):
probs = model.predict_proba([text]).ravel()
if probs.max() >= FLOOR:
return model.classes_[probs.argmax()]
return ask_llm(text)The gotcha is that you are training on the LLM’s own answers, so its mistakes become your model’s mistakes. You will recognise it when the small model agrees with the LLM almost every time and users still report bad routes. Step two is how you avoid it. A hand-checked sample is what tells you whether the labels were ever right.
A second symptom is a fallback that almost never fires. That usually means the model is overconfident on inputs unlike anything it trained on, so raise the floor.
My read: for narrow, repetitive decisions, the LLM belongs in two jobs only. It labels your first examples, and it handles the hard cases the small model refuses. If your decision model can’t beat the LLM on your own logged traffic after a retrain, your labels are the problem, not the approach.
The same logic of wrapping a small model in checks shows up in the Forge guardrails write-up. For where a cheap decision layer sits inside a larger agent pipeline, see the autonomous stack.
Want one of these a day in your inbox? Subscribe. One AI signal a day. 90 seconds. No fluff.