Evals Beat Vibes: Testing LLM Features Properly
Sep 2026 · 1 min read
Most teams building on language models start the same way: try a few prompts, it looks good, ship it. Then a prompt tweak fixes one case and quietly breaks three others.
Write the test set first
Collect real inputs, even 30 to 50 of them. For each, write down what a good answer must contain or avoid. This dataset is more valuable than any single prompt.
Mix cheap and expensive checks
Some things can be checked with plain code: is the output valid JSON, does it include the required field, is it under the length limit. Others need judgment, where a second model grading against a rubric is a reasonable compromise. Spot-check the grader against your own opinion.
Track it over time
Run the set on every prompt or model change. When a new model is released, you can answer "is it better for us?" in an afternoon instead of arguing about it.
Keep it honest
Evals only measure what you thought to include. Add every production failure to the set, so the same bug cannot come back twice.