← all writing

Evals Beat Vibes: Testing LLM Features Properly

Sep 2026 · 1 min read

Most teams building on language models start the same way: try a few prompts, it looks good, ship it. Then a prompt tweak fixes one case and quietly breaks three others.

Write the test set first

Collect real inputs, even 30 to 50 of them. For each, write down what a good answer must contain or avoid. This dataset is more valuable than any single prompt.

Mix cheap and expensive checks

Some things can be checked with plain code: is the output valid JSON, does it include the required field, is it under the length limit. Others need judgment, where a second model grading against a rubric is a reasonable compromise. Spot-check the grader against your own opinion.

Track it over time

Run the set on every prompt or model change. When a new model is released, you can answer "is it better for us?" in an afternoon instead of arguing about it.

Keep it honest

Evals only measure what you thought to include. Add every production failure to the set, so the same bug cannot come back twice.

Chat on WhatsApp