Skip to contentSkip to main content
Get Useful Answers from AI — a free microcourse with a reusable templateStart learning
TechlyUp
For developers

How to evaluate LLM outputs: test sets, rubrics, and automated checks

By TechlyUpUpdated 2 min readDevelopers and AI engineers

Quick answer

Evaluate LLM features like any other software: a representative test set, clear pass criteria, and automated runs on every change. Use exact checks where outputs are structured, rubric-based scoring (by people or a carefully validated model grader) for open-ended text, and track results over time. Evaluation is what lets you change prompts and models without guessing.

Build a representative test set

Sample real (or realistic synthetic) inputs across the variety you expect, including hard and adversarial cases. Label expected outcomes or acceptance criteria for each.

Choose the right check for each output

Match the check to the output type.

  1. Structured output: schema validation and exact-match fields.
  2. Classification: accuracy, precision/recall per class, confusion matrix.
  3. Extraction: field-level correctness against labelled data.
  4. Free text: rubric scores for correctness, completeness, groundedness, and tone.

Model-graded evaluation with care

Using a model to grade outputs scales well but must be validated against human judgements on a sample. Give the grader a specific rubric and examples, and watch for bias toward longer or more confident answers.

Make it continuous

Run evaluations in CI when prompts, models, or retrieval change. Keep a history so you can see regressions and improvements.

eval results — prompt v7 vs v8
schema_valid: 100% → 100%
category_accuracy: 91% → 94%
hard_cases: 12/20 → 15/20
cost_per_100: ₹X → ₹Y (fill from your own billing)

Evaluation mistakes to avoid

These make evaluation results misleading.

  1. Building the test set only from easy, typical cases.
  2. Changing the test set every time, so results aren't comparable.
  3. Trusting a model grader without checking it against human judgement.
  4. Reporting a single average that hides serious failures in important categories.

Worked example: evaluating a summariser

A team builds 40 test documents with human-written reference points: key facts that must appear and statements that must not. Each summary is scored on coverage of key facts, absence of unsupported claims, and length limits.

They validate a model grader on 20 summaries by comparing with two human reviewers, adjust the rubric wording until agreement is good, and then run it in CI. When a model update is released, they know within an hour whether quality moved — rather than finding out from users.

Try it yourself

Write a rubric with four criteria for a summarisation feature, score 20 outputs yourself, then compare with a model grader's scores.

Frequently asked questions

How big should an evaluation set be?

Start with dozens of well-chosen cases and grow it as you find failures in production.

Can I trust LLM-as-judge?

Only after checking its agreement with human reviewers on your task. Use it to scale, not to replace human judgement entirely.

What should I evaluate besides accuracy?

Latency, cost, safety behaviour, refusal rates, and consistency across runs.

Want a suggested next step for your situation?

Share a few details and someone from TechlyUp will get back to you. No automated sequences.

Sources and further reading

Examples are authored practice material, not measured learner outcomes. Tool behavior can change. Found an error? Contact TechlyUp with the page URL and correction.

Continue learning