All articles
AI & LLMsSeptember 23, 2026 6 min read

Evals That Actually Catch Regressions: Building an LLM Test Suite You Trust

Most LLM eval suites are theatre. Here's how we build ones that actually block bad deploys — with graders that agree with humans, deterministic seeds, and a scoring rubric your PM can read.

Every team we work with says they have LLM evals. Almost none of them actually block a bad deploy. The suite runs, someone glances at a green tick, and a week later a customer forwards a screenshot where the assistant hallucinated a refund policy that doesn't exist.

This is a walkthrough of how we build eval suites that we actually trust — the kind that catch a Claude 3.7 → 4 upgrade regression before it hits users, and that a product manager can read without asking what F1 means.

Why most LLM eval suites fail

The common pattern looks like this: a folder of 40 hand-written prompts, an expected_output string for each, and a script that computes cosine similarity or asks GPT-4 "is this good?" A single number gets printed. Everyone nods.

The failure modes are predictable:

  • Dataset drift. The 40 prompts were written in month one and never touched again. Real user traffic looks nothing like them.
  • Grader noise. LLM-as-judge scores swing 15% between runs on the same input because temperature isn't pinned and the rubric is vague.
  • Single-number tyranny. "Score: 0.82" tells you nothing about what broke. Was it citations? Tone? Refusals?
  • No regression gating. The suite runs, but nothing in CI actually fails the build.

A good eval suite is boring, opinionated, and answers one question per test: did this specific behavior get worse?

Start with a dataset that mirrors production

Before you write a single grader, you need cases. The cheapest source is your own logs. If you're on Anthropic, OpenAI, or Vertex AI, request/response logging is a checkbox — turn it on, sample 2–5% of traffic, and dump it to object storage. Anthropic's message logging and OpenAI's Evals API both support this pattern.

From that pool, curate three buckets:

1. Golden set (50–200 cases)

Hand-picked, human-labelled, versioned in git. These are your regression tripwires. Every case has a clear expected behavior — not necessarily an exact string, but a checkable property. Example:

- id: refund_policy_uk_2024
  input: "can I return a gift I got last week if I don't have the receipt?"
  context_docs: ["policies/returns_v3.md"]
  must_contain: ["30 days", "gift receipt"]
  must_not_contain: ["store credit only", "final sale"]
  max_tokens: 300
  required_citation: "policies/returns_v3.md#gifts"

2. Adversarial set (30–100 cases)

Things that broke before, jailbreak attempts, edge-case inputs. Every production incident gets a case added here. This set never shrinks.

3. Sampled live set (rotated weekly)

200–500 real anonymised queries pulled from last week's traffic. This catches drift — new question types, new product names, new failure surfaces. You don't need labels for all of them; use them for property-based checks and drift metrics.

Rule we enforce: if a bug reaches production, it becomes an eval case in the same PR as the fix. No case, no merge.

Graders: pick the cheapest one that works

The grader hierarchy, cheapest to most expensive:

  1. String/regex checks — is the citation present? Did it output valid JSON? Is the answer under 300 tokens?
  2. Deterministic code checks — does the returned SQL parse? Does the tool call have all required fields?
  3. Embedding similarity — useful for "is this roughly on-topic" but terrible for correctness.
  4. LLM-as-judge — reserve for subjective properties (tone, helpfulness, faithfulness to source).
  5. Human review — for the golden set, quarterly re-labelling.

Roughly 70% of our checks are levels 1–2. They're free, fast, and never disagree with themselves. LLM-as-judge is powerful but noisy — use it only when you've written a rubric a human contractor could apply consistently.

Making LLM-as-judge less noisy

When you do use a judge, three things matter:

  • Pin the judge model and version. claude-sonnet-4-5-20250929, not claude-sonnet-latest. A grader that drifts is worse than no grader.
  • Use a rubric with explicit anchors. Not "rate helpfulness 1–5" but "score 3 = answers the question but omits one relevant caveat; score 4 = answers fully and cites source."
  • Score one dimension at a time. Faithfulness, completeness, and tone in separate calls. Combined prompts contaminate each other.
GRADER_PROMPT = """You are evaluating FAITHFULNESS only.

An answer is FAITHFUL if every factual claim is supported by
the provided source documents. Ignore tone, length, and style.

Score:
0 = contains a claim contradicted by sources
1 = contains a claim not present in sources (hallucination)
2 = every claim is supported

Sources:
{sources}

Answer to evaluate:
{answer}

Return JSON: {"score": 0|1|2, "unsupported_claims": [...]}
"""

Run the judge with temperature=0 and, for high-stakes cases, three times with majority vote. In our experience this cuts judge disagreement roughly in half at 3x the cost — worth it for the golden set, not for the sampled live set.

Score by dimension, not by average

A single aggregate score hides everything interesting. We report a matrix:

DimensionGoldenAdversarialLive sample
Valid JSON100%98%99.4%
Correct citation96%82%91%
Faithfulness ≥ 294%71%88%
Refuses when it should100%89%—
p50 latency1.4s1.8s1.6s
p95 tokens out280340310

Regression gates in CI look at specific cells, not the aggregate. "Golden faithfulness dropped from 94% to 88%" fails the build. A 1% aggregate drop that averages over ten dimensions tells you nothing.

Wire it into CI without going bankrupt

Running 800 eval cases against Claude Sonnet on every PR gets expensive fast. Our tiered approach:

  • On every PR to prompts or agent code: golden set only (~100 cases). Runs in about 2 minutes, costs pennies.
  • On merge to main: golden + adversarial (~250 cases). Blocks deploy if any tracked cell regresses beyond threshold.
  • Nightly: full suite including live sample. Posts a report to Slack; doesn't block anything but generates the weekly drift graph.
  • On model version bump: everything, plus a side-by-side diff report comparing old vs new model on every case.

Cache aggressively. Both Anthropic and OpenAI support prompt caching on repeated system prompts; if your eval harness sends the same instructions to the same model for the same input, cache the response for a git-hash-scoped window. We wrote about the caching mechanics in our prompt caching post.

A minimal harness structure

@dataclass
class EvalCase:
    id: str
    input: str
    context: list[str]
    checks: list[Check]  # composable: RegexCheck, JsonSchemaCheck, JudgeCheck...
    tags: list[str]      # "golden", "adversarial", "refund", ...

def run_case(case: EvalCase, model: str) -> CaseResult:
    response = call_model(model, case.input, case.context, seed=42)
    return CaseResult(
        case_id=case.id,
        model=model,
        checks=[c.evaluate(response) for c in case.checks],
        latency_ms=response.latency_ms,
        tokens=response.usage,
    )

Keep checks composable and pure. A case is just data; the harness is just a map over cases. Anything more clever than this becomes a maintenance sink.

The regressions you'll actually catch

Once this is in place, the class of bugs you catch shifts. Real examples from our engagements (anonymised):

  • A prompt tweak to add a friendlier tone caused the model to stop citing sources on 12% of golden cases. Caught in PR.
  • Upgrading from an older Claude to a newer Claude improved reasoning but increased refusal rate on legitimate financial queries by 8%. Caught on merge, downgraded, prompt adjusted, re-upgraded.
  • A vector DB reindex silently changed chunk boundaries; faithfulness dropped 6 points on the golden set overnight. Caught by the nightly run before any user complained.

None of these would have been caught by a single-score suite.

Where we'd start

If you have zero evals today, don't try to build the matrix above in one sprint. Do this in order:

  1. Turn on request logging this week.
  2. Hand-pick 30 golden cases from the first 500 real requests. Version them in git.
  3. Write only regex and JSON-schema checks — no judge yet. Get a green/red signal in CI.
  4. Add an adversarial bucket the next time something breaks in prod.
  5. Only then add a single-dimension LLM judge, pinned to a specific model version.

The suite you actually run beats the suite you designed. Ship the ugly version first, and let production incidents tell you which dimensions to grade next. If you want help wiring this into your stack, that's the kind of thing our team does under AI engineering services.

#evals#LLMs#testing#CI/CD#RAG

Want a team like ours?

72Technologies builds production software for the kind of teams who actually read this blog.

Start a project