The agent testing problem and layered evaluation
Prerequisites: cost-management-governance
Traditional regression testing assumes the same input yields the same output every time. LLM-backed brokers break that assumption: the same request can produce differently worded - or even differently structured - responses across runs, with no code change. A credible strategy needs a layered evaluation model, not a single test type.
What you will learn
- Why exact-match assertions fail on generative behavior.
- How to separate what must never change from what is expected to vary.
- The four evaluation layers and when to run each.
Why agent testing is different
Because part of the system generates text, three things follow:
- Exact-match string assertions fail constantly on generative output, even when behavior is correct.
- Manual “looks fine to me” spot-checks do not scale and produce no repeatable pass/fail signal.
- A suite must split structural/deterministic behavior (must never change) from wording/phrasing (expected to vary) and evaluate each with the right method.
One question per layer
The four evaluation layers
Run the layers cheapest-first and treat each as a gate for the next.
| Layer | What it checks | Method | Cost |
|---|---|---|---|
| 1. Deterministic / structural | Schema/contract shape, routing/classification destination, tool invocation + parameters, guardrail boundaries | Exact assertions | Cheapest - run always |
| 2. Semantic similarity | Generative text conveys the same meaning as the golden answer | Embedding distance vs. a threshold | Low |
| 3. LLM-as-judge | Nuanced correctness against an explicit rubric (facts, completeness, tone, policy) | Separate high-capability model | High |
| 4. Human / adversarial | Highest-risk, ambiguous, out-of-scope, safety-sensitive cases | Sampled human review | Highest |
Layer 1 - Deterministic / structural (fast-fail gate)
Fast, free, and first. Validate the response shape (fields, types, required keys); assert the exact routing/classification destination (a router’s job is deterministic - test it as such, not with fuzzy matching); confirm the correct downstream tool was called with the correct parameters; and check the system stayed in scope (no out-of-domain routing, no unauthorized tool calls). If a case fails here, do not spend cost or latency on the layers below.
Layer 2 - Semantic similarity
For genuinely generative steps, do not assert wording. Convert the golden answer and the new output to embeddings and measure similarity. Set a threshold below which a case is flagged for review rather than auto-failed - this avoids false failures from paraphrasing while still catching real drift (wrong facts, missing information, tone changes).
Layer 3 - LLM-as-judge
Use a separate, high-capability model to grade output against a structured rubric with explicit criteria, not a vague “is this good?”. Keep the judge’s rubric and prompt versioned alongside the test suite - an ungoverned judge prompt is itself a source of regression. Reserve this layer for the cases that carry the most business or compliance risk, since it is the most expensive; it should never be the first or only line of defense.
Layer 4 - Human / adversarial review
Reserve human review for a sampled subset - especially adversarial and edge cases where an undetected failure is most costly. Track human-vs-judge agreement over time; when they diverge, retune the rubric or thresholds rather than patching the individual case.
Forward-looking
Where to go next
Next, build the foundation these layers grade against: Golden datasets, trace-to-test, and drift detection.