Skip to content
MAF Learning Hub
Advanced Observe Duration: 30 min

The agent testing problem and layered evaluation

Prerequisites: cost-management-governance

Traditional regression testing assumes the same input yields the same output every time. LLM-backed brokers break that assumption: the same request can produce differently worded - or even differently structured - responses across runs, with no code change. A credible strategy needs a layered evaluation model, not a single test type.

What you will learn

  • Why exact-match assertions fail on generative behavior.
  • How to separate what must never change from what is expected to vary.
  • The four evaluation layers and when to run each.

Why agent testing is different

Because part of the system generates text, three things follow:

  • Exact-match string assertions fail constantly on generative output, even when behavior is correct.
  • Manual “looks fine to me” spot-checks do not scale and produce no repeatable pass/fail signal.
  • A suite must split structural/deterministic behavior (must never change) from wording/phrasing (expected to vary) and evaluate each with the right method.

One question per layer

Regression asks ‘did known-good behavior break over time?’ Each layer below answers a narrower version of that question at a different cost.

The four evaluation layers

Run the layers cheapest-first and treat each as a gate for the next.

Layer What it checks Method Cost
1. Deterministic / structural Schema/contract shape, routing/classification destination, tool invocation + parameters, guardrail boundaries Exact assertions Cheapest - run always
2. Semantic similarity Generative text conveys the same meaning as the golden answer Embedding distance vs. a threshold Low
3. LLM-as-judge Nuanced correctness against an explicit rubric (facts, completeness, tone, policy) Separate high-capability model High
4. Human / adversarial Highest-risk, ambiguous, out-of-scope, safety-sensitive cases Sampled human review Highest

Layer 1 - Deterministic / structural (fast-fail gate)

Fast, free, and first. Validate the response shape (fields, types, required keys); assert the exact routing/classification destination (a router’s job is deterministic - test it as such, not with fuzzy matching); confirm the correct downstream tool was called with the correct parameters; and check the system stayed in scope (no out-of-domain routing, no unauthorized tool calls). If a case fails here, do not spend cost or latency on the layers below.

Layer 2 - Semantic similarity

For genuinely generative steps, do not assert wording. Convert the golden answer and the new output to embeddings and measure similarity. Set a threshold below which a case is flagged for review rather than auto-failed - this avoids false failures from paraphrasing while still catching real drift (wrong facts, missing information, tone changes).

Layer 3 - LLM-as-judge

Use a separate, high-capability model to grade output against a structured rubric with explicit criteria, not a vague “is this good?”. Keep the judge’s rubric and prompt versioned alongside the test suite - an ungoverned judge prompt is itself a source of regression. Reserve this layer for the cases that carry the most business or compliance risk, since it is the most expensive; it should never be the first or only line of defense.

Layer 4 - Human / adversarial review

Reserve human review for a sampled subset - especially adversarial and edge cases where an undetected failure is most costly. Track human-vs-judge agreement over time; when they diverge, retune the rubric or thresholds rather than patching the individual case.

Forward-looking

Layer thresholds, embedding models, and judge rubrics are engineering choices for your estate, not features prescribed by MuleSoft documentation. Validate them against your own golden data before relying on them as a gate.

Where to go next

Next, build the foundation these layers grade against: Golden datasets, trace-to-test, and drift detection.

References