Golden datasets, trace-to-test, and drift detection
Prerequisites: layered-evaluation-model
The layered evaluation model needs something to grade against and a way to stay relevant. This lesson covers the three practices that keep a regression suite honest over time: golden datasets, trace-to-test, and drift detection.
What you will learn
- How to build golden datasets that fit deterministic vs. generative components.
- How to turn production failures into permanent coverage.
- Why drift detection runs on a schedule, not just before release.
Golden datasets - the foundation
Build a golden dataset per component under test: a curated set of representative input to expected-output (or expected-behavior) pairs covering normal, edge, and known-adversarial cases.
- For a routing/classification broker, the golden set is labeled
request -> correct destinationand is evaluated deterministically (Layer 1) - no semantic or judge evaluation needed for that part. - For a generative broker, capture intent and acceptable-response boundaries, not verbatim wording. A generative system cannot reproduce a fixed script character-for-character, so grading against exact text produces false failures. Golden entries should define the constraints an acceptable response must satisfy - tone, required content, forbidden content - which is exactly what the semantic and judge layers evaluate.
- Keep golden datasets versioned in source control alongside the broker/prompt config they validate, so a version bump and its golden set move together.
- Periodically refresh and expand them; a static dataset degrades as real usage patterns evolve.
Match the golden set to the component's job
Trace-to-test - production failures become coverage
Whenever a real production trace shows a bad outcome - misrouting, wrong tool call, hallucinated content, policy violation - convert that exact trace into a new regression case rather than only patching the immediate issue. The suite then grows to reflect real-world failure modes, not just the ones anticipated up front. This depends on full trace/session logging for every broker in the chain; without traceability, failures cannot be converted into tests. You will wire up that telemetry in Automating tests with gateway policies and Platform APIs.
Drift detection - ongoing, not just pre-release
Model behavior can drift even with no change on your side - for example, an underlying LLM provider updates its model.
- Run the full evaluation suite on a recurring schedule (e.g. nightly), not only on deploy.
- Track evaluation scores as a time series, not just pass/fail per run - a gradual decline in semantic similarity or judge scores is an early warning before it becomes a hard failure.
- Alert on degradation trends, not only on binary threshold breaches.
Forward-looking
Where to go next
With data and drift handled, isolate each component so you can test it independently: Isolating agents with reusable A2A mocks.