Skip to content
MAF Learning Hub
Advanced Observe Duration: 30 min

Golden datasets, trace-to-test, and drift detection

Prerequisites: layered-evaluation-model

The layered evaluation model needs something to grade against and a way to stay relevant. This lesson covers the three practices that keep a regression suite honest over time: golden datasets, trace-to-test, and drift detection.

What you will learn

  • How to build golden datasets that fit deterministic vs. generative components.
  • How to turn production failures into permanent coverage.
  • Why drift detection runs on a schedule, not just before release.

Golden datasets - the foundation

Build a golden dataset per component under test: a curated set of representative input to expected-output (or expected-behavior) pairs covering normal, edge, and known-adversarial cases.

  • For a routing/classification broker, the golden set is labeled request -> correct destination and is evaluated deterministically (Layer 1) - no semantic or judge evaluation needed for that part.
  • For a generative broker, capture intent and acceptable-response boundaries, not verbatim wording. A generative system cannot reproduce a fixed script character-for-character, so grading against exact text produces false failures. Golden entries should define the constraints an acceptable response must satisfy - tone, required content, forbidden content - which is exactly what the semantic and judge layers evaluate.
  • Keep golden datasets versioned in source control alongside the broker/prompt config they validate, so a version bump and its golden set move together.
  • Periodically refresh and expand them; a static dataset degrades as real usage patterns evolve.

Match the golden set to the component's job

Deterministic components get exact-match golden entries; generative components get constraint-based ones. Using the wrong style is the most common source of false failures.

Trace-to-test - production failures become coverage

Whenever a real production trace shows a bad outcome - misrouting, wrong tool call, hallucinated content, policy violation - convert that exact trace into a new regression case rather than only patching the immediate issue. The suite then grows to reflect real-world failure modes, not just the ones anticipated up front. This depends on full trace/session logging for every broker in the chain; without traceability, failures cannot be converted into tests. You will wire up that telemetry in Automating tests with gateway policies and Platform APIs.

Drift detection - ongoing, not just pre-release

Model behavior can drift even with no change on your side - for example, an underlying LLM provider updates its model.

  • Run the full evaluation suite on a recurring schedule (e.g. nightly), not only on deploy.
  • Track evaluation scores as a time series, not just pass/fail per run - a gradual decline in semantic similarity or judge scores is an early warning before it becomes a hard failure.
  • Alert on degradation trends, not only on binary threshold breaches.

Forward-looking

Nightly cadence, score-trend alerting, and time-series retention are operational choices for your environment. The telemetry that feeds them is covered later using GA Anypoint Monitoring and Platform APIs.

Where to go next

With data and drift handled, isolate each component so you can test it independently: Isolating agents with reusable A2A mocks.

References