The multi-broker test pyramid
Prerequisites: isolating-agents-with-a2a-mocks
When a request passes through multiple brokers or agents before reaching an end user or system of record, no single test level is enough. The test pyramid layers unit, integration, and end-to-end coverage so each catches what the others cannot.
What you will learn
- The three levels of the multi-broker test pyramid and their cadence.
- The common gap: testing the broker API but never the true entry point.
- Where multi-turn state is validated versus per-request correctness.
The three levels
The pyramid is widest and cheapest at the bottom; run lower levels most often.
| Level | What it tests | Dependencies | Run cadence |
|---|---|---|---|
| Unit (per broker) | One broker in isolation | Downstream mocked (previous lesson) | Every change / every commit |
| Integration (broker-to-broker) | The real handoff between two adjacent brokers | Both brokers live | Every release of either component |
| End-to-end (full chain) | The whole path as a real user/system drives it, including any conversational front end | Everything live | Every release touching the chain, plus a scheduled cadence |
- Unit-level is the cheapest and fastest - each broker with downstream dependencies mocked. Run it most frequently.
- Integration exercises the real handoff between two adjacent brokers, catching contract mismatches that unit-level mocking cannot.
- End-to-end drives a request through the entire path exactly as a real user or system would, including any conversational front end in front of the chain. Run it on every release that touches any part of the chain, and on a schedule independent of releases.
The common gap: test the true entry point
A frequent mistake is testing the broker API directly (via a generic API client) but never exercising the true end-user entry point. If a conversational front end sits in front of the broker chain, E2E coverage must originate from that front end, not just from direct API calls to the first broker. Otherwise a whole class of integration issues - how the front end forms requests, how it handles multi-turn state - goes untested.
Direct API calls are not E2E
Task-based vs. context-managing components
Split responsibilities cleanly:
- A task-based broker (no multi-turn/context switching) is tested with independent single-task request/response pairs.
- The component that owns conversation state (typically the front end) is where E2E tests validate multi-turn behavior - separately from, and in addition to, the task-based broker’s per-request correctness.
Test each in isolation, then together
Where to go next
The pyramid proves behavior stays consistent. Next, prove the wire format and specs are correct in the first place: Protocol conformance and spec validation.