Isolating agents with reusable A2A mocks
Prerequisites: golden-datasets-and-drift
Each broker or agent in a chain should be testable in isolation, without every downstream dependency being live. This lesson shows how to build a reusable A2A mock, switch to it without a code change, and govern it across broker versions.
What you will learn
- Why a gateway policy cannot mock an A2A action target.
- How to build a config-driven, reusable A2A mock server.
- How to switch environments via an externalized target property.
- How task-based brokers and broker versions shape the test suite.
Why a gateway policy is not a mock
Mocking an A2A action target is not something a gateway policy can do. If a downstream call is configured directly as an A2A action in a broker’s Agent Script, an Omni/Flex Gateway policy (DataWeave transformation, schema validation, etc.) can validate or transform traffic on that connection, but it cannot redirect the call to a different target. There is no built-in mocking framework for this today.
Forward-looking
What actually works - a real A2A-compliant mock
Build the mock as a Mule application using the Anypoint Connector for A2A in server mode. It receives the same JSON-RPC calls the real dependency would (message/send, tasks/get, and so on) and returns canned responses in the same shape, looked up from a config-driven fixture set (a JSON/YAML file mapping request intent to response) so new scenarios are added without touching code. Deploy it as its own lightweight app.
Make it reusable across brokers: build one generic, configurable A2A mock server that impersonates different downstream agents by loading a different fixture set per route (for example /mock/agent-a, /mock/agent-b). Other brokers reuse the same app and just add their own fixture file and environment property.
# fixtures/agent-a.yaml - request intent -> canned A2A response
routes:
- intent: "check-order-status"
response:
kind: "task"
status: "completed"
artifacts:
- type: "text"
text: "Order 1234 shipped on 2026-02-01."
- intent: "unknown" # default / fallback scenario
response:
kind: "task"
status: "failed"
error: "No fixture matched the request intent."
# broker Agent Script config - externalize the downstream target
# DEV/TEST points at the mock; PROD points at the real dependency.
a2a.downstream.orders.url = ${A2A_ORDERS_URL}
# env: dev / test
A2A_ORDERS_URL = https://a2a-mock.internal/mock/agent-a
# env: prod
# A2A_ORDERS_URL = https://orders-agent.prod.example.com/a2a
Switch environments without a redeploy
Externalize the downstream target as an environment-scoped property in the broker’s Agent Script config rather than hardcoding the URL. Point the property at the mock app in DEV/TEST and at the real dependency in PROD. Switching targets is then a configuration change, not a code change or rebuild.
Isolate what you control from what you do not
Task-based brokers
If a broker is explicitly task-based and does not manage multi-turn conversation or context switching, structure its regression tests as independent single-task request/response pairs - no need to model conversation history at that layer. Any multi-turn or context-management behavior is the responsibility of whichever component owns it (typically a conversational front end), which end-to-end tests validate separately.
Version governance
Pin regression suites to the specific broker/runtime version in use. When a version upgrade is planned, run the existing golden dataset against the new version before cutover to measure the regression delta, then retire the old suite once migration completes. Do not maintain parallel suites for versions no longer in production - that adds maintenance cost without adding protection.
Where to go next
With components isolated, assemble them into a layered strategy: The multi-broker test pyramid.