8 min read Rocky Elsalaymeh
Eight turns, 8,000 tokens, 120 seconds
Most agent pitches sell autonomy as run until done. A loop with no hard stop is a billing incident. Here are the four caps Team-X enforces at v3.2.1, what happens when one trips, and where the limits leak.
Autonomy is the most overpromised word in AI. The pitch is that you hand the agent a goal and it runs until the job is done. Nobody mentions that “until done” is a condition the model gets to declare, while you pay for every turn it takes to declare it.
I took the opposite position when I built the question-answering path. As of v3.2.1, Team-X, an open-source, local-first desktop app for running AI-agent organizations, answers free-form questions with a ReAct loop that stops on whichever limit trips first: 8 tool turns, 8,000 tokens, or 120 seconds, with a 64-step ceiling behind them. When one trips, the run ends with a typed reason and no synthesized answer.
What actually stops the Team-X loop?
Four run-level caps and one per-tool timer, all defined in one file. The defaults sit in types.ts at v3.2.1, and the checks live in loop.ts.
| Budget | Default | Error reason | Ends the run as |
|---|---|---|---|
| Tool turns (one model call each) | 8 | budget_iterations | budget_exhausted |
| Emitted steps (plan, tool call, tool result, answer) | 64 | budget_steps | budget_exhausted |
| Tokens, prompt plus completion, summed | 8,000 | budget_tokens | budget_exhausted |
| Wall clock for the whole run | 120 s | budget_timeout | budget_exhausted |
| One tool call | 30 s | tool_timeout | failed |
Why two counters for turns and steps? Because the first version had one, and it lied. The 2026-05-07 audit found that with a default of 8 steps, each ReAct iteration consumed 3 (plan, tool call, tool result), so an operator who asked for 8 tool turns got only 2 or 3. The row is marked fixed on 2026-05-09.
The fix split the knob. maxIterations counts model calls and is the number an operator means by “turns”. maxSteps was raised from 8 to 64 and demoted to a safety net for one iteration that fans out into a flood of parallel tool calls.
Order matters too. At the top of every pass the loop checks cancel, then the clock, then turns, then steps, then tokens. Before each tool dispatch it checks the step ceiling again. Nothing gets to start another model call past a tripped cap.
What happens when a cap trips?
The run ends with an error step and no answer. Every budget exit in loop.ts goes through the same helper, emitErrorAndFinish, which appends a typed error step and returns. None of the five budget exit sites makes a final completion call to ask the model for a best-effort summary.
Here is the full termination vocabulary, copied from types.ts:
export type LoopErrorReason =
| 'budget_iterations'
| 'budget_steps'
| 'budget_tokens'
| 'budget_timeout'
| 'tool_call_invalid'
| 'tool_unknown'
| 'tool_threw'
| 'tool_timeout'
| 'provider_error'
| 'canceled';
The answer field on a finished run is documented as set only when the status is completed. A budget-exhausted run has none.
What the operator keeps is the trail. The service in agentic-loop-service.ts writes every step as a message row on a thread with the system-agent, so a tripped run leaves its plan, its tool calls, its tool results, and a final row shaped like [error:budget_tokens] plus the message. The run record is closed with status error, and an agentic.failed event carries the reason, tokens, and cost.
That is a trade, and I will name the cost. You paid for the tokens and you did not get a conclusion. I still prefer it to asking a model that has just run out of room to improvise a summary of evidence it could not finish reading.
Cancel works the same way. stop(runId) aborts the run’s controller, and the service coerces the final status to canceled. The cost of a canceled run is still recorded, because a stop button should not be a way to hide spend.
Does the loop ever write?
Directly, no. But “the loop is read-only” would be a false sentence at v3.2.1, so here is the exact map.
| Actor path | Tools in the registry | What they touch |
|---|---|---|
| System-agent (the default for palette questions) | Six read tools plus decompose_project, delegate_subtask, review_deliverable | Read: employees, tickets, projects, meetings, vault, audit events. Write-side: described below |
| System-copilot (the copilot service) | The same six read tools plus one insights query | Read only. The write-side set is carved out at the composition root |
The six read tools are query_employees, query_tickets, query_projects, query_meetings, query_vault, and query_events, listed in agentic-tools.ts. Each returns at most 50 rows with an explicit truncated marker, and the loop cuts any single observation at 8,000 characters.
The write-side three, in agentic-tools-write.ts, are built differently from what you might expect. decompose_project proposes a plan and emits plan.proposed. review_deliverable reads a finished ticket, makes one provider call, and emits review events. delegate_subtask writes a pending_delegations row that the operator must approve, and no ticket exists until that approval happens.
The only repository write I found in that file is the pending row. Whether that is enough separation is a fair argument. It is the separation that exists, and it is enforced by which tools are in the registry, not by a sentence in a prompt. The same argument applied to role specs is in A role is not a prompt. The copilot branch returns readSide plus one insights tool, and the system-agent branch returns readSide plus writeSide.
Why not let it run until it is done?
Because the people who build agents for a living say to stop them. Anthropic’s Building effective agents describes the agent loop and then says it is “also common to include stopping conditions (such as a maximum number of iterations) to maintain control.” Elsewhere it adds that “The autonomous nature of agents means higher costs, and the potential for compounding errors.”
The cost half is measurable. Anthropic’s engineering team reports that agents typically use about 4 times more tokens than chat interactions, and multi-agent systems about 15 times more. The same write-up says its simulations revealed a failure mode: “agents continuing when they already had sufficient results”.
The mechanism is old. The ReAct paper proposed letting a model “generate both reasoning traces and task-specific actions in an interleaved manner”. The paper gives the loop its shape. The stopping rule is something I had to add.
The reliability half is why a long leash is a bad bet. METR’s 2025 paper on long tasks defines a 50% time horizon, the length of task, measured in human time, that a model completes half the time, and puts Claude 3.7 Sonnet at “around 50 minutes”. A 50% rate is a coin flip. I do not want a coin flip with an open meter attached.
Security people name the same thing. The OWASP entry on unbounded consumption warns that attackers “exploit the cost-per-use model of cloud-based AI services”. The mitigations include limits: “Apply rate limiting and user quotas to restrict the number of requests a single source entity can make in a given time period.” A step cap is the same control applied to one request.
Other frameworks converge on it. Microsoft’s AutoGen documentation says plainly that a run can go on forever and ships termination conditions for it, including one that stops when a certain number of prompt or completion tokens are used. Vercel’s AI SDK 5 announcement says a request “runs for a single step by default” and lists a step limit among its stop conditions.
Where do the limits leak?
Three places, and I would rather list them than have you find them.
First, the checks run at the top of each pass. The loop adds up tokens after a completion lands, then compares at the start of the next pass. A call that crosses 8,000 tokens still finishes, still gets billed, and its tool calls still dispatch before the cap fires. The same is true of the clock: the elapsed time is compared at the top of the loop, and I found no timer in loop.ts that interrupts a completion already in flight. The caps bound the number of passes, not the length of one.
Second, the Settings screen and the loop are not connected at v3.2.1. The Settings repo clamps Max Steps to 1 to 32, Max Tokens to 512 to 64,000, and Timeout to 10 to 600 seconds, and the UI saves them. The service can take a getBudgets callback and otherwise falls back to the defaults. A search of the v3.2.1 tree finds getBudgets in the service, its tests, and an audit note, and not in the composition root that builds the service. So the numbers in this post are the effective budgets, and the saved values do not reach the loop.
Third, this is the inner fence. Before a run starts, the service asks the budget-governance layer whether the company may run an agentic job at all. Outside that sit your provider’s own rate and spend limits. I treat those as the last line of defense, not the design.
What can you do with this?
- Put a number on every loop you run. Turns, tokens, and wall clock, at minimum. If you cannot state the maximum bill of one run, you do not have a budget.
- Count in the operator’s unit. If “8” means tool turns to the person reading the dashboard, make that the binding cap and keep the raw step count as a backstop.
- Write down the exhaustion contract before you ship. Decide whether a tripped run returns a typed reason, a partial trail, or a best-effort answer. Mine returns the first two.
- Gate writes by registry, not by prompt. Leave write tools out of the registry for read-only actors, and park any irreversible action behind an approval row.
- Test each stop. The loop’s test file names a case for the token cap, the timeout, and the iteration cap. Break the cap in your own loop and confirm the test fails.
The loop and its budgets are documented in the agentic loop guide, and the code is at the v3.2.1 tag. If you want the weekly notes from the shop floor, Field Notes is where they land.
Frequently asked questions
What are the default budgets of the Team-X agentic loop?
In Team-X v3.2.1 the agentic loop defaults to 8 tool turns, a 64-step ceiling, 8,000 tokens summed across all model calls, and 120 seconds of wall clock. Each tool call also gets its own 30 second timer. Whichever limit trips first ends the run.
What does Team-X return when an agent loop runs out of budget?
Team-X ends the run with a typed error reason naming the cap that tripped, such as the token budget, and an exhausted status. It does not make a last model call for a best-effort answer, so no synthesized answer exists. The steps taken so far stay in the thread as messages.
Can the Team-X agentic loop change company data?
Not directly. Its six query tools only read. On the system-agent path it also gets three write-side tools, but the delegation tool only parks a pending delegation for operator approval. The copilot path gets the read tools plus one insights query, and no write-side tools.
A loop with no hard stop is a billing incident with a personality.