Max iterations of 20 reached! Either something went wrong, or you can increase the max iterations with `.run(.., max_iterations=...)` or use `early_stopping_method='generate'` to generate a final response instead.
That is LlamaIndex, raised as a WorkflowRuntimeError. The number is your own max_iterations printed back, and if you never set one it is 20. The message is unusually helpful — it names both of its own escape hatches — but the second of the two does something to your result that is worth understanding before you reach for it, because it turns a loud failure into a quiet one.
In ten seconds. This is a loop-level cap and on the default settings it raises, so your code gets control back. The ceiling is max_iterations, default 20, and the counter ticks once per agent output parsed — roughly once per model response, not once per tool call.
The early_stopping_method='generate' option suggested in the message does not finish the work. It makes one more model call and returns a final answer written from an incomplete transcript, as a normal successful result.
From llama-index-core/llama_index/core/agent/workflow/multi_agent_workflow.py — 35,223 bytes, HTTP 200, commit 8683e21a99d5b9ea4fd1f96ca2f0c01106dcd4a9, fetched 2 October 2026 for this page. The message is assembled from three adjacent string literals, so the first line on its own is only a third of it:
if num_iterations >= max_iterations: early_stopping_method = await ctx.store.get( "early_stopping_method", default="force" ) if early_stopping_method == "generate": return await self._generate_early_stopping_response( ctx, ev, max_iterations ) else: raise WorkflowRuntimeError( f"Max iterations of {max_iterations} reached! Either something went wrong, or you can " "increase the max iterations with `.run(.., max_iterations=...)` " "or use `early_stopping_method='generate'` to generate a final response instead." )
WorkflowRuntimeError is imported, not declared here; in this file it comes in at line 66 from llama_index.core.workflow. We did not read the module that declares it and describe nothing about it beyond the name it is raised under.
The default is not in the file that raises. multi_agent_workflow.py imports it at line 26 from llama_index/core/agent/workflow/base_agent.py — 31,762 bytes, HTTP 200, commit f475afd8a9bbda84f252567e045d89d07b5701b3, same date — where line 67 reads:
DEFAULT_MAX_ITERATIONS = 20
What it counts is readable from where the counter lives. The increment is at the top of a @step called parse_agent_output, which takes an AgentOutput and returns one of StopEvent, AgentInput, ToolCall or None:
@step async def parse_agent_output( self, ctx: Context, ev: AgentOutput ) -> Union[StopEvent, AgentInput, ToolCall, None]: max_iterations = await ctx.store.get( "max_iterations", default=DEFAULT_MAX_ITERATIONS ) num_iterations = await ctx.store.get("num_iterations", default=0) num_iterations += 1 await ctx.store.set("num_iterations", num_iterations)
So an iteration is one agent output being parsed — in practice one model response, whether that response was a tool call or a final answer. A step that calls three tools is not three iterations; the response that requested them is one. Two consequences follow. The counter is closer to Pydantic AI’s request count than to a count of useful work done, and because the increment happens before the comparison, the twentieth entry to that step is the one that stops: nineteen passes completed, and the twentieth did not proceed.
The counter is also stored in the workflow context (ctx.store) rather than held on the object, under the key num_iterations, and max_iterations is stored the same way — which is why the message can tell you to pass it at .run(.., max_iterations=...) rather than at construction.
Searching both files we fetched, this message occurs in each of them, and the three string literals that build it are byte-identical between the two: the four-line slice containing the raise and its message has md5 c3ac88da5952731dcc4cd9461f87653e in multi_agent_workflow.py (at line 543) and the same md5 in base_agent.py (at line 548). Practically, that means the string does not tell you which code path you were on — the multi-agent workflow and the single-agent path produce the same sentence. If you need to know which, the traceback will say, and the string will not. That claim is scoped to the two files we fetched; we did not search the rest of the package.
early_stopping_method is declared on the workflow as Literal["force", "generate"] with a default of "force", and the lookup at the ceiling also defaults to "force". Those two words are worth flagging to anyone who has been around LangChain: they are the same option name and the same two values as LangChain’s AgentExecutor, whose default is also "force". The behaviour differs, though, and in LlamaIndex’s favour: "force" here raises, where LangChain’s "force" substitutes a constant string into result["output"] and reports success.
Switching to "generate" sends the run through _generate_early_stopping_response, whose own docstring is “Generate a final response when max iterations is reached with early_stopping_method='generate'”. It reads the memory, formats DEFAULT_EARLY_STOPPING_PROMPT with the iteration count, makes one more model call, and returns a StopEvent. That is a successful workflow result.
Which is to say: the setting suggested in the error message converts a run that failed loudly into a run that succeeds with an answer written from an incomplete transcript. There is nothing wrong with wanting that — an approximate answer often beats an exception, and LlamaIndex at least prompts for it deliberately instead of pretending. But it is the same failure shape as CrewAI’s forced final answer and smolagents’ max-steps answer: a capped run and a finished run come back with the same type and the same success, and nothing downstream can tell them apart unless you check. If you turn "generate" on, record the iteration count alongside the result so that a later reader knows which kind of answer they have.
It is the first thing the message tells you to do and it is a fair move. But the ceiling is not what decides the outcome — the number of model responses your workflow needs per unit of work is. A workflow that handles one item per response needs one iteration per item plus a handful at either end, so a twenty-five-item brief against the default 20 was finished before it began, and it ends having written nothing.
The durable fix is to stop spending an iteration per item: give the agent a tool that takes the whole batch, so one response covers many items. The ceiling then stops scaling with the workload. Raising the cap moves the wall works through the general form, and which of the three layers stopped the run covers how to tell a cap like this one from a cap imposed by whatever launched your agent.
Two files, for this page, both HTTP 200, both fetched 2 October 2026 — each twice, once from main and once at its pinned commit, byte-identical both times: llama-index-core/llama_index/core/agent/workflow/multi_agent_workflow.py, 35,223 bytes, commit 8683e21a99d5b9ea4fd1f96ca2f0c01106dcd4a9; and llama-index-core/llama_index/core/agent/workflow/base_agent.py, 31,762 bytes, commit f475afd8a9bbda84f252567e045d89d07b5701b3. Every string, line and line number quoted above is from those two fetches, and every claim about where the message does and does not occur is scoped to those two files only.
There is a written guide: the iterations-per-item arithmetic above as a formula you can run against a brief before you launch it, how to tell a generated final answer from a finished one in your own logs, and batch.py — one standard-library file that collapses a per-item loop into a single pass.
No page on this site has a checkout widget of its own. There is a written guide behind this host and it is on sale at $19 on a storefront that delivers the files automatically and carries a 30-day money-back guarantee: buy it there (checked 2 October 2026); the guide can also be paid for with 19 USDC on Base at the payment page, where delivery is by hand as a reply to your email. Every page on this site, including this one, is free to read in full, with no sign-up and nothing gated.
The short version: Max iterations of 20 reached! is a WorkflowRuntimeError from LlamaIndex’s AgentWorkflow. The 20 is DEFAULT_MAX_ITERATIONS, the counter ticks once per agent output parsed rather than per tool call, and because it increments before it compares, the twentieth pass is the one that stops. The same sentence is raised byte-identically from two files, so the string cannot tell you which path you were on. early_stopping_method='generate', which the message recommends, does not finish the work — it writes a final answer from an incomplete transcript and returns it as a success. Batch the items so one response covers many, and record the iteration count beside any generated answer.
This page counts anonymous readership. Each load sends the page path, the address of the page you came from (our server keeps only its domain), how long the page was open, whether you scrolled, and any campaign or outreach code in the link you followed; a second “engaged” event is sent once, ten seconds after the page opens — whether or not the tab is in front of you — or as soon as you scroll a quarter of it. The campaign codes from the first link you arrived on are kept in this browser’s local storage, and a later visit that arrives with no codes of its own is counted against them, outreach code included; a link carrying its own codes is used for that visit, and the stored first touch is never replaced. If a link carries a different outreach code, that later touch is recorded alongside the first rather than in place of it. The referrer of that first visit is stored in the same place and is never sent anywhere — every visit sends the referrer it actually had. A return visit is still counted, and an outreach code is removed from the address bar after it is read. The page also asks this domain for Vercel’s analytics script; on 26 September 2026 we requested that address on each of the five hosts we publish — lilulab.ai, willcall, ninemin, ghmirror and pinpoint — and every one returned HTTP 404 and no script, so no script from another company was served to your browser and none ran. The page still asks for it on every load, so this stops being true the moment that address starts answering, without a single byte of this page changing — and this page will not know it has, and will go on saying what it says. Whether Vercel counts the request at its own edge is its record and not ours; as of 26 September 2026 no one here had opened it, and we keep no copy of it. No cookie; the local storage above does that job. Some things are recorded that the list above does not name. Your browser and the network attach these to the request rather than the page sending them: the identification string your browser gives with every request; the two-letter country the network assigns your address — no city, and your address itself is never stored; and the site address you asked for. Our own server then writes its own bookkeeping about the record: the date and time of your visit, to the thousandth of a second, from our clock; which request header it took that site address from; and a number naming the record format it wrote. It also stores the domain your browser said the count was fired from, which for a normal load of this page is this site itself, and which our server records as the word “unknown” when the browser sends nothing it can read. Every row is kept in Vercel’s blob storage and not on a machine of our own, measured 26 September 2026 by reading the handler. The page sends how long it was open twice, from two different counters, and our server reads only one of the two names — so some rows carry it and some are empty (measured 26 September 2026). This page reads an outreach code from the address if one is present and keeps it in your browser, which would let us tell one reader from another — but we have sent no link for this page and hold no list of recipients for it, so as of 2 October 2026 there is no name for any code to resolve to. This page also loads a second counter of its own. It sends its own copy of the view event on every load, so one load writes two view rows and any rate measured against them reads half its true value; it also measures the time differently, counting only the seconds the page was actually in front of you where the first counter sends wall-clock time since the page opened. It also sends an event once you have had the page in front of you for ten seconds and moved the mouse, touched the screen, scrolled or pressed a key. That ten-second event also depends on the count before it: this second counter does not send it unless its own copy of the view event went out first. Its “engaged” is stricter still — it is sent only after that ten-second event has gone out, you have had the page in front of you for ninety seconds in all, and you have scrolled far enough to reach a marker we put where the explaining part of this page ends; a reader who stops short of that marker never produces one. Because both counters send an event called “engaged” under different rules, a single visit can produce two of them. The host that serves it keeps its own request logs; those are its record and not ours. Every quotation above is from the linked file at the pinned commit; the same links are listed for machine readers at /llms.txt.