ninemin.lilulab.ai

“Max number of llm calls limit of `500` exceeded” — Google ADK’s model-call ceiling

Max number of llm calls limit of `500` exceeded

That is Google’s Agent Development Kit, raised as an LlmCallsLimitExceededError. The backticks around the number are in the source string, not added by us. The 500 is a real default and it is the one number on this page you did not choose: it lives in a module constant called _DEFAULT_MAX_LLM_CALLS. The counter does not count agent steps, tool calls or turns — it counts model calls about to be made, and it is checked immediately before the call goes out, so the call that trips it never happens.

In ten seconds. This is an invocation-level cap and it raises an ordinary Exception subclass, so your code can get control back. The ceiling is RunConfig.max_llm_calls, default 500, overridable by the ADK_MAX_LLM_CALLS environment variable.

The counter increments and then compares with >, so on the default a run completes 500 model calls and raises on the 501st attempt. A response produced by a before-model callback never reaches the counter, so cached and short-circuited responses are free.

The raise site, quoted in full

From invocation_context.py, lines 69–85 as fetched. This is the whole of _InvocationCostManager.increment_and_enforce_llm_calls_limit:

  def increment_and_enforce_llm_calls_limit(
      self, run_config: RunConfig | None
  ) -> None:
    """Increments _number_of_llm_calls and enforces the limit."""
    # We first increment the counter and then check the conditions.
    self._number_of_llm_calls += 1

    if (
        run_config
        and run_config.max_llm_calls > 0
        and self._number_of_llm_calls > run_config.max_llm_calls
    ):
      # We only enforce the limit if the limit is a positive number.
      raise LlmCallsLimitExceededError(
          "Max number of llm calls limit of"
          f" `{run_config.max_llm_calls}` exceeded"
      )

The string is built from two pieces — a plain literal and an f-string beginning with a space — which is why the rendered sentence has exactly one space before the backtick. If you are grepping your logs for it, Max number of llm calls limit of is the part that does not change.

Three conditions have to hold for the raise, and each one is a way for it not to fire. There has to be a run_config; max_llm_calls has to be greater than zero; and only then does the count get compared. The comment on the line above the raise says what the second one means in practice: We only enforce the limit if the limit is a positive number.

The exception class itself, line 54 of the same file, in its entirety:

class LlmCallsLimitExceededError(Exception):
  """Error thrown when the number of LLM calls exceed the limit."""

An ordinary Exception. except Exception catches it, and so does a handler written for this class by name.

Where the 500 comes from

From run_config.py, line 38 and lines 41–52 as fetched:

_DEFAULT_MAX_LLM_CALLS = 500


def _default_max_llm_calls() -> int:
  """Resolves the default max LLM calls limit from environment or fallback."""
  if env_val := os.getenv('ADK_MAX_LLM_CALLS'):
    try:
      return int(env_val)
    except ValueError:
      logger.warning(
          'Invalid value for ADK_MAX_LLM_CALLS env var: %s. Using default %s.',
          env_val,
          _DEFAULT_MAX_LLM_CALLS,
      )
  return _DEFAULT_MAX_LLM_CALLS

And the field it feeds, lines 264–266:

  max_llm_calls: int = Field(
      default_factory=_default_max_llm_calls,
      description=(
          'A limit on the total number of llm calls for a given run. Can be'
          ' overridden by ADK_MAX_LLM_CALLS environment variable.'
      ),
  )

Two consequences of that being a default_factory rather than a plain default. The environment is read when a RunConfig is constructed, not when the module is imported, so the value can differ between two configs built in the same process. And because the factory calls int(env_val), an ADK_MAX_LLM_CALLS that is not a number does not fail the run — it logs a warning and quietly gives you 500. If someone set that variable in a deployment and the cap is behaving as though they had not, that logger.warning is the place to look.

The validator, and the two ways the cap switches itself off

Lines 350–362 of the same file:

  @field_validator('max_llm_calls', mode='after')
  @classmethod
  def validate_max_llm_calls(cls, value: int) -> int:
    if value >= sys.maxsize:
      raise ValueError(f'max_llm_calls should be less than {sys.maxsize}.')
    elif value <= 0:
      logger.warning(
          'max_llm_calls is less than or equal to 0. This will result in'
          ' no enforcement on total number of llm calls that will be made for a'
          ' run. This may not be ideal, as this could result in a never'
          ' ending communication between the model and the agent in certain'
          ' cases.',
      )

This is the asymmetry worth knowing. Setting the cap absurdly high is rejected outright with a ValueError. Setting it to zero or a negative number is allowed, with only a log line, and it removes enforcement entirely — because the guard in the raise site requires max_llm_calls > 0 before it will compare anything. ADK’s own warning text names the outcome: a never ending communication between the model and the agent. So the one intuitive way to “turn off the limit for a moment” is the way that works silently, and a max_llm_calls=0 left behind in a config file is an uncapped run that looks configured.

What one llm call is, and when it is counted

The public method, invocation_context.py lines 499–510 as fetched, is a wrapper with nothing in it but the call and a docstring:

  def increment_llm_call_count(
      self,
  ) -> None:
    """Tracks number of llm calls made.

    Raises:
      LlmCallsLimitExceededError: If number of llm calls made exceed the set
        threshold.
    """
    self._invocation_cost_manager.increment_and_enforce_llm_calls_limit(
        self.run_config
    )

In the five files fetched for this page, that method has exactly one caller: flows/llm_flows/core/_model_call.py, line 223, inside _call_llm_with_tracing() — a generator nested in call_llm_async(), which begins at line 171 of that file. The three comment lines immediately above it say what the placement is for:

      # Check if we can make this llm call or not. If the current
      # call pushes the counter beyond the max set value, then the
      # execution is stopped right here, and exception is thrown.
      invocation_context.increment_llm_call_count()

Line 218 of that file resolves the model — llm = await flow._get_llm(invocation_context) — and the code that actually invokes it comes after line 223. So the counter is a budget on calls attempted, and the call that exceeds the budget is never sent. On the default you are billed for 500 model calls and the 501st raises before it costs anything.

One thing is not counted. Lines 190–205 of the same file run the before-model callback first, and if it returns a response the generator records the trace span, yields that response and returns — never reaching line 223. A response that came from a callback, a cache or a plugin instead of the model does not consume the budget. The counter is a count of real calls to the model, which is what you want if you are using it to bound cost, and is not a count of agent steps or of loop passes.

Its scope is the invocation. The counter lives on _InvocationCostManager, whose field, lines 282–288, is declared as a Pydantic PrivateAttr and documented as costs incurred “as a part of this invocation”. The docstring on _AbortState at lines 89–97 of the same file explains the mechanism for the neighbouring private attribute declared the same way: Pydantic model_copy() shallow-copies __pydantic_private__, so all derived contexts within the same Runner share the exact same instance. On that mechanism the budget is one budget for everything a single Runner invocation does, sub-agents included, rather than 500 each. That same docstring calls out cross-Runner sub-runs — it names AgentTool and nested Workflow node runners — as a separate case, and this page makes no claim about how the budget behaves across those.

Nothing in these files catches it

We fetched five files from this package. Across all five, at the commit in the provenance section, the identifier LlmCallsLimitExceededError occurs three times in total: the class definition at line 54 of invocation_context.py, the raise at line 82 of the same file, and the Raises: docstring at line 505 of the same file. There is no except clause for it in any of the five, including in runners.py, which is the largest of them at 89,740 bytes. That is a statement about those five files at that commit and nothing wider: the package has many files we did not read, and the absence of a handler in a file is not the absence of a handler. What it does mean is that if you want a capped run to do anything other than propagate, you are the one who has to write the handler.

Why raising max_llm_calls moves the wall

Nothing in the code above looks at what the calls achieved. The counter goes up by one per call attempted and the comparison is against your constant; there is no second condition, no measure of progress, no check that the agent is converging. A run that is making the same tool call with the same arguments forty times spends the budget at precisely the rate of a run that is working, because the only thing being measured is the count.

So doubling the cap buys exactly twice the model calls and the same sentence with a bigger number in the backticks. The useful version of the question is not “what should the cap be” but “how many model calls does one unit of my work cost” — and the answer to that is visible, because the counter counts the thing you are billed for. Divide 500 by your calls-per-item and you have the number of items one default invocation can hold. If that number is less than your brief, the cap is not the problem; the shape of the brief is.

The restructure that leaves finished work behind

There is a written guide: the calls-per-item arithmetic above as a formula you can run against a brief before you launch it, what to put inside the except block so a capped run still leaves its finished work behind, and batch.py — one standard-library file that collapses a per-item loop into a single pass.

No page on this site has a checkout widget of its own. There is a written guide behind this host and it is on sale at $19 on a storefront that delivers the files automatically and carries a 30-day money-back guarantee: buy it there (checked 2 October 2026); the guide can also be paid for with 19 USDC on Base at the payment page, where delivery is by hand as a reply to your email. Every page on this site, including this one, is free to read in full, with no sign-up and nothing gated.

Provenance

Five files, for this page, all HTTP 200, all fetched 2 October 2026, all pinned at the same commit — 2a2364fd71faeaa64bd5216e1cf593ede3654c25, the head of main in google/adk-python at the time of the fetch. Three are quoted above: src/google/adk/agents/invocation_context.py, 25,902 bytes; src/google/adk/agents/run_config.py, 14,617 bytes; and src/google/adk/flows/llm_flows/core/_model_call.py, 12,155 bytes. Two more were fetched and searched, and are the files the “nothing catches it” section above is scoped to along with the first three: src/google/adk/runners.py, 89,740 bytes; and src/google/adk/flows/llm_flows/base_llm_flow.py, 23,507 bytes. Every string, line and line number quoted above is from those five fetches, and every claim about what does and does not occur is scoped to those five files at that commit only. Line numbers are line numbers in the files as fetched and will move when the files do.

The short version: Max number of llm calls limit of `500` exceeded is an LlmCallsLimitExceededError from Google’s ADK, and it is an ordinary Exception that your code can catch. The 500 is _DEFAULT_MAX_LLM_CALLS, resolved through a factory that reads ADK_MAX_LLM_CALLS and falls back to 500 with a warning if that variable is not a number. The counter ticks once per model call about to be made — not per agent step, and not for a response that came from a before-model callback — and it increments before it compares with >, so the 501st attempt is the one that raises, before the call is sent. A cap of zero or less is accepted and switches enforcement off with only a log line. Catch the error by name, persist per event rather than per run, and batch the items so one call covers many.

Nearby

This page counts anonymous readership. Each load sends the page path, the address of the page you came from (our server keeps only its domain), how long the page was open, whether you scrolled, and any campaign or outreach code in the link you followed; a second “engaged” event is sent once, ten seconds after the page opens — whether or not the tab is in front of you — or as soon as you scroll a quarter of it. The campaign codes from the first link you arrived on are kept in this browser’s local storage, and a later visit that arrives with no codes of its own is counted against them, outreach code included; a link carrying its own codes is used for that visit, and the stored first touch is never replaced. If a link carries a different outreach code, that later touch is recorded alongside the first rather than in place of it. The referrer of that first visit is stored in the same place and is never sent anywhere — every visit sends the referrer it actually had. A return visit is still counted, and an outreach code is removed from the address bar after it is read. The page also asks this domain for Vercel’s analytics script; on 26 September 2026 we requested that address on each of the five hosts we publish — lilulab.ai, willcall, ninemin, ghmirror and pinpoint — and every one returned HTTP 404 and no script, so no script from another company was served to your browser and none ran. The page still asks for it on every load, so this stops being true the moment that address starts answering, without a single byte of this page changing — and this page will not know it has, and will go on saying what it says. Whether Vercel counts the request at its own edge is its record and not ours; as of 26 September 2026 no one here had opened it, and we keep no copy of it. No cookie; the local storage above does that job. Some things are recorded that the list above does not name. Your browser and the network attach these to the request rather than the page sending them: the identification string your browser gives with every request; the two-letter country the network assigns your address — no city, and your address itself is never stored; and the site address you asked for. Our own server then writes its own bookkeeping about the record: the date and time of your visit, to the thousandth of a second, from our clock; which request header it took that site address from; and a number naming the record format it wrote. It also stores the domain your browser said the count was fired from, which for a normal load of this page is this site itself, and which our server records as the word “unknown” when the browser sends nothing it can read. Every row is kept in Vercel’s blob storage and not on a machine of our own, measured 26 September 2026 by reading the handler. The page sends how long it was open twice, from two different counters, and our server reads only one of the two names — so some rows carry it and some are empty (measured 26 September 2026). This page reads an outreach code from the address if one is present and keeps it in your browser, which would let us tell one reader from another — but we have sent no link for this page and hold no list of recipients for it, so as of 2 October 2026 there is no name for any code to resolve to. This page also loads a second counter of its own. It sends its own copy of the view event on every load, so one load writes two view rows and any rate measured against them reads half its true value; it also measures the time differently, counting only the seconds the page was actually in front of you where the first counter sends wall-clock time since the page opened. It also sends an event once you have had the page in front of you for ten seconds and moved the mouse, touched the screen, scrolled or pressed a key. That ten-second event also depends on the count before it: this second counter does not send it unless its own copy of the view event went out first. Its “engaged” is stricter still — it is sent only after that ten-second event has gone out, you have had the page in front of you for ninety seconds in all, and you have scrolled far enough to reach a marker we put where the explaining part of this page ends; a reader who stops short of that marker never produces one. Because both counters send an event called “engaged” under different rules, a single visit can produce two of them. The host that serves it keeps its own request logs; those are its record and not ours. Every quotation above is from the linked file at the pinned commit; the same links are listed for machine readers at /llms.txt.