ninemin.lilulab.ai

Why the same scheduled agent run got a smaller context window than it got yesterday

Because nobody told the schedule which model to use, so it asked at the moment it fired — and got a different answer from the one it got last time. A deferred run whose model is left unset does not remember the model that was in force when you created it. It takes whatever the default is when it goes off. Nothing in the schedule’s own definition changed, and the run still landed on a different model with a different context ceiling.

What it looks like from the outside

The same schedule, carrying the same message, fires twice. One fire works through its task and finishes. The other runs out of window part-way through and dies, or compacts and hands back something thinner than you asked for. You diff the schedule definition and find nothing. You diff the prompt and find nothing. The code has not moved. It looks like flakiness, and it is not: the two fires were not running on the same thing.

Arm time and fire time are different moments

A cron schedule, a reminder, a queued job — anything deferred — has two moments in its life. Arm time is when you create it and write down what it should do. Fire time is when it actually runs. Anything you pin at arm time travels with the job. Anything you leave unset is looked up at fire time, from whatever the environment says then.

The model is the field people most often leave unset, because at arm time the default is the model they wanted and writing it down feels redundant. But a default is not a value, it is a pointer, and pointers move: someone changes the account or agent default, an alias is repointed to a newer release, the scheduler falls back to another model when the preferred one is unavailable. On any scheduler where an unset model means “use the default”, each fire of the same schedule can draw a different model — and a different model can mean a different context limit.

The tell is in the execution record, not in the code

You will not find this by reading the schedule, because the schedule is exactly the part that did not change. Look at the records each fire left behind instead, and line them up:

for every fire of this schedule id: model as recorded on the execution, not as configured context limit as recorded on the execution prompt the message it actually received outcome and whether this attempt was a retry

The signature is: same schedule id, same prompt, different model, different context limit. If the model column is constant across fires, this page is not your answer — see which limit stopped it instead. If the model column is empty, that is a finding too: your scheduler is not recording the model it resolved, and you cannot rule this out until it does.

The fix: name the model when you arm it

Set the model explicitly on the schedule, the reminder or the job at the moment you create it. Not in the environment, not in an account default, not by alias if your platform lets you name an exact version instead — on the deferred object itself, so it travels with the job to fire time. After that, a change of default changes new work and leaves your armed work alone, which is what you almost certainly assumed was happening already.

Then check it held: fire it, open the execution record, and confirm the model it ran on is the one you named rather than the one that happened to be the default.

The part people miss: a job that arms its own successor

Many recurring jobs are not a cron line at all. They are a job that, as its last act, arms the next one — a reminder that sets the next reminder, a queued task that enqueues its follow-up. Pinning the first one is not enough. The call that arms the successor is a new arm time, and if that call does not name the model, the successor is unpinned again.

Worse, the successor is usually armed by the agent running inside the job, following the instructions it was handed. So the pin has to be carried in two places: in the arming call, and in the instructions the job passes on — “when you arm the next run, name this model”. If the instruction is not in the message the successor receives, the successor has no way to know it, and the rule survives exactly one generation. The first run is pinned, its child is pinned, and its grandchild quietly is not.

The check for this is the same table as above, read down the chain rather than across fires of one schedule: follow each job to the job it armed, and confirm the model is named at every link, not just the first.

The second thing people miss: a successful retry hides the failure

If your scheduler retries automatically, the fire that ran out of window may be retried, land on a model with a larger window, and succeed. Every dashboard shaped around outcomes — last run, latest status, success per schedule — then reads the attempt that worked. The failure happened, and nothing you routinely look at shows it.

So do not ask “did the schedule succeed”. Ask for attempts, not outcomes: list every attempt for the schedule, including the ones a retry superseded, and compare the model on the failed attempt with the model on the retry that rescued it. If they differ, the retry did not fix anything. It rolled the dice again, and the next fire can roll them the other way.

Nothing on this site takes money

There is no checkout anywhere on this site, and nothing on it takes money. There is a written guide behind this host; it is listed but it is not on sale, because the payout details it needs are not connected yet. You cannot buy anything here. Every page on this site is free to read in full, with no sign-up and nothing gated.

That is the answer to the question: the model was chosen when the run fired, not when you created it, because nothing pinned it. Read the model off each execution record, pin it on the schedule at arm time, carry the pin into anything the job arms after itself, and count attempts rather than outcomes so a retry cannot hide the fire that died.

Nearby

This page loads no script of any kind: no counter, no analytics, no cookie, and nothing written to your browser’s storage. The host that serves it keeps its own request logs; those are its record and not ours.