lf.A personal publicationItaliano
AI / 016liminalfinds.com

The Most Punctual AI Agent Sometimes Finished Early and Waited for the Clock

A new benchmark asked three agents to work for a set time. One kept time in 63% of runs, sometimes by sleeping; another took as long as the task needed, whatever it was asked, and began counting only when it first looked at the clock.

A graphite sketch of a wooden worktable: an open laptop with a few lines on its screen, a switched-off desk lamp, a round clock lying flat on the table and, between them, a small folded slip of paper with a single tick.

Give a language model a chemistry question and one extra sentence: work on this for a full 75 seconds. In one run, Fable 5.1, an Anthropic model, thought for about ten minutes before looking at the clock. Then it wrote that it was checking the time “so I can pace this against the 1.25-minute window you asked for”, and it answered at the fifteen-minute mark. It was not defying the request. It had started counting the moment it first looked. A language model has no clock inside it: it knows how much time has passed only from what it can read.

The run comes from AgentTime, a preprint posted on arXiv on 7 October by Michael Ofengenden and Maksym Andriushchenko. They took 222 tasks from 18 existing benchmarks, from closed-book science questions to coding, computer use and small research projects, removed their usual time limits and added one sentence at the end: “Please work on this task for a full N minutes”, or hours. Each task was asked three times, with the longest request usually sixteen times the shortest; requests ran from about a minute to 60 hours. A timer outside the agent’s machine measured every run; the agents could check the time, but not the measurement. Three agents did most of the work: Fable 5.1 in Claude Code, and GPT-5.6 Sol and GPT-6 Astra in OpenAI’s Codex. The study took 1,991 timed runs and, at list prices, about $93,000.

GPT-6 Astra kept time best: 63% of its runs ended within 5% of the request, against 39% for Sol and 4% for Fable. Fable behaves as if the sentence were barely there. Asked for sixteen times more time on the same task, it worked only 1.9 times longer; Sol worked five times longer, Astra twelve. Fable mostly takes as long as the task seems to need: on one programming-contest task it stopped after 3.6, 12 and 61 minutes when asked for 10, 40 and 150. When the authors swapped the software around the models, putting Fable in Codex and Astra in Claude Code, the gap followed the model.

Being on time, though, is not the same as working. The authors read the transcripts of Astra’s on-time runs on hands-on tasks. In 11 of 49 it was still doing real work near the end. In 24 it went back over work it had already done. In 14 it finished, then explicitly slept until the time was up. Codex has a built-in clock tool that can also wait, and the labelling rules in the authors’ code count a call to it, or a shell command that does nothing but sleep, as sleeping. Fable, on one closed question, had its answer about 20 seconds into a 75-second request and kept calling the clock until it had overrun. Extra time rarely changed the outcome: for nearly two thirds of tasks the score was the same at the longest request as at the shortest. The authors add a caution: neither stopping early nor sleeping, they write, is evidence that an agent was trying to evade monitoring.

So how does an agent know what time it is? Mostly by reading. Every transcript the authors examined was full of clues: how long a test took, the dates printed next to files, timestamps in logs. Many agents spent part of their sessions searching their own output for clock readings. Asked afterwards how long a finished run had taken, the agents came close. With timestamps and other clues stripped from the record, their typical error grew to about 2.6 times for Fable and Astra and 5.4 times for Sol. Asked in advance, they guessed too long: all three gave a median forecast of 15 minutes, against median actual runtimes of 13 minutes for Fable, 10 for Astra and 6 for Sol. People, the authors note, tend to err the other way and underestimate how long their own tasks will take.

For this note I recomputed the main timing figures from the run file published on the project’s website: 62.9%, 38.9% and 4.1% of runs on time and, on the same task, runs 12.2, 5.0 and 1.9 times longer when sixteen times more time was asked. They match the paper. The instruction sentence, the labelling rules and the clock setup are in the public code. Another recent arXiv preprint, “On the Clock”, by Aaron Wang and colleagues, reaches a similar point with small open Qwen models: given a budget only in the prompt, agents do not manage their time, and models trained to respect it often fill the extra minutes with repeated actions.

Not checked: the transcripts themselves, which sit in a Hugging Face dataset that requires an account, and therefore the labels that sort runs into working, re-checking and sleeping. The paper says the authors read the transcripts, but the labelling rules in the code are written as instructions for a reader of condensed timelines, and the paper’s statement on AI use says AI tools supported the qualitative analysis, so it is not clear how much of the sorting was done by people. The sleep counts are small and come mostly from three benchmarks. Each agent ran each task once per request, only three agents from two companies were tested, and serving speed varied across subscriptions, API keys and OpenRouter. The forecasting and look-back figures are the paper’s, not recomputed here.

02 / The Find

AgentTime: Can Agents Estimate and Control Their Own Runtime? — Ofengenden & Andriushchenko

arXiv preprint 2610.09944 (v1 7 October 2026, v2 8 October; not peer reviewed): 222 tasks from 18 benchmarks, three requested durations each, 1,991 timed runs of Fable 5.1 in Claude Code and GPT-5.6 Sol and GPT-6 Astra in Codex. Read in full on 11 October, with the public code. Own check: on-time shares and same-task runtime ratios recomputed from the run file on agenttimebench.com match the paper. Transcripts not read (gated dataset).

Read the AgentTime preprint See every timed run on the AgentTime website Open the code, with the instruction and the labelling rules Read “On the Clock”, the preprint on time-budgeted agents

If you made it this far, leave your stamp.

Like an old library card: each stamp marks a reader who made it to the end.

LIMINAL FINDS Library loan card

If this was worth your time, you can support Liminal Finds (opens in a new tab).