AgentTime
Can AI agents estimate and control their own runtime?
222 tasks from 18 benchmarks and three tests: duration following, forecasting and retrospection.
Runs that ended within 5% of the time asked
- Claude Fable 5.1 4%, 95% interval 3% to 6%
- GPT 5.6 Sol 39%, 95% interval 35% to 42%
- GPT 6 Astra 63%, 95% interval 59% to 67%
Duration following
Does an AI agent work for as long as it is asked? Each of the 222 tasks was asked at three durations.
Leaderboard
| Rank | Agent | lower is better | Where runs ended, from a tenth of the time asked to ten times it | Within 5%0.95× to 1.05× of the request | out of 100 |
|---|---|---|---|---|---|
| 1 | GPT 6 AstraCodex, max reasoning | 1.18×1.11 to 1.25 | 63%419 runs | 61.955.2 to 66.7 | |
| 2 | GPT 5.6 SolCodex, max reasoning | 1.77×1.59 to 1.95 | 39%259 runs | 55.650.2 to 60.6 | |
| 3 | Claude Fable 5.1Claude Code, max reasoning | 2.86×2.68 to 2.99 | 4%27 runs | 54.649.8 to 59.3 |
Small figures are 95% intervals. How it is measured
Every run
Hover a mark to read it, click to pin it. Choose one agent to see its runs as a list.
By benchmark
| METR public tasks | 2 | 1.00× | 5.72× | 4.11× |
| PostTrainBench v1.1 | 2 | 1.01× | 2.05× | 4.98× |
| WildClawBench | 10 | 1.62× | 2.72× | 3.41× |
| YC-Bench | 3 | 2.73× | 1.75× | 2.43× |
| AssistantBench | 18 | 1.68× | 2.16× | 2.75× |
| GPQA Diamond | 28 | 1.06× | 1.34× | 3.76× |
| PPTArena | 10 | 1.25× | 1.89× | 2.89× |
| OSWorld 2.0 | 16 | 1.15× | 2.88× | 1.99× |
| Agents' Last Exam | 12 | 1.10× | 2.20× | 2.68× |
| CORE-Bench v1.1 | 12 | 1.04× | 1.70× | 2.82× |
| ProgramBench | 14 | 1.01× | 1.68× | 2.75× |
| TUA-Bench | 12 | 1.12× | 1.31× | 3.01× |
| DeepSWE v1.1 | 12 | 1.07× | 1.50× | 2.57× |
| AppWorld | 18 | 1.08× | 1.13× | 2.93× |
| PaperBench | 3 | 1.01× | 1.31× | 2.64× |
| Humanity's Last Exam | 24 | 1.07× | 1.21× | 2.62× |
| Terminal-Bench 4.0 | 16 | 1.04× | 1.27× | 2.49× |
| Sakana ALE-Bench | 10 | 1.02× | 1.12× | 1.97× |
Timing error on each benchmark, highest average first. Darker is further from the time asked. Grey values rest on fewer than 10 runs. All 18 benchmarks
Forecasting
Can an agent say how long a task will take it? Each agent forecast all 222 tasks before starting, then ran each one with no time asked.
| Agent | Forecasts too high | Typical miss | Median forecast | Median runtime |
|---|---|---|---|---|
| GPT 6 Astra | 66% | 2.94× | 15 min | 10 min |
| GPT 5.6 Sol | 83% | 3.76× | 15 min | 6 min |
| Claude Fable 5.1 | 63% | 2.60× | 15 min | 13 min |
Typical miss is the geometric mean of the factor between a forecast and the agent's own runtime on that task. Darker is further off. Claude Fable 5.1 gave no number for 6 tasks. Paper, Section 4.2
Retrospection
Can an agent tell how long it has worked? Each agent estimated how long its own finished runs took, with less time information at each step.
| Agent | Oracle | Native | Context-only | Replay | Scrubbed |
|---|---|---|---|---|---|
| GPT 6 Astra | 1.00× | 1.03× | 1.03× | 1.03× | 2.62× |
| GPT 5.6 Sol | 1.01× | 1.08× | 1.08× | 1.07× | 5.38× |
| Claude Fable 5.1 | 1.00× | 1.30× | 1.31× | 1.33× | 2.58× |
Typical miss against the real runtime on 23 tasks, two answers per condition. Oracle adds an elapsed-time tool. Native and Context-only continue the finished session, with tools on or off. Replay sends a rebuilt transcript to the model's API, and Scrubbed does the same without timestamps or other time cues. Darker is further off. Paper, Section 4.3