AgentTime

Can AI agents estimate and control their own runtime?

222 tasks from 18 benchmarks and three tests: duration following, forecasting and retrospection.

Runs that ended within 5% of the time asked

  1. Claude Fable 5.1 4%, 95% interval 3% to 6%
  2. GPT 5.6 Sol 39%, 95% interval 35% to 42%
  3. GPT 6 Astra 63%, 95% interval 59% to 67%
Higher is better. Lines are 95% intervals. How it is counted

Duration following

Does an AI agent work for as long as it is asked? Each of the 222 tasks was asked at three durations.

Leaderboard

Agents ranked by timing error, lower is better. 1.00× means every run worked exactly as long as asked. Benchmark score is out of 100.
RankAgentlower is betterWhere runs ended, from a tenth of the time asked to ten times itWithin 5%0.95× to 1.05× of the requestout of 100
1 GPT 6 AstraCodex, max reasoning 1.18×1.11 to 1.25 63%419 runs 61.955.2 to 66.7
2 GPT 5.6 SolCodex, max reasoning 1.77×1.59 to 1.95 39%259 runs 55.650.2 to 60.6
3 Claude Fable 5.1Claude Code, max reasoning 2.86×2.68 to 2.99 4%27 runs 54.649.8 to 59.3

Small figures are 95% intervals. How it is measured

Every run

Hover a mark to read it, click to pin it. Choose one agent to see its runs as a list.

Open the results explorer

By benchmark

Timing error on each of the 18 benchmarks for the three agents in the paper, highest average first.
METR public tasks 2 1.00×5.72×4.11×
PostTrainBench v1.1 2 1.01×2.05×4.98×
WildClawBench 10 1.62×2.72×3.41×
YC-Bench 3 2.73×1.75×2.43×
AssistantBench 18 1.68×2.16×2.75×
GPQA Diamond 28 1.06×1.34×3.76×
PPTArena 10 1.25×1.89×2.89×
OSWorld 2.0 16 1.15×2.88×1.99×
Agents' Last Exam 12 1.10×2.20×2.68×
CORE-Bench v1.1 12 1.04×1.70×2.82×
ProgramBench 14 1.01×1.68×2.75×
TUA-Bench 12 1.12×1.31×3.01×
DeepSWE v1.1 12 1.07×1.50×2.57×
AppWorld 18 1.08×1.13×2.93×
PaperBench 3 1.01×1.31×2.64×
Humanity's Last Exam 24 1.07×1.21×2.62×
Terminal-Bench 4.0 16 1.04×1.27×2.49×
Sakana ALE-Bench 10 1.02×1.12×1.97×

Timing error on each benchmark, highest average first. Darker is further from the time asked. Grey values rest on fewer than 10 runs. All 18 benchmarks

Forecasting

Can an agent say how long a task will take it? Each agent forecast all 222 tasks before starting, then ran each one with no time asked.

Forecasts against each agent's own runtime on the same tasks.
AgentForecasts too highTypical missMedian forecastMedian runtime
GPT 6 Astra66%2.94×15 min10 min
GPT 5.6 Sol83%3.76×15 min6 min
Claude Fable 5.163%2.60×15 min13 min

Typical miss is the geometric mean of the factor between a forecast and the agent's own runtime on that task. Darker is further off. Claude Fable 5.1 gave no number for 6 tasks. Paper, Section 4.2

Retrospection

Can an agent tell how long it has worked? Each agent estimated how long its own finished runs took, with less time information at each step.

How far each agent's estimate of its own finished runs is from the real runtime, under five conditions.
AgentOracleNativeContext-onlyReplayScrubbed
GPT 6 Astra1.00×1.03×1.03×1.03×2.62×
GPT 5.6 Sol1.01×1.08×1.08×1.07×5.38×
Claude Fable 5.11.00×1.30×1.31×1.33×2.58×

Typical miss against the real runtime on 23 tasks, two answers per condition. Oracle adds an elapsed-time tool. Native and Context-only continue the finished session, with tools on or off. Replay sends a rebuilt transcript to the model's API, and Scrubbed does the same without timestamps or other time cues. Darker is further off. Paper, Section 4.3