EXPLAINER
METR's Task Horizon, Explained
The single best public number for tracking real AI progress — and it's accelerating.
What it measures
METR (Model Evaluation & Threat Research) measures the time horizonof frontier AI systems: the length of a task — measured in how long a human expert would need — that an AI agent can complete on its own at a given success rate. A model with a "1-hour horizon at 50%" completes half of all hour-long expert tasks it's given, autonomously.
This is a better progress metric than any benchmark score, because benchmarks saturate (the standard coding benchmark, SWE-bench Verified, now sits at 96% — solved). Task length doesn't saturate: it just keeps growing, and it maps directly onto economic reality. A 10-minute-horizon AI is a tool. A one-week horizon is a coworker.
The numbers
From 2019 to 2025, the 50% horizon doubled roughly every 7 months— from about 4 seconds in 2019 to hours by 2025. METR's Time Horizon 1.1 update (January 2026) put the strongest measured models at roughly 16–20 hours at 50% (3–4 hours at 80%) — and found the doubling time has compressed to roughly 3–4 months in recent data.
Compounding at that rate: a two-day horizon within months, week-long autonomous work around mid-2027. That is the threshold where agentic labor stops being a demo and becomes an economic line item.
Why investors watch it
Task horizon is the leading indicator that connects capability to cash flows: it tells you when AI can absorb real jobs-worth of work, which drives software margins, labor exposure, and the payoff timing on the hundreds of billions in data-center capex. Prediction markets still price AGI in the 2030s; the horizon curve, extrapolated, argues meaningful automation years earlier. That disagreement — tracked weekly on our indicator board — is one of the most consequential open questions in markets.
The honest caveats
The measure is software-task-centric; physical work and messy organizational context aren't captured. Success at 50% is not reliability. And extrapolating any curve is a forecast, not a fact — which is why we log the extrapolation as a scored claim in our public ledger rather than asserting it.
Track it weekly, free
The Timeline Desk tracks the task horizon and nine other capability indicators — sourced, graded, and translated for allocators — every Monday.