MODELSMETR TASK HORIZON (50%)
▲16–20 hrs
doubling every ~4 months (was ~7) · as of 2026-01 (Time Horizon 1.1)
The single best public timeline metric: how long a human-expert task frontier agents can complete half the time. At the current doubling rate, week-long autonomous work arrives within ~12 months — the threshold where agentic labor becomes an economic line item.
METR Time Horizon 1.1 · AI Digest tracker
MODELSSWE-BENCH VERIFIED (TOP)
▲96%
saturated — signal moved to SWE-bench Pro (69.2%) · as of 2026-08-02
Frontier models have effectively solved the standard coding benchmark. Saturation itself is the signal: real-world coding automation is no longer capability-limited, it is deployment-limited.
SWE-bench Verified leaderboard · SWE-bench Pro
CAPITALPREDICTION-MARKET AGI ODDS
▲25% by 2029
50% by 2033; 'weak AGI' by end-2026 · as of 2026-02 (Metaculus)
The crowd's number — and the spread between it and insider claims (2027–28) is the tradeable disagreement this desk exists to referee.
Metaculus AGI question · 80,000 Hours timeline review
MODELSFRONTIER-LAB REVENUE RUN-RATE
SEED — VERIFY▲~$60B/yr
Anthropic ~60× YoY (from ~$1B) · as of 2026-07 (public statements)
Revenue is deployed capability. A 60× year means enterprises are paying for AI labor at scale now — not in a forecast.
Kokotajlo, Diary of a CEO (Jul 2026)
COMPUTECOMPUTE BUILDOUT
SEED — VERIFY▲$100B+ announced
multi-GW campuses under construction · as of 2026-H1
Capex is a hard-to-fake commitment to the timeline. Power contracted today is capability delivered 2027–2029 — the physical layer confirms or falsifies the software story.
Situational Awareness LP 13F ($13.7B AUM)
APPSROGUE-AGENT INCIDENT LEDGER
▲82% of firms
reported unexpected agent behavior, trailing 12mo · as of 2026-06 (Gravitee)
Our proprietary dataset-in-progress: every public incident of agents acting off-instruction (OpenClaw mass-deletion, Feb 2026, is entry #1). Incident frequency tracks real-world autonomy deployment — and prices the safety discount.
Gravitee survey · OpenClaw incident
MODELSAI SHARE OF AI RESEARCH
SEED — VERIFY▲rising
labs report majority-AI code authorship · as of 2026-H1
The recursive self-improvement indicator — the mechanism every fast-timeline scenario runs on. When the research loop closes, every other indicator accelerates together. Quantified series under construction.
AI 2027 scenario (mechanism)
COMPUTECOST PER TASK (CONSTANT CAPABILITY)
SEED — VERIFY▼~10× cheaper/yr
frontier-level output at commodity prices within ~18mo · as of 2026-H1
Capability tells you what is possible; cost tells you when it deploys. Falling cost at constant capability is what turns benchmarks into layoffs and margins.
Epoch AI (pricing analyses)
MODELSINSIDER-STATEMENT LEDGER
◆2027–2030
insider median vs. market's 2033 · as of 2026-08
Every named-insider timeline claim, timestamped and scored on resolution. Current spread: Kokotajlo 50% by ~2029-30; lab leadership privately 2027–28; Metaculus 50% by 2033. The 4–5 year disagreement is the largest mispricing candidate in any asset class.
Kokotajlo interview (primary source) · AI Futures Project update
LABORLABOR DISPLACEMENT SIGNAL
SEED — VERIFY◆US 4.2% / UK 5%
headline flat; entry-level exposed roles weakening · as of 2026-Q2
The lagging indicator everyone watches and misreads. Fast-timeline scenarios predict headline unemployment stays quiet until automation of research completes — so the tell is entry-level hiring in exposed categories, not the headline rate.
Kokotajlo interview (stats cited)