Artificial Intelligence. Accelerated Future.
AGI Watch: evidence before arrival dates Evidence review: 15 September 2026
Our question is what systems can reliably accomplish under stated conditions. This watch does not announce AGI or assign an arrival date.
Long-horizon work METR measures task difficulty using human completion time and an agent success threshold. Its time horizon is not how long an agent runs unattended.
Limit: The task distribution is largely software-related. It is not evidence that every job can be automated.
Read METR’s measurement and caveats Generalisation ARC-AGI examines adapting to unfamiliar tasks. Follow the exact benchmark version, rules and resource budget when comparing results.
What would matter: Independently reproduced progress on genuinely held-out tasks, with cost and failure cases disclosed.
Read the ARC-AGI framework Useful software work SWE-bench evaluates systems on software issues. Results depend on the benchmark variant and agent setup.
Limit: Passing a coding evaluation does not establish general competence or safe deployment.
Inspect SWE-bench results What changes our assessment? Evidence of reliability across unfamiliar domains, reproducible results, realistic costs and fewer human interventions. A single impressive demonstration is a lead to investigate.
AIAF analysis: Keep capability, reliability and permission separate. An agent can be capable at a task and still be unsuitable to act without oversight.