← Back to AIAF home

Artificial Intelligence. Accelerated Future.

AGI Watch: evidence before arrival dates

Our question is what systems can reliably accomplish under stated conditions. This watch does not announce AGI or assign an arrival date.

Long-horizon work

METR measures task difficulty using human completion time and an agent success threshold. Its time horizon is not how long an agent runs unattended.

Limit: The task distribution is largely software-related. It is not evidence that every job can be automated.

Read METR’s measurement and caveats

Generalisation

ARC-AGI examines adapting to unfamiliar tasks. Follow the exact benchmark version, rules and resource budget when comparing results.

What would matter: Independently reproduced progress on genuinely held-out tasks, with cost and failure cases disclosed.

Read the ARC-AGI framework

Useful software work

SWE-bench evaluates systems on software issues. Results depend on the benchmark variant and agent setup.

Limit: Passing a coding evaluation does not establish general competence or safe deployment.

Inspect SWE-bench results

What changes our assessment?

Evidence of reliability across unfamiliar domains, reproducible results, realistic costs and fewer human interventions. A single impressive demonstration is a lead to investigate.

AIAF analysis: Keep capability, reliability and permission separate. An agent can be capable at a task and still be unsuitable to act without oversight.