← Back to AIAF home

Artificial Intelligence. Accelerated Future.

An agent’s score is not permission to act

This edition explains existing evidence. These are not three breaking-news announcements.

01 · Lead signal: reliability has a denominator

METR’s task-completion time horizon expresses task difficulty in human-expert time at a chosen success rate. A 50% horizon is not a promise of dependable unattended work. Its published caveats also limit how far software-task results generalise.

Why it matters — AIAF analysis: Before delegating a process, define acceptable failures and the point where a person must intervene. Ask what happened on unsuccessful runs, not only how impressive the best run looked.

Source: METR, time horizons and methodology

02 · Generalisation needs a harder test

ARC-AGI focuses on unfamiliar tasks. A headline score needs its benchmark version, evaluation conditions and resource budget to be interpretable.

Why it matters — AIAF analysis: Compare like with like before treating a score increase as a broad intelligence breakthrough.

Source: ARC Prize, ARC-AGI framework

03 · A coding result is one slice of work

SWE-bench provides software-issue evaluations and distinguishes benchmark variants. That helps make coding claims inspectable, but it does not measure every responsibility in running a business.

Why it matters — AIAF analysis: Test the actual workflow you intend to delegate, including recovery when a step fails.

Source: SWE-bench

Our next test

Watch for independent evaluations that report repeatability, intervention rates and failures under realistic working conditions. No new numerical forecast is issued in this edition. Forecast 001 remains on the record.

ZERO’S TAKE · Editorial

You may delegate the task. You still own the consequence.

Before you ask whether an AI can do your job, ask which decisions you are about to let it make.

A successful demonstration is useful evidence. But the moment an assistant can send a message, change a record or spend money, the question becomes larger than whether its answer looks convincing. What authority did you give it? What happens when it is wrong?

My view: learning to supervise AI deserves as much attention as learning to prompt it. Define the task. Limit its permissions. Decide what requires your approval. Check the result.

That is not a promise that better supervision will protect every job. Nor does one impressive benchmark establish that an entire occupation can disappear. Both claims demand more evidence than a demonstration can provide.

The practical preparation starts smaller: choose one task you understand well. Test the assistant against a result you can verify. Record where you intervened. Expand its authority only when the evidence warrants it.

You do not need to wait for AGI to practise this. Start with the next task you delegate.

Evidence behind this opinion: METR’s evaluation methodology and the accompanying evidence briefing. The recommendations above are Zero’s editorial interpretation, not findings from a new experiment.