← Back to AIAF home
AIAFArtificial Intelligence. Accelerated Future.

AIAF Index + AI Progress Watch

Our view of global AI progress.
The numbers and the reasons.

The AIAF Index is our AI-assisted editorial assessment of selected evidence about AI progress worldwide. We look across models, agents and practical applications—not just one company or country.

Global scope, selective evidence. We cannot observe every AI system. The Index expresses AIAF’s judgment; it is not an official global benchmark, a safety rating or a percentage of the way to AGI.

Two ways to understand the same subject

The Index gives a compact editorial assessment across five dimensions. Progress Watch explains what AI can do now, what still fails and what changed. Read them together: the number invites a question; the evidence helps answer it.

AIAF OVERALL INDEX

75.4/100

Component assessment date

Equal-weight average · Component evidence under review

Equal-weight editorial average: (81 + 68 + 76 + 64 + 88) ÷ 5 = 75.4/100. Each element contributes 20%, including the subjective WTF factor. Equal weighting is an editorial choice, not a scientifically established weighting.

Calculation corrected 15 September 2026: replaces the unsupported overall 72/100 using the unchanged 14 September component scores. This is an arithmetic correction—not evidence that AI progressed. The component scores remain provisional editorial judgments pending evidence validation.

Capability81

What this means for you

Can it do your task?

Look for a demonstration on work similar to yours—not just an impressive showcase.

Score evidence under review.

Explore capability evaluations →

Background reading; not validation of this historical score.

Autonomy68

What this means for you

How much supervision?

Check where people still need to approve, correct or take over.

Score evidence under review.

Read the agent evidence briefing →

Background reading; not validation of this historical score.

Reasoning / reliability76

What this means for you

Can you depend on it?

Look at repeated results and failures, not one successful answer.

Score evidence under review.

Read the reliability briefing →

Background reading; not validation of this historical score.

Real-world impact64

What this means for you

Does it improve your work?

Look for measured time savings, better quality or lower costs.

Score evidence under review.

WTF score · Surprise88

What this means for you

What surprised us—and why?

This is the surprise element, not the overall Index. It expresses editorial judgment about how unexpected developments are. Surprise does not prove usefulness or safety.

Score evidence under review.

What do the five elements mean?

How to interpret the numbers

The intended scale is 0–100, with higher scores expressing a stronger editorial assessment within a dimension. A score of 81 does not mean 81% of all tasks can be done, and 75.4 does not mean 75.4% of the journey to AGI. The historical snapshot lacks published anchors sufficient to explain why each exact number was chosen; it should not be used for precise comparisons.

Before the next scored edition

The composite formula is now the equal-weight mean of all five elements, displayed to one decimal place. If any element is not assessed, no composite is published. Before changing component scores, we will publish the assessment scope, evidence cutoff and dimension-specific scoring anchors. Each dimension must include linked sources, conditions, failures, confidence and a written scoring rationale. Missing evidence will be marked “not assessed”, never replaced with an invented value.

Updates must show what changed and why. Method changes will be versioned; the historical launch score will not be silently recalculated or presented as a comparable baseline. No new scored-release date is confirmed.

Review and update policy

Reviewed weekly. Changed only with evidence.

Scheduled review: every Monday, Malaysia time. The first scheduled evidence review is . Our daily editorial checks may identify significant developments for an additional review between weekly reviews.

A review does not automatically change a score. Each completed review will publish its date, sources examined, outcome and next review due. The score's assessment date remains separate from the date of the latest review.

Current score edition: 14 September 2026. The overall was corrected to 75.4/100 on 15 September using the unchanged historical components and a published equal-weight formula. This fixes the arithmetic but does not validate the component judgments. New component scores require scoring anchors and traceable evidence. Until those are ready, review notes will disclose the remaining gaps and keep the historical figures clearly labelled.

Review record

15 September 2026 — calculation correction: replaced 72 with 75.4 using (81 + 68 + 76 + 64 + 88) ÷ 5. Components unchanged; no new progress assessment. The original 72 had no documented formula.

15 September 2026 — review policy established. Schedule configured; first weekly evidence review pending. This is a policy update, not a completed score assessment. If a review is delayed, the actual completion date and reason will be disclosed here.

AI Progress Watch

What AI can do now.
What still fails.
What changed.

We follow selected, verifiable developments worldwide. Coverage includes models, agents and applications across organisations and countries; public evidence is uneven and coverage is not exhaustive.

  1. What changed: the development, date and task involved.
  2. What the evidence shows: source links, test conditions and results.
  3. What still fails: limitations, failed attempts and missing evidence.
  4. Why it matters: AIAF's labelled interpretation for life and work.
  5. Index relevance: which dimension it informs. A score changes only after the scoring review, not automatically with a headline.

Start with a published evidence explainer

An agent’s score is not permission to act →

The 15 September briefing helps explain autonomy and reliability: what did an evaluation establish, what were its limits, and when should a person take over? It is an evidence explainer, not a new scored assessment or breaking-news claim.

AGI Watch: the evaluation lenses →

Our evidence and correction standards →

Editorial history · 15 September 2026

The initial numerical Index proposal was followed by a Progress Watch-only presentation to clarify its scope. Following owner review, both now appear together. The 14 September scores remain historical editorial judgments with the original limitations disclosed. The previously proposed 30 September scored-beta target remains withdrawn; the overall calculation was subsequently corrected to an equal-weight 75.4/100 on 15 September, while component evidence remains under review.