Skip to content
Wednesday, October 7, 2026
iInnovate MagSTARTUPS · INNOVATION · GADGETS · AI
AI

How AI Benchmarks Actually Work and Where They Mislead

AI benchmarks are standardized task suites with a named publisher and a public scoring method, and they supply most of the evidence behind claims that one model beats another. ARC-AGI-2, launched on March 24, 2025 by ARC Prize, reports that pure large language models score 0 percent on its…

Mei-Ling Chen · April 21, 2026 · 5 min read
ShareXFacebookLinkedInTelegramEmail
A researcher compares benchmark leaderboards across two monitors, pencil notes resting beside the keyboard in a warm off-white studio.
A researcher compares benchmark leaderboards across two monitors, pencil notes resting beside the keyboard in a warm off-white studio.

AI benchmarks are standardized task suites with a named publisher and a public scoring method, and they supply most of the evidence behind claims that one model beats another. ARC-AGI-2, launched on March 24, 2025 by ARC Prize, reports that pure large language models score 0 percent on its puzzle tasks, while every task has been solved by at least two humans in under two attempts.

What is an AI benchmark, and who publishes one?

A benchmark is a fixed set of tasks plus a scoring rule, run the same way for every system that attempts it. The publisher matters as much as the score: a result is only as trustworthy as the organization that maintains the tasks, keeps them uncontaminated, and publishes the method. Reputable publishers name the test, publish the tasks or a paper describing them, and update the suite when it saturates.

ARC Prize, the nonprofit behind the ARC-AGI series, describes ARC-AGI-2 as a benchmark "designed to stress-test the capabilities of state-of-the-art AI reasoning systems, provide useful signal on AGI progress, and inspire researchers to work on new ideas" on its official benchmark page. The tasks are visual puzzles that look like grids of colored shapes, and the scoring rule is blunt: a system either produces the correct completion or it does not.

Benchmarks differ in what they try to measure. Some test knowledge, some test reasoning, and some test whether a model can act usefully in a real environment. Reading a leaderboard without knowing what the tasks look like is the fastest way to be misled by a number.

Why does ARC-AGI-2 exist when ARC-AGI-1 already did?

ARC-AGI-1 endured five years of competitions and a reported 50,000x scale-up of base models with little progress until late 2024, when test-time adaptation methods changed the picture. Once frontier systems began scoring well, the benchmark stopped separating them. ARC Prize's answer was a harder successor, announced in the organization's launch post on March 24, 2025: "Pure LLMs score 0% on ARC-AGI-2, and public AI reasoning systems achieve only single-digit percentage scores."

The design goal is asymmetry. Tasks are easy for people and hard for machines, which is the opposite of most tests that reward memorized expertise. Every task in the set has been solved by at least two humans in under two attempts, according to the same post, which keeps the human baseline grounded in demonstration rather than estimation.

ARC-AGI-2 also changed what the leaderboard reports. Alongside accuracy, ARC Prize reports a cost axis, because a system that solves problems at a thousand times the human cost is a different engineering proposition from one that approaches human efficiency. A competition with $1,000,000 in prizes, hosted on Kaggle and opened alongside the benchmark, was set up to reward efficient systems rather than brute-force scale.

How does SWE-bench test real coding claims?

SWE-bench, described in a paper first submitted to arXiv on October 10, 2023, turns real software maintenance into a measurable task. The authors introduce "an evaluation framework consisting of 2,294 software engineering problems drawn from real GitHub issues and corresponding pull requests across 12 popular Python repositories," where a model receives a codebase and an issue description and must edit the code to resolve it, per the paper's abstract.

The tasks resist shortcuts. The abstract notes that resolving issues "frequently requires understanding and coordinating changes across multiple functions, classes, or even files simultaneously," which is closer to daily engineering work than answering quiz questions. A maintained family of leaderboards, including a Verified split reviewed by human engineers, publishes resolution rates and cost per resolved instance.

What problem is HELM trying to solve?

Not every benchmark is a single score. Holistic Evaluation of Language Models, published by Stanford researchers in a paper submitted on November 16, 2022, argues that "language models (LMs) are becoming the foundation for almost all major language technologies, but their capabilities, limitations, and risks are not well understood." HELM therefore evaluates across many scenarios and measures several metrics at once, including accuracy and calibration, in a framework the authors built for transparency, per the paper submitted on November 16, 2022.

The practical lesson for readers is that a model ranking depends on which scenarios and metrics were chosen. HELM makes those choices explicit and notes where coverage is thin, which is a discipline worth expecting from any leaderboard. The paper openly flags gaps in its own coverage, from question answering in neglected English dialects to metrics for trustworthiness, and that admission is a feature of the method rather than an embarrassment.

Multi-metric evaluation also changes what a "better model" means. A system can win on accuracy and lose on calibration, which means its confident answers are wrong at a predictable rate. Buyers who ship models into products care about that second number more than the first, and single-score leaderboards simply do not carry it.

BenchmarkPublisherTask typeScale, as documented
ARC-AGI-2ARC PrizeVisual reasoning puzzlesEvery task solved by 2+ humans; LLMs score 0%
SWE-benchAcademic authors, arXiv 2310.06770Real GitHub issue resolution2,294 problems, 12 Python repositories
HELMStanford researchers, arXiv 2211.09110Multi-scenario, multi-metric7 metrics across selected scenarios

How should a reader read a benchmark number?

A score means nothing without its context, and assembling that context takes about five minutes. The habit separates people who track model progress from people who retell marketing.

  1. Identify the publisher and confirm the test is named, public, and currently maintained.
  2. Check what the tasks actually look like, and whether they resemble the intended use case.
  3. Check the metric: accuracy, pass rate, resolution rate, and human-normalized scores are not interchangeable.
  4. Check the cost axis where one exists, since efficiency is part of the result, not a footnote.
  5. Check the date and the model versions, because leaderboards go stale quickly.

None of this requires inside information. Every fact cited in this piece sits on a publisher's own page, which is exactly where benchmark claims should live before anyone repeats them.

Sources

  1. ARC-AGI-2 — ARC Prize, Inc.
  2. Announcing ARC-AGI-2 and ARC Prize 2025 — ARC Prize, Inc.
  3. SWE-bench: Can Language Models Resolve Real-World GitHub Issues? — arXiv (Cornell University)
  4. Holistic Evaluation of Language Models — arXiv (Cornell University)

More from our brands

Part of the VUGA Network