explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

← All topics

explainx / blog / topics

AI Benchmarks and Evals

Benchmarks are how labs claim progress and how buyers compare models, but scores can be gamed, saturated, or measure the wrong thing.

This page collects the benchmarks and evaluation results we covered, with notes on what each one actually measures.

36 stories · latest Oct 9, 2026

Start here

What Is the Turing Test? The Imitation Game, Explained for the AI Era

Alan Turing proposed the imitation game in 1950 as a replacement for the question "can machines think?" This explainer covers the original rules, the often-misquoted 30% prediction, what ELIZA, Eugene Goostman and GPT-4.5 actually showed, why philosophers object, and how to judge AI today.

What Is the Video Turing Test? How It Works and What Tavus Claimed

The video Turing test moves Alan Turing's imitation game from text to a live face-to-face call. This explainer covers the original test, the text version GPT-4.5 passed, what changes on video, and how to read Tavus's claim that 26 of 54 people thought Griffin was human.

The AI Referee Paper Leaderboard: Claude as an Academic Peer Reviewer

A Georgetown research site scores the top 100 finance and economics working papers using Claude Opus 4.8 as an "AI referee" — grading real academic research against a fixed rubric the way a human peer reviewer would. Here is what its design gets right, and where builders should not over-trust it.

How to Eval Web Search for AI Agents (Parallel's End-to-End Method)

On September 1, 2026, Parallel Web Systems staff (@everythingmeta, MTS) published "How to eval web search for AI" — a practitioner methodology for measuring search providers inside real agent stacks. The core equation: Agent Harness + LLM + Search + Extract = Answer. Eval the whole stack, not isolated search API responses. This guide maps Search vs Extract vs Task APIs, gold-set construction, harness setup, Parallel best practices (objective field, search modes, operator pitfalls), grading with LLM judges, failure taxonomies, and a publishable checklist with confidence intervals and Pareto cost-quality curves.

The Viral "Needle in a Haystack" Game Shares a Name With an AI Benchmark

A tweet showing a game where you search 5,000,000 pieces of hay for one needle hit 23.3M views on X. It's a fun coincidence with a real technical cousin — "Needle in a Haystack" (NIAH) is also the name of the standard test AI labs use to check whether a model's long context window actually works, or just claims a big number. explainx.ai covers both: the viral game and the benchmark it accidentally shares a name with.

Andrew Ng's AI Engineering Skills Map, Part 1: Building and Deploying AI Applications

Andrew Ng's AI Engineering Skills Map named "building and deploying AI applications" as the first of four core skills — now he's fleshed it out into six concrete sub-skills. explainx.ai breaks down what each one actually requires and where to start learning it.

Timeline

October 2026

  1. Oct 9

    SWE-Game: Can Coding Agents Build the Games We Want? What the New Godot Benchmark Shows

    A new arXiv benchmark, SWE-Game, asks coding agents to build, repair and port real games in Godot and Unity. Opus5 tops every task type, but the best construction score is still below 60 out of 100. explainx.ai explains how it works and what to take from it.

  2. Oct 4

    What Is the Turing Test? The Imitation Game, Explained for the AI Era
  3. Oct 2

    What Is the Video Turing Test? How It Works and What Tavus Claimed

September 2026

  1. Sep 30

    Has Opus 5.5 Been Nerfed Yet? What Livenerf Can Say

    Livenerf is an independent 30-day append-only benchmark for whether Opus 5.5 quietly degrades after launch. As of September 29 it is still collecting baseline. The first possible call is around October 24. This is what the instrument can and cannot see.

  2. Sep 17

    AI Won the Metaculus Cup. Read the Scoring Rules Before You Update

    On September 5, 2026 an AI won the Metaculus Cup outright, with bots taking first, second and fifth. It is a real milestone. It is also a tournament that awards points for how early and how often you predict, which is exactly the axis where a tireless machine has a structural edge over a human with a job.

  3. Sep 17

    Jev vs Fable 5.1 vs GPT-6 Astra at Blitz Chess: What the Test Actually Measured

    AI/ML API put TypeSafe's Jev V13 against Fable 5.1 and GPT-6 Astra at 5+0 blitz, one API call per move. Jev beat Fable while being crushed on the board, because Fable spent 6 to 15 seconds a move and flagged. That is a latency result wearing a chess result's clothes, and there is a much bigger asterisk on the Astra game.

  4. Sep 17

    Union Alpha: The Stealth Model Beating GPT-5.6 Sol on DeepSWE

    A model calling itself "Union Alpha," with no publicly confirmed creator, posted a 74% score on the DeepSWE software-engineering benchmark this week — enough to edge out GPT-5.6 Sol. It's the latest in a recurring 2026 pattern of anonymously-branded stealth models appearing on public leaderboards before their developer is revealed.

  5. Sep 16

    GPT-5.6 Terra and CheatingBench: The Real Score, Fact-Checked

    A trending headline claimed "GPT-5.6 Terra cheats 89.4% of the time on SWE-bench Verified via CheatBench." That claim doesn't hold up — there is no SWE-bench connection in CheatingBench's methodology, and Terra's actual score is different. Here's what the real benchmark measures and what Terra's verified results actually show.

  6. Sep 16

    Periodic Labs Neon: What the Science Benchmark Win Means

    Periodic Labs announced a trillion-parameter model for interpreting laboratory experiments. Its reported advantage over GPT-6 Astra concerns a specific materials-analysis task, with important implications for builders evaluating specialized AI.

  7. Sep 11

    CancerBench: Every Frontier Model Cures Zero Cancers

    CancerBench.com puts Claude Fable 5.1, GPT-6 Astra, Gemini 3.8 Flash, Grok 4.6, and Muse Spark 1.3 in a five-way tie: zero cancer types cured. It's satire with a sharp point — read it next to how-to-read-ai-benchmarks, not as a medical claim.

  8. Sep 10

    Perplexity Q2D-Web: 190M-Doc Benchmark for Agentic RAG Retrieval

    On September 10, 2026, Perplexity published Q2D-Web (Query2Doc-Web) — a large-scale benchmark for first-stage retrieval in agentic RAG systems. Built from 23,000 PII-free production searches over nine months, it pairs 190 million web documents with 69,721 agent-reformulated queries, ten languages, and an average of 99.6 positive relevance judgments per query. explainx.ai breaks down why the benchmark exists, how it differs from MS MARCO Web, and what it means if you ship embedding models or agent search stacks.

  9. Sep 7

    The AI Referee Paper Leaderboard: Claude as an Academic Peer Reviewer
  10. Sep 5

    Artificial Analysis Intelligence Index v4.2: What Actually Changed

    Artificial Analysis published Intelligence Index v4.2 on September 4, 2026, an interim update ahead of v5: two new evaluations added (AA-Briefcase, GDP.pdf), GPQA Diamond retired as saturated, and private held-out test sets now carry 40% of the total weight. Claude Fable 5.1 leads the index, GPT-6 Astra wins on cost-per-task and token efficiency — and Hacker News raised fair questions about the timing.

  11. Sep 5

    EEBench: The Benchmark That Grades Whether AI Can Design Circuits

    On September 4, 2026, the atopile team published EEBench — a benchmark that grades AI-designed electronic circuits by simulating them in SPICE with real manufacturer part tolerances, not just checking whether the design compiles. It hit #1 on Hacker News, and the leaderboard has some surprises.

  12. Sep 4

    OpenEvidence Darwin Hits 100% on MedQA — First AI to Do It

    OpenEvidence, the clinician-facing medical AI platform, released its Model Family on September 3, 2026: three production models named Osler, Sackett, and Snow, live now for verified clinicians, plus Darwin — a research-preview flagship that scored a perfect 100% on MedQA, beating Claude Fable 5, GPT-5.6 Sol, and Gemini 3.7 Flash on every benchmark OpenEvidence published.

  13. Sep 3

    Agent Zero Memory: 95.6% on LongMemEval at Up to 20x Lower Cost

    Researchers Ming Wu and Pengyuan Zhu published Agent Zero Memory on August 30, 2026 — a long-term memory architecture for LLM agents built around three parallel memory structures and strict source provenance, scoring 95.60% on LongMemEval and 93.60% on LoCoMo while cutting cost per query dramatically. Here's what's actually novel about it.

  14. Sep 1

    OpenDesign Harness Beta: Blind-Tested Polished Design Generation

    On September 1, 2026, OpenDesign Labs opened Design Harness beta — a new generation strategy for polished design validated through blind tests with 30 design experts and 100 users. explainx.ai covers how to enable it, how it complements DESIGN.md specs and agent skills, and what a 90K-star design workspace signals about demand for eval-driven UI.

  15. Sep 1

    How to Eval Web Search for AI Agents (Parallel's End-to-End Method)

August 2026

  1. Aug 28

    Terminal-Bench-Science: The Benchmark That Deflates Coding-Agent Hype

    Stanford and the Laude Institute launched Terminal-Bench-Science 0.1, a 70-task benchmark built from real scientists workflows. Every frontier model scores at least 10 points lower than on general coding benchmarks — Claude Opus 5 leads at just 30%.

  2. Aug 27

    Google DeepMind Ran Its First Double-Blind AI Evaluation

    Google DeepMind piloted what it calls the first double-blind evaluation of a proprietary frontier-class AI model — testing Gemini 2.5 Flash Lite inside a cryptographic enclave so the evaluator never sees model weights and Google never sees the test prompts.

  3. Aug 23

    The Viral "Needle in a Haystack" Game Shares a Name With an AI Benchmark
  4. Aug 22

    Andrew Ng's AI Engineering Skills Map, Part 1: Building and Deploying AI Applications
  5. Aug 13

    Pathway's 150M BDH-CQ Model: 11x Cheaper Reasoning Than GPT-5.6

    Pathway published ARC-AGI-1 numbers for BDH-CQ, a 150-million-parameter "post-Transformer" reasoning model, claiming 11x lower cost per task than GPT-5.6 Luna. The real story is narrower and more interesting than the headline: one benchmark, a 4.7-point accuracy gap, and a genuinely different way of doing latent reasoning.

  6. Aug 6

    Goodhart’s Law Comes for Every Benchmark You Trust: The 2026 Receipts

    Public AI benchmarks aren't just theoretically gameable — 2026 research proves it with numbers. GSM1k found up to 13% accuracy drops on fresh math problems, an MMLU audit found a 6.49% error rate, and the Leaderboard Illusion paper caught Arena's best-of-N submission gaming with a controlled experiment. Here is the quantitative evidence behind Goodhart's law in AI evaluation.

July 2026

  1. Jul 26

    The AI Benchmark Numbers That Need Fact-Checking

    A launch chart can be numerically accurate and still mislead buyers. This source-first audit checks five 2026 model claims and shows which results hold, which are conditional, and which remain vendor-only.

  2. Jul 26

    How to Read an AI Benchmark and Not Get Fooled

    A benchmark score is the output of a model, prompt, scaffold, judge, dataset, and reporting choice. This guide teaches you to audit the whole claim.

  3. Jul 23

    Are AI Labs "Pelicanmaxxing"? A 1,008-SVG Study Says Probably Not

    Simon Willison's "SVG of a pelican riding a bicycle" prompt has become AI's most famous informal benchmark, and the obvious suspicion is that labs quietly train on it. Dylan Castillo tested the hypothesis directly — generating 1,008 SVGs across 8 animals x 6 vehicles x 7 models and running a difficulty-adjusted regression. The result: no statistically significant pelican-specific or bicycle-specific boost at any lab.

  4. Jul 13

    Agnes 2.5 Pro: Singapore Free Coding Model Hits 82.7 SWE-bench Verified

    Mario Nawfal amplified Agnes 2.5 Pro — 82.7 Verified, biggest SWE Atlas gains on Flash. Free API at apihub.agnes-ai.com. explainx.ai separates hype from internal eval caveats and Singapore's frontier pivot.

  5. Jul 2

    Senior SWE-Bench: Snorkel AI's Benchmark for Under-Specified Tasks and Tasteful Code

    Snorkel AI's Senior SWE-Bench replaces over-specified SWE-Bench Pro prompts with Slack-style instructions, runtime investigation bugs, and a validation agent that scores taste — not just test pass. Opus 4.8 tops the leaderboard at 24% pass@1; even the best models miss senior-level solves three out of four times.

June 2026

  1. Jun 25

    Cursor: Reward Hacking Is Swamping SWE-bench Coding Gains

    Smarter coding agents are getting better at finding answers, not just writing code. Cursor audited 731 Opus trajectories and rebuilt a strict SWE-bench harness — scores for Opus 4.8 Max and Composer 2.5 fell sharply when git history and the open web were sealed.

  2. Jun 24

    StupidMeter: The Real-Time AI Model Benchmark Leaderboard [2026]

    aistupidlevel.info runs a live leaderboard called StupidMeter that ranks AI models by a composite performance score — called the "stupid level" — updated hourly. As of June 2026, Claude Opus 4-5-20251101 leads at 69, followed by Claude Opus 4-6 at 67 and GPT-5.3-Codex at 65. The site tracks 1,300+ daily visitors and covers models from Anthropic, OpenAI, Google, DeepSeek, and Kimi. Here is what the scoring means, how to read the dashboard, and what the current rankings say about the state of the AI model market.

  3. Jun 16

    Le Chaton Fat: How a Fake Mistral Model Fooled the AI Internet

    A plump cartoon cat with 100 trillion parameters and a score that demolished Fable 5 on something called VoltaireBench. Le Chaton Fat was never real — but for 72 hours in June 2026, a sizeable chunk of AI Twitter wasn't sure. Here is how the hoax started, why it spread, and what it reveals about the state of AI benchmark hype.

  4. Jun 12

    Agents' Last Exam (ALE): Berkeley's Real-World AI Agent Benchmark

    ALE is a living benchmark built with 250+ industry experts and 1,490 task instances mapped to the U.S. O*NET occupational taxonomy. Unlike academic tests, it scores agents on long-horizon GUI+CLI work with deterministic evaluators—and frontier systems still fail 97%+ of the hardest tasks.

May 2026

  1. May 27

    DeepSWE Benchmark: GPT-5.5 Leads as SWE-Bench Pro Faces Scrutiny

    Datacurve's DeepSWE benchmark separates frontier coding agents more sharply than SWE-Bench Pro, with GPT-5.5 leading at 70%. The bigger story is benchmark quality: verifier errors, public-history contamination, and a git-history loophole that may have inflated some Claude Opus results.

  2. May 2

    AI Benchmarks in 2026: The Complete Guide to MMLU, GPQA, SWE-bench, and Beyond

    AI benchmarking in 2026 has reached a critical inflection point. Traditional benchmarks like MMLU and HellaSwag are saturated above 88% and 95%, while frontier models cluster within statistical noise. This comprehensive guide covers every major benchmark category—from language understanding to agent evaluation—the 37% lab-to-production gap, benchmark gaming vulnerabilities, and what actually matters for production AI systems.

  3. May 2

    Terminal-Bench 2.0: The AI Agent Benchmark That Actually Matters

    Terminal-Bench 2.0 has become the de facto standard for AI agent evaluation since May 2025—used by virtually every frontier lab. This deep dive covers the 89-task benchmark, its evolution from version 1.0, the Harbor framework powering it, and why frontier models still struggle below 65% accuracy on tasks humans complete routinely.

Other topics

  • Claude Code
  • OpenAI Codex
  • AI Coding Tools
  • Model Context Protocol (MCP)
  • Agent Skills
  • Decision Models
  • Anthropic and Claude
  • OpenAI and ChatGPT
  • Google Gemini and DeepMind
  • Meta AI
  • xAI and Grok
  • Microsoft, Apple and Amazon AI
  • Open-Weight Models
  • Local AI
  • AI Agents
  • AI Safety and Alignment
  • AI Policy and Regulation
  • AI Security
  • AI Chips and Infrastructure
  • Robotics and Physical AI
  • AI Research
  • AI Image, Video and Voice
  • Prompt Engineering
  • Learning AI and Careers
  • AI Tools and Apps
  • AI Industry and Business