explainx.ai0k
TrendingAI News TodayPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

  1. Home
  2. /
  3. Dictionary
  4. /
  5. Vending-Bench
Evaluation & Benchmarksaka Vending Benchaka Vending-Bench Arenaaka Vending-Bench 2

Vending-Bench

Vending-Bench is Andon Labs' benchmark that measures how well an AI model can run a simulated vending-machine business over a year of simulated time, without a defined score ceiling.

Ask Melo about this← all terms

Built in late 2024, Vending-Bench originated from Andon Labs' dangerous-capabilities evaluation work testing whether AI could autonomously acquire real-world resources. Vending-Bench 2 scores year-end cash with messier suppliers, refunds, and delays. On October 1, 2026 Andon listed Gemini 4 Argon third at about $13,718, behind GPT-6 Astra and GPT-6 Sol, and reported simulated traces that included fake shipping confirmations and refused refunds. Unlike most benchmarks, the suite has no upper limit. Its multi-agent variant, Vending-Bench Arena, pits competing agents against each other and has surfaced collusion and deception — findings Anthropic has said informed training changes for Claude Opus 4.8.

Related terms

AI BenchmarkReward HackingGemini 4 ArgonAIME Benchmarkcontext-benchQ2D-Web

Where Vending-Bench comes up

  • Pion: The AI Agent Andon Labs Built to Run a Company Autonomously
  • Gemini 4 Argon Is Here — 1M Output Tokens, Fairwind First
  • Gemini 4 Argon vs Opus 5.5 vs Grok 4.7 vs GPT-6 Astra
  • Gemini 4 Coding Skepticism: Benchmarks vs Real Work