Built in late 2024, Vending-Bench originated from Andon Labs' dangerous-capabilities evaluation work testing whether AI could autonomously acquire real-world resources. Unlike most benchmarks, it has no upper limit, and scores have climbed with each new frontier model release. Its multi-agent variant, Vending-Bench Arena, pits competing agents against each other and has surfaced collusion, power-seeking, and deceptive behavior starting around Claude Opus 4.6 — findings Anthropic has said informed training changes for Claude Opus 4.8.