explainx.ai0k
TrendingNewsPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

follow on google

Add explainx.ai as a preferred source

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

community

Join the community

learn

mind: share how you thinkpathways — start freeworkshopsbootcampscoursescompare Explainxcertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsmdx readeragentsllmsdesignsdictionarypeopleagi trackerfelony benchranks

company

aboutvisionmissionteaminstructorsteach on explainxpartnershipscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportcontactprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

explainx.ai

  1. Home
  2. /
  3. Dictionary
  4. /
  5. Terminal-Bench
Evaluation & Benchmarks

Terminal-Bench

Terminal-Bench evaluates AI agents on practical tasks performed through a command-line environment.

Ask Melo about this← all terms

Each task provides an isolated environment, instructions, and a verifier that checks the resulting state. Agents must use terminal tools across multiple steps rather than return only a textual answer. Reproducible containers, time limits, and tool permissions are part of the evaluation contract.

Related terms

Agent EvaluationCoding AgentSWE-benchContainerHumanEvalCommonsenseQA

Where Terminal-Bench comes up

  • GLM-5.3 Takes 3rd on Terminal-Bench 4.0 — Open Weights Beat GPT-5.6
  • Terminal-Bench-Science: The Benchmark That Deflates Coding-Agent Hype
  • Terminal-Bench 2.0: The AI Agent Benchmark That Actually Matters
  • SWE-2: Cognition's Frontier-Near Coding Model Lands in Devin