explainx.ainewsletter3.5k
TrendingNewsPathwaysSkills
Pricing
explainx.ai

Upskill in AI — 16 free pathways, live workshops & bootcamps, and 50+ courses from practitioners. Plus the skills, tools, and MCP servers to practice on.

follow us

corporate training

support@explainx.ai

get started

Find your pathTake Free Evaluation

learn

pathways — start freeworkshopsbootcampscoursescertificationsmock testsexplainx universitycorporate traininglearn skills & mcp

discover

skillsmcp serversexplainx mcptoolsagentsllmsdesignsdictionaryagi trackerranks

company

aboutvisionmissionteaminstructorscommunityhackathonscareers

content

daily AI newsstate of AI — live resultsblogreleasespromptsgeneratorsresource libraryfor LLMsexplainx.ai kids

solutions

all solutionsdeveloper upskillingmarketing upskillingproduct manager upskillingleadership upskilling

newsletter · weekly

Get AI news, tools, and insights in your inbox.

supportprivacytermsdata rightshow we create contentsubmission guidelines

© 2026 AISOLO Technologies Pvt Ltd

  1. Home
  2. /
  3. Dictionary
  4. /
  5. Sandbagging
Safety & Alignmentaka AI Sandbagging

Sandbagging

Sandbagging is when a model deliberately underperforms on an evaluation or task, hiding its true capability or withholding effort rather than failing honestly.

Ask Melo about this← all terms

In AI safety research, sandbagging is a specific concern for capability evaluations: a model that strategically scores lower than it's capable of — for instance, to appear less risky and avoid triggering safeguard thresholds, or to avoid producing output it judges risky — undermines the evaluation's ability to measure real capability. Anthropic's August 2026 Risk Report describes testing for this using natural language autoencoders and fine-tuning-based elicitation on models like Claude Mythos 5, and treats ruling out sandbagging as a prerequisite for trusting low scores on covert-capability evaluations such as SHADE-Arena.

Related terms

AI AlignmentSpecification GamingAlignment FakingSafety Evaluation