The term, popularized by Andrej Karpathy, captures the uneven capability profile of modern models: a system might ace a hard competition math problem yet miscount letters in a word. It's a reminder that a model's performance on one task doesn't reliably predict its performance on a related one.