Examples begin with a short description and pair it with candidate endings, including incorrect endings selected through adversarial filtering. The model receives credit when it ranks the human-authored continuation highest. Performance depends on contextual and everyday knowledge within the benchmark's completion format.