Problems are drawn from coding competitions and evaluated by executing submissions against tests. Release dates support contamination-aware comparisons by separating problems that appeared after a model's training cutoff when that cutoff is known. Execution environment, language, and sampling settings still affect results.