A model generates candidate implementations, which are executed against hidden functional tests in a sandbox. Evaluation commonly reports how often at least one of a set of samples passes, making sampling and execution settings important.
HumanEval is a code-generation benchmark built from programming problems with function signatures, descriptions, and tests.
A model generates candidate implementations, which are executed against hidden functional tests in a sandbox. Evaluation commonly reports how often at least one of a set of samples passes, making sampling and execution settings important.