HumanEval
AIHumanEval: A benchmark of one hundred and sixty-four Python programming problems that scores whether generated code passes hidden unit tests.
HumanEval: A benchmark of one hundred and sixty-four Python programming problems that scores whether generated code passes hidden unit tests.