Inter-annotator agreement
AIInter-annotator agreement: How consistently different human labellers give the same judgement, which bounds how reliable any human-scored benchmark can be.
Inter-annotator agreement: How consistently different human labellers give the same judgement, which bounds how reliable any human-scored benchmark can be.