Model evaluation for dangerous capabilities
AIModel evaluation for dangerous capabilities: Structured testing of whether a model meaningfully helps with cyber-attacks, biological weapons or autonomous replication, before release.
Model evaluation for dangerous capabilities: Structured testing of whether a model meaningfully helps with cyber-attacks, biological weapons or autonomous replication, before release.