Why AI Safety Testing Can't Be Trusted: MLCommons Exposes the Benchmark Problem
Industry consortium reveals that current jailbreak evaluations are non-reproducible, non-defensible, and useless for regulators
Tag
Industry consortium reveals that current jailbreak evaluations are non-reproducible, non-defensible, and useless for regulators