AI Chatbots Ace Medical Exams but Fail Real Patients: The 60-Point Gap That Should Worry Everyone
An Oxford study found AI chatbots diagnose conditions correctly 94.9% of the time on paper, but only 34.5% when talking to actual people. The implications for AI benchmarks extend far beyond medicine.