Longer Isn't Smarter: Google Research Shows Token Count Predicts Failure, Not Success
New research from Google and UVA reveals that longer AI reasoning traces actually correlate with wrong answers. The fix: measure how deeply the model thinks, not how much it writes.