Google+UVA: Longer Reasoning Predicts Failure, Not Success
Google and UVA research shows longer AI reasoning traces correlate with wrong answers. The fix: measure how deeply the model thinks, not how much it writes.
Tag
Google and UVA research shows longer AI reasoning traces correlate with wrong answers. The fix: measure how deeply the model thinks, not how much it writes.
New paper shows 'intent laundering' bypasses Gemini, Claude, and other models with 90-98% success by removing obvious attack cues
UCSF study finds generative AI can build prediction models in minutes that took human teams months, though only half the tested systems worked.
University of New Hampshire researchers used AI to scan 67,000 compounds and find alternatives to rare earth magnets critical for EVs and clean energy.
Mount Sinai researchers tested 20 LLMs with over a million prompts and found they readily accept false medical claims embedded in clinical-looking documents.
Researchers discovered that displaying an AI model's reasoning process creates a roadmap for attackers. OpenAI's o1 rejection rate dropped from 98% to under 2%.
An Emory study found that pairing clinical staff with AI tools improved accuracy in identifying eligible cancer patients without adding to workload.
Anthropic reports that explicitly permitting reward hacking prevents models from generalizing to sabotage and deception
AI systems mined Hubble and NEOWISE archives, surfacing 811 undocumented anomalies and 1.5 million previously uncataloged variable-object candidates.
UNH team built an AI that extracted magnetic data from 67,573 papers, identifying 25 new high-temperature magnets to replace rare earths in EVs.
Mass General Brigham's foundation model trained on ~49,000 brain MRIs outperforms specialized tools at predicting dementia, cancer survival, and mutations.
In a retrospective study of 6,401 cases, DeepRare's top diagnosis was correct 64.4% of the time, compared with 54.6% for specialists.
Research from ELLIS Alicante shows AI reasoning models can autonomously plan and execute attacks that bypass safety guardrails in nearly all other AI systems.
Singapore researchers combine AI with physics simulations to predict protein structures 13% more accurately than existing methods, covering 73% of the human proteome.
OpenScholar matches human expert citation accuracy while GPT-4o fabricates sources 78-90% of the time. The code, models, and 45 million paper corpus are all free to use.
Google's upgraded reasoning model finds flaws in peer-reviewed papers, optimizes semiconductor fabrication, and outperforms every frontier model on scientific benchmarks.