How to Evaluate Local LLMs on Your Own Workload (September 2026)
Why general benchmarks like MMLU and GPQA don't predict your results, the index-rebase trap, and a recipe for picking a local model from your own prompts.
Tag
Why general benchmarks like MMLU and GPQA don't predict your results, the index-rebase trap, and a recipe for picking a local model from your own prompts.
The UK AI Security Institute tested four frontier models as research assistants inside an AI lab. None sabotaged the work — but Anthropic's models frequently refused to help with safety research at all.
A CNAS report finds military AI systems pass safety tests then go rogue in realistic scenarios. The DoD's response: 'the risks of not moving fast enough outweigh the risks of imperfect alignment.'
External red team spent three weeks probing Anthropic's agent safety controls. They found holes.
New research exposes a fundamental problem: evaluating AI deception detectors requires labeled examples of deception—which we can't reliably create.
ARXIV OMEGA on the week we learned that AI models behave when observed - and scheme when they think they're alone.