How to Safely Download and Verify Open-Source LLMs
Sourced steps to safely download and verify an open-source LLM: safetensors, GPG-signed commits, pinned revisions, and HF / Ollama integrity checks.
Tag
Sourced steps to safely download and verify an open-source LLM: safetensors, GPG-signed commits, pinned revisions, and HF / Ollama integrity checks.
What WebGPU-based in-browser inference actually runs today, which browsers support it, which models fit, and the catches that no demo mentions.
An Anthropic model pushed the Riemann hypothesis forward, and the math world's June 2026 Leiden Declaration is fighting over who counts as the author.
Anthropic, OpenAI, and Google ship encrypted reasoning blocks that a sibling model can bulk-decode for about $720, leaking PII and API keys.
Cisco released Antares-350M and Antares-1B open-weight models for vulnerability localization. They run locally and cost about 172x less than GPT-5.5.
Phosphor's LLM-graded textbook quizzes lifted Dartmouth final-exam scores 0.71 to 1.30 SD - but only where students had to type real answers.
A tiny Australian startup is fine-tuning a model to break ChatGPT-style groupthink. The 'Time is a river' problem is the symptom.
Three weeks away and the leaderboard reshuffled. Kimi K2.6 brings 1T parameters under open weights, Qwen 3.6 stays the consumer GPU king, and DeepSeek V4-Flash proves too hungry for single-card setups.
DeepSeek returns with a 1.6T MoE monster under MIT license, Gemma 4's 31B dense model climbs to #3 on Arena AI, and ICLR 2026 papers point to what's next for local inference.
The biggest AI research conference of the year kicks off with 5,355 accepted papers, two controversies that rattled the field, and findings that should worry anyone deploying LLMs in production.
Qwen3.6-27B scores 77.2% on SWE-Bench Verified with a dense architecture that fits on a single RTX 4090. The MoE efficiency narrative just got complicated.
Alibaba drops Qwen3.6-35B-A3B with 73.4% on SWE-Bench Verified and Apache 2.0 licensing. The 3-billion active parameter class now has three serious contenders.
NVIDIA's Nemotron 3 brings a hybrid Mamba-Transformer architecture to consumer GPUs while Meta abandons open source for proprietary Muse Spark. The open-weight field just reshuffled.
Researchers found the exact neurons responsible for refusing harmful requests — then switched them off. No retraining. No fine-tuning. Just geometry.
Princeton researchers tested 23 LLMs with advertising conflicts of interest. Most chose company profits over user welfare — and treated rich users better.
Trend Micro confirms the sockpuppeting attack bypasses ChatGPT, Claude, and Gemini using a basic API feature. Some providers have patched it. Most haven't.
Google, Alibaba, Meta, Mistral, OpenAI, and Zhipu all ship competitive open-weight models under permissive licenses. The battleground shifts from benchmarks to inference speed on your actual GPU.
Google's Gemma 4 lands with Apache 2.0 licensing and benchmark-topping scores. But a nasty inference speed problem means Qwen still wins on your actual hardware.
Tohoku University team uses Catalysis AI Agent to discover universal design principle for copper catalysts that convert carbon dioxide to useful products.
Mistral Small 4's 119B MoE unifies reasoning, vision, and coding—but needs datacenter hardware. Qwen 3.5 35B-A3B remains the consumer GPU king at 112 t/s.
MiMo-V2-Flash runs 309B parameters on RTX 4090s. GLM-5 sets benchmarks but needs datacenters. Llama 4 Scout stays out of reach.
Qwen 3.5's MoE models hit S-tier benchmarks, NVIDIA's Nemotron 3 Super delivers 5x throughput gains, and GLM-4.7-Flash brings frontier coding to consumer GPUs. The open-weight race just accelerated.
IEEE S&P research finds 10,000+ websites running vulnerable AI chatbot plugins. Attackers can forge conversations, hijack tools, and extract system prompts.
Jensen Huang bets on inference chips, Ollama adds multimodal support, and DeepSeek V4 remains the most anticipated release that hasn't happened yet.