Liquid AI Lands on Setapp While iOS Private Relay Leaks Real IPs
Liquid AI's 2.6B model ships via Setapp for offline Mac agents, while iCloud Private Relay leaks real IPs through passkeys in the same 24 hours.
Tag
Liquid AI's 2.6B model ships via Setapp for offline Mac agents, while iCloud Private Relay leaks real IPs through passkeys in the same 24 hours.
Local embedding and reranking models for RAG, sized by VRAM. Real GGUF file sizes, licences, and the runtime gap that stops Ollama reranking.
Open image models sized for real GPUs, with the text encoder counted in and the licence checked. FLUX, Z-Image, Qwen-Image and SD3.5 compared.
Anthropic's July 8, 2026 policy trains on consumer chats unless you opt out, and Gemini keeps human-reviewed chats for up to three years.
Local AI on a 6GB GPU: GTX 1660, RTX 2060, RTX 3050 and laptop cards. Real weight sizes for chat, coding, vision, speech, translation and RAG.
Find your VRAM tier, then the current open-weight pick for chat, coding, vision, speech, translation or agents. Fourteen guides, one index.
Apple caps its on-device model at 4096 tokens per session, and Google's own tables put Gemma 4 E2B at 25.0 decode tokens/sec on an iPhone 17 Pro CPU.
Apple Silicon has no discrete VRAM, so tier guides mislead Mac owners. The real ceilings are bandwidth and the GPU-usable slice of unified memory.
Quantization formats compared with real file sizes and bits per weight. Why Q4 does not halve a model, and why low quants generate faster.
Alibaba announced 2.4T-parameter Qwen3.8-Max open weights and a 27B sibling that Unsloth says will run in 17GB of RAM or VRAM.
Two back-to-back merges add MiniMax Sparse Attention and a Qwen2.5-VL style vision tower to llama.cpp, but every existing MiniMax-M3 GGUF must be regenerated.
A new industry letter asks Washington to protect open-weight AI while separating legitimate distillation from alleged theft of closed models.
Cisco released Antares-350M and Antares-1B open-weight models for vulnerability localization. They run locally and cost about 172x less than GPT-5.5.
A leaked board email proposed a local GPT-3-class model in 2022. OpenAI's later gpt-oss release shows how that strategy changed.
A Go-based botnet is scanning exposed Ollama, ComfyUI, n8n, Open WebUI, Langflow, and Gradio instances for AWS keys and Kubernetes tokens, QiAnXin XLab says.
Hugging Face disclosed a July 2026 breach run end-to-end by an autonomous agent. The defender was an open-weight model.
Two new open-weight models point in opposite directions: Bonsai brings a 27B model toward phones, while Inkling targets customization at data-center scale.
VIDRAFT_LAB posts Ourbox-35B-JGOS to Hugging Face: 20 tok/s on an 8GB laptop GPU, ~17 tok/s on a CPU-only server, 86.4% on GPQA Diamond.
PhantaField's Sophon PFG-1 whitepaper claims ~95x Nvidia HBM4 bandwidth via monolithic 3D stacking. No silicon yet. Here's why it matters anyway.
Ditch GitHub Copilot's $19/month subscription. Set up Continue.dev with Ollama for private, local AI code completion in VS Code — zero data leaves your machine.
Three weeks away and the leaderboard reshuffled. Kimi K2.6 brings 1T parameters under open weights, Qwen 3.6 stays the consumer GPU king, and DeepSeek V4-Flash proves too hungry for single-card setups.
Ditch GitHub Copilot's $10/month subscription. Set up free, private AI code completion in VS Code using Continue.dev and Ollama — runs entirely on your hardware.
DeepSeek returns with a 1.6T MoE monster under MIT license, Gemma 4's 31B dense model climbs to #3 on Arena AI, and ICLR 2026 papers point to what's next for local inference.
Stop paying Midjourney $30 a month. Set up FLUX on your own hardware with ComfyUI and generate unlimited images with zero content filters and full privacy.