A 24/7 local AI server turns a one-time hardware purchase into a privacy-respecting chat and embedding endpoint that no third party can see. The trade-off is that you are now responsible for the power draw, the fan noise, the operating-system updates, and the storage budget for model weights. This primer walks through the three hardware shapes that cover almost every realistic home deployment, then through the operating-system and runtime choices that fit on top.
Pick the Hardware Shape That Matches the Workload
There are three shapes worth comparing directly, and the right one depends on which model sizes you intend to run.
The low-power mini PC class (an Intel N100 or N305 class board in a fanless or near-silent chassis, with no discrete GPU) is the right answer when the planned workload is small models on CPU. The 0.6B-parameter Qwen3 fits in under 1 GB on disk and works comfortably on this hardware; even a 4B-parameter model runs in CPU mode at usable speeds for short prompts. The trade-off is that anything larger than roughly 8B parameters on a CPU-only mini PC will feel slow for chat and painful for batch embedding jobs.
The used-workstation with a 3090 shape is the most popular middle ground. A used OEM workstation or pre-built PC equipped with a 24 GB RTX 3090 sits at the top of the practical consumer tier. The card’s published power draw is 350 W (NVIDIA GeForce RTX 3090 specifications) and it carries 24 GB of GDDR6X, which is enough VRAM to load 8B through 32B dense models at sensible quantizations and still leave headroom for KV cache and a parallel request. Ollama supports the RTX 3090 explicitly under the compute-capability 8.6 (GeForce RTX 30xx) bucket in its hardware support table (Ollama GPU support).
The rack GPU or eGPU shape is what readers running a 70B or a 235B mixture-of-experts model need. Two or three used RTX 3090s in a desktop chassis, or a single multi-GPU workstation board, is enough for the heavier Qwen3-30B-A3B and Qwen3-32B models. The cost rises fast: power, cooling, and noise all scale together, and a 70B-class model will not fit in the 24 GB that a single 3090 offers.
Choose the Operating System First
The OS choice should be made before the hardware, because the runtime you pick depends on it.
Ubuntu Server LTS is the default for almost every Ollama-on-Linux installation, and is what the Docker images assume. It has the broadest driver coverage for NVIDIA, AMD ROCm, and Intel GPUs.
Proxmox VE is the right answer when the AI server is also expected to host other services (a media server, Home Assistant, a DNS ad-blocker, a vault). The official installation documentation lists 1 GB of RAM as the minimum for evaluation, with additional memory required per guest, and requires a 64-bit x86 CPU with Intel VT or AMD-V for full virtualization (Proxmox VE installation documentation). The hardware-requirements page confirms the same baseline (Proxmox VE hardware requirements). Ollama then runs inside an Ubuntu or Debian VM, with the GPU passed through via PCIe passthrough. The catch is that GPU passthrough on consumer AMD and Intel cards is hit-or-miss, and NVIDIA consumer cards disable their UEFI when passed through on some boards, which means Proxmox is a great fit for office machines but not for a single-purpose AI box.
Unraid sits in the middle: it is a flash-boot OS designed for home servers, and it runs Docker and VMs in parallel. It is a popular choice for mixed-purpose home servers, though Ollama’s ROCm support on consumer cards is the same as on bare Ubuntu, so there is no GPU advantage either way.
Bare Windows is fine if the box is also a desktop you want to use for other things. The Ollama installer ships a Windows package and uses the same Vulkan backend where possible.
Run the Runtime: What Actually Goes on Top
For most readers the runtime is Ollama with Open WebUI as the chat frontend. The Ollama server binds to 127.0.0.1 port 11434 by default and reads OLLAMA_HOST, OLLAMA_MODELS, OLLAMA_KEEP_ALIVE, OLLAMA_NUM_PARALLEL, and OLLAMA_MAX_LOADED_MODELS for its runtime configuration (Ollama FAQ). OLLAMA_MAX_LOADED_MODELS defaults to 3 times the number of GPUs, or 3 for CPU-only inference, so leaving the default is usually fine.
Open WebUI runs as a container next to Ollama. Its Quick Start documentation provides a single-container docker run command using the ghcr.io/open-webui/open-webui:main image, exposing port 3000 on the host, and storing data in a Docker volume. It supports macOS, Linux (x86_64 and ARM64, including Raspberry Pi and NVIDIA DGX Spark), and Windows, and currently supports Python 3.11 and 3.12 (Open WebUI Quick Start). The Quick Start also documents a separate image variant, ghcr.io/open-webui/open-webui:ollama, that bundles Ollama inside the same container.
The trade-off the runtime page does not advertise is image size: the standard image is around 1.66 GB, while a slim variant ships at about 176 MB and is meant for the same use cases where you do not need the bundled document-parsing tooling.
Storage and Power Cost, Worked Out
Storage for model weights is the easiest line item to get right. A small Qwen3 tag for qwen3:0.6b is 523 MB on disk, the qwen3:1.7b tag is 1.4 GB, and the qwen3:4b tag is 2.5 GB (Ollama qwen3 tags). The same page lists qwen3:8b at 5.2 GB. A practical home library of small-to-mid models fits comfortably on a 256 GB SSD, with room for embeddings and chat history. A library that also pulls Qwen3-32B at 20 GB, Qwen3-30B-A3B at 19 GB, and Llama 3.3 70B at 43 GB needs closer to a terabyte of fast storage.
Power is the harder line item because it depends on duty cycle. The most recent EIA monthly average for US residential electricity is 18.31 cents per kilowatt-hour, for July 2026 (EIA Electric Power Monthly - Table 5.3). Worked against an RTX 3090 at its 350 W TDP running flat-out, that is roughly 6.4 cents per hour of pure inference (0.350 kW × $0.1831/kWh), or about $1.54 per day if the card is fully loaded around the clock. A mini PC in the 10-25 W range burns about 0.2 to 0.5 cents per hour, or roughly 4 to 11 cents per day, at the same rate. For most home deployments the model is loaded for chat and sits idle between prompts, so the real number sits between those extremes.
Internal Links Worth Reading First
Two articles already cover the pieces the primer does not. Local AI without a GPU covers which small models actually run on CPU and how to size RAM to fit them. Best used GPUs for local AI covers the budget side of the hardware question and pairs with the cluster’s Best GPUs for local AI for new-card pricing. Serve a local model to your household covers the network layer (Tailscale, MagicDNS, the Ollama host variable) once the server is up.
Bottom Line
A private home AI server is, at its core, an ordinary x86 box running Linux plus an Ollama container plus an Open WebUI container, sitting on the home network. The interesting choices are the shape of the hardware (low-power mini PC for small models and embedding, used 3090 workstation for chat at 30B-class, multi-GPU desktop for 70B-class), the operating system (Ubuntu for single-purpose, Proxmox for mixed-purpose), and the power budget (a mini PC costs roughly pennies per day to run, a 3090 workstation roughly a dollar per day at typical home chat duty cycles). Everything else is configuration.