Deploy a local AI server for $1,500 using consumer GPUs. Step-by-step guide on building with DeepSeek-R1, RTX 4090, tensor parallelism & cost ROI. In September 2023, to continue my exploration and research on Large Language Models (LLMs) outside of work, I assembled a dual RTX 4090 personal AI server. It has been running for nearly a year, and here are some observations: Noise: Placed under my desk, the fans can get quite loud under full. In today's AI-driven world, the ability to train AI models locally and perform fast inference on GPUs at an optimal cost is more important than ever. Building your own GPU server with an RTX 4090 or RTX 5090 — like the one described here — enables a high-performance eight-GPU setup running on PCIe. Can two RTX 4090s run Llama 70B or Qwen 72B? We break down the real performance of dual 4090 setups for local LLM inference — VRAM pooling, PCIe limitations, and what models you unlock. A single RTX 4090 is the best consumer GPU money can buy for local AI inference. This tutorial walks through constructing a dedicated inference machine for around $1,500, from component selection through a working OpenAI-compatible API endpoint. The goal is a reasonable configuration for running LLMs, like a quantized 70B llama2, or multiple smaller models in a crude Mixture of Experts layout. 24GB flagship that unlocks 70B+ models and top tokens/sec. Ready to assemble with standard tools.