local-llms, gpu, self-hosting, ai-development
Four GPUs, Two Weeks, and the Uncomfortable Truth About Local LLMs
What happens when you throw 96 GB of VRAM at open-source models, optimize every last flag, profile CUDA kernels, redesign system prompts, and still end up reaching for a cloud API. It started with a simple premise: why pay per token when you have the hardware? I have four