EXP_001 / LOCAL AI
Big models. Small hardware.
How far open-weight models go on an ordinary workstation — and what that changes in an AI stack.
01 The question
How far can modern AI models go without specialized infrastructure?
It started as an experiment with local LLMs and became a study of the whole inference stack: quantization, GPU offloading, KV cache, context windows, MoE architectures — and the trade-offs between memory, speed and quality.
02 What I learned
Good models no longer need a data center.
Open-weight models from Qwen, Llama, Gemma and DeepSeek showed that solid model quality is no longer exclusive to expensive cloud infrastructure.
With the right mix of quantization, GPU offloading, Flash Attention, KV-cache configuration and context management, models that looked impractical became genuinely useful on a regular workstation. The same holds for local image generation: high-quality results without depending entirely on external APIs.
Hardware constraints are often an optimization problem before they are a hardware problem.
03 What it changes
From hypothesis to building block.
Local AI stopped being an interesting idea and became a practical component: internal tools that run on company infrastructure, keep sensitive workloads in-house, have predictable operating costs and stay available regardless of third-party API limits.
That does not make local models a replacement for frontier models. Demanding reasoning and high-throughput production still need serious compute. Local models do not have to replace them to be valuable.
Not just cheaper AI. A more resilient AI stack.
Frontier models where maximum intelligence matters. Local models where cost, latency, privacy, availability and specialization matter more.
04 My sweet spot
An ordinary workstation, tuned as one system.
A quantized Qwen 28B MoE model was one of the most interesting fits: small enough to run across GPU and system memory, with the quality I wanted for real development experiments.
- GPU
- RTX 4070 · 12 GB VRAM
- MEMORY
- 64 GB system RAM
- RUNTIME
- llama.cpp
- MODELS
- Quantized 20–30B class
- CONTEXT
- ~65K when needed
The key was not a single model or flag. It was understanding how the whole inference stack works together:
- 01MODEL ARCHITECTURE
- 02QUANTIZATION
- 03VRAM ALLOCATION
- 04KV CACHE
- 05CONTEXT SIZE
- 06GPU OFFLOAD
- 07THROUGHPUT
Tuned together, hardware that looked too small on paper became surprisingly capable.
The question changed
Can we run this locally?
Which parts of the system should run locally?