p.PABLO CARDOZOBUSINESS × ENGINEERING
EN/ES
navigateopen
Enter the lab

EXP_001 / LOCAL AI

Big models. Small hardware.

How far open-weight models go on an ordinary workstation — and what that changes in an AI stack.

STATUS
FIELD NOTES / VERIFIED
RUNTIME
llama.cpp
GPU
RTX 4070 · 12 GB
CONTEXT
~65K

01 The question

How far can modern AI models go without specialized infrastructure?

It started as an experiment with local LLMs and became a study of the whole inference stack: quantization, GPU offloading, KV cache, context windows, MoE architectures — and the trade-offs between memory, speed and quality.

02 What I learned

Good models no longer need a data center.

Open-weight models from Qwen, Llama, Gemma and DeepSeek showed that solid model quality is no longer exclusive to expensive cloud infrastructure.

With the right mix of quantization, GPU offloading, Flash Attention, KV-cache configuration and context management, models that looked impractical became genuinely useful on a regular workstation. The same holds for local image generation: high-quality results without depending entirely on external APIs.

Hardware constraints are often an optimization problem before they are a hardware problem.
FIG. 01 / MEMORY FIT01 / 04 · DOES NOT FIT
Illustrative, not to scale. How a model that does not fit becomes one that runs.

03 What it changes

From hypothesis to building block.

Local AI stopped being an interesting idea and became a practical component: internal tools that run on company infrastructure, keep sensitive workloads in-house, have predictable operating costs and stay available regardless of third-party API limits.

That does not make local models a replacement for frontier models. Demanding reasoning and high-throughput production still need serious compute. Local models do not have to replace them to be valuable.

Not just cheaper AI. A more resilient AI stack.

Frontier models where maximum intelligence matters. Local models where cost, latency, privacy, availability and specialization matter more.
FIG. 02 / HYBRID STACKFRONTIER MODELS
Routing by what each workload needs, not by habit.

04 My sweet spot

An ordinary workstation, tuned as one system.

A quantized Qwen 28B MoE model was one of the most interesting fits: small enough to run across GPU and system memory, with the quality I wanted for real development experiments.

GPU
RTX 4070 · 12 GB VRAM
MEMORY
64 GB system RAM
RUNTIME
llama.cpp
MODELS
Quantized 20–30B class
CONTEXT
~65K when needed

The key was not a single model or flag. It was understanding how the whole inference stack works together:

FIG. 03 / INFERENCE STACKSTAGE 01
  1. 01MODEL ARCHITECTURE
  2. 02QUANTIZATION
  3. 03VRAM ALLOCATION
  4. 04KV CACHE
  5. 05CONTEXT SIZE
  6. 06GPU OFFLOAD
  7. 07THROUGHPUT
Each stage constrains the next. Tune one, re-check the rest.

Tuned together, hardware that looked too small on paper became surprisingly capable.

The question changed

Can we run this locally?

Which parts of the system should run locally?