There is one question nearly every SMB that reaches out to 5ac has asked this past year: "Can I run AI inside my own company, or am I stuck depending on external APIs?"
The short answer in 2026 is: you not only can — you should. Local AI has crossed its tipping point. What was once a hobbyist playground built around a single RTX 3090 is now the default infrastructure for the enterprise. And for the small and medium business, where every dollar of cost and every byte of customer data counts, on-prem is no longer a technical option. It is a strategic decision.
This article is written from a CTO's vantage point: which hardware to buy, which software to run, which models to pick, and — most of all — why running AI inside your own walls is a real competitive edge for a small company in 2026.
Thesis: In 2026, local AI is no longer "running a big model on a big machine." It is a complete stack — hardware, inference engine, open-weight models — mature enough that an affordable workstation can replace a swath of cloud APIs for an SMB's everyday workloads.
Part 1 — Hardware: two numbers decide everything
Before we talk models, one thing has to be clear: capacity decides what fits, bandwidth decides how fast it runs. These two numbers are independent, and conflating them is the most common mistake beginners make.
The only VRAM formula you need
The "model → VRAM" relationship only holds once you account for how the weights were quantized. The cleanest version:
VRAM (GB) ≈ Parameters (billions) × (effective bits per weight ÷ 8)
Three numbers worth memorizing:
- FP16 / BF16 → 16 bits → ~2 GB per 1B params
- FP8 / INT8 → 8 bits → ~1 GB per 1B params
- 4-bit quants → ~4 bits → ~0.5 GB per 1B params
GGUF sits in between depending on the scheme: Q6_K ~0.82 GB/1B, Q5_K ~0.69, Q4_K ~0.56, Q3_K ~0.43, Q2_K ~0.33. The deeper you compress, the cheaper it gets — and the more quality you lose. Math, code, JSON, and tool use degrade first.
Which model fits on which GPU
Translated into hardware people actually own:
| VRAM | FP16 | FP8 | 4-bit |
|---|---|---|---|
| 8 GB | ~3B | ~6-7B | ~12-13B |
| 12 GB | ~5B | ~10B | ~18-20B |
| 16 GB | ~7B | ~13B | ~25B |
| 24 GB (RTX 3090/4090) | ~10-12B | ~20B | ~35-40B |
| 48 GB | ~20-24B | ~40B | ~70-80B |
| 80 GB | ~35-40B | ~70B | ~140B-class |
This is why the RTX 3090 (24 GB) is still the local-AI gold standard in 2026: at 4-bit it runs the entire 35-40B class — powerful enough for most office work, coding, and agents. The RTX 5090 (32 GB) pushes further still. You do not need a data-center cluster to start.
The VRAM tax nobody talks about: weights are only part of the bill. The KV cache grows with context length and quietly eats memory at 32K or 128K tokens. Activations, batching, and concurrency (especially in agent workloads) multiply memory fast. Safe rule of thumb: add 10-30% extra VRAM for a comfortable run; add more for long context or high concurrency.
Bandwidth — capacity decides what fits, bandwidth decides how hard it breathes
A 128 GB machine can hold a very large model, but with low bandwidth it produces 3 tokens per second — like decoding through wet cement. Bandwidth is the hardware number that really decides perceived speed.
The 2026 landscape:
- 1.8 TB/s class — RTX PRO 6000 Blackwell, RTX 5090 → 1792 GB/s. Speed kings.
- 800 GB/s class — Mac Studio M3 Ultra → 819 GB/s, up to 512 GB unified memory.
- 450-650 GB/s — Mac Studio M4 Max (546), MacBook Pro M5 Max (460-614), Radeon AI PRO R9700 (640), Tenstorrent Blackhole (512). Serious workstation tier.
- 250-300 GB/s unified-memory — DGX Spark (273), Mac mini M4 Pro (273), Ryzen AI Max / Strix Halo (256). Where unified memory starts paying off.
- Thin-and-light — MacBook Air M5 (153), Snapdragon X Elite (135), Intel Lunar Lake (136). Fine for small models, assistants, edge workloads — not for 9B dense playgrounds or multi-agent.
Three thresholds to remember: below ~150 GB/s is thin-and-light; around 250-300 GB/s unified memory gets interesting; 450-650 GB/s is a serious workstation; above 800 GB/s is expensive and powerful.
Three pragmatic hardware choices for an SMB
- Hobbyist / experimentation: one RTX 3090 or 4090 (24 GB) → 35-40B model at 4-bit. Enough for chat, coding, mid-size RAG.
- Mac-first: Mac Studio M3 Ultra (819 GB/s, up to 512 GB) → fits very large models on one quiet box using MLX. Slower raw tokens/sec than discrete GPUs, but what fits, fits.
- Enterprise on-prem: DGX Spark (128 GB unified, 273 GB/s, full NVIDIA stack) → not a bandwidth monster, but a developer appliance: coherent memory plus the whole NVIDIA software stack, with NVFP4 support. This is exactly the direction 5acAI-Lab is deploying for Vietnamese SMBs.
The biggest hardware lesson: stop asking "which hardware is best" and ask "which bottleneck am I buying?" — capacity, bandwidth, or software.
Part 2 — Software/Inference: the engine follows strategy, not the other way around
The first rule of inference engines: you do not pick an engine first. You pick a hardware strategy, a workload shape, and a serving model — the engine follows.
An engine is not "the model." It is the traffic cop, memory manager, scheduler, cache accountant, parallelism planner, API surface — and sometimes the deployment framework. The right engine matches your memory hierarchy, interconnect, quantization format, latency and throughput targets, and operational maturity.
A one-page decision guide
| Situation | Engine to use |
|---|---|
| Laptop / edge / odd hardware | llama.cpp |
| Mac-first workflows | MLX / MLX-LM |
| Single RTX local inference | ExLlamaV2 |
| 2-4+ NVIDIA / CUDA GPUs | ExLlamaV3 |
| General production serving | vLLM |
| Long-context / MoE / routing | SGLang |
| NVIDIA max performance | TensorRT-LLM |
| Cluster orchestration | NVIDIA Dynamo |
Prefill and decode: two phases, two problems
Every LLM inference has two phases:
- Prefill reads the prompt and builds the initial KV cache. It is compute-intensive — the more you parallelize, the faster it goes.
- Decode generates one token at a time, repeatedly reading weights and KV cache. It is memory-bandwidth-bound — decode speed tracks bandwidth more than peak compute.
That distinction explains almost everything: short prompt + long answer → decode dominates → bandwidth and batching matter. Long prompt + short answer → prefill dominates → attention kernels and chunked prefill matter. Many users → scheduler quality matters. Long context → KV cache dominates → paged attention, KV quantization.
The main engines
- llama.cpp — the portability king. Runs on Apple Silicon (ARM NEON, Accelerate, Metal), x86 (AVX/AVX2/AVX512/AMX), RISC-V, CUDA, AMD HIP, Vulkan, CPU+GPU hybrid offload. llama-server offers OpenAI-compatible routes, Anthropic Messages compatibility, continuous batching, JSON schema, function calling, speculative decoding. Limitation: not for serious multi-node production (RPC backend is proof-of-concept).
- MLX / MLX-LM — the Apple Silicon weapon. Unified memory lets large models fit on a Mac that a 24 GB discrete GPU cannot run. Slower raw speed than discrete GPUs. Its server warns it is not recommended for production.
- ExLlamaV2 / V3 — consumer CUDA engines for one or 2-4+ GPUs, tuned for low-bit quantized inference. V3 adds EXL3 (QTIP-based), tensor/expert parallelism, TabbyAPI OpenAI-compatible server.
- vLLM — the default for open-source production serving. PagedAttention, continuous batching, chunked prefill, prefix caching, broad quantization support (FP8, GPTQ, AWQ, GGUF...). It still needs systems thinking: tune batching, context, parallelism.
- SGLang — vLLM's systems-brained cousin. RadixAttention prefix caching, prefill-decode disaggregation, strong with structured outputs, long context, MoE.
- TensorRT-LLM — maximum NVIDIA performance. FP8 (H100+) doubles performance and halves memory versus 16-bit. You trade portability for performance.
A warning worth remembering: many tutorials recommend Ollama because it is convenient. For production serving, stay away from it. Production means security, observability, backpressure, routing, autoscaling, and SLA behavior — none of which a casual local runner guarantees.
Do not choose on a single "tokens/s" number
A bad benchmark: "I got 180 tok/s." A good one records the exact model, weights/quant, engine version, hardware, input/output distribution, concurrency, and tracks TTFT, TPOT, p50/p95/p99, memory usage, cost per 1M tokens. The golden rule: never compare engines on single-user tokens per second; test your actual prompt and output distribution, with realistic concurrency.
Part 3 — Models: open-weight is now the frontier
The biggest shift of 2026 is on the model side. The "Llama vs. everything else" era is over. Open-weight is now an ecosystem choice: weights, license, tokenizer, template, quantization, runtime support, serving path. And Chinese labs are leading this race.
Four names worth tracking in 2026
- Kimi K3 (Moonshot AI) — the largest open-weight model out of China: 2.8T parameters, MoE 16/896 experts, 1M context, Intelligence Index 57 — the highest of any open-weight model. #1 on the Agent board, #2 on WebDev. Open weights landed July 27, 2026 (96 shards, ~1.56 TB). Custom license (not MIT).
- DeepSeek V4 — 284B Flash / 1.6T Pro, open source, 1M context, Hybrid Attention Architecture. Flash is the cheapest frontier-tier API model ($0.14/$0.28), Pro at $0.44/$0.87. Lindy.ai moved to DeepSeek V4 and improved performance on core use cases.
- GLM 5.2 (Z.ai) — within one percentage point of Opus 4.8 on agentic benchmarks at roughly one-fifth the cost. Strong for coding agents, long-horizon tasks, MoE, deployment.
- Qwen 3.5 / 3.6 — an open-weight family spanning the whole stack: small models for laptops, dense mid-size for workstations, MoE for multi-GPU, FP8, long context (27B dense up to 262K tokens), multilingual, coding, agentic. Qwen is a sensible default family when you want one ecosystem from laptop experiments to serious serving.
A note on model names: the 2026 market moves fast. At the time of writing, the verified releases are Qwen 3.5/3.6 and GLM 5.2. Always check the latest model card before choosing — but that does not change the overall picture: Chinese open models are delivering high performance at low cost.
MoE: do not fall for the total-parameter trap
Mixture-of-Experts confuses people. "8x7B" sounds like 56B, but only a subset of experts runs per token. So compute cost ≠ memory cost. Total parameters decide the memory footprint; active parameters decide speed. DeepSeek-V3 reports 671B total but only 37B active per token; Kimi K3 is 2.8T total, 104B active. Treat MoE like dense and you will misjudge badly — either over- or under-estimating memory.
Quantization: GGUF is not magic
GGUF is a container plus a quantization strategy optimized for llama.cpp-style inference and CPU+GPU hybrid setups. But those memory numbers only apply in that runtime — move to another framework and weights may be dequantized, pushing memory up sharply. "It fits in 6 GB" is runtime-specific truth, not universal truth.
The 2026 rule: FP16/BF16 is the quality baseline; Q8/INT8 is near-lossless but still large; Q4 is the default consumer sweet spot for chat and documents; Q3/Q2 only when you must fit a bigger model — and that is where math, code, structured output, and tool use degrade first. A smaller model at higher precision can beat a larger one crushed into too few bits. Do not worship parameter count.
Why an SMB should run local
When open models reach this quality at this cost, the reasons to run local become obvious:
- Predictable cost: pay once for hardware, not pay-as-you-go token bills that can surprise you.
- Privacy: customer data, financials, code never leave your walls.
- No rate limits: no throttling, no sudden rule changes.
- No vendor lock-in: you own the weights; swap models when the market improves.
Part 4 — Why on-prem is the right call for SMB and enterprise
Now the story leaves the hobbyist and enters the business. For a Vietnamese SMB — 800,000+ businesses, over 95% of them SMB, average IT budget of $5,000-30,000 a year, AI spend still low ($500-3,000 a year) — every dollar and every decision counts. And on-prem AI answers those pains directly.
Four strategic reasons
- Data sovereignty: more SMBs do not want to send customer and financial data to foreign clouds. GDPR, PDPA (personal data protection law) push on-prem AI. For Vietnamese firms that already prefer open-source, this is a natural advantage.
- Operating cost: token costs fall 10x every 18 months, but an AI agent team running around the clock on APIs can balloon into an unpredictable bill. Run local, and cost becomes a predictable fixed line item.
- Reliability and control: no dependence on a provider's uptime, price changes, or export controls. You are your own ops team — harder, but with absolute control.
- Agentic workloads: multi-agent and tool use create high concurrency, which makes KV cache and latency major problems on APIs. Local lets you control batching, context, and the cost of each agent loop.
Where 5acAI-Lab fits
5ac.vn is built on two pillars: One-Person Company and Local AI Engine. 5acAI-Lab is the Local AI piece: an on-prem engine built around DGX Spark — 128 GB unified memory, 273 GB/s, the full NVIDIA stack with NVFP4 — designed so Vietnamese small and medium businesses can run agentic AI inside their own walls.
We are not claiming on-prem is the answer to everything. For work that needs the absolute best model quality and you lack the hardware to match, a hosted API is still the right tool. But for the majority of SMBs — where private data, stable cost, and autonomy matter more than a few benchmark points — on-prem is already the smart default.
"Once you understand your VRAM math, bandwidth tier, and engine, you stop asking 'can I run this model?' and start asking 'how do I want to run it?' — that is when you start designing systems instead of guessing."
Conclusion: local AI in 2026 is the default, not the exception
The journey from a hobbyist RTX 3090 to enterprise on-prem is no longer a long one. The same VRAM math, the same inference engines, the same open-weight models — just a different scale. What changed in 2026 is maturity: hardware cheap enough, engines good enough, open models strong enough.
For an SMB, the question is no longer "should I run local AI." It is "where do I start, with which model, on which hardware, and for which workload first." Start small: an RTX 3090 for experimentation, a matching engine, an open model that fits. Once you see the value, scale into serious on-prem.
At 5acAI-Lab we are building exactly that path for Vietnamese businesses. If you are weighing whether to run AI inside your own walls, and want to understand which stack fits your scale and your data, reach out to 5acAI-Lab — we walk with you from your first GPU to a complete on-prem infrastructure.
Next step: Read the "Local LLMs From Zero to Hero" series (Vietnamese translation on the 5ac blog) to master GPU Memory Math, Memory Bandwidth, Inference Engines, LLMs 101, LLM Engineering Projects, and Evals/Benchmarks. That is the solid foundation before you buy your first piece of hardware.