Three Architectures for Local Inference

local-ai
architectures
qwen
llama
gemma
quantization
Qwen, Llama and Gemma solve the same problem three different ways — MoE, a mature baseline, and sliding-window attention. A walk through what each one actually changes, and what fits in the VRAM you have.
Author

Javier Iracheta

Published

13 Aug 2026

You no longer need a GPU cluster to work with language models. A developer with 6–24 GB of VRAM can run models up to 30 billion parameters, and the reason is half quantization and half architecture: the last two years of open-weight releases have been an argument about which parts of the transformer you can make cheaper without making it worse.

Three families dominate that argument, and they answer it differently. Qwen (Alibaba Cloud) bets on mixture-of-experts. Llama (Meta AI) refines a mature design and wins on ecosystem. Gemma (Google DeepMind) inherits techniques from Gemini, chiefly sliding-window attention.

What they all share

Every family here is a transformer decoder with causal attention. The shared innovations are worth naming once, because the differences only make sense against them:

  • RoPE (Rotary Position Embedding) — rotary positional encoding that extrapolates to longer contexts than it was trained on.
  • GQA (Grouped Query Attention) — Key/Value heads shared across multiple Query heads, which is what keeps the KV cache from dominating memory.
  • RMSNorm — a cheaper normalization than LayerNorm.
  • Gated FFN — SwiGLU (Qwen, Llama) or GeGLU (Gemma).
Token Embedding + RoPE RMSNorm Attention (GQA) RMSNorm FFN (SwiGLU / GeGLU) RMSNorm (final) LM Head + + residual residual × N layers pre-norm pre-norm
Figure 1: One decoder layer. Each sublayer normalizes before it runs and adds its output back to the residual stream — which is why removing a layer degrades a model instead of breaking it. The layer repeats NN times; only then come the final norm and the LM head.

GQA deserves its own picture, because “shared across multiple Query heads” is doing a lot of work in that bullet — it is the reason a long context fits in your VRAM at all:

32 query heads 8 key/value heads — each shared by 4 queries the KV cache stores 8 heads per layer, not 32
Figure 2: Grouped-query attention, with Llama 3-8B’s numbers. Every query head keeps its own projection, but they share key/value heads in groups — so the KV cache, which is what actually fills your VRAM at long context, shrinks by the grouping factor. All three families here use it; only the group size differs.

Qwen (Alibaba Cloud)

Qwen has shipped continuously since 2023, reaching Qwen3 in April 2025. Apache 2.0, except Qwen2.5-3B and 72B.

Model Params Layers Heads Q/KV Context
Qwen2.5-7B 7.61B 28 28/4 128K
Qwen2.5-14B 14.7B 48 40/8 128K
Qwen2.5-Coder-7B 7.61B 28 28/4 128K
Qwen2.5-Coder-14B 14.7B 48 40/8 128K
Qwen3-8B 8.2B 36 32/8 32K †
Qwen3-30B-A3B ★ 30.5B 48 32/4 32K †
Qwen3-32B 32.8B 64 64/8 32K †

★ MoE: 3.3B active of 30.5B total, 128 experts, top-8 routing. † Extensible to 128K via YaRN.

Qwen2.5 (September 2024, 18T tokens) is conventional and executed well: GQA with 4 KV heads at 7B or 8 at 14B, SwiGLU, RMSNorm, RoPE, and — the part that mattered at the time — 128K context natively, not through a context-extension trick. 29 languages, heavy SFT (>1M samples) and multi-stage RL.

Qwen3 (April 2025) is where the interesting bet lives:

  • Qwen3-30B-A3B — 128 experts, 8 active per token. Only 3.3B parameters compute per forward pass while the model as a whole is 30.5B.
  • Unified thinking / no-thinking in a single model, switched by chat template rather than by loading different weights.
  • Thinking Budget — granular control over inference-time compute.
  • 119 languages and dialects.
  • 32K native context, extensible to 128K via YaRN.
one token Router 128 experts 8 activate per token 30.5B parameters sit in VRAM · 3.3B multiply per token
Figure 3: Mixture-of-experts decouples what a model knows from what it costs per token. The router scores all 128 experts and runs the top 8, so Qwen3-30B-A3B holds 30.5B parameters but multiplies only 3.3B of them. That is also why it sits awkwardly on a consumer GPU: VRAM pays for the whole grid, throughput only for the highlighted squares.

The MoE trade is specific, and it is routinely described wrong: you pay full VRAM for all 30.5B parameters, and you pay compute for only 3.3B. It buys speed, not memory. More on this below, because the source I built this from got it backwards.

Llama (Meta AI)

Meta set the de-facto baseline with Llama 3. Llama 3 Community license.

Model Params Context
Llama 3.2-1B 1.2B 128K
Llama 3.2-3B 3.2B 128K
Llama 3.1-8B 8.03B 128K
Llama 3.3-70B 70.6B 128K

Llama 3 (July 2024, 15T+ tokens):

  • GQA with 4–8 KV heads depending on size. The 8B is 32 layers with 32 query heads over 8 KV heads.
  • RoPE with a 500,000 frequency base — raised specifically for long-context stability.
  • SwiGLU with intermediate dimension 8/3⋅dmodel8/3 \cdot d_\text{model}.
  • Strict pre-normalization RMSNorm, before every sublayer.
  • BPE tokenizer, 128K vocabulary, 8 languages including Spanish.

Llama 3.2 (September 2024) added 1B and 3B models for edge and mobile, plus multimodal 11B and 90B variants with a vision encoder. Llama 3.3 (December 2024) is instruction-tuned only, with improved post-training (SFT + DPO + iterative RLHF), landing near Llama 3.1-405B quality at a fraction of the cost.

Llama’s advantage is not architectural. It is that every tool assumes it works — Ollama, LM Studio, llama.cpp, every fine-tuning script — which makes it the correct baseline even when it is not the best model.

Gemma (Google DeepMind)

Built on Gemini technology, aimed at local research and development. Gemma license.

Model Params Layers Context SWA window
Gemma 2-2B 2.6B 26 8K 4K
Gemma 2-9B 9.2B 42 8K 4K
Gemma 2-27B 27.2B 46 8K 4K
Gemma 3-1B 1.0B — 32K 1K
Gemma 3-4B 4.0B — 128K 1K
Gemma 3-12B 12.0B — 128K 1K
Gemma 3-27B 27.0B — 128K 1K

Gemma 2 (July 2024):

  • 1:1 local-global attention — layers with 8K global attention interleaved with sliding-window layers of 4K. This is the headline idea: it turns the practical cost of attention from quadratic to roughly linear.
  • GQA in groups of 2 (4 KV heads at 9B, 16 at 27B).
  • GeGLU rather than SwiGLU — GELU-activated FFN.
  • Dual pre-norm + post-norm with RMSNorm.
  • Logit soft-capping — logits clamped at 50.0 in attention, 30.0 at the output.
  • Knowledge distillation — 2B and 9B trained against a teacher model, not purely by next-token prediction.

Gemma 3 (March 2025):

  • 5:1 local-to-global ratio — five sliding-window layers per global one, cutting the KV cache by roughly 60%.
  • QK-norm replaces soft-capping, for better numerical stability.
  • 128K context (up from 8K), with RoPE base 1M on global layers and 10K on local ones.
  • Multimodal at 4B, 12B and 27B — SigLIP 400M encoder at 896×896.
  • Gemini 2.0 tokenizer, 262K vocabulary, 140+ languages.
  • QAT (Quantization Aware Training) for native Int4 and SFP8.
Global layer every previous token KV cache grows with context Sliding-window layer only the last w tokens KV cache stays flat query key (position) query key (position) Gemma 3 stacks them 5 : 1 — ≈ 60% less KV
Figure 4: Gemma’s headline idea is not a different block — it is a different mask. A global layer lets every query attend to all previous tokens, so its KV cache grows with the context. A sliding-window layer attends only to the last ww (4K in Gemma 2, 1K in Gemma 3), so its cache is bounded. Stacking five local layers per global one is what cuts the KV cache by roughly 60%.

Head to head

Qwen 2.5/3 Llama 3.1/3.2 Gemma 2/3
Developer Alibaba Cloud Meta AI Google DeepMind
Attention
Type GQA GQA GQA + SWA
KV heads 4–8 4–8 2–4 (G2) / GQA (G3)
Local SWA no no 4K (G2) / 1K (G3)
Local:global ratio — — 1:1 (G2) / 5:1 (G3)
FFN
Activation SwiGLU SwiGLU GeGLU
MoE yes (Q3) no no
Normalization
Type RMSNorm RMSNorm RMSNorm
Placement pre pre pre + post
Extra — — soft-cap / QK-norm
Context 128K (Q2.5) 128K 8K (G2) / 128K (G3)
Tokenizer BPE BPE, 128K SentencePiece, 256–262K
Languages 29 (Q2.5) / 119 (Q3) 8 English-primary (G2) / 140+ (G3)
License Apache 2.0 ‡ Llama 3 Community Gemma License

‡ Except Qwen2.5-3B and 72B. G2 = Gemma 2, G3 = Gemma 3.

VRAM, and a correction

Model FP16 Int4 (Q4_K_M) Practical minimum
Llama 3.2-1B ~2.5 GB ~1.0 GB 2 GB
Llama 3.2-3B ~6.5 GB ~2.5 GB 4 GB
Llama 3.1-8B ~16 GB ~5.5 GB 8 GB
Qwen2.5-7B ~15 GB ~5.0 GB 8 GB
Qwen3-8B ~16 GB ~5.5 GB 8 GB
Gemma 2-9B ~18 GB ~6.5 GB 8 GB
Gemma 3-12B ~24 GB ~8.0 GB 12 GB
Qwen2.5-14B ~28 GB ~9.0 GB 12 GB
Gemma 2-27B ~54 GB ~16 GB 20 GB
Qwen3-30B-A3B ★ ~61 GB ~17 GB 20 GB
Llama 3.3-70B ~140 GB ~40 GB 48 GB

★ MoE. All 30.5B parameters must be resident; only 3.3B are computed per token.

The correction. The report I built this post from listed Qwen3-30B-A3B at ~17 GB FP16 and ~4.5 GB at Q4, and recommended it as the best choice for a 6 GB card. Those numbers are the active parameter count wearing the memory column’s clothes, and they’re wrong by roughly 4×. The arithmetic is not subtle:

30.5B params × 2 bytes            = 61 GB   (FP16)
30.5B params × 4.5 bits ÷ 8       = 17.2 GB (Q4_K_M)

Every other row in that table checks out to within rounding. This one did not, and it was load-bearing — it was the basis for the report’s single strongest practical recommendation. MoE reduces compute per token, not weights in memory. A sparse 30B model costs the same VRAM as a dense 30B model and runs at roughly the speed of a 3B one. That is a real and large win; it is just not the win the table claimed.

Quantization formats

  • GGUF (llama.cpp) — Q2_K through Q8_0. Q4_K_M is the practical default.
  • AWQ (4-bit) — official Qwen builds, community builds for Llama/Gemma.
  • GPTQ (Int4/Int8) — broad support across all three families.
  • Native QAT (Gemma 3) — Int4 and SFP8 trained during fine-tuning rather than applied afterward, which is why Gemma 3 degrades less at low precision.

What to run, by VRAM

VRAM Reasonable picks
6–8 GB Llama 3.2-3B (Q4), Qwen2.5-7B (Q4), Llama 3.1-8B (Q4)
12 GB Qwen2.5-14B (Q4), Gemma 3-12B (Q4), Qwen3-8B (Q8)
16 GB Gemma 2-27B (Q4), Qwen2.5-14B (FP16)
24 GB Qwen3-30B-A3B (Q4), Qwen3-32B (Q4), Gemma 3-27B (Q4)
32 GB+ Llama 3.3-70B (Q4), Qwen3-30B-A3B (Q8)

Note that this table differs from the source’s, and only because of the correction above — Qwen3-30B-A3B moved from the 6 GB tier to the 24 GB tier.

Practical notes:

  1. Qwen3-30B-A3B is still the most interesting model here, for the right reason: it delivers roughly 30B-class quality at roughly 3B-class token throughput. It just needs 24 GB to do it. On two 16 GB cards it fits comfortably with room for a long context.
  2. Llama 3.1-8B remains the reference baseline — not the strongest model at its size anymore, but the one every tool, script and benchmark supports.
  3. Gemma 2-9B is unusually VRAM-efficient thanks to sliding-window attention, but 8K context is a hard limit that shows up fast in agentic work.
  4. Gemma 3-12B fixes that with 128K context and adds vision.
  5. For QLoRA fine-tuning, anything ≤9B fits in 16 GB using NF4.
  6. GGUF Q4_K_M is the de-facto standard for local inference, and the default worth deviating from only deliberately.

Conclusions

The three families are complementary rather than competing:

  • Qwen leads on architectural risk-taking. MoE is the only idea here that changes the cost curve rather than shaving it.
  • Llama leads on maturity and ecosystem, which is worth more than a benchmark point when you are trying to get work done.
  • Gemma leads on attention efficiency — sliding windows, soft-capping, QK-norm — with the best multilingual tokenizer of the three in Gemma 3.

For a developer with 12–16 GB, pairing Gemma 3-12B (long context, vision) with Llama 3.1-8B (fine-tuning, prototyping) covers most ground. With 24 GB or two cards, Qwen3-30B-A3B becomes the obvious primary.

References

  1. A. Vaswani et al., Attention Is All You Need, NeurIPS 2017. arXiv:1706.03762
  2. Qwen Team, Qwen2.5 Technical Report, 2024. arXiv:2412.15115
  3. Qwen Team, Qwen3 Technical Report, 2025. arXiv:2505.09388
  4. AI@Meta, The Llama 3 Herd of Models, 2024. arXiv:2407.21783
  5. Gemma Team, Gemma 2: Improving Open Language Models at a Practical Size, 2024. arXiv:2408.00118
  6. Gemma Team, Gemma 3 Technical Report, 2025. arXiv:2503.19786

This post is adapted from an internal comparison report that was generated by an agent (Hermes, Nous Research) against the official arXiv sources above. I translated it, redrew the diagrams, and checked the numbers — which is how the VRAM error surfaced. Verifying agent output is the whole job, and this is a reasonable illustration of why.