My Hardware

local-ai
hardware
pcie
homelab
Two identical 16 GB GPUs, and one of them is worth a quarter of the other. The parts list is the boring half of this post — the interesting half is what the PCIe topology of a consumer board does to a dual-GPU inference box.
Author

Javier Iracheta

Published

13 Aug 2026

Every “my local AI rig” post is a parts list. Parts lists are the least interesting thing about a machine, because they describe what you bought, not what you got.

What I got is a box with two identical GPUs where one of them is worth about a quarter of the other, for reasons that have nothing to do with the cards and everything to do with where the motherboard let me plug them in. That’s the part worth writing down.

The parts list, quickly

Board MSI MPG B550 GAMING PLUS (MS-7C56), BIOS 1.J0 (March 2025)
CPU AMD Ryzen 7 5700G — 8 cores / 16 threads, 16 MB L3, boosting to ~4.67 GHz
GPU 0 NVIDIA GeForce RTX 5060 Ti, 16 GB, 180 W cap
GPU 1 NVIDIA GeForce RTX 5060 Ti, 16 GB, 180 W cap
iGPU Radeon Vega (Cezanne), integrated in the 5700G
RAM 58.7 GiB visible to the OS
Storage Crucial BX500 1 TB, SATA
Network Realtek RTL8111/8168 gigabit
OS Ubuntu 24.04.4 LTS, kernel 7.0.0-28
Driver NVIDIA 580.159.03, CUDA 13.0

Two 16 GB cards gives 31.9 GiB of VRAM, which is the number that decides what this machine can and can’t run. It’s enough for a 24B–35B model at 4-bit with a 32k context, and not enough for anything that wants to be a 70B at a quantization you’d defend in public.

The 5700G is doing far less work than the price of an 8-core suggests. On an inference box the CPU’s job is to feed the GPUs and stay out of the way, and this one does that. Its lanes, on the other hand, are the whole story.

The finding: one of these GPUs is on a different network

nvidia-smi will happily report two identical cards and let you assume they’re interchangeable. They aren’t. The link each one negotiated:

Card’s max link Negotiated Theoretical bandwidth
GPU 0 PCIe 5.0 x8 Gen 3 x8 ~7.9 GB/s
GPU 1 PCIe 5.0 x8 Gen 3 x4 ~3.9 GB/s

Both cards are natively PCIe 5.0 x8 parts. At Gen 5 that link would be ~31.5 GB/s. GPU 0 is running at 25% of what the card can do, and GPU 1 at 12.5%.

The PCIe tree explains why, and it’s worth reading the topology rather than trusting the slot labels on the board:

flowchart TD
    CPU["Ryzen 7 5700G<br>PCIe Gen 3 only"]
    GPU0["GPU 0<br>Gen 3 x8 — ~7.9 GB/s"]
    CHIP["B550 chipset"]
    GPU1["GPU 1<br>Gen 3 x4 — ~3.9 GB/s"]
    NIC["Realtek NIC<br>1 GbE"]

    CPU -- "00:01.1" --> GPU0
    CPU -- "00:02.1" --> CHIP
    CHIP --> GPU1
    CHIP --> NIC
Figure 1: PCIe topology of the board. GPU 0 hangs off a CPU root port; GPU 1 reaches the CPU through the B550 chipset, sharing that uplink with the onboard NIC.

GPU 0 hangs off a root port on the CPU itself — 00:01.1 straight to bus 10. GPU 1 does not: the path is 00:02.1 → 16:00.2 → 20:00.0 → bus 21, out to the B550 chipset, through a downstream switch port, and only then to the card. And look at what shares that switch — 20:09.0 → bus 2a is the onboard network controller. GPU 1 and the NIC are on the same chipset uplink.

Two separate things are going on here, and it took me a while to stop conflating them:

Gen 3 instead of Gen 5 is the CPU’s fault. The 5700G is a Cezanne APU, and Cezanne’s PCIe controller is Gen 3. It doesn’t matter that the board is B550 and advertises Gen 4, or that the cards are Gen 5. The lowest common denominator on the link wins, and the APU sets it. A Zen 3 non-APU — a 5800X in the same socket, same board — would negotiate Gen 4 on the primary slot and roughly double GPU 0’s link. That’s a CPU swap, not a rebuild.

x4 instead of x8 is the board’s fault. The second full-length slot on this class of board is chipset-fed and electrically x4. There is no BIOS setting that fixes that; the lanes physically aren’t there.

So the two problems have two different solutions, and only one of them is cheap.

Why this matters less than it sounds, and more than I expected

The reflex is to panic about PCIe bandwidth. Usually that reflex is wrong. Once a model’s layers are resident in VRAM, inference reads weights at the card’s memory bandwidth — 448 GB/s on a 5060 Ti — and the PCIe link is idle. It carries the prompt in and the tokens out, which is nothing.

The link stops being idle in exactly two situations, and this machine hits both.

Loading. Every model start pushes the whole file across the bus. A 14 GB model over GPU 1’s ~3.9 GB/s link is several seconds of pure waiting before anything happens, and that’s the optimistic number — see the storage section, because the disk gets there first.

Layer splitting. This is the one that actually costs throughput. With --split-mode layer, the activations for each token have to cross from the last layer on GPU 0 to the first layer on GPU 1 and back. That’s a small transfer, but it happens per token, synchronously, and it is latency-bound rather than bandwidth-bound. Sending it over a chipset-attached x4 link — which adds a switch hop on top of the narrow width — means every token pays for the topology.

That’s why the server that’s running on this box right now looks like this:

llama-server -m Devstral-Small-2-24B-Instruct-2512-UD-Q4_K_XL.gguf \
  -ngl 99 -dev CUDA0,CUDA1 -sm layer -ts 88,12 -mg 0 \
  -c 32768 --jinja -fa on -ctk q8_0 -ctv q8_0 -t 8

-ts 88,12. Two identical 16 GB cards, and the tensor split is 88/12. That is not a typo and it is not laziness. A 14 GB model at Q4 nearly fits in one card; pushing the last few layers onto GPU 1 buys the VRAM headroom needed for a 32k KV cache, and nothing more. Every layer beyond that would add a per-token round trip over the slow link for capacity I don’t need.

-mg 0 puts the main device — the one holding the KV cache and doing the scratch work — on the fast card. That’s the same decision stated a second way.

Note also -ctk q8_0 -ctv q8_0: the KV cache is quantized to 8-bit. At 32k context that roughly halves the cache footprint versus fp16, which is what makes the 88/12 split viable at all. VRAM pressure and PCIe topology are the same constraint wearing two hats.

The honest summary: this is a 1.5-GPU machine. The second card contributes VRAM capacity, and contributes it cheaply. It does not contribute proportional throughput, and any benchmark I publish from this box has to say so.

Storage is the real bottleneck, and it isn’t close

One SATA SSD. 915 GiB usable, 475 GiB used. The model store alone:

Model On disk
qwen3-coder-next-80b-iq4 36 GB
qwen3-coder-next-80b 27 GB
muse-glimmer-30b 22 GB
qwen3.6-35b-a3b 21 GB
ornith-1.0-35b 20 GB
qwen3.6-27b / qwen3-coder-30b 17 GB each
gemma4-26b-a4b 16 GB
devstral-24b 14 GB
gpt-oss-20b 13 GB
Total model store 207 GB

A SATA link tops out around 550 MB/s, and the BX500 is a DRAM-less drive that doesn’t sustain its own peak on long sequential reads. Loading the 36 GB model is a minute-plus of disk, every time, and it’s the disk — not the x4 PCIe link — that sets that number. Page cache absorbs it on the second load, which is why 58.7 GiB of RAM is more useful here than it looks: with 49 GiB currently sitting in buff/cache, a recently-used model comes back from memory instead of from SATA.

For a benchmark harness that swaps models between runs, this is the single biggest quality-of-life problem on the machine, and the cheapest to fix. An NVMe drive in an M.2 slot would cut cold model loads by roughly 5–7×. It’s the upgrade I’d do first, ahead of anything involving GPUs.

The 5 GiB that went missing

64 GB installed, 58.7 GiB visible. The gap is mostly the iGPU: Cezanne’s Radeon Vega carves a UMA block out of system RAM at boot, and the kernel never sees it.

The machine is headless. That graphics block is being reserved for a display nobody is looking at. Dropping the UMA frame buffer to its minimum in BIOS should return most of it — I haven’t done it yet, and I’m noting it here mainly so I stop rediscovering it. (I couldn’t read the DIMM configuration or memory speed directly; dmidecode needs root and I ran this inventory unprivileged. So I know how much memory the OS sees, not how many sticks or at what timing.)

What else this box is doing while it infers

This is not a dedicated inference appliance, which is worth being explicit about before anyone reads a latency number off it. It’s currently running 26 containers: three Postgres instances, several application stacks behind a Caddy edge, Plex, Open WebUI, Portainer, and the Hugo container serving a local mirror of this blog.

Load average at the time of writing: 0.68 over one minute, on 16 threads. The containers are mostly idle mostly of the time, and the CPU has room. But “mostly” is not “always”, and a benchmark run that happens to coincide with a Plex transcode is not a clean measurement. Any timing I publish from this machine needs to say what else was running, or it isn’t reproducible.

The upgrade order

Written down so I argue with a list instead of with an impulse:

  1. NVMe boot/model drive. Biggest real-world improvement per peso. Fixes cold loads, fixes the harness swap time, fixes nothing about inference throughput — and that’s fine, because inference throughput isn’t what’s annoying me.
  2. Drop the iGPU UMA reservation. Free, five minutes in BIOS, returns several GiB of page cache to a machine that would rather cache models.
  3. A non-APU CPU. A Zen 3 with a real PCIe Gen 4 controller would double GPU 0’s link and lift GPU 1’s from Gen 3 x4 to Gen 4 x4. Same socket, same board, same RAM. It would also cost me the iGPU, which on a headless box is not a loss.
  4. A different board. This is the only thing that fixes x4, and it’s the only item on the list that is really a rebuild. Not worth it until layer splitting becomes the bottleneck I actually care about — which, at 88/12, it isn’t yet.

Notice that the two GPUs don’t appear anywhere on that list. The cards are the part of this machine I’d change last. That’s usually the sign that the parts list was never the interesting question.