Scoring a local model is easy. Scoring it well enough to rank one above another is a different problem — and my previous instrument could not resolve its own top two. Here is what I rebuilt and why.
Two identical 16 GB GPUs, and one of them is worth a quarter of the other. The parts list is the boring half of this post — the interesting half is what the PCIe topology of a consumer board does to a dual-GPU inference box.
Qwen, Llama and Gemma solve the same problem three different ways — MoE, a mature baseline, and sliding-window attention. A walk through what each one actually changes, and what fits in the VRAM you have.
It started as a merge of three tools I was tired of maintaining separately. It turned into a research assistant, and the reason is that a hypothesis and a requirement are the same shape — something stated, refined, and eventually resolved.
How to build a Hugo site into a multi-stage image and serve it with Nginx, without shipping the compiler inside the final image.