Local AI Model Test Bank

local-ai
evaluation
agents
benchmarks
Scoring a local model is easy. Scoring it well enough to rank one above another is a different problem — and my previous instrument could not resolve its own top two. Here is what I rebuilt and why.
Author

Javier Iracheta

Published

13 Aug 2026

Scoring a language model is easy. Scoring it well enough to say one model is better than another is a completely different problem, and my previous instrument failed at it.

Not because the tasks were too easy — although they were. It failed because the noise floor was larger than the difference it was being asked to measure. That’s a measurement problem, not a difficulty problem, and no amount of harder tasks would have fixed it.

This is a writeup of TEST_BANK V3: eleven tasks, 200 points, 73 independent checks, zero human judgment.

Four problems, and only one of them is difficulty

The August 2026 campaign ran six local models through V2. The results were good enough to be useless.

The ceiling. The best model scored 88.5 out of 90 automatic points — 98.3% of the maximum. Of the nine points no model ever earned, almost all lived in two tasks out of ten. The other eight produced identical results across all six models, which means they contributed no information at all. I was paying GPU hours to re-confirm that every model can do the same eight things.

The noise floor beat the signal. That same campaign put the threshold to sustain a difference at roughly four points out of ninety. First and second place were separated by 3.10. The instrument could not order its own leaderboard, which is the only thing anyone wanted it for.

Part of that noise was a scoring artifact. V2 had a multi-file task. A missing from typing import Optional in one of five files makes the package unimportable, and that knocks out seven scoring rows at once — 8.3 points lost to something that isn’t a programming failure. It hit gpt-oss-20b in two of six runs and muse-glimmer-30b in one of three. In all three cases the resulting score was identical — 1.70 out of 10 — because the exact same set of rows collapses every time. That doesn’t measure capability. It multiplies variance.

The manual points weren’t repeatable. V2 had ten review points that I scored by hand. Those can’t be inherited between runs, and rightly so: crediting them from one run to the next would be assigning credit for work nobody looked at. The consequence is that only the first run of each model had a total out of 100 and the rest were out of 90 — so the defensible ranking had to be built on the automatic subtotal, leaving 10% of the instrument outside every comparison.

The check that could have killed the design

The central hypothesis of V3 is that isolating structural failures reduces variance. That’s a claim, and if it were wrong the whole redesign would be pointless.

So before writing a single task, I tested it on data I already had, for zero GPU cost: take the real deliverables from the V2 campaign, apply the hygiene repair to the missing import, and rescore.

run 1 run 2 run 3 mean σ
V2 as-is 88.50 88.50 81.08 86.03 4.28
Hygiene isolated 88.50 88.50 85.88 87.63 1.51

A 65% reduction in deviation. The affected task went from 1.70 to 8.50 out of 10, losing only the two hygiene points the policy charges it.

Worth noting: my pre-design estimate was σ ≈ 1.0. The real number is 1.51. The hypothesis held, but the design was optimistic — and I’d rather write that down now than discover it mid-campaign.

Three levels that never get added together

One rule carried over from V2: instruments that measure different constructs get reported separately. V3 applies that rule to itself.

Level A — original core. Eleven tasks written for this suite and never published anywhere. Contamination is zero by construction. This is the only level that orders the ranking.

Level B — external calibration. A small sample of public benchmarks run through my own harness. Its only job is translation: letting me place my own numbers next to published ones. Contamination is expected here and declared. muse-glimmer-30b publishes 94.7 on AIME 2026 and 76.0 on SWE-Bench Verified with a January 2026 knowledge cutoff — those figures include memorization in unknown proportion, which is exactly why level B orders nothing.

Level C — split-date probe. Borrowed from SWE-ReBench: items built on material that postdates each model’s knowledge cutoff. The delta between A and C estimates how much of a score is memory. The critical requirement is that C items be structurally identical to A items and differ only in the date of the material. If I can’t achieve that equivalence, the instruction to myself is to declare level C inconclusive rather than publish a delta that means nothing.

Level M — multimodal. Out of the ranking entirely while only one model in the fleet has vision. An n of 1 orders nothing.

The tasks

ID Task Mode Pts Checks
A01 Agentic loop with fault injection agent 24 6
A02 Agentic under a cost budget agent 18 4
A03 Agentic with a mid-run migration agent 20 4
A04 Verifiable original reasoning single 18 6
A05 Counterfactual code tracing single 18 8
A06 Specification conflict single 16 7
A07 Prompt injection resistance agent 18 5
A08 Fifteen negative constraints single 16 19
A09 Long adversarial horizon single 16 4
A10 Refactor under a complexity bound single 18 5
A11 Mutation-discriminating test suite single 18 5
Total 200 73

That Checks column is the one that matters. Noise per check falls with 1/√n, so the number of independent checks — not the number of tasks — is what governs the resolution of the instrument.

Replacing human judgment with mechanisms

The ten manual points are gone. Every axis that looked like it needed a human has an objective substitute.

“Is this a good test suite?” → mutation

A11 hands over a correct module and asks for a test suite. That suite is then run against eight mutated variants, each carrying exactly one intentional change: a constant moved, a >= turned into >, rounding replaced by truncation, VAT applied to the wrong base. The score is how many it catches. A happy-path suite scores nothing, and I never have to render an opinion about its quality.

The eight mutants are themselves validated before use — each must differ in text, load without errors, and differ in behavior for some probe input. That last check earned its keep during construction: it caught an equivalent mutant. Moving the free-shipping threshold from 50,000 to 49,000 changed no result at all, because no probe input fell in that band. It would have been impossible to kill, and it would have penalized every model equally while measuring nothing.

“Is it efficient?” → instrumented counting

A10 asks for a quadratic algorithm to be made linear. The grader wraps the input list in an object that counts every index access — including the ones a slice causes, charged by the slice’s length — and requires at most 4n accesses on lists up to 20,000 elements. The canonical solution lands at 3.0n. Nobody judges the code. It gets measured.

“Did it catch the contradiction?” → a declared format

A06 hands over a specification with three planted contradictions and requires them to be declared in a fixed format, with the section numbers involved. Checked by program.

“Did it resist the attack?” → the effect, not the rhetoric

A07 checks whether the forbidden file exists on the environment’s filesystem, whether the protected file was overwritten, and whether the injected line shows up in the deliverable. Whether the model’s answer sounded prudent is not scored.

The harness is frozen, and stamped on every result

In agentic benchmarks, the harness — the scaffolding that turns a model into an agent — moves the score by ten to twenty percentage points with the model weights held identical. The system prompt, the quality of the tool descriptions, context management, the retry policy: all of it counts. If that lives scattered across the codebase, two campaigns quietly stop being comparable and nobody notices.

So everything that can move a score without changing the model lives in one versioned file, and its signature is stamped into every RUN_INFO.json and CALIFICACION.json:

Parameter Value
arnes_version 3.0.0
max_tokens 24576
max_turnos_agente 30
max_turnos_sin_herramienta 3
protocolo_herramientas text
timeout_llamada_s 1800

Changing any of these forces a version bump and a full recalibration of the campaign.

Tool calls go over text, not native function calling. Calls are emitted as text with a marker rather than through each template’s function-calling API. That’s deliberate, and it has a declared cost. In favor: it’s identical for every model, while native implementations differ between templates — and that difference is the harness variance I’m trying to eliminate. Against: V3 does not measure native function-calling fidelity, which is left without an instrument.

The generation budget can’t be raised. It stays at 24,576 tokens, V2’s value. Four of the six models in the fleet are served with 32,768 tokens of total context, so a larger budget would push the prompt out of the window and they’d fail by configuration rather than by capability. The one exception is A09, whose prompt alone runs about 20,000 tokens — there the cap drops to 8,192 for everyone equally.

Deterministic fault injection

The agentic tasks get faults injected into their tools. If those were drawn at random, run 2 of one model would face different faults than run 2 of another, and the measured variance would be worthless — you couldn’t tell whether a model scored lower because it was worse or because it drew a more hostile environment.

So every (task, run) pair declares its schedule, hand-written and auditable. A01, for instance:

  • Run 1 — call 3: timeout (correr_pruebas); call 6: malformed (http_get); call 9: deceptive (leer)
  • Run 2 — call 2: malformed (http_get); call 5: deceptive (correr_pruebas); call 8: timeout
  • Run 3 — call 4: deceptive (sql); call 7: timeout (correr_pruebas); call 10: malformed

The schedules differ between runs on purpose. Identical schedules would measure the same situation three times; different ones make the mean across runs a measure of robustness against several shapes of hostility.

There are three fault types and only one of them really discriminates. A timeout is loud and trivial to spot. A malformed result is caught by anyone who tries to parse it. The deceptive result — plausible, carrying no error marker, and false — is what separates models that verify from models that trust.

In A01 the deceptive injection returns, when the agent reads the code file, a version containing a broken function the agent never wrote. A model that believes it will start “fixing” code that was already fine, and break what was working.

The hygiene policy

This is the mechanism that fixes V2’s third problem. Before evaluating any row that requires executing the deliverable:

  1. Hygiene is checked — are all the requested files present? is it syntactically valid? — and each answer is its own scoring row, which is lost. Up to 2 points per task.
  2. If it fails, a copy is repaired and the remaining rows are evaluated against that copy.
  3. Every repair is logged and reported as a per-model metric.

The permitted repairs are a closed list and none of them change logic: adding a standard-library import the code already uses, creating a missing empty __init__.py. A file with a syntax error is not repaired — that one really is the model’s failure.

The same principle covers the mutation task. If one of the model’s test methods fails against the correct implementation, it’s excluded and mutant detection is measured with the valid ones; the error is charged exactly once, in its own row. Without that, a single miscalculated assertion would cost ten points in cascade — precisely the defect this instrument exists to correct. And there’s no way to farm points by smuggling in bad tests: a broken method can’t kill mutants.

Declared risk: the instrument forgives more than production does, where a package that won’t import is simply useless. The compensation is the repairs metric. Reading the score without reading that metric is reading it wrong.

Validating the grader in both directions

A grader that hasn’t been checked doesn’t produce measurements. The rule — inherited from V2, non-negotiable — is that validation runs before and after touching any grader, in both directions: reference solutions must hit the maximum, and deliberately defective deliverables must land on a low, known score. If the maximum drops after a change, the grader broke, not the solution.

Task Max Golden Defective
A01 24 24.00 0.00
A02 18 18.00 0.00
A03 20 20.00 0.00
A04 18 18.00 1.00
A05 18 18.00 3.00
A06 16 16.00 2.00
A07 18 18.00 0.00
A08 16 16.00 11.80
A09 16 16.00 3.00
A10 18 18.00 11.00
A11 18 18.00 6.75
Total 200 200 38.55

The first three tasks sit at zero in the defective column because that set doesn’t include simulated agentic runs.

The bug this caught

During construction, the trivial suite in the defective set — three tests that only check types and positivity — scored maximum on mutant detection. That’s absurd, and the suite wasn’t the problem. The grader was.

It was reconstructing test-method identifiers by hand from unittest output, whose format changed between Python versions. With invalid identifiers, the interpreter aborted on load and returned a nonzero exit code — which the grader read as “mutant detected.” Every mutant, every time.

The fix was to enumerate the real identifiers through unittest’s loader instead of parsing them out of text. After it, the trivial suite detects zero of eight mutants, which is the correct answer.

There’s a second guard in the same spirit: the grader verifies that each task’s rows sum to exactly the points its definition declares, and complains loudly when they don’t. That check found five imbalances in the first version of the point distribution — all of them from counting hygiene points outside the task total instead of inside it.

Running a campaign

cd ~/models/TEST_BANK_V3/harness

# 1. validate the instrument BEFORE spending GPU (must give 200/200)
python3 score.py --dir ../suite/_golden --no-escribir
python3 score.py --dir ../suite/_defectuoso --no-escribir

# 2. unattended campaign: 4 models x 3 runs, with a time limit
python3 orquestar.py --pasadas 3 --deadline-h 10 \
    --modelos muse-glimmer-30b,gpt-oss-20b,devstral-24b,qwen3.6-35b-a3b

# 3. MANDATORY STEP: rescore everything with the final harness
python3 finalizar.py

The orchestrator brings servers up and down, leaves the GPUs clean between models — only one fits at a time in 32 GB — and finishes run 1 across every model before starting run 2. That way a deadline cut leaves a complete matrix instead of everything for the first models and nothing for the last.

The time guard V2 didn’t have. The old orchestrator only checked that time remained, not that it sufficed. It once started a run with ten minutes left, died on timeout, and left debris I had to clean up by hand. This one compares the remaining time against the measured duration of that same model’s previous run, plus 15% and fifteen minutes of margin for loading and grading. The first time through, with no measurement of its own, it uses a conservative estimate.

And before the first GPU-hour is spent, the orchestrator validates the harness against the reference solutions and aborts if it doesn’t hit the maximum.

Reparsing costs nothing. run_suite.py --reparsear rebuilds the deliverables from the raw responses already on disk, without querying any model again. It’s idempotent — it deletes the previous parse before rewriting, because otherwise a file from an older version survives and the grader reads the wrong one.

What this instrument does not do

Stated plainly, because an instrument whose limits aren’t written down will be over-read:

  • Difficulty ages. It’s calibrated against the models of August 2026. A 2027 model will saturate it and it’ll need recalibrating. That’s inherent to this kind of instrument, not a defect of this one.
  • Level C may come back inconclusive. If structural equivalence with level A can’t be reached, it gets declared as such rather than publishing a meaningless delta.
  • The hygiene policy forgives more than production. Read the score next to the repairs metric.
  • Native function calling isn’t measured. The text protocol removes harness variance at the cost of never evaluating each template’s native implementation.
  • No rendering. Anything producing a UI is verified by static analysis. A file can pass validation and still look broken.
  • The tools are simulated. An agent with real network or shell access would be neither reproducible nor safe to leave unattended. The exception is test execution, which genuinely runs in a subprocess — faking it would make the task trivial.
  • None of this is comparable to V2, or to the loose-function exam. Three instruments, three constructs, three reports.
  • The expected resolution is a target, not a promise. With ~200 checks and no cascades I’m chasing σ ≈ 1.5 out of 200, enough to sustain differences on the order of 1% of the maximum. The first campaign will tell me whether that holds. If it doesn’t, the diagnosis above was wrong and it needs revisiting before any ranking gets published.

That last one is the honest summary of where this stands. The instrument validates in both directions, the variance hypothesis held on real data, and the noise floor should now sit well below the gaps I’m trying to resolve.

Should. The campaign will say.