What AgentX Is For

agentx
agents
research
mcp
dotnet
It started as a merge of three tools I was tired of maintaining separately. It turned into a research assistant, and the reason is that a hypothesis and a requirement are the same shape — something stated, refined, and eventually resolved.
Author

Javier Iracheta

Published

13 Aug 2026

AgentX began as housekeeping. I had three tools I’d built at different times for different reasons — a test management system, an agentic Kanban board, and a diagram versioner — and each one had its own database, its own login, and its own half-finished API. Merging them was supposed to be a weekend of yak shaving.

It became something else, and the turn happened at a specific place: when I went to model a hypothesis and realized I had already written that class.

The housekeeping part

The platform is a modular monolith — ASP.NET Core, EF Core on PostgreSQL, an Angular front end, one database, one identity core, one MCP server. Seven feature modules now hang off a shared kernel that owns users, projects, memberships, roles, personal access tokens, audit, and an in-process event bus.

Module What it holds
Builder Requirements, test cases, cycles, executions, defects, plans, coverage
Tracker Phases, tasks with claim/lease, dependency DAG, sessions, ideas, decisions, risks
Visualizer Views (wireframe / system design / database) with append-only versions and rollback
Experimenter Hypotheses, protocols, subjects, runs, measurements, aggregates
Writer The manuscript: a file tree, LaTeX compilation, bibliography
Sync Jira and Zephyr Scale connectors, field mapping, credential storage
Core Identity, tenancy, RBAC, passports, audit feed, event bus

The consolidation goal was narrow and it was met: one token, every module. An agx_pat_ bearer authenticates an agent against Builder, Tracker, Visualizer, Experimenter and Writer over both HTTP and MCP. Tenancy is a header (X-Project), the effective role is min(tokenRole, membership), and anything cross-tenant returns a uniform 404 rather than a 403 — because a 403 confirms the resource exists.

That’s the boring half. Here’s the half worth writing about.

A hypothesis is a requirement

When I started the Experimenter module I expected to design a new domain. I didn’t. The comment I ended up leaving on the entity says it better than a paraphrase would:

A claim under investigation. Shaped after Requirement — markdown body, project-configurable status, parent/child nesting — because a hypothesis is the same kind of object: something stated, refined, and eventually resolved.

The same thing happened one class down. A protocol — what you do, in what order, under what preconditions — is a test case. It even needed the same escape hatch: TestCase already carried an AutomationRef/AutomationType pair binding it to an external runner, which is exactly what the constraint the harness pushes, we never run it requires.

So HYP-1 behaves like REQ-1 and EXP-1 behaves like TC-1, deliberately. Test management, it turns out, is applied epistemology with worse branding: you state a claim, you design a procedure that could falsify it, you run the procedure, you record what happened, and you keep the trail. QA has been doing science badly-named for forty years.

What a hypothesis needed that a requirement didn’t is an edge to another hypothesis:

supersedes | refutes | supports | relates

supersedes and refutes require a note explaining why — the service rejects them without one. That constraint exists because of a specific failure: two contradicting papers ended up in the same tree with nothing connecting them, sitting side by side, both marked done. A status field can’t express “this one killed that one.” An edge can.

The two problems Experimenter actually solves

Everything else in the module is CRUD. These two are the reason it exists.

What is a subject, exactly?

A Subject is the thing under test — a model, a configuration, an agent, or a composite of several. The naive design gives each one a row and calls the row its identity. That breaks immediately: two people describing the same model with the same sampling settings create two subjects, and the results never join.

So identity is computed, not assigned. CompositionHasher derives a CompositionHash from what the subject is made of — recursively, for composite subjects — under a normalization rule whose version is baked into the hash input:

RuleVersion = "axh1"

Two writers describing the same model land on the same subject. Change one agent in a seven-agent matrix and you land on a different one. And because the rule version is inside the hash, a future change to normalization can’t silently collide with hashes written under the old one. The config that feeds it is contractually free of anything environmental — no endpoints, no hosts, no paths — so where you ran it doesn’t change what it was.

“Not comparable” is a useless answer

This is my favorite thing in the codebase, and it came out of being wrong.

When you compare two runs, the obvious API returns a boolean. I wrote that, and it was worthless — because not comparable tells you nothing you can act on. So compare returns a report of what differs, and each difference carries a severity:

Severity Meaning
blocking Changes what was measured. Comparing across it describes the setup, not the subject.
varied This is what the comparison is about — the independent variable.
notable Worth stating in a result, but doesn’t invalidate the comparison.

The varied tier was added after a real conduction embarrassed the design: two replicates of one treatment came back “not comparable” because their replicate labels differed. The verdict was useless at exactly the point the module was supposed to earn its keep. A factor that differs is the point of an experiment, not a contaminant.

The failure this addresses isn’t that people forget to hold factors constant. It’s that past a certain number of factors, nobody can hold in their head which ones were held. The canonical example is a published head-to-head that compared a best-of-fifteen run against a single run of a different system, on a different model and a different provider, and concluded “parity.” Every fact needed to catch that was present in the write-up. No single line said the two things were not the same kind of thing.

That line is what the module produces.

Writer, and the part where the loop closes

Writer holds the manuscript as a tree of nodes — folders, text files, binary assets — with the LaTeX compile handled by texc, a small Python sidecar container. Around that sits what you’d expect: a BibTeX parser, DOI import through Crossref, cite-key generation, an outline extractor, text statistics, a LaTeX log parser that turns compiler vomit into structured diagnostics, table formatting, and a freeze operation for when a chapter is done.

But two of the node kinds aren’t files:

Node kind What it pins
FigureRef A specific Visualizer ViewVersion
TableRef A specific Experimenter aggregate

Both are resolved into generated LaTeX at compile time, against the pinned version.

That’s the whole thesis of the platform in one mechanism. A number in the compiled PDF is not a number somebody typed into a table. It traces back through the aggregate, to the measurements, to the run, to the subject that produced it and the instrument that scored it — and the diagram next to it traces to a specific version of a specific view. Change the upstream data and the table is stale until you refresh it, explicitly, with an endpoint that exists for exactly that (/api/manuscript/tables/refresh-all).

The chain the platform is built to preserve:

flowchart LR
    H[Hypothesis] --> E[Experiment] --> S[Subject] --> R[Run]
    R --> M[Measurement] --> A[Aggregate] --> T[Table] --> P[PDF]
    T -- "supports / refutes" --> H
Figure 1: The chain the platform keeps attached. A number in the compiled PDF traces back through the aggregate to the run and the subject that produced it.

Nothing in that chain is novel on its own. Keeping all of it attached, in one database, under one identity model, is the point.

Agents are first-class, not bolted on

The provenance model doesn’t treat an agent as a user with a funny name. A Run records:

public string ActorKind { get; set; } = ActorKinds.Human;
public Guid?  ActorUserId { get; set; }
public string? ActorAgentRef { get; set; }
public string? ActorAgentVersion { get; set; }

Who produced this run — a person, or an agent, and which version of that agent. Six months from now, when a result looks off, “an agent did it” is not an answer. The agent’s version is a factor like any other, and the comparability report will tell you if it moved between two runs you’re about to compare.

The platform also carries an authoring surface for the agents themselves — /operations, /actions, /workflows, each with a publish step — so the catalog of what an agent is allowed to do is versioned data inside the system rather than prompt text scattered across repositories.

What it isn’t

The intent is a research assistant. Some of that intent is still intent.

It does not run experiments. The harness pushes results in; AgentX never invokes a runner. That’s a deliberate boundary — the thing that records results shouldn’t also be the thing that produces them — but it means a run only exists here if something else put it here.

The docs are behind the code, badly. The README documents phases M0–M5 and three modules. There are seven. The agent guide lists four MCP tool groups and none of them are Experimenter or Writer. Two design documents referenced by name in code comments — Lab_module.md, Experimenter_schema.md — are not in the repository at all. I know why this happened (the last forty commits are almost entirely Writer and Experimenter, and prose is the first thing to slip), but a platform whose whole argument is keep the trail attached should not have dangling references in its own comments.

The legacy cutover isn’t finished. The per-module ETL is built and idempotent. The identity migration it depends on — deduping users, projects and memberships across three source databases — is still a runbook, not code.

Why build this instead of using something

The honest answer is that I kept needing a specific join that no single tool offers: from a claim, to the procedure that tested it, to the run that executed the procedure, to the number in the document that reports it. Jira has the first two thirds. A notebook has the middle. A reference manager has the end. Nobody has the join, and the join is the thing that makes a result auditable six months later — which is the only property of a result that survives contact with anyone who wasn’t in the room.

That’s also why the test-management heritage stopped being embarrassing about halfway through. I spent twenty-something years building systems whose entire job is proving that a claim about software was checked, by a stated procedure, with the evidence still attached. Pointing that at research instead of at release notes required renaming three classes.