Roger Mendoza

NULL is a deliverable: auditing 113 research findings

A burst of AI-assisted research produced 88 findings in two days. Auditing them against their own rules found that the filter never ran — and that the audit was the most useful finding of all.

Dictyon-net keeps a research ledger: every lead anyone touched, with a verdict and a reason. ADOPT means it changed the project. QUEUE means real but deferred, with the reason. NULL means tested and doesn't apply — and the ledger's header is explicit that a NULL with a reason is a deliverable, not a failure. The only unrecoverable waste is a dead end someone walked and didn't write down.

Leads climb rungs, from "named" to "mechanism stated" to "computed" to "run on hardware". The design intent is that most leads die early. A process that promotes everything is just being credulous with extra steps.

The burst

In mid-September, AI agents working in parallel landed 88 findings in about two days — network coding, fountain codes, swarm intelligence, digital twins, blockchain trust, quantum key distribution and much more. Each had a hypothesis, a measurement script and a verdict. It looked like an extraordinary amount of progress.

So I audited all 113 findings in the tree against the ledger's own rules.

What the audit found

Verdicts across 113 research findingsADOPT44QUEUE40NULL21OPEN2unparsed6
74% promoted, 19% killed, and not one lead died at the first two rungs. The audit then found the 44 ADOPTs had changed zero lines of code. Source: finding 115.

The verdicts were inverted. 74% of findings were promoted (ADOPT or QUEUE) and 19% killed. Not one died at the first two rungs. A ladder that nothing falls off isn't filtering anything.

44 ADOPTs changed zero lines of code. The ledger defines ADOPT as "changed the project". A git diff over the whole burst touched the research and measurement folders and nothing else — no spec, no reference code, no test, no roadmap entry. By the ledger's own definition, the correct ADOPT count was zero.

40 of 101 measurement scripts didn't run in a clean checkout. Thirty-four imported an optional maths library without guarding the import; four contained absolute paths from the machine that wrote them. A number nobody else can regenerate is an assertion with a filename attached.

Fourteen findings re-derived earlier findings. Network coding was investigated four separate times with the same script and the same conclusion. Several findings said so themselves, and were filed anyway.

Why it happened

Nobody did anything malicious, and the individual scripts were real: where they ran, their numbers matched the findings that cited them. The failure was structural. Generating findings in parallel is cheap; filtering them is the expensive part, and the filter — the step that says "this doesn't apply here, retire it" — was skipped because each agent only saw its own lead.

Breadth only pays if the filter sits at the back end. Without it, volume looks like progress.

What changed

The audit's corrections were applied, and the fourteen re-derivations were reclassified as NULL as independent results. Two rules follow directly from it:

  • "Changed the project" means a diff outside the research folder, or it isn't ADOPT.
  • A measurement only counts if it runs in a clean checkout.

The general lesson

This applies well beyond radio research. Any time AI makes it cheap to produce analyses, reports or findings, the bottleneck moves to verification. The questions I now ask of any batch of AI-generated work are the same ones the audit asked:

  1. Did anything actually change because of this?
  2. Can someone else re-run it from scratch?
  3. How many of these are the same result counted twice?
  4. What fraction was rejected? If the answer is "almost none", the filter isn't running.

The most valuable finding in that burst wasn't any of the 88. It was the one that checked them. More in the Academic Theory paper.