Part 1 of 4  ·  2026

The Verification Gap

Every breakthrough in AI for science rests on a checker that is not a model. Comparing claims across fields has never had one.


In 2019 a consortium of laboratories across three continents set out to do something that sounds almost trivially simple. Two leading theories of consciousness had been arguing past each other for two decades. The task was to find the places where the two theories actually disagreed, stated clearly enough that an experiment could tell them apart.

It took about a year of expert workshops to agree on three such points.

Three. That year involved the founders of both theories, funding from a major philanthropy, and a preregistration process built so that neither camp could move the goalposts once results came in. When the experiments finally ran, the work was genuinely valuable. But look at the cost of the step before the experiments. A year of the most qualified people in a field, sitting in rooms, to produce a list of three disagreements.

That step has a name nobody uses, because nobody has built anything to do it. Two bodies of work describe something. They use different vocabularies. Are they saying the same thing in different words, or genuinely different things, and if genuinely different, could any experiment tell?

This work is done by hand, constantly, in places where the stakes are high. A biotechnology company does it before committing tens of millions to one explanation of a disease over another. An investor does it before backing a company whose entire value rests on a mechanism. A doctoral student does it every time they write the paragraph explaining how their work differs from the work it resembles. In each case the process takes weeks or months, the output is one qualified person's judgment with a bibliography attached, and nobody downstream can check it, repeat it, or find out what was considered and rejected along the way.

That was the situation before language models arrived. What they changed is not quite what most people describe.

Finding became free. Trusting became impossible.

Ask any capable model to connect two fields and it will give you something fluent and specific within seconds. Ask again tomorrow and you will get something different. Ask whether the answer is correct and it will tell you, confidently, in exactly the register it uses for everything else.

Generating cross domain connections now costs nothing. Verifying them is harder than it has ever been, because the supply of plausible unverified material has gone up by orders of magnitude while the cost of producing more approaches zero. Citation hallucination in deployed research systems runs somewhere between eleven and fifty seven percent depending on which system and which measurement you trust. The best automatic attribution classifiers reach around eighty percent macro F1. Eighty percent is a respectable research result and a useless guarantee.

So the bottleneck moved. It stopped being "can we find a connection" and became "can anyone check the connection we found."

This is not a minority view. Yuan Cao, who spent years at Google DeepMind working on the Gemini models and now runs an AI for science company, said in 2026 that the real bottleneck is not compute, it is verification. He means it about experiments, and about experiments he is right. I mean it about claims, one layer earlier. The two readings are complementary rather than competing. Before you spend a laboratory on a question, something ought to establish that the question is real.

Every breakthrough has a checker, and the checker is never the model

Strip the announcements from the last two years of AI for science and one pattern shows up with almost boring consistency.

AI reached medal level performance on International Mathematical Olympiad problems. The system that did it proposes proof steps with a language model and verifies every one of them in Lean, a formal proof assistant that is deterministic, exhaustive, and completely indifferent to how persuasive a step sounds. DeepMind's own framing credits the verifier. A result in 2026 sharpened the point considerably: a stock language model with no additional training, placed inside a harness with a Lean verifier, outperformed a purpose trained specialist model, taking solve rate from under ten percent to around seventy. The harness beat the specialisation.

AI found a faster algorithm for a problem where the human record had stood since 1969. That system mutates candidate programs with a model and scores them by running them. A program either executes faster or it does not, and there is no argument to be had with a benchmark.

AI now runs autonomous research campaigns lasting half a day, reading well over a thousand papers per run, with roughly seventy nine percent of its statements traceable and accurate. Its credibility does not come from a cleverer model. It comes from a structured world model that code updates after every step, and a discipline in which every claim links to a specific source or code cell.

Same shape in all three. The model proposes, something that is not a model disposes. Where a domain has a verifier, coupling generation to it produces the breakthrough. Where a domain has none, the system inherits the generator's failure mode, which is fluent, confident, unreproducible output.

The most useful evidence for this comes from the people with the strongest incentive to claim otherwise. Google's AI Co-Scientist is a seven agent system that debates and ranks its own hypotheses through Elo tournaments, and it independently re-derived a finding that had taken a laboratory a decade. Its own paper states that self-evaluation is unreliable. The team that built the largest model-judging-model system in existence published its failure mode in the paper announcing it. Anyone building in this space should read that sentence twice.

The domain with no verifier at all

Mathematics has Lean. Algorithm discovery has execution. Empirical science has, eventually and expensively, the laboratory.

Comparing claims across vocabularies has nothing. There is no proof assistant for "do these two frameworks describe the same structure." There is no execution test for "is this correspondence real or a coincidence of metaphor."

And here is the part that took me longest to see, because it is the reason this cannot be solved the way the other domains were solved. No downstream laboratory will ever settle it. Nobody is going to run an experiment to determine whether two theories share an abstract structure. In this domain the verdict is not a hypothesis awaiting a test. The verdict is the product.

Which means the check has to happen inside the system, or it never happens at all.

That is the gap. Not a gap in generation, where a great deal of talent and capital is currently pointed, but in the layer that decides whether what was generated holds. Between literature search, which finds documents, and expert judgment, which costs a specialist's month and produces an unauditable opinion, there has been nothing.

Why the judge cannot be a language model

The obvious objection is that reasoning models keep improving, so the judging problem will dissolve on its own. I think that is wrong, and the reasons have nothing to do with capability. There are three, and each one on its own is disqualifying.

The first is reproducibility. A reasoning model produces different outputs across runs. Temperature, sampling, a silent version upgrade behind an interface. For an instrument whose entire product is a trustworthy verdict, "usually the same answer" is not a weaker form of the right property, it is the absence of it. The first question any serious evaluator asks is whether the same input gives the same output twice, and a system that has to answer "approximately" has already lost the conversation.

The second is incorruptibility. Language models are trained toward helpfulness, and the failure mode of this particular domain is false positives. Ask a capable model whether two frameworks correspond and it will find you a correspondence, because finding one is the helpful act. A more capable reasoning model simply reasons more elegantly toward yes. Determinism is what makes it possible to publish rejections at all, and publishing rejections is the only credible way to demonstrate that a system can say no.

The third is auditability. A deterministic verdict comes apart under inspection. You can show which claims matched, on which typed fields, under which rule, and precisely why the next grade up was refused. A neural verdict is a forward pass. It can be accompanied by an explanation, but that explanation is generated afterwards and is not the thing that produced the answer. That is an oracle asking for faith, and science has enough of those.

None of this is an argument that models are bad. Models are extraordinary at the part of this problem they are actually suited to, which is reading.

What we built

The architecture follows from the split, and its commitments are easy to state even though the machinery underneath is not.

A model reads. It turns a passage into a typed claim: what is being asserted, about what, in which direction, at what scale, and what would measure it. Reading natural language is the part models do better than any alternative, and this is the only place they are used.

Then the models are switched off and deterministic code does the judging. Each body of work becomes a typed directed graph, and matching is subgraph isomorphism under type constraints rather than similarity search, which places the design in the neurosymbolic family and specifically in the LLM-Modulo pattern that Subbarao Kambhampati argues for: the model proposes, an external verifier disposes. Every match is graded on a four rung ladder, from formal isomorphism at the top down to thematic resonance at the bottom, which is quarantined rather than celebrated. A guard named for M. H. A. Newman's 1928 objection enforces the types, because untyped structure is trivially satisfiable and a single defect edge vetoes a correspondence rather than being averaged away. A further layer then asks, of each genuine disagreement, whether both sides make predictions in the same observable space, at the same physical scale, reachable by instruments that exist. The answer is testable now, testable only under a stated unproven assumption, or not testable with current technology, and the reason travels with it. This is falsifiability in Popper's sense given a computational form, checked against a versioned instrument capability map rather than argued in prose.

Three commitments make this an instrument rather than a demonstration.

Every verdict carries a ceiling. A grade is not a score, it is a statement about how far the correspondence licenses you to go, written down explicitly, so that nobody can use the top of the ladder to justify a conclusion only the bottom supports.

Every claim resolves to a sentence. Not a document, not a page. Provenance is a visible surface in the product rather than a footnote, because a citation you cannot open in one click is a citation you cannot check.

Every rejection is published. Each match the instrument refused, with the structural reason it failed. This is the part competitors are least likely to copy and the fastest way to demonstrate which population a project belongs to.

There is one further guard, and it sits exactly where the design is most exposed. The judging is deterministic, but its inputs come from a model, and a wrong input produces a wrong verdict reproducibly. So claims destined for matching are extracted under a constrained schema and then independently re-derived by a second model of a different lineage, a cross model agreement gate where disagreement blocks a claim rather than being quietly resolved. That turns extraction error from an unknown into a measured, gated quantity. It does not eliminate it, and the difference between those two statements appears on the page rather than in a footnote.

The test that mattered

An architecture is a proposal until it is tested against something it cannot have memorised.

We ran the instrument across six theories of consciousness, with the consortium's published comparison papers held out and the exclusion enforced by a script whose pass state is displayed beside the results. Working from the theories' own texts, the system recovered the location and connectivity disagreements that the year of expert workshops had preregistered. It recovered the timing disagreement only partially, and that partial recovery turned out to be the more interesting result, because the missing component appears nowhere in the theory's published work. It was negotiated between scientists in a room. The instrument separates what a theory says from what its champions agreed to, which is a distinction no prior tool has been able to compute.

It also found something the answer key did not contain. It surfaced a disagreement about substrate, whether a perfect simulation would be conscious, graded it not testable with current technology, and attached the reason. The human collaboration could not preregister that one precisely because no experiment reaches it. A system that had merely absorbed the answer key could not have contained more than the answer key.

Then it did something I had not anticipated when we started. In January 2020 a workshop convened the founders of two theories, in a room that included a Nobel laureate, to find a single experiment that could test one against the other. It found none, and the failure is documented. Our instrument, working blind, computes why. Beneath maximally different vocabularies the two theories agree on the field's deepest question, since both deny that consciousness is computation, so functional experiments cannot separate them even in principle. And where they genuinely disagree, they disagree at incompatible physical scales, brain networks against molecules, with no shared measurement space. The historical record and the computed verdict sit beside each other, and they match.

The whole comparison regenerates from source in about seven seconds under a content hash, with twenty one rejections published and eighteen computed disagreements left standing as public predictions to be scored when results appear.

What it costs

An honest account of an architecture includes what it gives up, and this one gives up two things.

It has lower recall than a frontier reasoning model would. Typed matching misses subtle analogies that a capable model would catch, and those misses are logged rather than quietly dropped. The trade was deliberate. In a field where the dominant failure is confident nonsense, precision with receipts is worth more than recall with vibes.

And it needs someone to author the claim vocabulary for each new domain. A learned world model generalises on its own. An authored one does not. That is the genuine advantage held by systems that learn their representations, and it is the real price of this approach, measured in weeks of expert time rather than in compute.

Both appear here for the same reason the rejections are published. An instrument that cannot describe its own limits is not an instrument.

Where the work stands

The engine is in active build, gated behind an adversarial calibration benchmark that includes deliberately convincing fakes it is required to refuse. The consciousness comparison described above is live and inspectable. The engine code will be released open source and the graded corpus published as an open dataset, because a verification layer nobody can audit is a contradiction in terms.

The larger claim I want to make is not about this project. It is that the shape of the opportunity in AI for science has been widely misread. Enormous capital and talent are pointed at generation, at proposing more hypotheses, more molecules, more designs, faster and more cheaply. That work is real and I am not arguing against it. But every one of those systems is now producing candidate claims at a rate that far exceeds anyone's ability to check them, and the field's own leading systems say so in their own papers.

The generation side will keep getting better. The interesting problems have moved to the verification side, and in the specific business of working out what different fields are actually claiming, that side has been empty.

That is the gap, and it is what this instrument is for.