How the Judge Actually Works
Inside a deterministic verifier for scientific claims. Five quality gates, typed graph construction, the Newman guard, subgraph isomorphism, a grading ladder that refuses to produce a score, and a testability check against instruments that exist.
An instrument that cannot be inspected is not an instrument. It is an oracle, and science has enough of those.
So this essay opens the machine. What follows is the mechanism inside The Consilience: a system that takes claims from different fields, written in different vocabularies, sometimes centuries apart, and decides whether two of them are structurally the same, genuinely different, or not comparable at all.
The short version of the design is that a language model reads and deterministic code judges. That single sentence is the architecture. Everything below is how it is actually done, and where it can be attacked.
Why the judging layer is not a model
Before the mechanism, the constraint that shapes all of it.
The obvious way to build this system is to hand two passages to a capable reasoning model and ask whether they correspond. I tried a version of that first and it failed, badly and instructively, producing several thousand confident false positives. The failure was not a tuning problem. Similarity and structural correspondence are different relations, and a system optimised for one will systematically mistake it for the other.
But there are three reasons the judging layer stays out of the model even now that reasoning models are far better than they were.
The first is reproducibility. A model produces different outputs across runs, because of sampling, temperature, or a silent version change behind an interface. For an instrument whose product is a verdict, "usually the same answer" is not a weaker form of the right property. It is the absence of it.
The second is what I would call incorruptibility. Language models are trained toward helpfulness, and the failure mode of this domain is the false positive. Ask a model whether two frameworks correspond and it will find you a correspondence, because finding one is the helpful act. A more capable model simply reasons more elegantly toward yes.
The third is auditability. A deterministic verdict decomposes. You can show which claims matched, on which typed fields, under which rule, and why the next grade up was refused. A neural verdict is a forward pass, and any explanation attached to it was generated afterwards rather than being the thing that produced the answer.
This places the system in the neurosymbolic family, and specifically in the pattern Subbarao Kambhampati calls LLM-Modulo: the model proposes, an external verifier disposes, and nothing on a claim that matters reaches a user without passing a check that is not a language model. His group's results are the empirical backing. Models fail block stacking problems that classical planners solved in the 1970s, and performance collapses further when the object names are obfuscated, which shows the apparent reasoning was pattern recall rather than planning.
Five gates before the judge sees anything
A deterministic judge running on bad input produces bad verdicts deterministically, which is worse than a probabilistic system because the error looks authoritative. So before any claim reaches the judge it passes five independent gates.
Gate one: schema constrained extraction
A model reads a passage and fills in a claim card with a fixed shape. What kind of claim it is, whether it asserts or denies, in which causal direction, at what physical scale, what instrument would measure it, and the exact source sentence.
The output is constrained to that schema at the decoding level and then validated field by field after it arrives. If the claim kind is not one of the permitted types, rejected. If the commitment direction is neither asserts nor denies, rejected. Missing required field, rejected. Two retries, then the passage is flagged for human review and excluded from matching.
This is an ordinary validation layer and it catches the most obvious failures: malformed output, invented field values, structural nonsense. Cheap, unglamorous, and it does a great deal of work.
Gate two: cross model agreement
Every claim card destined for matching is independently re-extracted by a second model of a different training lineage. Not a different size of the same model. A model built by different people, on different data, with different architectural choices.
Both read the same passage. Both fill in the same schema. Code then compares the two cards field by field. Agreement on the typed fields makes the claim eligible for matching. Disagreement on any structural field flags the claim and excludes it.
The reasoning is specific. Models from one family share training data and therefore share blind spots. If one misreads a causal direction, its sibling is likely to misread it the same way. A model from a different lineage has different blind spots, so independent agreement between the two carries information.
An honest caveat belongs here, because it limits what this gate can claim. A study across more than three hundred and fifty models found that when two models both get a question wrong, they converge on the same wrong answer roughly sixty percent of the time. Cross model agreement is correlated evidence rather than independent confirmation. That is exactly why it is gate two of five rather than the only gate. It does not certify that an extraction is correct. It surfaces the places where automated reading is not trustworthy.
Gate three: adversarial calibration in continuous integration
A frozen set of passages with known correct extractions runs through the pipeline on every code change. Some are real claims that must extract correctly. Some are deliberately convincing fakes that must fail.
If a code change or a model update causes a previously correct extraction to break, or causes a known fake to start passing, the test fails and the change is blocked. This is regression testing applied to epistemics, and the fakes have to be written by a person, because their adversarial value comes from anticipating failure modes the models cannot anticipate about themselves.
Gate four: human review as a publication gate
Before a claim enters the public canon, a reviewer sees the source text beside the extracted card and approves, rejects, or revises. Only approved claims carry the status of signed.
One design commitment underneath this matters more than it looks. Reviewer approvals are a publication gate, not a training signal. They do not feed back into any model, tune retrieval, or build a personalised knowledge graph. An earlier version of this system used confirmations as training data and it was removed for a precise reason: if the reviewer's biases train the system, the system reflects those biases back, and you have built a confirmation machine wearing the costume of an instrument.
Gate five: trust tier labelling
Every claim that ships carries a label stating how far it was verified. Signed for human reviewed and approved. Machine extracted for claims that passed gates one and two without human review. User amended where someone edited a claim, with the original preserved.
Nothing ships unlabelled, and the label cannot be removed or hidden, including in exports.
Turning claims into wiring diagrams
After the gates you have a set of validated claim cards. A list of cards is not a theory, though. A theory is a set of commitments that stand in relations to each other, so the next step converts the list into a graph.
Each claim becomes a node. Each relationship between claims becomes a typed, directed edge. The result is a small directed graph representing the structural shape of the theory.
A concrete example makes this clearer than any definition. Suppose a theory commits to the following: quantum collapse causes consciousness, consciousness operates at microtubule scale, and microtubule scale effects are measurable by quantum coherence detection. The nodes are quantum collapse, consciousness, microtubule scale, and quantum coherence detection. The edges carry types: causes, operates at, measurable by.
Now a second theory says: orchestrated objective reduction generates experiential states at cytoskeletal scale, detectable by decoherence timing measurements. Almost no shared vocabulary. The same wiring diagram.
This is the step that makes vocabulary independence possible, and it is where the design sits in the lineage of Dedre Gentner's structure mapping theory of analogy, which argued that analogical reasoning operates over relational structure rather than surface features. That idea is forty years old in cognitive science. What is new here is applying it to scientific claims with types strict enough to support a refusal.
The Newman guard, or why types are not optional
In 1928 the mathematician M. H. A. Newman raised an objection in the philosophy of science that has a precise consequence for anyone building structural comparison tools. If you define structure using untyped relations, meaning you say A is related to B without specifying how, then any collection of things with the right number of elements automatically shares structure with any other collection of the same size. Structure without types is trivially satisfiable. Everything matches everything.
This is not an edge case. It is the default failure mode of every tool claiming to find deep structural patterns across fields. If your matching lets a causes edge pair with an is located in edge because both are merely relations, a theory of consciousness will match a theory of plate tectonics. The match is real in graph theoretic terms and meaningless in every other sense.
The Newman guard is the rule that prevents this. Every edge in a correspondence must match in type. If a mapping requires pairing causes with is substrate of, the match fails at that point, and the failure is hard rather than soft. A single defect edge, one place where types do not align, vetoes the correspondence down to the lowest rung the remaining match supports. No averaging. No eighty percent of edges matched so it is mostly right.
This one rule is why the system can publish rejections. A tool without typed matching cannot reject anything with confidence, because it has no principled basis for saying two things are structurally different. The Newman guard is exactly that basis.
The matching: subgraph isomorphism and maximum common subgraph
Two theories are now typed directed graphs. The question is whether one structure fits inside the other.
The algorithm that answers it is VF2, from Vento and Foggia, published in 2004, shipped in networkx, free and open source. Subgraph isomorphism asks whether every node and edge of graph A maps onto a subset of graph B while preserving all types and directions. If it does, A's structure lives entirely inside B.
VF2 is not novel and I want to be explicit about that. It is a textbook algorithm. It was not invented for this project and it was not modified. It was selected because the problem, do these two typed structures correspond, is precisely the problem it was designed to solve. The innovation is not in the algorithm. It is in the schema the algorithm runs on, and in the fact that nobody had built that schema for scientific claims before.
When the match is not complete, a second computation finds the largest piece that does match, the maximum common subgraph. If a theory has eight nodes and the largest cleanly mapping subgraph has six, the ratio tells you how much structure corresponds and the two unmatched nodes tell you precisely where the theories part company. Those unmatched nodes become the disagreements the system surfaces, and they are the scientifically valuable output.
The combination of the two is what separates "these theories fully correspond" from "these theories partially correspond, and here is exactly where they diverge."
Four rungs, not a score
Matching tells you whether two structures correspond and how much of them does. That is not enough, because a bare number invites misuse. A score of 0.73 becomes "probably corresponds" in the hands of someone who never asks what the missing 0.27 contained. It might contain the only part that mattered.
So the system assigns one of four rungs, and each carries an explicit ceiling stating what the correspondence does and does not license.
Formal isomorphism. A complete match. Every node maps, every typed edge maps with the same type. Two bodies of work describing the same thing in different words, in the way matrix mechanics and wave mechanics turned out to be equivalent formulations of quantum mechanics. Rare. The ceiling licenses treating results from one as evidence about the other, because structurally they are the same theory.
Partial mapping. No complete match, but substantial overlap. The system reports what matched, what did not, and the ratio. This is the most common useful verdict. The matched part is genuine correspondence and the unmatched part is where the real disagreement lives, which is often the more valuable finding. The ceiling holds conclusions to the matched substructure only.
Generative analogy. Real but thin overlap. Worth investigating as a direction for new work, carrying no evidential weight. The ceiling says this is a lead rather than a finding.
Thematic resonance, which is quarantined rather than celebrated. Similar sounding vocabulary, no structural correspondence under typed matching. This is where quantum consciousness meets quantum computing. Both use the word. They commit to entirely different structures. The ceiling states plainly that this is a vocabulary coincidence and should not be cited as evidence of a connection.
The grading is computed by pure functions from the matching output. Same graph pair, same rung, same ceiling, every time. No model, no probability, no temperature.
Decidability, or whether anyone can actually settle it
A disagreement between two theories is only scientifically useful if someone can settle it. The matching surfaces where two theories diverge. The decidability check asks whether any instrument alive today can reach that divergence.
Three possible answers and nothing else.
Testable now. Both theories predict in the same observable space, at the same physical scale, and an instrument that currently exists reaches that scale with sufficient resolution. The system names the instrument class and the measurement type. This is the most valuable thing the system produces, because it is a disagreement someone could go and settle with equipment that already exists.
Testable under a stated assumption. The disagreement becomes testable if a specific unproven assumption holds. The assumption is named and travels with the verdict. A reader who finds it plausible has a testable prediction. A reader who does not finds the assumption itself interesting, because it is now explicit where it was previously buried.
Not testable with current technology. The theories disagree at incompatible physical scales, or in an observable space no instrument reaches. The system states why.
The check is a lookup against a versioned instrument capability map, which is structured data rather than a model. It records what each class of instrument can measure, at what spatial and temporal resolution, at what physical scale. Functional magnetic resonance imaging reaches whole brain network scale at millimetre resolution. Patch clamp electrophysiology reaches single neuron scale at millisecond resolution. No current instrument measures quantum coherence at intracellular scale in living tissue. The disagreement states where the theories diverge, the map states what instruments reach, and the verdict falls out of the comparison.
The map is versioned because instruments improve. A disagreement that is not testable under one version may become testable under a later one when a new technique reaches a previously unreachable scale. The old verdict is not deleted when that happens. It was correct under the capability that existed. Both verdicts coexist with their version numbers and both stay citable.
This is the oldest idea in the philosophy of science given a computational form. Karl Popper defined falsifiability in 1934 and philosophers have worked on what makes a claim empirically testable ever since. Instrument ontologies exist: the OBO Foundry catalogues biomedical instruments, and space agencies maintain capability databases for missions. Feasibility assessment for individual claims exists as an active benchmark task: Matter of Fact (Jansen et al., EMNLP 2025) tests whether individual claims in materials science are feasible using language models over retrieved literature, and HARPA scores hypotheses with a testability driven reward model. These systems ask whether one hypothesis is feasible, scored by a learned model or a probability estimate. What I have not found is a system that asks the different question: whether the divergence between two theoretical frameworks is settleable by a discriminating experiment, computed as a deterministic lookup against structured instrument data, versioned so that old verdicts remain valid when instruments improve. The object is different, a disagreement rather than a single claim. The method is different, a structured lookup rather than a learned model. The output is different, a three way verdict naming an instrument and a reason rather than a probability. If someone has built it and I have missed it, I would like to know, because that person is a collaborator rather than a competitor.
Building the map is the slowest and least glamorous work in the project. No database exists saying which instrument reaches which scale measuring which observable. That knowledge lives in experimentalists' heads across thousands of laboratories and nobody has written it down in queryable form. It is also why the system expands domain by domain rather than paper by paper. The consciousness work ran because the map held entries for neuroscience instruments. Upload a materials science paper tomorrow and the matching still runs and the grading still assigns a rung, but the decidability layer returns nothing, because it has no entries for electron microscopes or X ray diffraction. Each new domain costs weeks of expert time. That is the real price of expansion, and it is also why the asset compounds.
Provenance, or how you check the checker
A verdict is only as trustworthy as your ability to take it apart. Every output ships with a receipt stating how it was produced.
This is not a bibliography. A bibliography says see Smith 2024. Provenance here records which exact sentences each claim came from, which model versions performed the extraction, which ontology version defined the relation types, which judge version computed the rung, and which capability map version the decidability check used. One click from any verdict to the source sentence. Not the paper. Not the page. The sentence.
Alongside it sits a content hash computed over the entire output. Run the same comparison again with the same inputs, ontology version and judge version, and the hash is identical. Change anything and the hash changes. This is how determinism is demonstrated rather than asserted, and in continuous integration a test runs the same comparison twice and blocks any change where the hashes diverge. That catches silent nondeterminism, which is the dangerous kind, because nobody notices it until a verdict is challenged.
The versioning principle extends to verdicts themselves. When the ontology improves and a verdict changes as a result, the old verdict is not deleted. It was correct under the rules that produced it. Both coexist, and the difference between them is itself informative, because it shows exactly what the ontology change affected.
In most AI research tools the output is a forward pass through a model. You cannot trace how it was produced, cannot verify you would get it again, and if the model is updated behind the interface your old results may quietly become unreproducible. An instrument that cannot be recalibrated against its own past results is not an instrument.
Published rejections, which is not a technology at all
Every other tool in this space is built to show what it found. This one publishes what it refused.
When the matching determines that two claims do not correspond, because subgraph isomorphism fails, or the common subgraph is trivial, or the Newman guard vetoes on a defect edge, the rejection is not discarded. It appears on screen beside the confirmed matches with the structural reason it failed. In the consciousness comparison, twenty one rejections are visible, each stating which claims were compared and precisely why the match was refused.
This is the hardest decision in the design and it has nothing to do with engineering.
Publishing rejections makes the product look worse in a demonstration. Every competitor shows a wall of discovered connections. A first time viewer here sees twenty one things the system said no to and might reasonably conclude it does not find much. But a buyer who has been burned, the company that committed tens of millions to the wrong mechanism, the investor who backed a spurious correspondence, the researcher whose published work rested on a false analogy, sees twenty one things the system refused to fake. That is the buyer who pays.
The reason competitors will not copy this is not technical. Any tool could publish its rejections. The reason is incentives. Products in this category are measured on engagement, recall, and the number of things discovered, and publishing refusals reduces all three. Anyone proposing the feature inside such a company would be told, correctly, that it hurts the metrics.
That single decision determines the entire customer segmentation. The buyer who wants to see how much was discovered will never want this instrument. The buyer who needs to know what was considered and rejected, with the reason, is currently unserved by everything on the market.
What this does not do
It does not generate hypotheses, design molecules, run simulations, or predict experimental outcomes. It does not say which theory is correct, only where they differ and whether the difference is reachable.
Its recall is lower than a frontier reasoning model's would be. Typed matching misses subtle analogies a capable model would catch, and those misses are logged rather than quietly dropped. That trade was deliberate. In a field where the dominant failure is confident nonsense, precision with receipts beats recall with fluency.
And it requires an authored claim vocabulary for each new domain. A learned world model generalises on its own. An authored ontology does not. That is the genuine advantage of systems that learn their representations, and it is the real cost of this approach, measured in weeks of expert time rather than compute.
That is the machine. Five gates protecting the judge's inputs, typed graphs that keep structural commitments while discarding vocabulary, a guard named for a mathematician who died before computers existed, two graph algorithms that are older than most of the people using them, a grading ladder that refuses to emit a number, a testability check against instruments that actually exist, provenance that makes every verdict reproducible, and a decision to publish refusals that is not a technology at all.
Enough, I hope, for anyone who wants to inspect it, challenge it, or tell me where it breaks. That last one is the most useful thing you could do, and I mean it.