How the engine earns trust
The model reads. The code judges. They never share a substrate.
The verification pipeline
Extract→Validate→Graph→Match→Guard→Grade→Decidability→Provenance→Rejections
Extract
A language model reads each passage and fills a typed claim card with fixed fields: claim kind, commitment direction, causal direction, physical scale, measurable observable, and the exact source sentence. Causal direction is a first class field because it catches what embeddings lose. Two theories can use the same concepts with opposite causal arrows and only typed extraction captures the difference.
Validate
Five independent gates before any claim reaches the judge. Schema validation against the ontology. Cross model agreement from a second model of different training lineage, where disagreement excludes the claim rather than silently passing it. An adversarial calibration benchmark seeded with deliberately convincing fakes that must fail. Human review as a publication gate, never a training signal. Trust tier labelling on every claim: signed, machine extracted, or user amended, non removable.
Graph
Validated claims for each theory become a small directed graph. Each claim is a node, each relationship is a typed edge. The graph captures structural commitments independent of vocabulary. Two theories using entirely different words produce the same graph shape if they commit to the same relationships.
Match
Typed subgraph isomorphism checks whether one theory's structure maps onto another's, preserving all node and edge types. When the complete match fails, partial matching finds the largest substructure two theories share. The matched portion is the genuine correspondence. The unmatched portion names the exact structural disagreements. Both algorithms are deterministic. Same inputs, same output, same content hash.
Guard
Multiple guard mechanisms operate during matching. Typed edge constraints require every mapping to match in relation type. A vocabulary dependency test checks whether correspondences survive when shared terms are stripped. Provenance checks flag self referential evidence chains. One structural mismatch in the edge types vetoes the entire correspondence to the lowest supportable grade.
Grade
Four rungs computed by pure functions. Formal isomorphism: complete structural match. Partial mapping: substantial overlap with exact disagreements identified. Generative analogy: thin but real overlap, suggestive not evidential. Thematic resonance: vocabulary overlap only, quarantined. Each rung carries a machine readable ceiling stating what the correspondence does not license. A second axis runs orthogonal: falsifiable, metaphysical, unsupported, or contested. The two axes never collapse.
Decidability
Each genuine disagreement is checked against a versioned instrument capability map recording which instruments reach which scales measuring which observables. Three verdicts: testable now with the instrument named, testable under a stated assumption with the assumption named, or not testable with the reason named. The map is versioned. When instruments improve, past verdicts stay valid under their original version.
Provenance
Every verdict carries a full manifest: source sentences, model versions, ontology version, judge code version, capability map version. The comparison regenerates from source under a content hash. A continuous integration test runs the judge twice on identical inputs and asserts identical output.
Rejections
Every failed match published alongside confirmations, with the structural reason it failed.
Why rejections matter
Any tool can show you what it found. We show you what we found and what we refused. Twenty one rejected pairings visible in the demo, each with the structural reason it failed. This is not a limitation. It is the product.
Questions, answered plainly
What is a graded correspondence?
A claim that two passages from different fields share underlying structure, carrying the typed mapping, a strength rung, a demarcation label, provenance to both sources, and a reviewer signature.
Why not just use ChatGPT or Claude for this research?
Ask a chatbot whether two theories from different fields describe the same structure, and it will say yes, eloquently. Ask it again tomorrow and it may say yes differently, or no. It cannot tell you which parts of its answer are documented, which are interpolation, and which are its training data’s favorite poetry. It will never hand you a verdict it is willing to be wrong about, because it holds no verdicts at all: only fluency. A language model is a brilliant reader and a tireless explainer, and we use one for exactly those tasks. But research needs something a conversation cannot give: a claim that stands still. A graded correspondence does not change when you rephrase the question. It carries the same rung tomorrow, cites the same passages, shows the same reasoning in the audit view, and names the human who signed it. A chatbot is a voice. This is a record. Science is built on records.
How is this different from other AI for science systems?
Hypothesis engines such as Google’s Co-Scientist, SciAgents, Kosmos, and ResearchAgent generate ideas with language models and evaluate them with language models, inside domains rich in public data. Verifier coupled systems such as AlphaProof and AlphaEvolve achieve real rigor, but only where a formal verifier exists: mathematics and code. Literature tools such as Elicit, Consensus, and Semantic Scholar search and summarize what is written; they do not judge structure. Notebook tools such as NotebookLM ground answers in your sources but grade nothing. The Consilience occupies the space none of them touch: cross vocabulary structural correspondence, judged by a deterministic engine, graded on two axes, in domains that have no verifier.
Is this like The Consilience Project or the E.O. Wilson book?
No affiliation. Wilson’s 1998 book named the dream of unified knowledge; The Consilience Project (Daniel Schmachtenberger) works on civilizational sensemaking. We build an instrument: the graded hypothesis engine.
Can I see why a verdict was reached?
Yes. One click opens the audit view: both structure graphs, the mapping, why the next rung up was denied, the ontology version, the judging parameters, the reviewer decision. Built for the reviewer who wants us to be wrong.
What did you deliberately not build?
No fine tuning: nothing a researcher ingests ever enters model weights, a privacy guarantee and a judgment guarantee in one. No end to end learned judge: a verdict no one can inspect is not a scientific instrument. No engagement metrics in the loop: the grader answers to the calibration benchmark, never to what users would prefer to hear.