How a hypothesis engine earns trust

Why nothing existing can do this

Embeddings match words. “Quantum vacuum” and “Sunyata” never meet in vector space, yet describe structurally similar dynamics. Word search cannot cross a vocabulary wall.

Language models cannot judge their own output. Today’s hypothesis engines generate with a model and evaluate with a model. The model that produces a seductive false connection is the same model asked whether it is real.

Nothing grades strength. The gap between formal isomorphism and poetic resemblance is the gap between a finding and a slogan. No tool measures it.

Five engineering principles

The full method will be published with the engine’s open source release. The principles are permanent.

01
Structure before words.
Every passage is reduced to a representation that is invariant to vocabulary before any matching occurs. Correspondence is computed where words no longer exist. This is what lets a physics paper and a Sanskrit commentary meet at all.
02
Generation and judgment never share a substrate.
Language models propose, extract, and explain. Whether two structures actually correspond is decided by a deterministic engine whose parameters are derived from a versioned ontology, frozen per version, and published in the audit view. The model that imagines a connection is structurally incapable of approving it.
03
Two axes, never one score.
The strength of a correspondence and the testability of a claim are measured independently and never collapsed. A framework can be honestly falsifiable yet structurally empty; a correspondence can be structurally tight yet untestable. Any system that outputs one number is hiding one of these.
04
Calibrated against seduction.
The grader is gated behind a benchmark seeded with false correspondences built to be convincing, and it is promoted only when it rejects them. An instrument tested only on successes has never been tested.
05
Provenance complete, or unpublished.
Every published claim resolves to source passages, the ontology and engine version that judged it, and a named human signature. Anything that cannot show its full chain does not ship.

Two commitments, in one breath: nothing you ingest ever trains a model, and every rejected pairing is published, with its reason, in a public negative results ledger. A grader that never says no has a yes worth nothing.

What you will see when the engine ships

A correspondence page
Two passages, side by side, from two traditions that never met. Between them, the structural bridge. Above them, the verdict: a rung badge and a testability label. Below, every source cited to the page. Each correspondence is a permanent public page with its own stable ID: linkable, citable, exportable with a DOI.
The Explorer
Browse the full graded canon. Filter by field, tradition, grade, or source. Follow a tradition and receive new approved correspondences as they land.
Ask
Type a question: “Is the quantum vacuum the same as Sunyata?” The answer comes back grounded in the canon, with citations inline and the honest grade attached: what is documented, what is analogy, what is resemblance only, and why.
The audit view
One click on any verdict opens everything beneath it: both structure diagrams, the mapping, why the next rung up was denied, the ontology version, and the reviewer’s signature. Nothing to take on faith.
The negative results ledger
A public page of pairings that failed, each with the specific mismatch that failed it. The record science usually loses, kept.
Grade cards and overlays
Any correspondence exports as a shareable card or a video overlay, sources and grade included, built for the classroom, the talk, and the feed.

The output: a new scientific artifact

The graded correspondence. Stable ID. Version stamp of the ontology and grader that judged it. Full provenance to both source passages. Reviewer signature and date. Exportable with a DOI. Citable in a paper today, resolvable in five years to exactly what was graded under which ruleset. Cross domain knowledge becomes compoundable, and the growing canon of graded correspondences is an open dataset any group can build on: a community resource, not a walled product.

One gap, claimed precisely

The Consilience works where today’s AI for science cannot: domains with no formal verifier and vocabularies that share nothing. We build no data platforms, no trial infrastructure, no clinical tools, no pooled statistics. Our strongest outputs hand off to the verification layer; we never enter it.

Questions, answered plainly

What is a graded correspondence?
A claim that two passages from different fields share underlying structure, carrying the typed mapping, a strength rung, a demarcation label, provenance to both sources, and a reviewer signature.
Why not just use ChatGPT or Claude for this research?
Ask a chatbot whether the quantum vacuum and Sunyata describe the same reality, and it will say yes, eloquently. Ask it again tomorrow and it may say yes differently, or no. It cannot tell you which parts of its answer are documented, which are interpolation, and which are its training data’s favorite poetry. It will never hand you a verdict it is willing to be wrong about, because it holds no verdicts at all: only fluency. A language model is a brilliant reader and a tireless explainer, and we use one for exactly those tasks. But research needs something a conversation cannot give: a claim that stands still. A graded correspondence does not change when you rephrase the question. It carries the same rung tomorrow, cites the same passages, shows the same reasoning in the audit view, and names the human who signed it. A chatbot is a voice. This is a record. Science is built on records.
How is this different from other AI for science systems?
Hypothesis engines such as Google’s Co-Scientist, SciAgents, Kosmos, and ResearchAgent generate ideas with language models and evaluate them with language models, inside domains rich in public data. Verifier coupled systems such as AlphaProof and AlphaEvolve achieve real rigor, but only where a formal verifier exists: mathematics and code. Literature tools such as Elicit, Consensus, and Semantic Scholar search and summarize what is written; they do not judge structure. Notebook tools such as NotebookLM ground answers in your sources but grade nothing. The Consilience occupies the space none of them touch: cross vocabulary structural correspondence, judged by a deterministic engine, graded on two axes, in domains that have no verifier.
Is this like The Consilience Project or the E.O. Wilson book?
No affiliation. Wilson’s 1998 book named the dream of unified knowledge; The Consilience Project (Daniel Schmachtenberger) works on civilizational sensemaking. We build an instrument: the graded hypothesis engine.
Does it prove spirituality is scientific?
No. The instrument grades honestly in both directions, and many beloved claims grade as thematic resonance only. We publish that. What survives means something precisely because the grader rejects what does not.
Can I see why a verdict was reached?
Yes. One click opens the audit view: both structure graphs, the mapping, why the next rung up was denied, the ontology version, the judging parameters, the reviewer decision. Built for the reviewer who wants us to be wrong.
What did you deliberately not build?
No fine tuning: nothing a researcher ingests ever enters model weights, a privacy guarantee and a judgment guarantee in one. No end to end learned judge: a verdict no one can inspect is not a scientific instrument. No engagement metrics in the loop: the grader answers to the calibration benchmark, never to what users would prefer to hear.