Why This Tool Exists
The correspondences that matter most are invisible to every tool we have, because vocabulary is exactly what they do not share.
There is a moment that keeps happening across the sciences, and nobody has built the infrastructure for it.
A computer scientist develops an algorithm for how a system learns from delayed rewards. Decades later, a neuroscientist discovers that dopamine neurons in the midbrain fire in exactly that pattern. The mathematical structure is the same. The vocabulary shares nothing. Nobody noticed for years, because the two fields do not read each other's journals. When the connection was finally made, it became one of the most important bridges in computational neuroscience, the link between temporal difference learning and reward prediction error.
That bridge was built by accident. By a person who happened to work across both fields.
The second law of thermodynamics says entropy increases in a closed system, that nothing persists without energy sustaining it. The Buddhist doctrine of impermanence says the same thing: all conditioned phenomena are in constant flux, nothing endures without causes maintaining it. The mathematical form is identical. The vocabularies were developed on different continents, in different centuries, for entirely different purposes. The correspondence is structural, not metaphorical. And no tool alive today can find it.
This pattern repeats across every pair of fields that happen to describe the same underlying structure using different words. It is not rare. It is the norm. What is rare is someone noticing, because every tool we have searches by vocabulary, and vocabulary is exactly what these correspondences do not share.
The problem
Every scientific literature tool in existence, Elicit, Semantic Scholar, Consensus, Google Scholar, NotebookLM, operates on words. You search for terms. You retrieve documents containing those terms. You receive summaries in those terms. If two fields describe the same structure using entirely different terminology, no keyword search and no embedding will pair them. The wall is not one of access. It is one of representation.
Embeddings get closer. They cluster by meaning rather than by exact words. But they fail on this task in a specific, measurable way. We built an embedding based system and tested it. It produced 6,341 mappings across a research corpus. The structural audit found systematic false positives wherever vocabulary overlapped without structural correspondence, and systematic false negatives wherever structural correspondence existed without shared vocabulary. Cosine similarity and structural correspondence are different mathematical relations. The system was computing the wrong thing efficiently.
Language models get closer still, and fail in a way that is harder to see. A frontier model asked to compare two theories produces a fluent, plausible answer. Ask again tomorrow, different answer. The output reads like insight. It is a forward pass through a network trained to be helpful, and helpfulness training makes the model reason toward yes. The more capable the model, the more convincing the false positive. There is no receipt. No reproducibility. No way to check whether the correspondence is real or whether the model is pattern matching on vocabulary, just with better prose.
The field's own results confirm this. Kambhampati showed at ICML 2024 that autonomous language model plans are correct about twelve percent of the time. AlphaProof solved mathematical olympiad problems not because the model became smarter but because they coupled it to Lean, a formal verifier that told the model when it was wrong. The pattern is the same everywhere the results hold: a model generates, a non model checker verifies. Mathematics has Lean. Code has execution. Cross domain scientific correspondence has had nothing.
What we built
The Consilience is the missing checker.
Language models read papers and extract typed claim cards: what kind of claim, which direction the causality runs, at what physical scale, what instrument would measure it, and the exact source sentence. Output is constrained, validated, and cross checked by a second model of different training lineage. Disagreement excludes the claim from matching.
Then deterministic code takes over. Typed directed graphs capture the structural commitments of each theory, independent of vocabulary. Subgraph isomorphism with type constraints checks whether one theory's structure maps onto another's. A graded ladder assigns one of four rungs, each with an explicit ceiling stating what the verdict does not license. A versioned instrument capability map checks whether each genuine disagreement is testable with equipment that exists today. Published rejections travel alongside the confirmations, with the structural reason attached.
Same inputs, same output, every time. Verifiable by content hash.
We blind tested it against the COGITATE adversarial collaboration, which spent eleven laboratories and nineteen months comparing two major theories of consciousness through structured expert deliberation. The engine recovered the same structural disagreements from theory texts alone, and surfaced one the experts could not resolve, with the specific reason it is untestable attached. One person built it. It took weeks. It required zero expert panels.
Why now
Two shifts are converging that make this tool both possible and necessary at the same moment.
The first is that analytical intelligence is becoming infrastructure. Jensen Huang frames it as the third force to reshape civilisation after electricity and the internet. When cognitive capability becomes abundant, the bottleneck moves from generating answers to verifying them, from producing hypotheses to knowing which ones are structurally sound and testable. The verification gap becomes the constraint. This tool is built for that gap.
The second is that cross domain research is no longer optional. The hardest problems in science, consciousness, climate, disease mechanisms, intelligence itself, sit at the intersection of fields that have developed independent vocabularies for the same underlying phenomena. Solving them requires crossing vocabulary walls that have stood for decades. The people who will make the breakthroughs are the people who can see that two fields are describing the same structure, and who have an instrument to verify that recognition rather than relying on intuition alone.
The engine is subject agnostic. It serves any field where the domain vocabulary has been authored. Consciousness science was first because the founder had domain access and because the COGITATE adversarial collaboration provided a preregistered answer key to validate against. Biopharma mechanism comparison is next, where the same structural comparison problem exists and real budget lines already fund manual versions of it.
The deeper why
There is a personal reason this tool exists, and it matters because it shaped the design.
The founders of quantum mechanics, Schrodinger, Heisenberg, Bohr, Bohm, all encountered something remarkable. The formal results they were deriving in physics resonated structurally with contemplative traditions that had been mapping consciousness for centuries. Schrodinger kept the Upanishads on his desk. Heisenberg said quantum theory would not look ridiculous to people who had read Vedanta. Bohm spent a decade in dialogue with Krishnamurti and it changed his physics. These were not casual interests. These were instances of the same structural recognition appearing in two places that shared no vocabulary.
That kind of recognition is accelerating. Mental health as a global crisis is driving millions toward meditation and contemplative practice. Neuroscience research on these practices has moved from the margins to the centre of the field. And as artificial intelligence makes the analytical mind abundant, a growing number of people are discovering through direct experience that there is something beyond analysis, something that contemplative traditions have mapped with extraordinary precision but that scientific vocabulary has barely begun to describe.
When someone has that experience, a precise structural recognition that crosses the boundary between what science describes and what contemplative practice reveals, they face a problem. No tool can tell them whether the recognition is structurally real or whether their mind is projecting similarity onto unrelated things. Embeddings find vocabulary matches. Language models produce eloquent confirmations. Neither can show you where the correspondence actually breaks. Neither can tell you whether the disagreement is testable.
I know this because I had that experience. I have a physics degree and an MBA. I spent five years building blockchain infrastructure from zero, working as the first hire at the Arbitrum Foundation and reporting to the board. I left to study and build. During that period I spent sustained time with contemplative practice and with the scientific literature on consciousness. I encountered, through direct phenomenological experience, something that the analytical mind could not reduce but that multiple traditions had clearly mapped. The recognition was structural, not doctrinal.
The question that followed was precise: is this real, meaning structurally grounded, or is my mind projecting similarity onto unrelated things?
No tool could answer it. So I built the instrument.
Not to prove that any tradition is right. Not to confirm what I experienced. To build something that can distinguish a real structural correspondence from a vocabulary coincidence, with receipts that anyone can check. If the correspondence is real, the instrument shows it and names the evidence. If it is not, the instrument rejects it and publishes the reason. The answer is the answer, not what the user hoped for.
Why nobody else is building this
The vocabulary problem is invisible from inside any single discipline. A neuroscientist does not know that their model of reward prediction matches a computer scientist's model of temporal difference learning, because they never read each other's papers and the words share nothing. The problem is obvious only from the outside.
Recall sells. Precision does not. Every tool in this space is measured by how much it finds. Publishing rejections hurts the demo. But it is the only thing that builds trust with someone who has been burned by a false positive, and in biopharma and scientific diligence, false positives cost millions.
The moat is authored, not computed. The typed vocabulary that makes matching meaningful, the adversarial test set that keeps the system honest, the instrument capability map that checks testability, these are hand authored by domain experts. Weeks of work per field. No shortcut. No way to prompt a model into producing them. That is why the asset compounds and nobody can replicate it quickly.
The long term
Every breakthrough that connected distant fields did so through a representation where structure is explicit and correspondence is checkable. What a system that helps researchers at exactly that step would look like is the question driving this work.
The engine today handles the structured comparison. The next layer is category theoretic matching, where a functor preserves relational structure as a mathematical guarantee rather than a typed approximation. That layer does not yet exist in usable form. Our typed matching is the practical version. If the formal version arrives, our matching layer is the natural replacement candidate.
Until then, the engine works. It finds real structure across vocabulary walls that have stood for decades. It tells you where theories agree, where they disagree, and whether anyone can test the disagreement. It publishes what it refuses. And it gives the same answer every time you ask.
Nobody else is building this. Because nobody else sits at the intersection of the direct experience that makes the need obvious, the technical depth to build the instrument, and the willingness to let the instrument say no.