One Engine, Four Instruments
The same verification pipeline serves a free tool for researchers, a decision artifact for biotech, a diligence report for investors, and a sourcing desk for creators. What changes between them is smaller than anyone expects.
A question I get from technical people, usually about four minutes into a conversation, is whether four products from one founder means four half built things.
It is a fair suspicion and it has a precise answer. There is one pipeline. A vertical is a configuration of it rather than a codebase. Exactly four things change between the products and everything else is shared. What follows is what those four things are, what each instrument does for the person holding it, and where each one is genuinely better than the alternative and where it is not.
What the engine does
Give it two or more bodies of claim bearing text. A model reads each passage into a typed claim under schema constrained decoding: what is asserted, about what, in which direction, at what scale, and what would measure it. Claims destined for matching pass a cross model agreement gate, independently re-derived by a second model of a different lineage, where disagreement blocks a claim rather than being quietly resolved. Then the models are switched off. Each body of work becomes a typed directed graph, matching is VF2 subgraph isomorphism under type constraints with maximum common subgraph scoring for partial matches, a Newman guard rejects untyped or defect edge matches, and every match is graded on a four rung ladder from formal isomorphism down to thematic resonance, with an explicit ceiling stating what the correspondence does not license. A final layer asks, of each genuine disagreement, whether both sides predict in a shared observable space, at the same physical scale, reachable by instruments that exist. Every claim resolves to a source sentence. Every refused match is published with the structural reason it failed.
Same inputs, same verdict, regenerated under a content hash in about seven seconds.
The four things that vary
Who selects the inputs. In the consciousness work I selected six theories. In the biopharma instrument a company names its own mechanistic thesis and its rivals. In diligence a fund points at a target company and the frameworks competing with it. For creators it is their own library alongside a shared canon.
Whose instruments define testable. This is the variable people underestimate. Testability is computed against a versioned instrument capability map, which is structured data rather than a model. For a research question that map is the field's instruments. For a biotech workspace it is extended with that company's own assays, so a verdict of testable now means testable by them, in their laboratory, this quarter. For diligence it narrows further, to the milestones the financing round is actually paying for.
What the output looks like. Four interactive panels for a researcher. A board ready landscape for a biotech. A per deal report for a fund. A sourced script with citation overlays for a creator.
Which claim vocabulary is loaded. The ontology is native to mind and brain sciences. A new field needs a vocabulary pack, which extends the ontology rather than forking the engine.
Nothing else changes. Ingestion, extraction, the agreement gate, graph construction, matching, grading, ceilings, receipts, rejections, tenancy, metering, evaluation. All shared.
The researcher's tool, and why it is free
A researcher selects two to five papers by identifier or upload. They get the claim maps side by side, the agreements in the middle including whether the two even make predictions in the same observable space, the fork ledger with each genuine divergence graded and marked for testability, the rejections tray, and a manifest with a regenerate button that demonstrates determinism in front of them.
This is the work behind every related work section, every response to a reviewer asking how your contribution differs from someone else's, every choice between two theoretical framings at the start of a project.
Where it is better than what exists: Elicit, Consensus, SciSpace, scite and their peers find, screen, summarise and rank documents. They are good at that and I use them. None of them compares what the selected sources structurally commit to, none produces a testability verdict, and none publishes what it refused to conclude. As of late 2026 that layer is empty.
Where it is not: this will not find you papers. It compares things you already chose. And structural comparison is an occasional task rather than a daily chore, which is exactly why it is free. Researchers pay for chore removal, and I would rather say so than build a subscription on an unevidenced willingness to pay.
The mechanism landscape for biotech
A company developing a therapy has a mechanistic thesis. This target, this pathway, this account of why the disease behaves as it does. Two or three rival accounts exist, defended by serious people, and choosing between them is the most expensive decision the company makes.
The landscape takes their thesis and the rivals and returns what everyone agrees on, where the genuine forks are, which forks their own assays can settle now, which would need instruments that do not yet exist, an experiment sketch for each settleable fork, and the full rejection log.
Where it is better than what exists: today this is bought as a consulting engagement at fifty to two hundred and fifty thousand dollars, taking two to three months, arriving as a static document whose data collection ended weeks before delivery. It cannot be re-run, audited, or updated when new papers land. The engine version is reproducible, carries a receipt under every claim, refreshes as literature arrives, and publishes what it refused to conclude. Against the knowledge graph platforms the distinction is sharper. They tell you what the evidence says about a link. This tells you where rival explanations genuinely disagree and which disagreement an experiment can settle.
Where it is not: it does not design molecules, run simulations, predict outcomes or nominate targets. It has nothing to say about which mechanism is correct, and it will not tell a chief scientific officer their thesis is weaker than a rival's, because that is neither a claim the instrument is entitled to make nor a product anyone would buy.
Diligence, and the question only this layer answers
A fund's scientific partner has a live deal and an investment committee meeting in three weeks. The company's entire value rests on a mechanistic story.
Two tiers exist, and the split is about trust rather than pricing. The first uses public sources only, the company's published papers and preprints against its rivals, delivered in five working days. No data room, therefore no confidentiality negotiation, therefore no barrier for a supplier nobody has heard of yet. The second adds the data room and does the thing that makes the product distinctive.
That thing is testability mapped onto the company's own stated milestones. Not "could an experiment settle this fork" but "is the experiment that settles this fork inside the plan this round is paying for?" A crux that is decidable in principle but not inside the runway is a different investment from one that resolves before the next raise, and nothing on the market computes that distinction.
Where it is better than what exists: expert network calls run five hundred to two thousand dollars an hour with annual minimums, and produce informed opinions with no receipts and no reproducibility. Diligence consultancies do excellent work over two weeks of senior time. Both stay useful, and the report is a complement rather than a replacement. The map tells you which three questions to actually ask the expert.
Where it is not: no investment recommendation, ever. No invest, no pass, no score. The report is the analyst's instrument and the judgment stays with the partner, which protects them in their own committee and is the reason the artifact is buyable at all.
The creator desk, which inverts the design
This one is shaped differently and is worth understanding precisely because it breaks the pattern.
Creators working at the intersection of science and contemplative traditions have a specific problem. The genuinely interesting connections require reading across fields, their audience has learned to distrust the genre, and the available tools produce fluent unsourced scripts that get them fairly criticised. What they need is sourced material and receipts, not a verdict on their idea.
So the assignment changes. Models find, using retrieval trained on this project's own signed correspondences and its own historical false positives as hard negatives. Deterministic code ranks candidates and supplies the plain words reason two passages match, but never vetoes. A person signs every published connection, with a label and an explicit statement of where the correspondence stops. And in drafting, the fact carrying sentences are assembled from verbatim spans of the cited passage rather than written freely, with free prose surviving only in an opening and closing that are constrained so they cannot assert anything.
That last mechanism is the most interesting engineering result the project has produced. It makes citation hallucination structurally impossible rather than statistically unlikely, because a receipt is a database row rather than generated text and the fact carrying language is assembled rather than composed. Thirteen of thirteen tested bridges, zero violations.
Where it is better than what exists: video tools script, voice and assemble from nothing, with no corpus, no citations and no honesty about confidence. Nothing in that market grades a connection or shows the passage behind it.
Where it is not: it does not make video, and it should not. The assembly layer is already solved by tools creators pay for.
What is shared, and why that matters more than what differs
Every instrument above runs the same intake, the same claim extraction against the same frozen ontology, the same immutability rule on source segments, the same receipts as database rows discipline, the same tenancy isolation, the same metering and cost controls, and the same evaluation suites.
The practical consequence is that a new vertical is weeks of vocabulary authoring rather than months of engineering, and improvements to the shared layers reach every product at once. When extraction quality improves for biotech, the researcher tool improves the same day.
The strategic consequence is what I would say to anyone weighing whether four products from one founder is a red flag. It would be, if they were four products. They are four reading surfaces on one instrument.
The line that runs through all of them
The model reads and finds. Code decides and grounds. A person signs.
Which of those three does the deciding varies. Deterministic code decides in the science instruments, because the output is a verdict and a verdict has to be reproducible and open to inspection. A named person decides in the creator instrument, because the output is a public claim and accountability cannot be delegated to a function.
In none of them does a language model get the last word. That is the architecture, and the rest is configuration.
The honest constraint
The ontology is native to mind and brain sciences, which is where it was authored and calibrated. Consciousness science and neuroscience run today. Oncology, immunology, materials, anything else: the core loop is domain neutral and the domain specific content is small, so a new field is an authored vocabulary pack with its own calibration slice rather than a rebuild. But authoring it takes a domain expert several weeks, and I would rather say that plainly than have it discovered during a pilot.
That is the real cost of building a verifier rather than a generator. A learned world model generalises on its own. An authored one does not. The trade is precision and auditability against reach, and in a field whose dominant failure is confident nonsense, I think it is the right trade.