Part 2 of 4  ·  2026

Choosing an Architecture

A survey of the serious alternatives to transformer based reasoning, and an account of which pieces I took, which I ranked first and could not build, and which I declined.


Most architecture posts are advocacy dressed as survey. The author walks through six approaches, finds fatal flaws in five of them, and discovers that the sixth happens to be the one they built.

This is not that, and I want to establish it early. I ranked one approach first on merit and did not build it, because one person could not. I took working ideas from a system designed by someone who now leads technology at a company operating in adjacent territory. I declined several directions that are probably correct and simply belong to a different problem than mine. And the architecture I did build carries a weakness that a learned system does not have.

The field is small and the people in it are serious. The fastest way to have a useful conversation is to put the decision record on the table.

The problem, stated once

Two bodies of work describe something. They use different vocabularies. Are they claiming the same structure, genuinely different structures, or things that cannot be compared at all? And where they genuinely differ, could any experiment tell them apart?

This is not literature search, which finds documents. Not summarisation, which compresses them. Not knowledge graph construction, which encodes entities and relations inside one field's ontology. It is adjudication, and today it is done by qualified people in rooms, over months, producing an opinion nobody downstream can audit.

Every architecture below was assessed against that problem. Several of them are better than mine at problems that are not mine.

Gary Marcus and the neurosymbolic case

The longest standing argument in the field. Language models are pattern recognisers rather than reasoners, and scaling will not repair what is an architectural limitation rather than a data one. The future is hybrid: neural front end for perception and language, symbolic back end for reasoning that carries guarantees.

Marcus is often read as a critic rather than a proposer, which does him a disservice. His requirements are specific. The hybrid has to be one system rather than two sitting next to each other, prior knowledge has to be encoded structurally rather than learned from scratch, and formal reasoning techniques have to be present rather than approximated.

This is not an alternative I rejected. It is the family my system belongs to. Neural reading, symbolic judgment, one pipeline. Naming it correctly matters, because it moves a conversation from "so you wrote some rules" to a research tradition with thirty years of literature behind it.

Subbarao Kambhampati and external verification

The sharpest empirical case, and the one I lean on hardest. Kambhampati describes language models as non-veridical memory systems, and he has results rather than rhetoric behind it. Models fail block stacking problems that classical planners solved trivially in the 1970s, and performance degrades further when the block names are obfuscated, which shows the apparent reasoning was pattern recall wearing the costume of deliberation.

His proposal is not to discard the model. Keep it for what it is genuinely good at, which is language and retrieval and intuition, and pair it always with an external verifier. Nothing on a claim that matters reaches a user without passing a check that is not a language model.

I accepted this as the governing pattern of the whole system. It is also the answer to the objection I hear most often, which is that reasoning models keep improving so the judging problem will resolve itself. The failures Kambhampati documents are structural. They do not dissolve with scale. They get more fluent.

Yoshua Bengio and the generator estimator split

Bengio's Scientist AI programme starts from a different complaint. Imitation trained models treat all text as truth and will reproduce whatever pattern fits, including the false ones. His alternative trains a system to explain why people say what they say, with the true state of the world as a latent variable to be inferred rather than asserted.

Two mechanisms matter to anyone building in science. The first is contextualisation, which splits training data into communication acts and verified facts rather than treating a tweet and a replicated experiment as the same kind of object. The second is architectural: a creative generator proposes, a neutral estimator scores, and the two cannot be the same network.

I accepted the separation in principle and adapted it in practice. The generator estimator split is the same commitment as generate then verify, arrived at from a different direction, and it is the strongest academic citation available for the design. Bengio's contextualisation maps directly onto the source tiering I use, where a peer reviewed paper, a contemplative text, and an unverified web page carry different evidential weight, and the lowest tier can seed a hypothesis but can never appear in an evidence chain.

What I did not adopt is the estimator as a neural network emitting Bayesian posteriors. Mine is deterministic code emitting a grade. A probabilistic estimator still drifts across runs, and still cannot break its verdict into steps a reviewer can attack.

Bernhard Schölkopf and invariance

Observations are generated by a few low dimensional causal variables projected into high dimensional data. Current machine learning learns the projection rather than the generators, which is why models break under distribution shift and cannot reason about interventions. The principles are independent causal mechanisms, disentanglement, and invariance across environments.

This framework fits my problem more precisely than anything else on this list, and it took me a while to notice why. The claim that two fields describe the same structure is formally a claim that they share a causal generator, where the field is the environment and the vocabulary is the projection. Cross vocabulary matching is causal disentanglement applied to disciplinary context.

I adopted it as framing and as an evaluation standard rather than as machinery. There is no formal disentanglement in my system and no identifiability proof. What I took is the standard it implies: a claim's structural representation should be identical whether the claim is written in physics, in Vedānta, or in phenomenology. My calibration set tests exactly that invariance, including against deliberately convincing fakes that are required to fail it.

Yann LeCun, concept space, and the thing I built and killed

LeCun argues that text is a serialisation of reasoning rather than reasoning itself, so systems trained to predict tokens learn the shadow rather than the object. JEPA predicts in latent representation space instead.

The variant that matters most here is the Large Concept Model work from Meta FAIR in December 2024. Process sequences of sentence embeddings rather than tokens, train on concept prediction, reason in a space where the same idea across two hundred languages lands on the same point, and decode to words only at output.

I want to be exact about this one, because it is the closest published relative of something I built and then destroyed. My first architecture translated each passage into a neutral vocabulary and embedded the translation. That is concept level reasoning at the application layer, done with prompting instead of a trained latent space. Meta was doing the same thing properly, at the foundation model layer.

It produced six thousand three hundred and forty one connections before I killed it, and the reason is worth stating plainly because it is the single most important thing I learned. Similarity and structural correspondence are different relations. Two texts can share every word and no structure, which is how numerology works. Two texts can share no vocabulary at all and be formally identical, which is why temporal difference learning in computer science and dopamine reward prediction error in the midbrain are the same object described twice, and no embedding would ever have paired those papers. The system generated false positives and missed true positives simultaneously, for the same reason, and no amount of tuning addresses a system computing the wrong relation efficiently.

So JEPA and concept models were studied and not adopted, for two reasons of different weight. The small one is buildability, since closing the gap means training a foundation model. The large one is that they change how a model reads, not who judges. If a JEPA class architecture replaced transformers entirely tomorrow, I would swap my reader and keep my judge, because the reason my judge is not a model has nothing to do with the model's architecture.

There is an irony worth recording. The component of my first system with the strongest frontier pedigree, the concept space layer, was eventually retired for a reason unconnected to any of this. The neutral translation was written through an individual user's interpretive lens, which makes results incomparable between users and makes a shared citable public corpus impossible. It died of a product requirement rather than a technical one.

The six thousand wrong answers were not deleted. They are the best collection of convincing but false correspondences in this domain that exists anywhere, and they now teach the candidate finding layer what a plausible looking fake looks like. You only get that asset by building the wrong thing first and being honest about it.

Symbolica, and the deferral that matters most

Symbolica's bet is that reasoning should be modelled as structure preserving maps between formal systems rather than as pattern matching in vector space. A functor preserves relational structure: what causes what, what implies what, what excludes what.

I ranked this first. Not first among things I rejected. First on merit against my specific problem, and I want to be precise about why. Category theory is quite literally the mathematics of two different descriptions of the same thing. If a quantum description and a contemplative description of a phenomenon genuinely point at one structure, they should be related by a functor. My old concept space layer asserted this informally. A categorical layer would make structural equivalence verifiable rather than asserted, which is the entire point of my product.

My own architecture review put the verdict this way: beautiful, but not productionisable in 2026 outside Symbolica itself, so borrow the concept without the heavy machinery. There was no library to adopt and the only people building it were a pre product company with thirty three million dollars.

What I borrowed instead was the concept in pragmatic form. Structured claim objects with typed predicates, directed edges and causal links, extracted before any embedding happens, so that matching filters for structural compatibility rather than for similarity. That is the typed claim my system runs on today, and the categorical layer remains the stated upgrade path rather than a road not taken.

This deferral is my honest answer to anyone whose work sits in that space, including teams layering neurosymbolic mathematical abstractions over a world model. It is not an approach I rejected. It is an approach I ranked first and could not afford, and the distance between those two sentences is the whole of my respect for the direction.

GFlowNets, program synthesis, active inference

Three more, each declined for a reason worth recording.

Bengio's GFlowNets sample from a learned distribution over structured objects proportional to reward, producing a diverse set of high reward candidates rather than one mode collapsed answer. Deferred, never built, but the diagnosis survived without the mechanism. Researchers do want the space rather than one answer, and my expression of that instinct lives on the verification side as a ledger of live disagreements between frameworks rather than a tidy resolution.

François Chollet argues that intelligence is skill acquisition efficiency rather than accumulated skill, and proposes deep learning guided program synthesis where the network guesses which programs are worth trying and classical search verifies them. Deferred as a later option. The relevance is to generation rather than adjudication, and Ndea is pre product on a three to five year research horizon, which is the right horizon for the problem and the wrong one for a solo build.

Karl Friston's active inference and Tomer Ullman's Bayesian theory acquisition both formalise something my early system approximated informally. Each user had a lens, an explicit interpretive prior they selected, and each confirmation they gave was evidence. Ullman is the exact formal model of that. Friston's version would have made a failed prediction the most valuable signal in the product. Both were designed and then dropped, because they belong to the lens and the lens turned out not to be the product. I record them because the reason for dropping them was strategic rather than technical, and that is a different kind of decision which deserves a different label.

The deployed systems, and what I took from each

The research programmes above are mostly pre product. The systems below are running, and they taught me more per hour of reading.

AlphaProof and the Lean harness. A model proposes proof steps and Lean verifies every one. A 2026 result sharpened the lesson considerably: a stock language model with no additional training, placed in an agentic harness with a Lean verifier, outperformed a purpose trained specialist, taking solve rate from under ten percent to around seventy. The harness beat the specialisation. I took this as the anchoring pattern, with the caveat recorded that mathematics has a total verifier and the empirical sciences do not, so it is an inspiration rather than a template.

AlphaEvolve. Model mutates candidate programs, execution scores them, and a matrix multiplication record standing since 1969 falls. Corroboration in a second domain that wherever a verifier exists, coupling generation to it is where the result comes from.

Kosmos. Twelve hour autonomous research campaigns, over a thousand papers per run, roughly seventy nine percent statement accuracy. Its credibility mechanism is not a better model, it is a structured world model that code updates after each step plus a traceability discipline linking every claim to a source. I took that directly. A world model exists in my schema with provenance on every row. The difference is authorship of its contents. Their agents write beliefs into it. My code writes proofs of match.

Google's AI Co-Scientist. Seven agents, Elo tournaments, and an independent re-derivation of a finding that took a laboratory a decade. Its own paper states that self-evaluation is unreliable. I did not adopt it as architecture, I use it as evidence. The team that built the largest model-judging-model system in existence documented the failure mode in the paper announcing it, and that citation does more work defending my design than any argument of mine.

SciAgents, from Markus Buehler's group at MIT. An ontological knowledge graph of roughly thirty three thousand nodes built from about a thousand papers, with a multi agent team reasoning over sampled paths between two concepts. An Ontologist owns the concept and relation definitions, a Scientist drafts a hypothesis, another expands it, a Critic reviews. My own landscape research names it the single closest architecture to mine.

I took two things from it. The Ontologist role, which became my fixed ontology, so that concept and relation definitions are owned by one authored artifact rather than improvised per query. And the two concept path query, which became my bridge mechanism, with their contrast between random and shortest paths serving as a ready made novelty control.

I declined two things. In SciAgents the agents reason over the sampled path and produce the hypothesis, which puts judgment inside the model. And structurally, SciAgents works inside one domain, bioinspired materials, where everyone shares a vocabulary. My research note on it is exact: their mechanism is mine minus the vocabulary stripping. They did not need that step. I cannot work without it.

What I deliberately did not copy is the autonomous laboratory stack, meaning Lila Sciences, Periodic Labs, Sakana's end to end paper writers and their peers. Not because the work is weak. Because their verifier is physical reality and mine can never be. Nobody will run an experiment to determine whether two theories share an abstract structure. In my domain the verdict is the product rather than a hypothesis awaiting a laboratory.

What I built, and where each piece came from

A model reads and emits a typed claim under schema constrained decoding: what is asserted, about what, in which direction, at what scale, and what would measure it. That lineage is Marcus and my own architecture review.

Claims destined for matching are independently re-derived by a second model of different lineage, a cross model agreement gate where disagreement blocks the claim rather than being quietly resolved. That is my own addition, and it sits exactly where the design is most exposed. It is also correlated evidence rather than independent confirmation, since models that both err converge on the same error more often than chance, which is why it is one gate among several rather than the only one.

Each body of work becomes a typed directed graph, and matching is VF2 subgraph isomorphism under node and edge type constraints, with maximum common subgraph scoring where the match is partial. Neither algorithm is novel and neither was modified. VF2 dates to 2004 and ships in networkx. The lineage here is Kambhampati's external verifier, Gentner's structure mapping theory of analogy, and the pragmatic residue of the categorical direction.

Every match is graded on a four rung ladder from formal isomorphism down to thematic resonance, which is quarantined rather than celebrated, and every grade carries an explicit ceiling stating what the correspondence does not license. Mine, and the thing I would defend hardest.

A guard named for M. H. A. Newman's 1928 objection handles the deepest technical risk in the whole design. Newman pointed out that any two systems share some structure once you stop typing the relations, which makes unconstrained structural claims trivially satisfiable. Typed relations are the answer, and a separate mechanism handles the case where a correspondence leans on a shared word rather than on shared structure.

A decidability layer then asks, of each genuine disagreement, whether both sides predict in a shared observable space, at the same scale, reachable by instruments that exist. It is a deterministic lookup against a versioned instrument capability map, which makes Popper's falsifiability criterion computable rather than rhetorical. Testable now, testable under a stated assumption, or not testable today, with the reason attached. Mine, and the piece with the clearest value to anyone making a decision that costs money.

Every verdict resolves to source sentences, regenerates under a content hash, and every refused match is published with the structural reason it failed.

The weakness

A learned world model generalises across domains on its own. An authored ontology does not. Someone has to write the claim vocabulary for each new field, and that cost is measured in weeks of expert time rather than in compute.

This is the genuine advantage held by systems that learn their representations, and I would rather state it here than have it discovered later. The bet is that in a field whose dominant failure mode is confident nonsense, precision with receipts is worth more than recall with fluency, and that the authoring cost is a real price rather than a fatal one.

The second weakness follows from the same trade. My judge has lower recall than a frontier reasoning model would. It misses subtle analogies a capable model would catch, and those misses are logged rather than hidden.

The one sentence difference

In the systems I admire most, symbolic structure feeds the model's reasoning and the model produces the answer. In mine, the model produces typed claims and something that is not a model produces the answer.

A model never gets the last word. Either deterministic code decides, or a named person signs. That is the whole architecture, and everything above is the record of how I got there.

If you are building in this space and think I have a piece of this wrong, I would rather hear it than not.