Research programme
One ever-larger model is a hypothesis, not a conclusion.
Project Octoryn Cortex asks whether general capability must live inside a single scaled model, or whether it can be formed as bounded cognitive domains and composed through a governed protocol. We are publishing the contracts, the budget and the conditions under which we will call it a failure — before the first run, not after the results.
Where the work stands
Contracts exist. Nothing has been trained.
The repository holds versioned domain manifests, an immutable exchange protocol, a memory state machine, dataset and run manifests, safety interlocks and an append-only audit chain. It holds no training run, no model weights, no benchmark result and no inference service. Every operational feature flag is off, and the diagnostic tool fails if one is switched on while the project remains in foundation mode.
The question
Scaling works. That is not the same as it being the only thing that works.
Capability keeps arriving from larger models trained on more data, and nothing here disputes that. But "it works" has quietly become "it is the only route," and that second claim has never been tested against a serious alternative under a matched budget. Three narrower questions are answerable.
When has a domain actually formed?
A prompted role is not a domain. A retrieval index is not a domain. We require an explicit boundary, its own lifecycle and versions, a separate memory namespace, release gates, and the ability to abstain — and then we have to show the difference is measurable.
What does composition actually cost?
Routing, protocol overhead and cross-domain exchange are usually invisible in reported results. They are counted here as first-class costs, alongside FLOPs, active parameters and latency, because an architecture that wins only by hiding its overhead has not won.
What would prove the thesis wrong?
If a merged model matches the domain network on capability, retention and update cost under the same budget, the thesis fails. That comparison is pre-registered, and the outcome table is written before any arm is run.
The thesis
Formation happens in layers. Each one has to earn the next.
Capability is built as a sequence, not a single pass. Each layer is a separate research question requiring separate authorisation, and none of them inherits evidence — or approval — from the layer below it.
A K–12 educational progression
The foundation is formed along a developmental learning sequence rather than a scraped corpus. The ordering of educational material is treated as a property being learned, not as an artefact of where the data came from — which is why curriculum stage, subject and level have to be evidence-backed rather than inferred from folder names.
Discipline material at university-major level
A cognitive domain is intended to sit at roughly the granularity of a university major. Biology and Medical are the first pair precisely because they are adjacent: close enough that a single merged model is a genuinely strong control, rather than a straw baseline chosen to be beaten.
An organisation's own material
A formed domain plus an organisation's training material is intended to yield a role-scoped capability. This layer is why abstention, declared scope and evidence traceability outrank aggregate accuracy throughout: something relied on inside a role has to be able to decline, cite and escalate. It is out of scope until layers 1 and 2 produce evidence.
Read the full statement of goals, non-goals and open problems on Goals and scope.
Approach
Domains propose. The protocol carries the doubt.
A domain never gets the last word by being fluent. It emits a claim with its evidence, scope, assumptions, confidence and known contradictions attached, and the receiving domain is not permitted to silently broaden it. Fluency is not traceability.
Exact versions, or no run
Every domain names its exact foundation, prerequisites and model artefacts. A run is bound to immutable inputs, budgets, outputs and audit evidence before it may start, and changing any of them produces a new run identity rather than a quiet amendment.
Memory is not training
Experience moves from raw observation to approved consolidation through explicit stages with human authorisation. Raw events never reach parameters. An agent that silently learns from what it just saw cannot be audited afterwards.
Abstention is a first-class outcome
Domains are scored on whether they decline and escalate correctly, not only on whether they answer correctly. The router is scored separately from the answering domain, so routing failures are never reported as reasoning successes.
Provenance before content
The source snapshot was profiled from paths, sizes and hashes without opening a document. Selection is not approval, classification is not extraction, and extraction is not permission to train — each is a separate gate with named approvers.
Position
A confident answer outside its boundary is the failure, not the feature.
The systems this programme is aiming at are meant to be relied on inside real work, where the cost of a fluent wrong answer is borne by someone who had no way to check it. That makes the interesting engineering question not how often a model is right, but whether it knows the edge of what it was formed to know — and says so.
Which is also why the governance came first here. The manifests, the audit chain and the two-person authorisation gates are not compliance theatre bolted on after a result; they are the reason a result would be worth believing. A programme that cannot say exactly which bytes, code and budget produced a number has not produced a finding.
We are not claiming AGI, and we are not claiming this architecture wins. We are claiming the comparison is worth running honestly, and building the apparatus that makes an honest answer possible either way.
Nothing here has been trained
There is no model, no checkpoint and no benchmark result behind this page. What exists is the contract layer that a future run must satisfy. Any capability claim about Project Octoryn Cortex — from us or from anyone else — is unsupported until a run manifest and its evidence are published.
Read next