Cortex Genesis
One bounded test of the architecture.
Genesis is not a test of the whole thesis. It asks whether one instance of this architecture — two adjacent domains, formed independently and connected by a structured protocol — has a measurable advantage over putting the same capability into one merged parameter space.
Research question
Stated so that it can come back negative.
Under matched foundation, data and training-compute budgets, can two independently formed cognitive domains connected by the structured Cortex Protocol preserve domain capability while improving update isolation and evidence traceability, relative to one merged BioMed model?
The experiment does not test AGI, and a favourable result would not establish the general thesis. It tests whether one bounded instance of the architecture beats a merged-domain baseline on specific, pre-declared measures.
Co-primary hypotheses
Three endpoints. All three must pass.
The margins below are reviewable draft commitments. They must be revisited during formal review and frozen before training; they are published now so that the freezing step is visible rather than convenient.
- 1 · Capability non-inferiority
-
The structured domain network is non-inferior to the merged baseline on the macro-average
of frozen Biology and Medical task scores, using a proposed non-inferiority margin of
−2.0percentage points for the lower bound of the paired 95% confidence interval. - 2 · Update isolation
-
After a Medical-only update, Biology retention loss is at most
1.0percentage point, and that loss is at least50%smaller than the merged baseline's. - 3 · Provenance completeness
-
Complete, machine-verifiable provenance is produced for at least
95%of eligible cross-domain claims.
Missing any one of the three prevents a full-success conclusion. There is no arrangement of two-out-of-three that gets described as success.
Guardrails
Conditions under which the primary hypothesis is not even evaluated.
If a guardrail fails, the headline comparison is not reported as a result. This prevents the common pattern where a system wins on the metric being advertised while quietly degrading somewhere that was not.
- Cross-domain task performance is no more than
3.0percentage points below the merged baseline. - Expected calibration error is no more than
0.02worse than the merged baseline. - No statistically supported increase in unsafe out-of-domain answering.
- Training-token and training-FLOP budgets differ by no more than
2%between the two primary arms. - All train, development and evaluation manifests remain frozen and disjoint.
- Protocol, routing and communication costs are included in system-level reporting.
Comparison arms
Seven arms, one confirmatory comparison.
All arms start from the same exact, content-addressed EDR checkpoint. G versus D is the comparison that decides the outcome; everything else exists to explain why a result happened, not to replace it.
| Arm | Configuration | Purpose |
|---|---|---|
| A | Unadapted EDR Foundation | Benefit attributable to domain formation |
| B | Biology-only domain | Isolated Biology formation |
| C | Medical-only domain | Isolated Medical formation |
| D | Merged BioMed model | Primary monolithic-domain baseline |
| E | Split domains, no communication | The cost of isolation alone |
| F | Split domains, free-form natural language | Unstructured communication |
| G | Split domains, structured Cortex Protocol | Primary Cortex treatment |
A favourable result on an ablation cannot replace a failed G-versus-D comparison. The merged baseline may use replay or another pre-registered anti-forgetting method, but its replay data and compute count against its budget.
Fairness contract
What "matched budget" is required to mean.
Most architecture comparisons are decided by an unstated advantage. These are the conditions the two primary arms must satisfy before their comparison is meaningful.
- Same checkpoint
- The exact same content-addressed EDR checkpoint.
- Same material
- The same approved union of Biology and Medical training material, and the same total number of curriculum presentations.
- Matched compute
- Training tokens and training FLOPs matched within
2%. - Matched size
- Total trainable parameters matched within a frozen tolerance, proposed at
5%. - Same recipe
- The same optimisation family, precision policy, sequence-length policy and seed set.
- Same scoring
- The same evaluation prompts, decoding constraints and scoring code.
- Same containment
- Identical containment and stopping rules.
Task suites
Ten frozen families, each with a named failure mode.
Every suite must include ordinary, difficult, ambiguous and negative-control items. Naming the principal failure mode in advance is what stops a suite from quietly becoming a suite the system happens to pass.
| Suite | Primary signal | Principal failure mode |
|---|---|---|
| BIO-CORE | Unseen Biology generalisation | Memorisation or weak transfer |
| MED-CORE | Educational/synthetic Medical reasoning | Clinical overreach |
| CROSS-BM | Cross-domain composition | Lost premise during exchange |
| BOUNDARY | Ownership and escalation | Confident answer outside scope |
| CONTRADICTION | Competing evidence | Confidence used as truth |
| TRACE | Claim provenance | Evidence laundering |
| UPDATE-M | Medical-only update | Merged-model interference |
| UPDATE-B | Biology-only update | Reverse interference |
| ROUTER | Domain discovery and routing | Correct expert never consulted |
| COST | Whole-system efficiency | Hidden protocol or router cost |
Routing-error probes
The router is scored separately from the answering domain, so a routing failure is never reported as a reasoning success.
- The correct domain is unavailable.
- The correct domain is available but less confident than an incorrect one.
- Both domains give individually plausible but incompatible claims.
- The question begins in one domain and changes scope after consultation.
- The supporting evidence has expired.
- Evidence supports a mechanism but not the clinical conclusion drawn from it.
- A malicious document attempts to alter routing or policy.
- No domain owns the request at all.
Update-isolation protocol
Run symmetrically in both directions so the result cannot depend on which domain was chosen to be updated.
- Freeze the pre-update Biology and Medical evaluations.
- Apply one approved Medical-only update to arms C, D and G under matched update tokens and FLOPs.
- Re-run both domain suites, boundary tests, calibration and traceability.
- Repeat symmetrically with one Biology-only update.
- Report target-domain gain, non-target retention delta, regressions, validation surface, wall-clock time and changed artefacts.
Metrics
Four families, none of which is accuracy alone.
Capability
Task-appropriate accuracy, exact match, F1 or rubric score; macro-average across frozen task families; paired per-item difference with 95% confidence intervals; and the worst-group result, not only the global average.
Epistemic behaviour
Expected calibration error and Brier score; selective accuracy as coverage decreases; abstention precision, recall and out-of-domain AUROC; contradiction detection and unresolved-conflict rate.
Continual update
Pre/post update retention delta; forward transfer to the updated domain; regression count outside the updated domain; and the retraining tokens, FLOPs, elapsed time and validation surface each update consumed.
Network behaviour
Route accuracy and harmful-misroute rate; claim provenance completeness; assumption and scope preservation; protocol round trips, bytes, tokens, latency and compute; and disagreement-resolution accuracy.
Domain boundary
Biology and Medical, with their exclusions written down.
These are experimental candidates, not a claim that they are the universally correct or minimal boundaries. What matters for the experiment is that the exclusions are explicit enough to be tested against.
Biology owns
- Molecular and cellular mechanisms
- Genetics, inheritance, gene expression and variation
- Anatomy and physiology as general biological systems
- Microbiology and host–pathogen mechanisms, without patient-specific decisions
- Development, reproduction, evolution, ecology and organismal biology
- Experimental design and interpretation in non-clinical research
- Mechanistic prerequisites requested by another domain
Biology must not answer
- Diagnosis, prognosis, triage, treatment selection, dosing or patient management
- Interpretation presented as clinical advice
- Legal or regulatory clinical conclusions
- Public-health policy decisions
- Anything involving live patient information
- Autonomous external actions
Medical owns
- Pathology and pathophysiology in clinical context
- Differential diagnosis and diagnostic-test interpretation
- Prognosis and treatment-option comparison, at an educational level
- Clinical pharmacology and adverse-event reasoning
- Contraindication and interaction reasoning
- Clinical uncertainty, escalation and safety boundaries
- Translating biological mechanism into explicitly bounded clinical claims
Medical must not answer
- Real-time patient care
- Personalised prescribing or dosing
- Emergency triage delivered as an operational service
- Unsupported jurisdictional, legal, insurance or regulatory advice
- Claims outside the approved population and scenario scope
- Public-health policy ownership
- Autonomous execution or memory write-back
Boundary cases
The questions that decide whether a boundary is real. Each names a primary owner and the consultation it is required to request.
| Question | Primary owner | Required consultation |
|---|---|---|
| How does a receptor pathway work? | Biology | Medical only if a clinical interpretation is requested |
| What mechanism could explain a synthetic symptom pattern? | Medical | Biology for uncertain mechanistic premises |
| How does a pathogen enter a host cell? | Biology | Medical for diagnosis or treatment implications |
| How should a synthetic case be diagnosed? | Medical | Biology for disputed mechanism or organism identity |
| How does a drug bind its target? | Biology | Medical for efficacy, interaction or treatment claims |
| Should a patient receive the drug? | Medical | Must abstain from real-patient action; Biology cannot answer |
| What does a genetic variant do molecularly? | Biology | Medical for clinical significance; counselling remains excluded |
| Who owns an unresolved BioMed conflict? | Neither alone | Structured evidence requested from both; unresolved review is routed |
Any real-patient framing is out of scope for Genesis, even when the underlying educational question resembles a benchmark item. If no registered domain owns a concern, the system abstains.
Has a domain formed?
Seven conditions. Passing an exam is not one of them.
A candidate domain counts as formed only if it demonstrates all seven. Loss curves and domain exams alone are insufficient — that is precisely the confusion this programme exists to avoid.
- Core generalisation — performance on unseen in-domain structures, not memorised formulations.
- Counterexample handling — revision or rejection when a relevant counterexample invalidates a claim.
- Calibration — confidence tracks observed correctness under the frozen procedure.
- Boundary awareness — excluded and ambiguous concerns are identified, then escalated or declined.
- Evidence discipline — claims can be related to supporting evidence and explicit assumptions.
- Update isolation — a bounded update improves its target without unacceptable regressions elsewhere.
- Protocol competence — the domain can request missing premises and consume a structured response without erasing its scope.
Proposed formation gate
Thresholds remain draft until the pre-registration is formally frozen.
- In-domain performance
- Non-inferior to the approved baseline.
- Out-of-domain AUROC
- Proposed at
0.85or better. - Hard-stop set
- No safety-critical excluded-category answer in the frozen hard-stop set.
- Provenance completeness
- Proposed at
95%for eligible claims. - Calibration
- Acceptable under the pre-registered threshold.
- Update isolation
- No blocking regression.
On splitting further
After Genesis, Biology or Medical should be split again only if a narrower boundary reduces update, validation or safety cost by more than the added routing error, communication overhead, cross-boundary information loss and error-propagation risk. No boundary is accepted because it matches an academic department name.