Foundation 0. No training run has been authorised. Nothing on this site reports an experimental result.

Plan and roadmap

A sequence of gates, not a schedule to be executed.

The first cycle takes an accepted raw snapshot through governed cleaning to a bounded 200M-parameter pilot. Each milestone has a hard spend ceiling and an exit gate. No milestone authorises the next one, and unused budget does not carry forward.

Cycle objective

Three questions the first cycle answers.

Can the source become a dataset?

Can the accepted source be converted into a reproducible, reviewable native-text dataset without depending on a NAS or on any artefact from the prior project?

Can a tokenizer represent it?

Does a newly trained tokenizer represent the approved Chinese/English educational material with acceptable fidelity and efficiency across curriculum stages?

Can a bounded pilot learn from it?

Can a 200M-parameter pilot train on the frozen dataset with valid checkpoints, complete evaluation evidence and measured cost?

The cycle ends with a single go / hold / stop decision about a separately budgeted 500M foundation pilot. It does not include Cortex Genesis or any domain formation.

Fixed scope

What is bound before anything runs.

Source snapshot
snapshot:education-cn-raw-20260730496,175 exact S3 object versions, 2,510,613,958,218 accepted bytes.
Immutability
Raw stays immutable in S3 and is never written to. Derived artefacts reference exact parent digests.
Clean lineage
Prior code, tokenizers, checkpoints, dataset derivatives and benchmark claims are excluded as inputs. This means starting again from governed original material, not training directly on raw files.
Modality
Native text and structure only. Audio and video are never decoded, sampled, transcribed or embedded; they keep a terminal excluded disposition in reconciliation evidence.
Deferred
OCR and image understanding are outside the first-cycle budget. Scanned-only pages remain explicit deferred or unsupported records rather than silently empty text.
Project mode
foundation throughout, with every operational feature flag false.
Budget envelope
A$4,000 including GST for first-cycle AWS infrastructure. Human legal, privacy, cultural and quality-review labour is not included.

Milestones

Nine gates from snapshot to decision.

Each ceiling below is a hard planning limit on that milestone alone. Elapsed time is dominated by human review and AWS quota, not GPU runtime — a realistic first-cycle range is four to seven weeks, excluding unresolved rights and privacy decisions.

M0 — Scope and execution-contract freeze

A$0

Freeze the audio/video exclusion rule, the named data owner and two independent approvers, immutable storage prefixes, the reviewed container image digest, per-job spend and wall-clock limits, and the deterministic Batch concurrency, retry and capacity limits.

Exit: every identity, prefix, digest, limit, reviewer and stop procedure is exact. Readiness still reports execution disabled, but no planning ambiguity remains.

M1 — 512-record format canary

A$100

Run the already-selected canary — 512 records covering all 212 populated extension/size strata — using version-specific S3 reads. Verify signatures, active-content risk, encryption, corruption, unsupported formats and extractor routing.

Exit: all 512 inputs reconcile to a terminal disposition; raw has zero writes; every result binds its exact source version and pipeline digest; no audio/video body is decoded.

M2 — Bounded native-text extraction pilot

A$250

Deterministically select a representative pilot across DOCX, PPTX, native-text PDF, plain text, XML/HTML/RTF/CSV and selected legacy Office strata. Exercise malware scanning, bounded container handling, extraction, normalization, lineage, resume and reconciliation.

Exit: one terminal disposition per input; repeated deterministic stages produce matching digests; resource ceilings are measured; failure and quarantine samples get human review.

M3 — Full eligible-source cleaning

A$1,200

Process every eligible native-text object through malware screening, bounded extraction, deterministic normalization, initial language/subject/level labelling and immutable derived shards.

Exit: the full accepted inventory reconciles exactly; every object has an explicit terminal result; nothing is overwritten; output-volume and curriculum distributions are published.

M4 — Deduplication, risk review and dataset freeze

A$600

Exact, normalized-text and bounded near-duplicate clustering; development and evaluation families frozen; overlap screened; rights, privacy, cultural and quality review completed. Deduplication weights or selects representatives and never deletes lineage.

Exit: train/dev/eval manifest digests are disjoint; unresolved contamination is zero; restricted or unknown-rights items are absent from approved partitions; exact record, byte, unit and token counts are frozen.

M5 — Tokenizer bake-off and approved shards

A$250

Train and compare a 48K SentencePiece unigram tokenizer against a 48K byte-level BPE, using only the approved training partition. Measure Chinese and English fertility by education stage, mathematical and scientific symbol round-trip fidelity, byte fallback and token-count stability.

Exit: one candidate selected from frozen evidence; trainer, config, seed, input manifest and artefact digests exact; development and evaluation content never inspected; shard order and token counts reproduce.

M6 — 50M synthetic smoke

A$100 · ≤4 h

A randomly initialised model on deterministic synthetic data only, structurally prohibited from referencing the research dataset freeze. This validates packaging, BF16 execution, logging, heartbeats, finite-loss stops, checkpoint publication, resume and teardown.

Exit: no research data was readable by the job; checkpoint/resume preserves run lineage; emergency stop and zero-capacity teardown are demonstrated; no training-quality claim is made. GPU launch requires a new explicit instruction.

M7 — 200M approved-data pilot

A$600 · ≤24 h

Train the smallest selected dense model, not exceeding 200M parameters, on a frozen approved curriculum slice using one H100. Record throughput, utilisation, memory, loss, held-out progression, restart behaviour and exact cost.

Exit: all inputs, code, environment, image, tokenizer, seed and outputs immutable; loss finite; curriculum-retention evidence complete; no leakage or unapproved input; all billable compute returns to zero capacity.

M8 — Pilot review and scale decision

A$300

Reconcile budget, artefacts, quality, failures, utilisation and effective tokens per dollar against the pre-registered promotion criteria.

Exit: publish exactly one decision — go (prepare a separately budgeted 500M proposal), hold (name the defect to correct before replay), or stop (preserve artefacts and close the cycle). M8 cannot launch 500M training.

Plus A$600 of project contingency, storage and bounded retries. Total ceiling A$4,000. The proposed GPU target is a single p5.4xlarge with one H100 80 GB; the earlier eight-GPU drafts are disabled and are not eligible for this cycle unrevised.

Scaling ladder

Model size is selected from measurement, not assumed.

The first authorised family is a conventional dense, decoder-only Transformer. Mixture of Experts, domain routing, adapters, memory write-back and multi-domain communication are deliberately excluded from the foundation baseline — if the foundation itself contained hidden routing, later comparisons against independently formed domains would be uninterpretable.

Architecture
Dense decoder-only Transformer with RMSNorm, SwiGLU, RoPE, grouped-query attention and tied embeddings.
Size envelopes
Synthetic smoke ≤ 50M; data pilot ≤ 200M; foundation pilot ≤ 500M. These are selection ceilings, not claimed parameter counts — the smallest candidate meeting the frozen quality and systems gates must be selected, and a larger one requires measured evidence.
Context
2K for synthetic packaging tests, then a provisional 4K for data and foundation pilots.
Tokenizer
Trained only from the approved training partition. Must preserve Chinese without depending on whitespace, retain mathematics, scientific notation and code fragments, define deterministic Unicode normalization, and never inspect the frozen evaluation partition.
Still open
Vocabulary size, parameter count, layer/width/head dimensions, positional encoding details, precision and optimiser, curriculum mixture and stage schedule, and the token, FLOP, time and dollar budgets — all pending extraction and token profiling.

The ladder runs pipeline smoke → data pilot → foundation pilot → scale decision → EDR Foundation candidate. No step automatically authorises the next; each has its own immutable run manifest, budget, stop conditions and two approvers.

Readiness gates

What must pass before an experiment may even be proposed.

Draft documents existing is not evidence that a gate has passed. The machine-readable readiness report returns not_ready, which is the correct state during preparation; it becomes ready only from referenced authorised evidence, never from the existence of files.

G0 — Contract freeze

  • Versioned domain manifest exists
  • Protocol claims carry evidence, assumptions, scope, confidence, uncertainty, contradiction and temporal validity
  • Memory has no raw-to-consolidated shortcut
  • Forge plans are declarative and require dataset approval
  • Foundation mode prevents operational features
  • Protocol and manifest 0.1.0 reviewed by all four owner roles

G1 — Scientific pre-registration

  • Falsifiable primary hypothesis defined
  • Parameter and compute budgets frozen for both primary arms
  • Domain, cross-domain, unknown-boundary and traceability tasks frozen
  • Negative controls and routing-error probes defined
  • Success, partial success, null and stop conditions defined before training

G2 — Domain boundary

  • Biology ownership, exclusions and escalation approved
  • Medical ownership, jurisdiction, population and high-stakes prohibitions approved
  • Evidence that a domain has formed is defined
  • Out-of-domain and abstention tests defined
  • Conflict resolution for shared questions defined

G3 — Data and checkpoint governance

  • Exact EDR checkpoint selected with provenance and licence
  • Every dataset slice inventoried with an immutable digest
  • Licence, privacy, consent, cultural and contamination review complete
  • Train/evaluation separation and leakage checks approved

Transfer, reconciliation, one-way checksum and the deterministic inventory are accepted, and a canary is selected — but version binding, cleaning, rights review and partition freeze remain incomplete. This does not pass G3.

G4 — Operational containment

  • Isolated artefact namespace and budget
  • Reproducible environment and seed policy
  • Compute scheduling and emergency stop
  • No production routing or external actions
  • Signed run manifest and append-only audit sink

The AWS boundary, single-node-first ladder, checkpoint policy, budget envelope and emergency-stop sequence exist in draft. No compute has been provisioned. This does not pass G4.

G5 — Authorisation

  • Formation plan frozen
  • At least two named approvers
  • Every dataset marked approved for training
  • Training flag enabled only in an experiment-specific configuration
  • Explicit user authorisation recorded in the exact phase-gate revision

Explicitly absent today

No model or tokenizer. No dataset or download script. No optimiser or training loop. No inference endpoint. No benchmark scores. No experiment-tracking integration. No scheduled job. No active domain. No automatic memory consolidation. The diagnostic tool passing confirms this configuration is internally safe; it does not mean any gate has passed.

After this cycle

Deliberately outside the current envelope.

No downstream milestone inherits authorisation, unused budget or approval from this roadmap.

  • EDR 500M foundation pilot — only after an M8 go decision.
  • EDR Foundation candidate — scale selected from measured token volume, quality and cost rather than assumed in advance.
  • Cortex Genesis dataset formation — freeze Biology and Medical boundaries and matched data.
  • Cortex Genesis experiment — the comparison described on Experiment design.
  • Octoryn Cognitive Network pilot — only if Genesis evidence supports domain composition.