NRT-ATOM-SCHEMA-01 — Criterion Atom Schema (v1.2 — RATIFIED)¶
Status: v1.2 RATIFIED by Tim, 2026-08-06; v1.2.1 erratum nodded by Tim same day (R2 input correction per census Q4.1 6b6efe86… — no other change). This document is the schema of record for the criterion atom layer; on repo filing it enters the ADR log per CANONICAL-DECISION DISCIPLINE. The remaining census items (Q4.1 integrity, Q6 provenance, Q5 KN outliers) do not touch it. v1.2 folds in: the measured section-code inputs (86 titles, generation by co-occurrence), the disposition taxonomy, the four-way structure model, parse-by-shape, source keying on item_id, and the LICENSING generation correction.
Author: Claude (Fable seat), 2026-08-06.
Provenance: [C] = measured read-only against rto-nrt-db by Claude this session. [A] = Alex's census figures, gate-reviewed on artefact bytes (digests verified), not independently re-derived. The BSBCMM211 demo parse (70 atoms, conservation-checked both polarities) is the reference artefact for the target shape.
1. The question this answers¶
What is the permanent, citable address of a performance criterion — such that it survives supersession and renumbering, still answers an auditor who cites the published label, degrades gracefully for items with no published number, and does not invalidate the immutable KN enrichment already keyed to it? [2026-09-15, NAME-RETIRE-01: the prototype's KN corpus is retired; the address stands on register fact alone]
2. Identity design¶
2.1 Atoms are release-scoped. The atom is the item as published in release N of unit X. It is immutable: it never changes when a later release renumbers, rewords, splits or merges. Substrate confirms release-scoping is native: tga_nrt_content_items keys content by code and release — 77,713 code-release pairs across 61,486 distinct codes [C].
2.2 Address grammar.
{unit_code}:R{release}:{doc}:{section}:{slot_type}.{ordinal}
e.g. BSBCMM211:R1:UOC:EPC:performance_criterion.3
- doc: UOC (unit of competency) | AR (assessment requirements). Derived from the source bundle type: register Default → UOC; Assesment Requirements (the register's own spelling) → AR. Gen-1 releases publish a single document and are UOC throughout.
- section: canonical section code (§3)
- ordinal: 1-based position within (section, slot_type), assigned in document order
Logical sections (census Q4.2 amendment [A]): the citation address above is human-facing and unchanged, but (code, release, title) does not identify a single source row — 30,563 collisions, of which 30,297 are benign bundle-type projection and 266 are residue on 26 code-release pairs [A]. Therefore: all source rows belonging to one logical section are concatenated in (sequence, item_id) order before parsing, and ordinals run across the concatenation. A logical section never maps ambiguously; the concatenation order is deterministic down to the register's own primary key.
2.3 Published label. Kept verbatim as an opaque string (published_label, nullable). It is display and audit vocabulary, never the key, never sorted numerically. Items published without a number (PE/KE/AC items, FS descriptors) carry published_label = NULL and are cited by position — which is the sector's existing convention ("knowledge evidence, third dot point").
2.4 Content hash. Every atom carries text_hash (sha256 of normalised text). Used for cross-release identity candidates and dedup QA (KN binding validation, §5, is retired [2026-09-15, NAME-RETIRE-01]). Never the citable address.
2.5 Continuity is NOT identity. Cross-release links (identical / reworded / split-into / merged-from / no-successor) are a separate assertion layer with provenance, designed in its own brief. Nothing in this schema depends on it; the atoms table never needs re-parsing when the continuity theory changes.
2.6 Parent links. PCs carry the address of their element; FS descriptors carry the address of their skill.
2.7 Provenance classes — the atoms table holds register-published bytes only. Two classes exist in this system: register fact (content as published on the National Register) and computed derivative (anything generated — KN/Knowledge Navigator output, regulatory_context, classifications, summaries, embeddings). The atoms table admits only the first class. Computed content never enters it — not as rows, not as columns — regardless of quality or model. Derivatives live outside the fact layer and link in by atom address. This is the content-level analogue of the HARD SEPARATION RULE, and it is what prevents computed output masquerading as fact.
2.8 Source keying and duplicate handling (census Q4.2 amendment [A]). The parse is keyed on the register's item_id — unique across all 784,392 content rows [A], the register's own opaque identifier; the citation address is derived, never the key. Every atom carries the item_id, bundle_id, and sequence of its source row as provenance. Byte-identical same-address source rows (256 [A]) collapse by content_sha256 with all contributing item_ids recorded, and are filed as a mirror defect. Differing-content collisions (10, enumerated in the Q4.2 artefact [A]) are parsed in full, both versions, under a contested-source flag — nothing discarded; adjudication belongs to a later single-call register re-query brief.
3. Section model¶
3.1 Sections arrive pre-cut. The register delivers each unit as titled section rows in tga_nrt_content_items (784,392 rows [C]). The parser never discovers section boundaries; it parses within sections.
3.2 Dual naming. Every atom carries the canonical section code and the verbatim TGA section title of its source row. Title variants exist ("Unit sector" vs "Unit sector(s)") — canonicalisation is a lookup, verbatim is the record.
3.3 The section-code table — measured, complete (v1.2 [A]). Source artefacts: the Q2b unit-only inventory (3434406e… — counts authoritative, generation column VOID per 0111221d…) and the co-occurrence generation assignment (d00a50da…). The unit-linked population (680,071 rows, 49,894 codes) carries exactly 86 distinct titles: 59 gen-1 · 13 gen-2 · 3 BOTH · 11 NEITHER — the NEITHER all explained (1 title on withheld unit-releases that publish nothing else; 10 foreign training-package front matter on two deleted codes).
- Canonical code rule: deterministic slug from the verbatim title, with an explicit override table for the load-bearing sections (EPC, PRECONTENT, APP, APPUNIT, SECTOR, FS, PE, KE, AC, LICENSING, MODHIST, MAP, LINKS, DESCRIPTOR, RSK, RANGE, ROC, EVGUIDE, EMPSKILLS, PREREQ, COREQ, CONFIDENTIAL, …). The full 86-row table is generated from the two source artefacts at parser-build time; the verbatim TGA title is stored on every atom regardless (§3.2).
- Generation is recorded per atom from its unit-release context, never forced per title. The three BOTH titles take their generation from the release they sit in; NEITHER is recorded as such.
- Near-identical names are never merged without structural agreement:
Range statement(gen-1, 2.6% prose) andRange of conditions(gen-2, 48.6% prose) are distinct codes [A]. - Measured warnings inherited from the census: the naming convention runs backwards from intuition (
Unit sector(s)is gen-1,Unit sectoris gen-2);pre-contentis entirely gen-1 and is descriptor-furniture except for 3 enumerated EPC-shaped exceptions parsed by shape.
3.4 LICENSING is a first-class slot — and the standalone section is a gen-1 convention (corrected v1.2 [A]). The register carries "Licensing/regulatory information" as its own section for 28,272 unit-linked rows (38,621 corpus-wide [C]), and that standalone section co-occurs with gen-1 markers 27,909 : 361 [A]. In gen-2 the licensing statement is prose inside the Application section (measured across the enriched cohort this session [C]). The slot therefore exists wherever the register publishes the section; the derived licensing classification (below) ranges over both locations — standalone LICENSING atoms and APP atoms carrying the licensing sentence. Its atoms carry the verbatim text only, per §2.7. The classification (none-standard | none-occupational | applies-or-varies | regulatory-other | absent) is itself a computed derivative: it lives in the derived layer keyed by atom address, never as a column on the atom. Measured distribution over the 15,128 enriched-cohort Application texts: 7,505 / 1,787 / 4,493 / 1,246 / 97 [C]. applies-or-varies is the attachment point for the outside-TGA regulator layer. units.regulatory_context (54,488 rows [C]) is suspected computed content of the same class as KN — quarantined from this design pending the census provenance audit (CENSUS-02 Q6); it is not an input here.
3.5 Nothing is discarded. An unrecognised section title or heading creates a slot and a flag, never a drop. Stems, column descriptors, modification-history rows are all atoms.
3.6 Disposition taxonomy (v1.2 [A]). Every atom carries a disposition, measured into existence by the census over 781,792 VET-linked content rows:
- content — parseable substance (83.68%).
- asserted-empty — a sentinel stating the section is empty:
Not Applicable,Not Supplied,Not Available,nil,none,n/a, … (15.35%). - withheld —
Not For Public Access./Confidential content: the register asserts content exists and is not served. Never conflated with empty — an atom recorded empty where the register said withheld is a false statement about the register. Measured: withheld is a strict subset ofrestricted_access = 1(36/36 flagged, 0 exceptions [A]). - redirect — the section's content is a pointer to a sibling section (
Refer to …, 0.72%). Carries a resolved link to the sibling atom address (5,613 of 5,613 covered forms resolve, zero dangling [A]) or unresolved-flagged (the class exists: the register publishesrefer to unit descriptpr.[A]). Never an error, never a drop. - empty-markup — markup with no text after tag-strip (0.24%).
Sentinel and furniture predicates are allow-lists that fail open into flagged atoms. Anything unmatched is kept and flagged for human review, never guessed at and never dropped — the guard names the known set; the unknown lands in the visible lane. The taxonomy is measured not to extend to accredited components (2,554 rows: one title, zero withheld, zero redirect [A]).
3.7 Structure model (v1.2 [A]). Four structure classes, measured per section: table, table+list (a list nested in a table cell), list, prose. The parser implements a path for each, chosen by shape, per §3.8:
- EPC has zero pure-list rows (table 50.3% · table+list 49.3% · prose 0.4% of 58,539 [A]). No list-only EPC branch exists; the in-cell question is whether criteria are
<li>or<p>. The 255 prose rows are measured empty-or-sentinel — units that publish no criteria — plus 8 withheld. - Gen-1 sections are dominated by table+list nesting (Evidence guide 87.6%, RSK 84.0%, Range statement 82.7% [A]). Nesting is the norm, not an edge case.
- Foundation Skills requires a prose path: 50.6% of FS rows are prose in the register itself [A].
- Numbering, measured over 380,635 EPC label cells [A]: decimal 81.88% · integer 17.29% (element numbers — a parser that assumes every label cell is a criterion label mis-assigns one in six) · alpha 0.82% (3,136 cells — the population that makes
published_labelan opaque string) · roman 0.
3.8 Parse by shape, never by title. Two independent evidence classes force this: 3 full EPC tables published under the pre-content title, and modification-history prose published in the same section [A]. Section title selects the expected shape; the observed shape decides the parse path, and a mismatch parses by shape and flags.
4. What the atoms table is NOT¶
- No currency columns. Current/superseded/deleted is resolved at read time by joining to
tga_training_componentsvia the canonical fragments (PITH-CURRENCY-EXPOSURE-01). A status column on atoms would be a writerless cache — the defect that brief exists to remove, recreated at 1M rows. - Not fed from
unitscontent columns. Those columns are the earlier flattened extraction: structure destroyed where it existed (the FS table survives as an unpaired word list — though for the 50.6% of FS rows that are prose in the register, the structure was never there to destroy [A]), 15,128 units only, single release only [C]. They are display-grade. Engine-grade parsing readstga_nrt_content_items.contentonly (v1.2.1 [A: Q4.1]):r2_keyis NULL table-wide — the R2 carriage never completed — and must not be read as a fallback. The 1,239 no-content rows have no recovery path in the mirror: 1,193 are accredited rows the register publishes empty; the 46 VET rows are a named, enumerable set for a later register-fetch brief.content_sha256is the integrity instrument (proven both polarities, 25/25 sampled hashes verify [A]) and is re-verified per unit-release at parse time.
5. KN linkage — RETIRED 2026-09-15 (was: presentation-layer only)¶
The prototype's KN corpus is retired (NAME-RETIRE-01, 2026-09-15; Tim's ruling of 2026-09-09; ADR-079): no surface reads kn_* or regulatory_context, and the Knowledge Navigator name now means the AI/RAG companion. [2026-09-15, NAME-RETIRE-01] This section is kept as the record of the linkage the prototype used.
Epistemic status: KN (Knowledge Navigator) is a computed derivative — commentary generated by a low-capability model against unparsed content. It is crib notes for humans, not fact, and under §2.7 it has no role in the fact layer, the parse design, or any machine-processing path. The Time Machine's §5.2 framed KN as a potential constraint on the parse; with the provenance now stated, the correct posture is the reverse: the parse is designed entirely from register fact, and KN links in afterwards for display, or doesn't.
Measured facts (they make the display linkage easy, nothing more): kn_criteria_reasoning is structured JSON referencing criteria by published number AND verbatim text — criterion_id and criterion_text present in 15,102 of 15,127 populated rows [C]; the 25 exceptions are enumerated in the census.
Linkage rule (display only): a KN entry attaches to an atom when criterion_id = published_label and normalised criterion_text matches the atom text. Either-only matches are flagged. An unbound KN entry degrades a UI panel, never a pipeline. Mismatch flags are evidence about KN quality, not about the parse. The corpus is retired and no longer governed [2026-09-15, NAME-RETIRE-01]. This schema is safe either way, because §2.7 keeps it out of the fact layer regardless of where it sits.
Caution: KN's element_number is an integer; atom labels are opaque strings. Any join goes string-side, never integer-side.
Fence status, measured on the working tree 2026-08-06 [C]: workspace units/[code] route excludes KN by construction; internal-api strips kn_* from at least two response paths; ADAPTER-01 dry-run excludes regulatory_context from its predicate with a filed deviation. KN readers are presentation surfaces (workspace/site UnitDrawer, admin compute page) plus the generation scripts. The census (Q6) completes this inventory; anything found reading computed columns in a processing path is a defect to file.
6. Verification requirements (build gates for the parser brief)¶
- Conservation per unit-release: normalised parsed-atom text reconciles against normalised source text, both directions, zero loss; parser-added separators enumerated. Disagreements are a list, not a rate.
- Instruments proven on both polarities before trusted (pass on known-good, fail on known-bad).
- Every count names its population. A join count is not a row count [A: measured 2-row join inflation inside a census figure]. Text predicates are case-insensitive or justify why not [A: three casing-variant rows escaped a case-sensitive sweep].
- Structural invariants checked independently of the parser: every PC has a parent element; ordinals gapless per (section, slot_type); labels unique per release where present.
- No derived column in a measured table without marking it as different in kind [A: the Q2b generation column, superseded].
7. Open items (v1.2: NOTHING BLOCKS FREEZE)¶
Every v1.0/v1.1 blocking item is delivered and gated: the section-title inventory and code table (§3.3), the coverage gaps classified (25,373 non-current units without content; zero current gaps except the enumerated 80 [A]), the structure and numbering sweeps (§3.7), the duplicate diagnosis and source keying (§2.8), the disposition taxonomy (§3.6).
Remaining census items, none schema-blocking: Q4.1 (no-content rows + R2 verification — parser input integrity, gates the parser run, not the schema), Q6 (computed-column provenance — gates ADR-KN-RELOCATION-01), Q5 (KN outliers — presentation-layer). Post-ratification work in its own briefs: the parser build (with §6 as its gates), the continuity/equivalence edge layer (§2.5), the 80-unit drain backfill, the 46 VET no-content rows register fetch, the register re-query adjudication of the 10 contested-source collisions, the 256-duplicate mirror-defect trace, and the one-row register typo (descriptpr) recorded as a register erratum.