TGA-DATA-MODEL-02 — TGA's own structure, TGA's own names, and which database each table lives in¶
Canonical. Filed from outputs/TGA-DATA-MODEL-02-2026-09-02.md (sha256 feaf3fe540f56327a80688092e8567c28d5ab64e08a27eeb80d7fe055a2de852 / 31883 bytes) on 2026-09-03. Tapped by Tim 2026-09-02.
Status: TAPPED by Tim, 2026-09-02 — law for every brief on the TGA data from this date. Open item: §12.3 (RTO key) held pending the ninety-code query in BRIEF-TGA-MODEL-01.
Revised to v3 2026-09-04 under BRIEF-TGA-CANON-01 463680ba8666186f (ADR-082 six databases · ADR-083 the storage convention · ADR-085 the ledger as the measure of completeness). Every figure in this revision carries its as-at and the artefact it was read from.
Corrected 2026-09-11 under BRIEF-CANON-CORRECTIONS-02 (6369158c: §1 the reason the estate builds locally; §3 the 0116 row's counts and header claim; §4 od_contacts searchable; §6 release_current current in the mirror's sense only), BRIEF-CANON-CORRECTIONS-03 65d96495ac55d348 (§2.2's Parent training package and Classifications rows, which said two landed families were not stored) and BRIEF-CANON-CORRECTIONS-04 53471f32ef494ac7 (§2.1's usage-recommendation row, which said a populated column did not exist; §7's tga-nrt size, restated at the register of record; §9 item 2, closed as landed; §12's release-count finding, re-measured in a dated bracket).
Corrected 2026-09-15 under BRIEF-CANON-HYGIENE-01 7b5afbd0b8dd273f (Gate 1 verdict 64da23ee310544f4, R-HYG-4): §1's tga-rto row — scope sits in tga-rto today, measured on the register rto store; §1's tga-scope row — scope_entry's figure re-measured.
Corrected 2026-09-16 under BRIEF-COUNT-RECONCILE-01 c950f18b0ebbad92 (Gate 1 verdict 11b3fefb6eb7e00f, Gate 2 verdict f5c040415da0d692): §1's count sentence now reads eight — seven built and one proposed; tga-nrt-content-legacy (ADR-087 §A) and tga-cvig (Proposed, ADR-091) gain §1 rows; the tga-history row is dated to its LANE-Z-03C build and names the unruled rebuild question; §8's signed line names the estate's databases rather than a stale six. The 2026-09-04 revision line above is historical and stays.
Drafted: 2026-09-02, advisory seat. Supersedes TGA-DATA-MODEL-01 whole
(its glossary is withdrawn; its database list was missing history and the RTO side).
Revised 2026-09-02 (Tim): "pots" → databases; the per-edition database sets are withdrawn in favour of the diff model (§1.1).
Rule of this document: every table and column is named as training.gov.au names it — the label on
the page, or the field in the API response where the page has no label. Nothing is renamed. The only
words that are ours are edition (the fetch-date stamp on a row — §1.1) and the _norm join
columns (a lower-cased copy of a code, beside the published one).
Control set: Tim's screen caps of training.gov.au (21 Aug 2026), UEEAS0009_Complete_R2.pdf,
the BSB40920 API field inventory (docs/docs/ops/tga-api-field-inventory.md), and the local
Edition-2 stores queried read-only today.
0 · The seam, as TGA draws it¶
The site's search box has one toggle: NRT | RTO. That is the top-level split and it is the whole architecture.
NRT — training products RTO — training organisations
training package registered training organisation
└ qualification ── units of competency ├ summary · contacts · addresses
└ skill set ├ registration · registration manager
└ unit of competency ├ regulatory decisions
accredited course ── accredited units └ SCOPE ──────────────┐
│
┌────────────────────────────────────────────┘
▼
scope = the join, and it runs both ways
qualification page → "Find RTOs" (which RTOs deliver this)
RTO page → Qualifications / Skill sets / (what this RTO delivers)
Units / Courses tabs
Every page carries Show history. History is not a separate thing bolted on; it is the same tables with every dated row kept — releases with dates, legal names with start and end, registration periods, mapping edges with dates. "Current" is a filter on history, never a different store.
1 · The databases¶
Eight D1 databases, one R2 bucket — eight built (RULING-SIX-DATABASES-01, Tim 2026-09-04, on six; tga-nrt-content-legacy the seventh per ADR-087 §A; tga-cvig the eighth per ADR-091, built and certified under CVIG-STORE-01 2026-09-16). The store cells below cite
docs/docs/ops/store-register.md, the instruments' read list: the register is what ledger-02-gate1-v7.py
and ledger-02-emit-v7.py actually open, and this prose and that table must agree — the instruments STOP
if they do not (BRIEF-LEDGER-V7, R-V7-1). One set, not one per snapshot. Named after TGA's own split. The count moved from four on two measurements: tga-rto came in at 6.708 GB against an estimate of 1–1.5 GB, of which scope_entry and its six indexes are 96.3% (Lane Z Gate 3, 2026-09-03); and tga-nrt-content at all 36 section types projects to 6.720 GB against an estimate of 1.5–2 GB (CANON-01 Gate 1 §3, 2026-09-03).
| database | holds | TGA page(s) it serves | today, on the Mac |
|---|---|---|---|
tga-nrt |
every training product: code, title, type, usage recommendation, releases, mapping, parent training package, classifications, companion-volume listings. What exists, and its history. | Training package · Qualification · Unit · Skill set · Accredited course — all tabs except the section bodies | TODAY — ~/mirror/tga-nrt-05/20260914T052706Z/tga-nrt.sqlite 767f89ac3afb6ffc · 932,077,568 B (CHILD-APPEND-01, 2026-09-14: the 263's level-2 children and the 175 releases' level-2 bodies, held by CHILD-FETCH-01 under the children ledger 40ad6e4a3672a3d2, landed at ed6-2026-09-14-childappend01 — component_classification +482 · taxonomy_industry_sector +4 · taxonomy_occupation +41 · component_recognition_manager +261 · release_training_package_usage +175 · _release +175 · release_asset +134 · release_external_link +33 · release_link +35 · release_content_bundle +64 · release +175 at ed6, the 175 ed5 rows stamped by join (§6) · qualification_specialisation +9, body-driven (§13a) · edition +1. R-CA1-7 applied first: taxonomy_industry_sector.description nullable (§13a), and the DDL of record outputs/ingest-append-01/tga-nrt-v3.sql carries it (now 213a21f7629eb8e4). Built twice byte-identical into a copy proven digest-equal to its predecessor before any write; every other table byte-identical by manifest; verified by a second instrument that never imported a loader, 80 of 80 checks (outputs/child-append-01/CHILD-APPEND-01-GATE-4-20260914T072602Z.md afd36fcb96a68980). Cite 767f89ac3afb6ffc from here on.) · predecessor RETAINED, untouched: ~/mirror/tga-nrt-05/20260914T031157Z/tga-nrt.sqlite bb3167d140d369c7 · 928,993,280 B (register row nrt-p4; NRT-APPEND-01, 2026-09-14: the 263 components the ledger lacked — the 158 slash codes set aside on 2026-08-17 (C-0914-05) and 105 additions — landed from their held L1 bodies at ed5-2026-09-14-nrtappend01: component +263 · component_currency_period +201 · release +175 · supersession_edge +42 · component_parent +175; no level-2 body is held for the 175 releases (specialisations_state no-body), so their children are owed by a fetch (CHILD-FETCH-01). Built twice byte-identical into a copy proven digest-equal to its predecessor before any write; every pre-existing table byte-identical to the predecessor by manifest; verified by a second instrument that never imported a loader, 87 of 87 checks (outputs/nrt-append-01/NRT-APPEND-01-GATE-4-20260914T041918Z.md 05787d2ba59f82e1). §13 F-NA-2. superseded by CHILD-APPEND-01) · predecessor RETAINED, untouched: ~/mirror/tga-nrt-04/20260912T091900Z/tga-nrt.sqlite 8b84a2e98745b474 · 928,423,936 B (register row nrt-p3; INGEST-02, 2026-09-12: three new tables — component_recognition_manager 123,537 · release_training_package_usage 134,112 · release_training_package_usage_release 500,453 — plus the seven ed1 fetch-list releases' children at ed4-2026-09-12-ingest02 (release_asset 17 · release_content_bundle 4 · release_external_link 2 · release_link 3) and 936 build_finding rows. Built into a copy proven digest-equal to its predecessor before any write; every count re-derived by a second instrument that never imported the loader; orphan queries 0/0/0 with a planted control; foreign_key_check 0. superseded by NRT-APPEND-01) · predecessor RETAINED, untouched: ~/mirror/tga-nrt-03/20260911T033219Z/tga-nrt.sqlite 4190424cbe933e5a · 872,943,616 B, under outputs/ingest-append-01/tga-nrt-v3.sql 0254084cafd61922 plus outputs/vocab-01/vocab-01-ddl-v1.sql 7c5fcf86d8a47bfb (VOCAB-BUILD-01, 2026-09-11 — five additive tables: vocab_nrt_classification_scheme 6 · vocab_nrt_classification_value 3,569 · vocab_classification_purpose 2 · application_metadata 1 · component_restriction 2, built into a copy of 66ad976acfda6dc8 proven equal before any write, its 23 pre-existing tables byte-identical by two instruments. Before it, INGEST-APPEND-02, 2026-09-08, 66ad976acfda6dc8 — an append of the replacer's rows at ed3, the DDL unmoved. In both builds the file's byte count did not change, because the new rows filled free pages, so the byte count is not a check here — the digest is) · of record per store-register.md row nrt (predecessors nrt-p4, nrt-p3, nrt-p1, nrt-p2) |
tga-nrt-content |
the content of each release, section by section, as addressable rows: elements, performance criteria, evidence, prerequisites, packaging rules, and the rest of the 36 section types. Plus the section index that says which section of which bundle each row came from. | Qualification details · Unit details · Skill set details · Course details — the body of the page and the downloadable PDF | ~/mirror/tga-stores-2026-09-04/tga-nrt-content.sqlite 3.806 GB, 16 of 36 types (81af99aa1f184209, Q3-APPLY-01 Gate 5, 2026-09-04). Lane Q's 0116/0126 rows applied insert-only: reference 49,117 → 670,851, section_parse 506,995 → 525,370; reference 0126 = 29,212 by content_type_code (29,128 from q4 + 84 from q4b, both 0126). The 2026-09-03 file (24dd0063bdf83a4f, 3.328 GB) is superseded, retained, deletion deferred to the close of Q5 and LANE-W-02, which assert its digest. Projected 6.720 GB at all 36 types, 4.764 GB under §1.2 — §7 · TODAY — the pair is SPLIT across two directories: main ~/mirror/tga-content-v3/PARSEOPEN01-20260914T090939Z/tga-nrt-content.sqlite b93666be2931a824 · 6,809,022,464 B (store register row content; PARSE-OPEN-01, 2026-09-14 — the four named-open families (lane-w · content-0000 · reference-0116 · skill-set) lifted from their registered pipelines and landed on the 87 open sections and the 5 skill-set partials, into a copy of 2b5fda3894f6b56c proven byte-identical before any write: section_parse +97 · block +214 · reference +1,032 · foundation_skill +48 · range_of_conditions +55 · skillset_unit +27 · section_search +81; vocab_edition +2 (ed4 2026-09-13, ed6 2026-09-14) and the nine *_current views admit every registered edition (F-LG-11 closed, §3); pipeline_dependency re-pointed to register row nrt 767f89ac3afb6ffc. open_no_parse_row 87 → 0; every pre-existing row byte-identical; built twice byte-identical; full integrity_check ok, FK 0; verified by a second instrument, 46 of 46 (outputs/parse-open-01/PARSE-OPEN-01-GATE-4-20260914T095939Z.md 9c490f831fbaafcb). Cite b93666be2931a824 from here on.) · predecessor RETAINED, untouched: ~/mirror/tga-content-v3/CHILDAPPEND01-20260914T063655Z/tga-nrt-content.sqlite 2b5fda3894f6b56c · 6,806,659,072 B (register row content-p6; CHILD-APPEND-01, 2026-09-14 — the 64 bundles held for 35 of the 175 releases appended at ed6-2026-09-14-childappend01 into a copy of 233169dddbff85d2 proven byte-identical before any write: section +377 · section_parse +313 · block +1,925 · criterion +996 · element +169 · reference +22 · modification_history +70 · licensing_determination +35 · section_search +209 for the new addresses by the R-LA-4 rule; parse rows for the six proven families, 64 sections with a section row and no parse row (§3: 87 open store-wide), legacy-type rows 0; every pre-existing row byte-identical; FK 0, integrity_check ok; built twice, identical on every new row but licensing's clock column; verified by a second instrument, 80 of 80 (outputs/child-append-01/CHILD-APPEND-01-GATE-4-20260914T072602Z.md afd36fcb96a68980). ⚠ The nine *_current views admit none of the rows appended since ed1 (§3, F-LG-11). superseded by PARSE-OPEN-01) · predecessor RETAINED, untouched: ~/mirror/tga-content-v3/CPROV01-20260914T012806Z/tga-nrt-content.sqlite 233169dddbff85d2 · 6,799,839,232 B (register row content-p5; CONTENT-PROVENANCE-01, 2026-09-14 — the pipeline register applied to a copy of 0f68a6988bc9290d proven byte-identical before any write: pipeline_register 158 rows, pipeline_dependency 1 row, parser_of_record gains superseded_by_register with its 51 rows unchanged, and section_search is rebuilt whole on its key of record (bundle_id_norm, content_type_code, section_index, body_sha256) from of-record and co-record block labels — 397,187 → 415,201 rows, FTS = table (R-LA-4); the 27 untouched tables multiset-equal to the predecessor, FK 0, integrity_check ok, built twice byte-identical, verified by a second instrument that never imported the builder (Gate 4 695b10891bc9cdc1, 18/18) — §3. DDL of record: v3.5 and asset-text-01-ddl-v1.sql below plus the two tables and one column created by outputs/content-provenance-01/content-provenance-01-gate3-successor-v3.py 285020ef4db08245. superseded by CHILD-APPEND-01) · predecessor RETAINED, untouched: ~/mirror/tga-content-v3/CAPPEND01-20260913T042812Z/tga-nrt-content.sqlite 0f68a6988bc9290d · 6,794,579,968 B (register row content-p4; CONTENT-APPEND-01, 2026-09-13 — the 22 held bundles appended: section +118 · section_parse +95 · block +418 · element +24 · criterion +152 · reference +6 · modification_history +79 · licensing_determination +15 · skillset_unit +0; parse rows only where the pipeline is proven whole — §3. Pre-existing rows byte-identical to the predecessor on all nine tables, FK 0, integrity_check ok; Gate 4 fa534cb9b2f75539) · legacy is ~/mirror/tga-content-v3/LEGAPPEND01-20260914T074832Z/tga-nrt-content-legacy.sqlite 9e92baebbda1b5f7 · 3,483,017,216 B (register row content-legacy; LEGACY-APPEND-01 v2, 2026-09-14 — the 22's six legacy sections landed, §3; outputs/legacy-append-01/LEGACY-APPEND-01-V2-GATE-4-20260914T090149Z.md a8792459025e90c4. Cite 9e92baebbda1b5f7 from here on. Predecessor RETAINED, untouched: ~/mirror/tga-content-v3/APPLYPOP02-20260907T232701Z/tga-nrt-content-legacy.sqlite 253c2bfd7b4769b4, register row content-legacy-p1) — no 0011 section lives there · DDL of record unchanged: outputs/apply-pop-02/tga-nrt-content-v3.5.sql d9cfe8e482b803a1 plus outputs/lane-a/asset-text-01-ddl-v1.sql 47610d818046518a (two additive tables, one index). Before it: ASSETTEXT01's main 2eecdc96e4e37acf (ASSET-TEXT-BUILD-01, 2026-09-11; register row content-p3), superseded-retained · of record per store-register.md row content + content-legacy (predecessors content-p6, content-p5, content-p4, content-p3, content-p1, content-p2; legacy predecessor content-legacy-p1) |
tga-nrt-content-legacy |
the legacy template family (§3), split from tga-nrt-content on content family, never on status (ADR-087 §A), held as legacy_block with the block model as evidence_item. A product addresses its material by address, never by database — no consumer branches on which of the pair holds a section; the static resolver turns an address into a database at read time. |
the same section-body surfaces as tga-nrt-content, resolved by address at read time |
~/mirror/tga-content-v3/LEGAPPEND01-20260914T074832Z/tga-nrt-content-legacy.sqlite 9e92baebbda1b5f7 · 3,483,017,216 B (LEGACY-APPEND-01 v2, 2026-09-14) — 171,383 sections (§3); predecessor RETAINED, untouched: 253c2bfd7b4769b4 (APPLY-POP-02, 2026-09-07) · of record per store-register.md row content-legacy (predecessor content-legacy-p1) |
tga-rto |
every organisation: summary, contacts, addresses, registrations, registration managers, regulatory decisions, restrictions, codes, CRICOS. Scope sits here today — scope_entry (6,526,345 rows, measured 2026-09-15 on 6956114f01106a89), scope_entry_org (21,513), views scope_fact / scope_fact_current; tga-scope holds delivery_notification only; the move to tga-scope is Z2b, not done. [2026-09-15, CANON-HYGIENE-01] |
RTO — every tab except the scope tabs | TODAY — ~/mirror/tga-rto-01/20260915T020659Z/tga-rto.sqlite 52154a2d6efdd7bb · 6,000,332,800 B (FIND-RTOS-01, 2026-09-15: two views over scope_fact — find_rtos, the declared read of §5, and find_rtos_current — and no table changed; every table byte-identical to its predecessor by manifest; built twice byte-identical; Gate 4 outputs/find-rtos-01/FIND-RTOS-01-GATE-4-20260915T034151Z.md 2f0415837da4be2f; no second write — R-FR-2's index withdrawn at ATTRIB-01's close. DDL of record gains outputs/find-rtos-01/find-rtos-01-views-v1.sql c4a1a0f2503525c8. §5.2. Cite 52154a2d6efdd7bb from here on.) · predecessor RETAINED, untouched: ~/mirror/tga-rto-01/20260914T031157Z/tga-rto.sqlite 6956114f01106a89 · 6,000,328,704 B (register row rto-p5; NRT-APPEND-01, 2026-09-14: the 117 organisations C-SWEEP-05 found absent, landed from their held bodies — organisation +117, od_contacts +1,071, od_addresses +275 and every other od_* family at ed5-2026-09-14-nrtappend01, registration_manager re-derived whole — and a new table, scope_entry_org, 21,513 rows, with the views scope_fact / scope_fact_current (§5.1). scope_entry untouched, byte-identical. DDL of record: outputs/lift-audit-01/tga-rto-v2.sql e84041e60ccae17e plus outputs/nrt-append-01/NRT-APPEND-01-GATE-2-20260914T024718Z-scope_entry_org.sql 3b3b461d1663387b. Built twice byte-identical; 84 of 85 pre-existing tables byte-identical by manifest, registration_manager the named re-derived exception; 87 of 87 checks (outputs/nrt-append-01/NRT-APPEND-01-GATE-4-20260914T041918Z.md 05787d2ba59f82e1). §5 F-NA-1. superseded by FIND-RTOS-01) · predecessor RETAINED, untouched: ~/mirror/tga-rto-01/20260913T025348Z/tga-rto.sqlite a75c7118659e125f · 5,978,304,512 B (register row rto-p4; LIFT-AUDIT-01, 2026-09-13: scope_entry 6,526,322 → 6,526,345 — the 24 real rows Lane Z's lift collapsed for organisation 52583 restored from their ledgered ed1 delivery bodies, and the one row no ledgered body supports left out by filter, written whole to build_finding as orphan-source-row. ux_scope_entry_fact is rebuilt on the lift key of record (organisation_id_norm, component_nrt_id_norm, element_index, row_sha256) — R-LA-4; the other five indexes and all 17 views byte-identical. DDL of record outputs/lift-audit-01/tga-rto-v2.sql e84041e60ccae17e. Three makes byte-identical. superseded by NRT-APPEND-01) · predecessor RETAINED, untouched: ~/mirror/tga-rto-01/20260912T091900Z/tga-rto.sqlite 71b334e0ab7b2620 · 6,707,904,512 B (that predecessor: INGEST-02, 2026-09-12: scope_entry 6,526,172 → 6,526,322, the 150 rows CUFDIG402A's truncated ed1 body could not give, at capture anchor 2026-09-07T05:09:12+00:00; mapped by lane-z-02-build-v1.py's own rule, five row_sha256 re-derived by hand at Gate 4. superseded by LIFT-AUDIT-01) · predecessor RETAINED, untouched: lifted from e2-rto-01 into ~/mirror/tga-rto-01/v1/tga-rto.sqlite (e93515bf7ed061d4, 6.708 GB including scope; Lane Z Gate 3, 2026-09-03) · of record per store-register.md row rto (predecessors rto-p5, rto-p4, rto-p3, rto-p1, rto-p2) |
tga-nrt-packaging |
the qualification's published unit list and TGA's derived unit grid: packaging_rule · packaging_block · packaging_group · packaging_column · packaging_unit · packaging_unit_cell, plus unitgrid_fetch · unitgrid_unit · unitgrid_unit_link. |
Qualification details — the packaging-rules table; the grid as cross-check | ~/mirror/tga-nrt-packaging/v1/tga-nrt-packaging.sqlite 0.868 GB (f04ff17c8e2bb640, Q3-APPLY-01 Gate 3, 2026-09-04) — the sixth database, built and gate-certified. ⚠ Lane Q's delta ~/mirror/tga-packaging-01/v2/delta.sqlite (c74061d962e78fac, 1.346 GB, with the four _current correlation indexes in) is the build input, not the database; it stood in this cell as a proxy until the database existed. v1 of the delta is deleted, gone-is-gone · of record per store-register.md row packaging |
tga-scope |
the join, and it runs both ways — scope_entry (6,526,345 rows) [2026-09-15, CANON-HYGIENE-01], the find_rtos read, and delivery notifications, which are RTO × product facts. |
RTO Scope overview / Qualifications / Skill sets / Units / Courses, and the product page's Find RTOs | built 2026-09-11 (BRIEF-DNH-BUILD-01) with one table, delivery_notification — 1,975,641 rows, RTO × product delivery notifications from the dnh family: TODAY — ~/mirror/tga-scope/20260912T091900Z/tga-scope.sqlite 734cfa9c3ba1230f · 493,576,192 B (INGEST-02, 2026-09-12: delivery_notification 1,975,641 → 1,975,705, the 9 bodies TGA served but never delivered plus the one transient 417, at capture anchor 2026-09-11T09:52:13+00:00; 200 served-absence findings beside them. Cite 734cfa9c3ba1230f from here on.) · predecessor RETAINED, untouched: ~/mirror/tga-scope/20260911T045028Z/tga-scope.sqlite fe6f150ea8dcda0a · 493,522,944 B, DDL of record outputs/dnh-01/dnh-01-ddl-v1.sql 1e1b7dfbf690f27e; scope_entry still sits inside tga-rto — Z2b moves it · of record per store-register.md row scope (predecessor scope-p1) |
tga-history |
the derived, recomputable timeline: every dated event across NRT and RTO in sequence. Built at LANE-Z-03C (2026-09-04) from the five databases then in the estate — before tga-nrt-content-legacy was split out (ADR-087, 2026-09-05; its earliest store is APPLY-POP-02, 2026-09-07) and before tga-cvig was proposed (ADR-091; built and registered 2026-09-16, CVIG-STORE-01). Which of them a rebuild reads is not yet ruled. Rebuilt after each pull; never hand-edited. |
Show history, Compare, and anything that asks "what changed between these dates" | tga-history-01 on the Product account (852 MB); the local twin exists — ~/mirror/tga-history-01/v2/tga-history-gate3-build3.sqlite df7425aed29071a7, deterministic across 24 tables (LANE-Z-03C close, 2026-09-04). The "no local twin yet" wording was stale and is corrected here · of record per store-register.md row history |
tga-cvig |
the companion-volume derivation: the rows extracted from the companion-volume files — codes, titles, and the pathway relationships between them — held in their own container so a later custodian-level restriction is a detach rather than surgery on a store of record (ADR-091, licence containment). | Companion volume — the files listed against a training package release | ~/mirror/tga-cvig/20260916T105323Z/tga-cvig.sqlite 6313f93878101ee2, 267,530,240 B — thirteen tables from staging plus cvig_build, 575 provenance rows each carrying its tga_file_uri (CVIG-STORE-01, 2026-09-16; built twice, byte-identical) · of record per store-register.md row cvig |
tga-bodies (R2) |
the verbatim JSON and HTML as fetched — provenance. Never read at runtime once ingested. | — | the mirror pith-assets-01 (8.3 GB); section bodies also duplicated inside e2-content-01.content today (2 GB, to leave D1 per Tim's ruling) |
⚠ The ledger's edition of record, against these cells (INGEST-02 Gate 5, 2026-09-12). -ed5 6dca4e12876cb7d7 is the edition of record; it predates INGEST-02 and reads held-bodies 3; -ed6, which will read in for tpusage and recognitionmanager, is owed by BRIEF-LEDGER-V7 and is the next emission. The v6 instrument cannot emit it: it hard-codes the store paths and pins their digests, and derives from counts frozen in its Gate 1 store, so it cannot see the successors this section now names.
Why the estate builds locally (Tim, 2026-09-09; first written here 2026-09-11). The stores in the table above are built and served on the Mac because Cloudflare D1 read and write costs blew out at this scale. The Mac mirror is the copy of record; the cloud accounts are not. The Prototype account stays up only for the landing page and the domains, and Claude and Alex are being disconnected from it (Tim, 2026-09-09 — time machine v60.0 d2a1e8d9892121bc, Tim's rulings — 2026-09-09). The practice is older than its written reason: this table's last column has read today, on the Mac since the model was filed (TGA-DATA-MODEL-02, tapped 2026-09-02).
1.0a Runtime location — where each store also lives now (PROMOTE-02, 2026-09-19)¶
The "today, on the Mac" column above is unchanged and still names the store of record. What
follows is a second axis, not a replacement: every store has been copied up to Postgres on
db-01.bne and certified there. The SQLite file on the Mac remains the store of record; the box
holds a loaded copy. A store moves when the Mac file moves — the register's digest is still the
identity — and the box is re-loaded from it, never the other way round.
| store of record | runtime schema | tables | rows | size | host |
|---|---|---|---|---|---|
tga-cvig |
library.cvig |
14 | 123,991 | 108 MB | db-01.bne |
tga-nrt |
library.nrt |
31 | 3,093,118 | 832 MB | db-01.bne |
tga-scope |
library.scope |
5 | 1,975,923 | 434 MB | db-01.bne |
tga-nrt-packaging |
library.packaging |
18 | 2,020,573 | 647 MB | db-01.bne |
tga-rto |
library.rto |
55 | 6,971,556 | 5,113 MB | db-01.bne |
tga-nrt-content (main) |
library.content |
30 | 9,253,168 | 4,311 MB | db-01.bne |
tga-nrt-content-legacy |
library.content-legacy |
26 | 4,287,174 | 2,201 MB | db-01.bne |
tga-history |
library.history |
24 | 6,919,812 | 2,029 MB | db-01.bne |
| eight schemas | database library |
203 | 34,645,315 | 15 GB | 4 vCPU · 8 GB · 100 GB NVMe, Brisbane |
Table names are unchanged; one schema per store; no type was altered on the way up — a copy, not
a redesign. FTS5 virtual tables and their shadow tables did not transfer (Postgres has its own
full-text) and are listed, not loaded; so are SQLite's sqlite_stat1/sqlite_stat4 planner
statistics on tga-history. No schema carries a primary key or uniqueness constraint from an
implicit sqlite_autoindex_* — 183 of those are listed, not built, by standing ruling of
2026-09-19; row identity is in the data, and a consumer must not assume a PK exists. The 283
explicit indexes did transfer, and the far-side proof asserts that count per schema against the
source.
Certified on the far side by three exact count instruments and a per-table content fingerprint
against the SQLite source, and sealed so that library_ro reads and cannot write:
outputs/promote-02/promote-02-farside-proof.tsv and the Gate 5 close. The host, its fences, the
certificate authority and the Hyperdrive path are docs/docs/ops/binary-lane-db-01.md.
1.1 How the databases are kept current — the diff model (Tim, 2026-09-02)¶
There is one set of these databases. A new pull of TGA is compared body-by-body against what is held, by digest: every row already carries the SHA-256 of the body it was parsed from, so the comparison is a lookup. Only a body whose digest changed is re-parsed. Its new rows are written beside the old ones, stamped with the new fetch date and body digest; the old rows keep their dates (the history rule — nothing dated is ever overwritten). A body no longer served is an absence, found by list comparison, and recorded as such — never inferred from a digest.
no-body is a mirror state inside the vocabulary of published states (E3, TGA-MODEL-03 G2-5,
2026-09-05), carried with its least-sure: the column then holds two populations — what TGA
published, and what our mirror happens to hold — and the note column says which. Until a reader
can tell those apart from the row alone, no-body is read as a statement about the mirror, never as a
statement about TGA.
The three-state operating clause (E2, TGA-MODEL-03 G2-2, 2026-09-05). A three-state column is
added where TGA publishes the empty form at scale, on a row a product-page read returns — not
where the empty form is rare and no surface asks for it. The fact is never lost either way: the
bodies and raw_index hold what was served, so a state not modelled is still recoverable. The clause
is a rule about columns, never about evidence.
"Edition" is therefore a stamp, not a store: the fetch date on a row ("as at 2026-08-28"). It says when TGA said this; it does not create a separate copy of the databases.
One loader rule decides whether this holds: insert on change only. scope_entry today carries
the capture time in its key, so an unchanged re-pull would insert all 6.5 M rows again. Under the diff
model a scope row is written only when its content differs from the row held.
Why scope lives beside the RTO and not inside it (revised 2026-09-04). On TGA a scope entry is
a fact about a registration — it has the RTO's code, a start date, an end date and a status — and
the product it names is a reference by code. Both lookups are one query on the same table: filter by
RTO code for the RTO's tabs, filter by product code for Find RTOs. Titles come from tga-nrt on a
second query, keyed on code, which is how the site itself does it (the Find RTOs tab is a separate
fetch). None of that requires it to share a database with the organisation tables, and the measured
sizes say it must not: in the lifted tga-rto, scope_entry and its six indexes are 6.458 GB of
the 6.708 GB store — 96.3% — and the six indexes alone are 3.126 GB, more than every od_* table,
the whole organisation table and every vocabulary put together (Lane Z Gate 3 addendum B,
2026-09-03). Scope grows with the number of registrations; the organisation tables grow with the
number of RTOs. They are two populations with two growth curves and they get two databases.
1.2 The storage convention for our own columns (ADR-083, 2026-09-04)¶
Published values are stored verbatim, always. Every column that holds something TGA published keeps TGA's own representation — a GUID published as a 36-character string is stored as that string. Nothing in §8's "every column is TGA's label or API field" moves.
Our own columns change representation:
| column class | today | from each store's next DDL of record |
|---|---|---|
the _norm join copies of GUIDs (release_id_norm, bundle_id_norm, organisation_id_norm, component_nrt_id_norm) |
TEXT(36) |
BLOB(16) |
the digests (row_sha256, record_sha256, body_sha256, section_sha256) |
TEXT(64) hex |
BLOB(32) |
⚠ record_sha256 holds the source-object digest — the body the row was read from — and not a digest of the record. The name invites the wrong query; tga-rto.scope_entry carries e2-rto-01's source_object_sha256 under it (C-0912-04) |
Measured, not assumed (CANON-01 Gate 1 and Gate 1b, 2026-09-03/04):
- On a
scope_entry-shaped row — every column a key or a digest — the saving is 44.07% locally and 44.40% on D1. The two stores agree to within thirteen 4 KB pages across 53 MB: D1's storage accounting is SQLite's. - On the content store of record, whose rows carry prose as well as keys, the same convention saves 29.1% (3.328 GB → 2.359 GB). The convention pays most where rows are all key, which is scope.
- The fact-key lookup is index-served in both representations: medians 0.679 ms (hex/text)
against 0.668 ms (BLOB) over 100,000 rows on D1,
rows_read: 1on every run.
⚠ D1 returns a BLOB as an array of byte integers, not as hex and not as base64 — measured
"[1, 35, 69, 103, …]". So a reader takes hex() in SQL or converts the array in JS. A consumer
that assumes a hex string gets an array and no error. Rendered as hex only at the read; the stored
form is binary.
Applies from each store's next DDL of record. Existing stores migrate in their own briefs —
Z2b for tga-rto/tga-scope, the Q3 apply as tga-nrt-packaging's v3 DDL, -03 for the tables §13
names. No store is migrated in place by this document.
The join keys are TGA's identifiers and nothing else: product code (and id, TGA's UUID),
release id, RTO code. Every row also carries its fetch date and body digest.
2 · NRT — the training products¶
trainingPackageGroup and qualificationSpecialisation are TrainingComponentType values TGA
models and serves, and the mirror holds no component of either. 99,761 L1 register bodies walked,
0 of each (control: unit 56,725), and component.type_code holds neither. They appear only as
attributes inside org-training-packages bodies — a render surface. No table is owed, and none
is until a body of that type is held; FACET-ROUNDTRIP-01 asks TGA's own facet. (TAIL-RECON-01 C,
C-0913-16)
Component status — the derivation of record (R-FACET-G3-1, 2026-09-14). No table holds a component's status. VET
(unit, qualification, skill set, training package): component.usage_recommendation, with superseded split by the
component's own supersession edges as predecessor (maps_to_code_norm = code_norm) — supersededEquivalent only where every
such edge is is_equivalent = 1; a mixed set or no edge reads supersededNonEquivalent (H-S1b). Accredited (course,
unit/module): the L1 body's own status, served verbatim on all 38,739 accredited bodies and landed in no column. Checked against
TGA's listing of 2026-09-14: 149 held components differ, every one named as movement since capture
(outputs/sweep-01/SWEEP-01-GATE-3-20260914T010043Z-C-STATUS-MOVED.tsv 1a07678fbb0b3a9c).
Facet populations include test artefacts (R-FACET-G3-2). A round-trip's "ours" is component_current whole —
is_test_artefact stated beside, never subtracted: TGA counts its 259 training packages with the test artefact.
Recognition manager — TGA's code for our short name, settled by data (LISTING-SWEEP-01 Gate 3, every component carrying both
agreeing, outputs/sweep-01/SWEEP-01-GATE-3-20260914T010043Z-D-RECOGNITION-MANAGER-MAP.tsv 90ac0c60eca808ec):
SWMC 20 · ASQA 12 · VRQA 01 · WA TAC 04 · QLD DET 02 · ACT DET 07 · SA DFEEST 03 · NSW VETAB 08 · Tas TQA 05 ·
NT DET 06. component_recognition_manager holds names, not codes, and od_registration_managers spells 02 QLD DETE and has no
SWMC — neither is a join key without this map.
2.1 What every NRT product page carries (the header block, all types)¶
| page label | API field | our column today |
|---|---|---|
| Code | code |
component.code |
| Title | title |
component.title |
| (type) | type / componentType |
component.type_code |
| Usage recommendation · date | usageRecommendation, usageRecommendationLabel + change date |
component.usage_recommendation + usage_recommendation_label in tga-nrt, populated on 87,107 component rows of 125,848 — superseded 45,088 · deleted 24,006 · current 18,013 — and NULL on exactly the 38,741 accredited rows (accreditedCourse 19,439 · accreditedUnit 19,302), which carry no usage recommendation; 0 NULL on any other type. Measured at the register of record 4190424cbe933e5a, read-only, at CANON-CORRECTIONS-04 Gate 1 d1f3348709322d10 (the queries are printed there); second instrument, LEDGER-02's Gate 1 run (LEDGER-02-GATE-1-20260911T010159Z-RUN.md, "populated 87,107"). Landed by TGA-MODEL-03 (DDL of record 37f735f5b941571f). Two populations, not crossed: 87,107 counts component rows; the L1 register bodies carry it for 87,099 distinct codes. Until 2026-09-11 this cell said the column did not exist (CANON-CORRECTIONS-04). §13 |
| Release · Current · date | releases[].releaseNumber, .releaseDate, .currency, .currencyChangeDate, .id |
release.release_number/date/currency, release_id |
| Release · which body each row came from | the L1 register body's releases[] entry, and (from 2026-09-07) the releases endpoint body that changed it |
release.source_object_sha256 is the L1 REGISTER body this row was parsed from — which is what release.source_lane says on all 103,349 rows, none of which reads releases. release.refresh_object_sha256 is the refresh body that caused the row (INGEST-APPEND-01 Phase 2). NULL on every ed1-2026-08-28 row means NOT RECORDED — the 08-28 build overlaid nine columns from a refresh body whose digest it did not keep — never 'none' (ABSENT-NOT-DEFAULT-01) |
| Show history | the full releases[] array |
all releases kept — release has one row per release |
| Download | the PDF = the release's bundles rendered in order | §3 |
⚠ The 87,107 above is a ROW count. component is UNIQUE (code_norm, source_lane), so those rows
cover 87,099 distinct codes; the difference of 8 is the eight codes held on both training-L1 and
unit-L1 — AURRTE010 CPCPPS5001 FBPHVB3005 ICTPRG555 MSFFL2038 PMAOPS223 SFIFSH301
UEERA0019. The 87,099 quoted elsewhere in this file counts CODES. Both figures are true of their own
population and the difference is not a gap. (TAIL-RECON-01 A, C-0913-15)
2.2 Qualification¶
Tabs: Qualification details · Units of competency · Summary · Find RTOs.
Summary tab (from the ACM10121 screen cap — every table, every column, verbatim):
| table on page | columns on page | API field | our table today |
|---|---|---|---|
| Mapping | Mapping · Notes · Date | mappingInformation[]: code, mapsToCode, mapsToTitle, mapsToId, isEquivalent, date, notes. Direction: code is the successor, mapsToCode the predecessor (audited 2026-04-16) |
supersession_edge |
| Releases (+ Advanced compare) | Compare · Release · Release date | releases[] |
release |
| Number of required units | Core · Elective · Total | /releases/{n}/unitgrid → isEssential per unit; counts derived |
qualification_specialisation.packaging_core/elective (partial) |
| Parent training package | Code · Title · Release | parent.code, parent.title, parent.id |
component_parent in tga-nrt — child · parent · link state, one row per component whose body carries a parent (87,107 rows): 86,848 resolved (unit 75,135 · qualification 8,034 · skillSet 3,679) + 259 code-absent (every trainingPackage body carries a parent with no code key), 0 orphans, 0 bodies missing of 125,848; accredited courses and units carry no parent at all (38,741). Landed by TGA-MODEL-03-QUALS-01, 2026-09-07 (close 814aa6cc3993bceb, f0cfc91d; the by-type split from its Gate 1 df7ed22f99bd37a0); the register of record is §1's. The release range is not a field — no body names one, and parent carries only {code, id, title} (the same Gate 1); it is a derivation, held under R-M3 for a later brief. Until 2026-09-11 this cell said the table did not exist (CANON-CORRECTIONS-03). |
| Classifications | Scheme · Code · Classification value · Start date · End date | classifications[] (ANZSCO, ASCED field, ASCED level), taxonomy.industrySectors[], taxonomy.occupations[] |
stored in tga-nrt as component_classification (eleven keys, §13) and the two taxonomy tables taxonomy_industry_sector · taxonomy_occupation. Landed by TGA-MODEL-03, 2026-09-05, build-verified at its Gate 3 28c2f4b266bf871b (store 3c5c2458e864d90f): 280,883 · 8,262 · 14,022 rows. LEDGER-02 Gate 1 5da0c3e8366e4ba9 measured the same three counts on 2026-09-11; that report names no store digest for these counts. Rows, not bodies: §13's 1,244 and 3,103 are taxonomy-* body counts. Until 2026-09-11 this cell said classifications were not stored (CANON-CORRECTIONS-03). |
| Companion volumes | Title · Filetype · Status · Created date · Release | the companion-volume listing on the parent package | Lane F's file register (outputs/lane-f/), not yet a table |
Qualification details tab — corrected 2026-09-02 against UEE63020_R4.pdf and the store.
A qualification is one bundle (Default), and its sections are whatever that bundle carries —
not a fixed list. Measured over 12,902 qualification releases: Description 0001 10,117 ·
Packaging rules 0116 10,077 · Modification history 0012 10,076 · Entry requirements 0110
9,592 · Licensing 0011 4,901 (the other half carry the licensing sentence inside the
description, as UEE63020 does) · Employability skills summary 0109 4,620 · Pathways 0117 4,619
· General 0000 989 · Pre-requisites, Workplace requirements, Credit arrangements — a handful.
The PDF then appends Qualification Mapping Information (from Mapping) and Links (the
package). UEE63020 R4 is exactly: 0012 · 0001 · 0110 · 0116 · mapping · links.
The packaging rules ARE the unit list. There is no separate "unit grid" in the document.
0116 is one table: a rule-prose row (the weighting-point totals, the per-group ranges, the
imported-unit rule, the prerequisite-asterisk rule), then a group header — Core units,
Group A: Imported and common elective units, and Group B through Group E, each
General elective units (corrected 2026-09-04 against the PDF; the earlier reading gave that
wording to Group B alone) — each followed by unit rows of code (as <ntr-tcref data-nrt-code
data-nrt-title data-nrt-type>) · title as published (the * prerequisite marker is in the title
cell, not the ref) · Weighting Points.
A group header has three markings, and two of them are not table rows (measured over the 68,283 groups Lane Q's parser built, 2026-09-03):
| marking | groups |
|---|---|
| a table row carrying the column label | 40,252 |
a paragraph immediately before the table — bold-in-<p> 21,550 · plain <p> 5,773 · heading tag 708 |
28,031 |
| a paragraph inside a packed cell — seen by neither predicate above, not yet parsed (Q3b) | unmeasured; 198 units are ungrouped because of it |
The column set is the package's, not TGA's: 843 of 10,077 packaging-rules sections carry a
weighting-points column; the rest carry code · title only, or other headers. The predicate is
named because the figure depends on it: 843 is Lane Q's normalised cell census over 0116,
and it includes four sections whose header is not a text node. A strict match on the parsed
column header instead returns 809 (CANON-01 Gate 1 §2.2) — the same corpus, a narrower
predicate, not a contradiction. The 844 this document carried until 2026-09-04 had no
measurement artefact behind it and is withdrawn. Columns are captured as published headers,
never assumed. /releases/{n}/unitgrid (isEssential) is TGA's derived view
of the same list and is the cross-check, not the source. "Number of required units" on the site is
derived from it too. The grid is held as three derived-class tables in tga-nrt-packaging —
unitgrid_fetch (12,902 rows), unitgrid_unit (510,406), unitgrid_unit_link (570,641) — and its
refresh is time-based, never event-triggered: grid bodies churn on unit currency, not on
qualification releases (274 changes across 92 grids in six days, zero membership movement, every one
current → superseded; measured at Lane Q's Q2 walk, filed at outputs/lane-q/LANE-Q-02-CLOSE-2026-09-03.md d62b0406730b3a85 §2 finding 5 and §3). A refresh keyed on "this qualification has a new
release" would have caught none of them.
⚠ packaging_grid_relation.equal means equal COUNTS, not equal membership. Over the 10,020
releases holding both a packaging table and a non-empty grid: count-equality 6,153 (61.41%),
set-equality 5,763 (57.51%). Reading equal as "the grid and the table agree" is wrong 390
times. The relation carries both, each named (Lane Q Gate 4 §5, 2026-09-03).
The unit row, and what is not a unit. One <ntr-tcref> per row is the unit the row is about.
A later tcref in the same row is a prerequisite, not a second unit, where the column header, the
∟ / (Note pre-requ cue, or the unit's own 0121 section says so. The positional cue — "a tcref
after the first is a prerequisite because of where it sits" — is retired: a witness in the data
(the 0121 section) replaces it, and the positional form was wrong on rows that list two units
side by side.
⚠ packaging_unit is not unique on the unit. 535,043 rows over 500,803 distinct
(release, section, unit) — 34,240 repeats (measured 2026-09-04). A qualification may list the
same unit in more than one group, and a key that assumes otherwise loses rows.
Entry requirements have the same shape as prerequisites: <ntr-tcref> items in lists with
bare and/or paragraphs between them, plus free-text alternatives ("a current Electrical Fitter
Occupational License…"). Same token model.
Find RTOs tab — the scope join, product side. → tga-rto.scope, §4.
2.3 Unit of competency¶
Tabs: Unit details · Summary · Find RTOs. Header and Summary as 2.1/2.2 (Mapping renders as "Unit Mapping Information" in the PDF). Unit details = the sections, §3: Modification history, Application, Pre-requisite unit, Competency field, Unit sector, Elements and performance criteria, Foundation skills, Range of conditions; then the Assessment Requirements bundle: Performance evidence, Knowledge evidence, Assessment conditions.
2.4 Skill set · Accredited course · Training package¶
Same header, same Summary shape.
Skill set — checked against UEESS00210_R1.pdf and the store, 2026-09-02. One bundle
(Default), eight sections, and the PDF is those eight in order with nothing appended: Modification
history 0012 · Description 0001 · Pathways information 0117 · Licensing/regulatory information
0011 · Skill set requirements 0126 · Target group 0127 · Suggested words for statement of
attainment 0128 · Custom content section 0000 ("Not applicable." — declared absent, stored
verbatim). The most uniform type in the register: of 5,461 skill-set releases, 5,446–5,454 carry
each of the first seven; 2,145 carry the custom section. Two things to hold onto: the mapping
sentence lives inside the modification history ("This Skill Set replaces and is equivalent to
<ntr-tcref>UEESS00136…") as well as in the register's mapping edge — the PDF has no separate
mapping block; and Skill set requirements is the same table shape as packaging rules — a rule
row ("A total of 1 unit of competency must be attained", the asterisk rule) then unit rows of
<ntr-tcref> code · title-as-published with *. One table model serves both.
Accredited course
details carry 0105/0106/0201–0203 and the accredited-unit list. Training package summary carries
0010 Imprint, 0002 Overview, the qualification/skill-set/unit lists, and the companion volumes.
Each gets its own column-by-column table like 2.2, transcribed from its screen cap, as an addendum;
none of the labels are invented here.
3 · NRT content — the section bodies as rows¶
content_family covers all 36 types (A8, Gate 2c verdict 04cb260c579234b3 §2, 2026-09-05).
Every one of the 36 content type codes in the table below carries a content_family; none is left
unassigned. general = 0000 plus the eight small types. 0103 is prose, not a block model.
0001 in block has one parser of record: prose e5f2cacc781cc4f3, on every component type (R-CP-0001-01,
Tim 2026-09-14). A skill-set parse of 0001 is retained, not current: eb6efd624ea861f0+b28da335ad85453f's
6,650 block and 5,461 section_parse rows stay in the store and are never indexed.
THE PIPELINE REGISTER IS THE REGISTER OF RECORD (CONTENT-PROVENANCE-01, 2026-09-14). A row in this store is the
product of a pipeline, not of a parser alone, so the content main carries pipeline_register: one row per
(store, content type, table, parser_version), naming the parser, the writer, the loader, the entry point and its
arguments, and the substitutions a replay needs, with the row count, a currency verdict and its evidence. 158 rows:
carried 51 · added 40 · resolves-in-pair 34 · retained 15 · added-routed 13 · co-record 5; none
second-label, none ruling-owed. co-record is a second label of record at the same type and table, indexed
beside the first — at 0126, reference-0116's 147 synthesised declared-absent blocks sit beside skill-set's, and all
147 are absence beside absence (Gate 4). retained rows are kept, never current and never indexed.
parser_of_record stays as the per-(type, table) pointer, its rows unchanged, with superseded_by_register =
'pipeline_register'. The register is written by the register instrument, never by a loader (PIPELINE-REGISTER-01).
section_search.release_currency is derived from register row nrt, live rows only (superseded_at IS NULL),
for the whole column (R-CP-G3-1); pipeline_dependency names that dependency — on b93666be2931a824 it names register row nrt 767f89ac3afb6ffc (415,491 rows · 75,138 releases), re-pointed by PARSE-OPEN-01 from 8b84a2e98745b474, the replaced row kept whole in build_finding (the table has no history column). It replaces a copy from E2
2eff8cadfe218c21, which was stale on 14 releases replaced 2026-09-07 and blind to 14 appended at ed4.
A PARSER CHANGE IS A RE-POINTING (APPLY-POP-02 v2, 2026-09-08). When a family's parser is
replaced, the new parser's rows land at their own parser_version label and parser_of_record is
re-pointed to it; the incumbent rows are RETAINED, NEVER CURRENT, and are NOT stamped.
superseded_at means a body TGA no longer serves — it is not the column for "we parse this
differently now", and using it that way would say the source withdrew something it did not.
AN APPEND WRITES AT THE TYPE'S REGISTERED LABEL (CONTENT-APPEND-01, 2026-09-13). Sections appended to
an existing store are parsed by each type's parser_of_record, at that label, by the pipeline that produced
the store's own rows — parser and loader, proven whole by replay against the store before a row is
written. A section whose pipeline is not proven whole carries a section row and no parse row. It is
counted open by the ledger's open_no_parse_row column and is never given a typed state:
declared-absent means the parser of record ran and TGA's own text declares the thing absent, and a section
nobody parsed is not that. An absent row is the honest record of an unrun pipeline. CA-01 appended 118
sections: 95 carry parse rows at five proven pipelines plus licensing's 0011; 23 carry a section row
and no parse row (0000 0112 0116 0123 0126 0127 0128), owed by producer — CONTENT-APPEND-02 was folded into
CHILD-APPEND-01 (R-CA1-5) and retired as a name. 87 open after CHILD-APPEND-01 (23 + 64) (2026-09-14; Gate 4
outputs/child-append-01/CHILD-APPEND-01-GATE-4-20260914T072602Z.md afd36fcb96a68980). CHILD-APPEND-01 appended 377 sections at ed6 with parse rows for the six families
proven by replay (prose · head · 0012 · 0121 · evidence · licensing) and named four open, each at the one line a
fixture run would have had to change — a whole-corpus count assertion every time (R-CA1-8): lane-w
lane-w-01b-build-v1.py:198 (0112 6 · 0123 23) · content-0000 content-0000-01-build-v1.py:341 (0201 25) ·
reference-0116 q3-apply-01-gate4-step2-guard.py:164, its positive control (0116 4 · 0126 2) · skill-set
apply-pop-02-build-v1.py:97 (0126–0128 6; its 35 0011 sections carry licensing's parse row and no skill-set
block — a partial, not open). The same four families hold CA-01's 23. PARSE-OPEN-01 (2026-09-14): the four families lifted (their registered labels resolved to files by digest, ≥ 20 held sections per family reproduced), 87 → 0 open, 5 partials → 0; Gate 4 9c490f831fbaafcb 46/46.
THE 22's SIX LEGACY SECTIONS ARE OWED TO THE LEGACY STORE (CONTENT-APPEND-01, R-CA-G2-10). 0100 0108
0111 0119 0124 0125 — one bundle, one release, one body digest — are held in bodies and owed to the leg
store by LEGACY-APPEND-01. The legacy family is fully built there (171,383 sections under
787ea843799f6a48+a2ec60bf0b9d69d9); route-absent is retired for them, because a declared absence naming
a parser that exists and has run would be a typed state that is false. Landed by LEGACY-APPEND-01 v2, 2026-09-14 — 6 section · 6 section_parse · 181 block · 2 reference · 6 section_search at ed4-2026-09-13-append01; Gate 4 a8792459025e90c4, 22/22.
THE LEGACY FAMILY IS THREE STAGES (LEGACY-APPEND-01 v2, standing line). Emitter 787ea843799f6a48 · builder a2ec60bf0b9d69d9 · the assembly d96a76ec02c6d561: the builder writes a delta and the assembly's copy_pair carries it into the leg store, adding markup_state and synthesising or re-kinding declared-absent rows. An append lifts all three by declaration and proves them on held rows before it lands anything (26/26 sections at Gate 1). corpus_frequency is build-grained (R-LG-6): an append stamps its own rows over the existing corpus plus itself and never re-stamps a written row. The leg store's nine *_current views admit every registered edition (R-LG-12) — their MAX(as_at_date) clause was a one-edition assumption, and with a second edition it would have emptied block_current for every existing row (after the append block_current = 3,735,221). vocab_edition: the ed4-2026-09-13-append01 row's as_at_date 2026-09-13 is taken from the edition's own name, CONTENT-APPEND-01's records not carrying it (R-LG-11).
section_search IS REBUILT WHOLE FROM THE REGISTER (CONTENT-PROVENANCE-01 Gate 3, 2026-09-14). One row per
block address under each type's of-record and co-record block labels, retained labels excluded; the text is the
address's non-NULL block texts joined in (parser_version, item_index) order. 397,187 → 415,201 rows, the gap on
indexed types 18,016 → 0, rows outside the indexed set 2 → 0, section_search_fts = table; re-derived by an
independent instrument at Gate 4 (695b10891bc9cdc1, 18/18, a 200-address sample by its own rule). This closes
CONTENT-APPEND-01's known-stale finding (17,999 of 415,186 addresses unindexed; SEARCH_SEL in
content-v3-01-assemble-v8.py :262-266 folded retained rows in). Appended sections with block rows are indexed;
the 23 with no parse row have nothing to index. Search coverage is now a function of block under the register —
rebuilt, never patched.
F-LG-11 — closed on the main by PARSE-OPEN-01: vocab_edition rows for ed4 and ed6, the nine views admit every registered edition (C-0914-09 → law) (2026-09-14; b93666be2931a824, Gate 4 outputs/parse-open-01/PARSE-OPEN-01-GATE-4-20260914T095939Z.md 9c490f831fbaafcb). Each view's DDL is the predecessor's with the one e.as_at_date = (SELECT MAX(as_at_date) FROM vocab_edition) clause replaced by e.edition IS NOT NULL, AND t.superseded_at IS NULL kept; each view's count after = its count before + exactly the ed4/ed6 rows its parser_of_record join admits (block_current +2,521 of 2,557 — the 36 under labels parser_of_record does not hold stay out). Standing, as law: a content store's *_current views admit every edition vocab_edition registers, and every appended edition is registered with its own date in the build that appends it. Was: F-LG-11 — the nine content *_current views admit only the editions vocab_edition holds at MAX(as_at_date), so
every row appended since ed1 is outside them (standing line; measured at CHILD-APPEND-01 Gate 4, outputs/child-append-01/CHILD-APPEND-01-GATE-4-20260914T072602Z.md afd36fcb96a68980;
candidate C-0914-09). block_current · criterion_current · element_current · foundation_skill_current ·
licensing_determination_current · modification_history_current · range_of_conditions_current · reference_current ·
skillset_unit_current each join vocab_edition and keep e.as_at_date = (SELECT MAX(as_at_date) FROM vocab_edition),
and vocab_edition holds only the two ed1 spellings. On 2b5fda3894f6b56c none of the nine holds an ed4 or an ed6
row — block_current alone leaves out 418 ed4 and 1,925 ed6 rows — and their counts are identical on its
predecessor. section_parse_current has no edition join and is not affected. A read through these views does not
see CONTENT-APPEND-01's or CHILD-APPEND-01's rows. Fix in PARSE-OPEN-01 (vocab_edition rows for the appended
editions, the nine views re-pointed; R-CA1-13: not applied to 2b5fda3894f6b56c, which Gate 4 verified as built).
The acceptance for such a move is stated per content type on the _current view, because a type
ruled to another parser cannot appear in the difference at all: at APPLY-POP-02 v2 the five
re-pointed pairs moved 0011 +1,862/0 · 0126 +151/139 · 0127 +48/47 · 0128 +3/2, 0001 0/0
(ruled to the prose parser, its 6,650 new blocks retained and never current), and the other 25
content types 0/0, measured.
⚠ A re-pointing is invisible to the schema parity guard. The DDL of a re-pointing differs only
in parser_of_record seed rows, so the guard — which compares schema — is green on the new store
and green when shadowed to the old one. Measured at APPLY-POP-02 v2's Gate 5. A guard for this
class of move does not exist yet. The family is decided at parse by what the body is, and it is the column
ADR-087's split reads.
ALL-RELEASES-01 (Tim, 2026-09-02): the parse ranges over every release of every component,
not the current one. Every bundle of every release the mirror holds is parsed and loaded, with its
release id on every row. "Current" is a read-time filter, never a parse-time filter. Any loader that
takes releases[current], contentBundles[0] or the latest release is a defect.
A rule lives in a view, a fact in a table (E4, Lane Q Gate 3, 2026-09-05 — a build that returned
125,836). lane_of_record holds R1's type→lane mapping, which is a rule and therefore a view.
R1-a's per-code tiebreak is a fact about a code and belongs in component_current. Where a rule is
stored as a table it drifts from the logic that should have derived it; where a fact is computed in a
view it cannot be pointed at.
Consequence: unit_parse is keyed on unit_code_norm alone today (current release only). Its key
becomes (unit_code_norm, release_id_norm) before the cloud load, and element /
criterion re-key with it — a local re-run now, a migration later.
This is the part that was built lane-by-lane without a top-level model. The model is TGA's:
release ─ has bundles (units: "Default" 0000 and "Assesment Requirements" 0013 — TGA's spelling)
bundle ─ ordered sections, each typed by contentTypeCode (36 in the corpus)
section ─ parsed into addressable rows in ONE table per section family
The section index (one row per section; TGA's fields): bundleId, bundleTypeCode,
bundleTypeName, releaseId, sequence, contentTypeCode, contentType, title, format,
plus our body_sha256 and r2_key. No body column — bodies are in R2.
The row tables, one per section family, keyed on releaseId + contentTypeCode + array position,
every row carrying its fetch date, parser_version, body_sha256:
| section family (contentType) | table | rows per section | state |
|---|---|---|---|
0118 PerformanceCriteria |
element → criterion |
element: number · title; criterion: number · text | done — element 227,630 / criterion 1,028,001 (2026-09-03) |
0120 0113 0104 PerformanceEvidence · KnowledgeEvidence · AssessmentConditions |
evidence_item |
one per published block (stem, list item with depth, prose, table cell with coordinates) | done — 1,325,227, zero lost blocks (2026-09-03) |
0011 / 0103 licensing statement |
licensing_determination, read through licensing_determination_by_release |
sentence + determination. The view returns distinct sentences per release (distinct on the sentence digest) and is the only surface a model reads — the store keeps every copy at its own address, nothing ranked (RULING-LICENSING-DUPLICATE-SECTION-01, Tim 2026-09-03) | 97,047 done; backfill OPEN — 34,927 releases / 34,949 sections, the largest single open figure in the store. -02b |
0126 0127 0128 skill set |
skillset_block, skillset_unit |
0126 is the packaging-rules table shape (rule row + unit rows); 0127/0128 are prose |
done — to be re-keyed onto the shared packaging_group/packaging_unit model when that lands, not rebuilt |
0012 ModificationHistory |
modification_history |
one per row of the table: text verbatim; release number parsed beside it | done — 155,901 (2026-09-03). -02b: 4,486 decimal lifts in release_number_parsed, 79 distinct — training-package versions in a column named for the release number |
0121 Pre-Requisites |
reference (the universal table, §12.6) |
one per <ntr-tcref>: data-nrt-code, data-nrt-title, data-nrt-type (TGA's own attribute names), plus the and/or tokens in published order |
done — reference 49,117 (2026-09-03); the 495 0121 prose sections are -02b, so this row is not closed |
0123 RangeOfConditions |
range_of_conditions |
lead-in prose row, then variable × values rows | to build |
0112 FoundationSkills |
foundation_skill |
sentence, or skill × description rows | to build |
0103 Application · 0200 CompetencyField · 0102 UnitSelector · 0001 Description · 0110 EntryRequirements · 0117 PathwaysInformation · 0107 CreditArrangements |
section_text |
ordered paragraphs | done — 416,139 (2026-09-03) |
0116 PackagingRules |
six tables in tga-nrt-packaging: packaging_rule (15,531) · packaging_block (117,105) · packaging_group (70,874) · packaging_column (35,875) · packaging_unit (563,487) · packaging_unit_cell (113,982) — as loaded, v2 + q4, store f04ff17c8e2bb640 (Q3-APPLY-01 Gate 3 21f88d56d7a4b887; the v2 + q4 split in its STOP report 1cfcac001380104f). Until 2026-09-11 this cell carried Lane Q's v2 delta alone, with no population named |
rule prose as ordered blocks; one packaging_group per group header in any of its three markings (§2.2) with its published column headers; one packaging_unit per row: code · title as published (incl. *); each extra column's value is a packaging_unit_cell at its col_index, and the cell does not carry its header — packaging_unit_cell.header is NULL on 75,123 of 113,982 cells (65.9%), across 2,220 of the 3,182 qualification-releases that have cells (RELEASE-CURRENT-CONTRACT-01 Gate 1 243ef6a0967cb144, window/main-02 @ 120099fa). The published header is held in packaging_column, not on the cell: a read reaches it through that table or reads by position (the join key is not stated here — the cell carries no table_index or group_position, which packaging_column does). A read that selects by the header Weighting Points finds 5 of 15,007 qualification-releases (QUAL-SERVE-01 close d603676ff010d77b, window/main-02 @ 6c5337a9). Core/elective is the group; unitgrid is the cross-check |
built and verified (Lane Q Gate 3/4, 2026-09-03); applied as tga-nrt-packaging's v3 DDL of record cfc8b4b860c0c7dc (Q3-APPLY-01 Gate 3 21f88d56d7a4b887, PASS, 2026-09-04), under §1.2 |
legacy template — 0100 0119 0124 0125 0111 0108 0109 |
legacy_block |
block model as evidence_item |
built — 171,383 sections in the legacy store under 787ea843799f6a48+a2ec60bf0b9d69d9 (lane-w-01b + lane-q-05-build-v2.py), measured by CA-01's legacy probe 2026-09-13; "to build" is retired by measurement (C-0913-09) |
0000 General and the rest |
general_block |
block model | to build — 24,112 sections, 280.7 MB of body, 18% of all published content bytes and the single largest unparsed type (§7) |
Owed on reference, and it is not merely a nicety: text_as_published. The column holds the visible text inside the <ntr-tcref> element, verbatim, beside the attribute. TGA's data-nrt-title and the text inside the same tag differ — measured on 0121, where the attribute says "…for electrical installations" and the visible text says "…for general electrical installations". It is also the home of the 69,081 prerequisite title texts Lane Q's residual sweep holds today (Q Gate 4 §9, 2026-09-03). packaging_unit.title_as_published already does this for 0116; the column lands on reference in the Q3 apply's v3 DDL and the 0116 rows take it too.
Owed on 0118: the table-header measurement. Four blocks the round trip found unstored — Elements · Performance criteria and their two explanatory sentences. Measure whether the header text is one closed set across all 58,537 0118 sections: if it is, a vocabulary row plus a per-section pointer; if not, rows. -02b.
The 36 codes, measured 2026-09-03 by walking every bundle body in the mirror (~/mirror/tga-content-ed1-2026-08-28/content/bundle, 105,911 files, 0 unreadable): 781,817 sections over exactly 36 content type codes, 1,547,782,711 bytes of published content. That section count reproduces section in the store of record and the completeness ledger's §4 independently. Per-code section counts and body bytes: outputs/canon-01/canon-01-bodybytes-v1.json 32aa1e9b68730fa7. Bytes per section range from 118 (0102) to 25,211 (0105) — 213× — so no per-section mean is a safe basis for anything (§7).
Round trip (the test): render tga-nrt + tga-nrt-content for UEEAS0009 Release 2 in bundle
order; diff against the PDF text with cover, headers and footers removed. Pass = identical. "Unit
Mapping Information" comes from Mapping; "Links" is the parent package's VETNet URL, stored once
per package.
3.1 Three identity rules for bodies, sections and types (CANON-PASS-01, 2026-09-04)¶
A digest identifies a body; only the address identifies a section. Two sections may carry the same body and are still two sections; the address — bundle, release, content type, index — is what distinguishes them, and a join keyed on the digest alone silently merges them.
A body digest identifies a body, not a section — and not a content type. Four of the eight small
types are 0000 bodies republished under another code, so a per-type body count double-counts them
silently. Every type-scoped figure is counted from the store, never from a per-body record that
holds the type it was first met under. (Lane W, W-03 census, 2026-09-04: this defect produced three
separate wrong figures in one instrument before a total failed to close.)
A title is a label on a section, not a description of its content. Where a 0000 body carries
its own <th>, the two disagree on 1,285 bodies — a build that keys on title alone mis-routes
them.
4 · RTO — the training organisations¶
Tabs: Summary · Contacts · Addresses · Scope overview · Qualifications · Skill sets · Units · Courses.
Summary tab (from the 0573 screen cap, verbatim labels):
| table on page | columns on page | our table today |
|---|---|---|
| (header) | Code · Legal name · Status · Registration manager · Show history · Download | organisation (code, id, status, rto_status) |
| Summary | Code · Legal name · Business name(s) · Status · ABN · RTO type · Web address | organisation + od_legal_names + od_trading_names + od_legal_names_abns + od_web_addresses + od_classifications (RTO type) |
| History (Code / Legal name / Web address / Business name / RTO type) | value · Start date · End date | the same tables — every dated row kept (od_codes, od_legal_names, od_web_addresses, od_trading_names, od_classifications) |
| Registration | Registration manager · Initial registration date · Start date · End date · Legal authority · Exerciser | od_registrations, od_registration_managers |
| History (Registration manager / Registration / Renewal application received / School registration / Higher education registration) | Description · Start date · End date | od_registration_managers, od_registrations, od_roles |
| Regulatory decisions · Regulatory decision history | (decision, dates, status) | od_regulatory_decisions + _decision_events + _decision_mappings + _reverse_decision_mappings |
Contacts → od_contacts. Addresses → od_addresses. Restrictions (shown on scope) →
od_restrictions. CRICOS → od_cricos_codes. Column labels for these tabs are transcribed from
the Contacts / Addresses screen caps in the addendum, not invented here.
The RTO side, as lifted (Lane Z Gate 3, 2026-09-03). In the source e2-rto-01 every od_*
table was (org_code, element_index, element_json, source_object_sha256) — the API's JSON element
stored whole in one column. That lift is done: the seventeen od_* families are columns in
tga-rto, 407,915 of 407,915 rows reconstructed losslessly from the rows as stored.
Four things the lift established, each of which the model had wrong or unstated:
parent_indexis on exactly four tables — the nested ones:od_legal_names_abns,od_regulatory_decisions_decision_events,od_regulatory_decisions_decision_mappings,od_regulatory_decisions_reverse_decision_mappings.- ⚠
source_object_sha256is the source store's column, not the lifted store's.tga-rtocarriesrow_sha256(the element's own content) andrecord_sha256(the organisation body). Namingsource_object_sha256in a query againsttga-rtoselects nothing. od_contactsisverbatim— lifted into columns, held exactly as TGA publishes it, and never enriched, merged or joined to any other fact this house holds about a person. Its person fields are indexable and searchable (Tim, 2026-09-11). Its use is entry 1 of OPEN-TAPS-01 (standing-rules.md). Superseded 2026-09-11, recorded and not deleted: the rule of 2026-09-03 made itverbatim · personal, lifted with no index and no full-text search on any person field; Tim cancelled it on 2026-09-11, "for now".- Within-organisation duplicate elements are TGA's, not the lift's — one organisation serving the same element twice: ABNs 3,859 · decision events 342 · addresses 25 · contacts 12 · restrictions 2, and zero on the other nine tables. A digest collision across organisations is the digest doing its job and is not a duplicate.
od_cricos_codes holds 0 rows, and the emptiness is the held fact: TGA serves the endpoint
empty. (ABSENT-NOT-DEFAULT-01 — an empty served list is a fact, not a gap.)
5 · Scope — the seam, both directions¶
Scope overview / Qualifications / Skill sets / Units / Courses tabs on the RTO page and Find RTOs on the product page are the same table read from either end.
| on page | field | our column today |
|---|---|---|
| RTO | code |
scope_entry.org_code |
| Code · Title (of the product) | nrtId / code → title from tga-nrt |
scope_entry.component_nrt_id (title by join) |
| Extent · Start date · End date · Status (delivery / assessment; current / expired) | the scope entry's own fields | confirmed as columns — scope_start_date, scope_end_date, status / status_label, and the extent as extent_label. ⚠ See the note below on extent |
| Restrictions | od_restrictions |
od_restrictions |
| Show history | every dated scope row kept. source_captured_at is a column and is NOT in the key — it is a per-fetch stamp with 89k+ values, so keying on it makes every re-pull a new row by construction |
scope_entry 6,526,172 rows (from 6,526,449 in the source; 277 duplicate fact tuples collapsed by the key, all 12 NULL-end-date pairs among them) |
⚠ extent the field is served, and extent the column is not. TGA declares both an extent
code and an extent_label. Measured over all 6,526,449 source rows: extent is NULL on 100.00% of
them and extent_label on 0.00% — two values, Deliver and assess 6,461,437 and Assess only
64,735. So the extent is published, in full, on every row; it is the code column TGA never
populates. Z2 dropped extent rather than carry a column that reads as a field which exists;
anything filtering or grouping on extent gets one silent bucket. A future document must not
read this as "we do not hold the extent" — we hold it on every row.
The fact-tuple key is (organisation_id_norm, component_nrt_id_norm, extent_label,
scope_start_date, IFNULL(scope_end_date,''), status). The IFNULL is load-bearing: twelve pairs
differ only in a NULL end date, and a plain unique constraint would have kept both without saying
so.
The six denormalised organisation columns stay out of scope_entry (Lane Z Gate 4 §2.1,
2026-09-03). The falsifier measured the alternative: find_rtos — LEFT JOIN with correlated scalar
subqueries — returns 43.2 ms median on the largest product in the register (5,461 scope rows),
index-served throughout, against a 50 ms wall. And it found two real defects in the naive join on
the way: an inner join to organisation silently drops the 77 unresolved organisations, and the
_current LEFT JOINs fan out where an organisation carries more than one current legal name — the
two errors nearly cancelled, so a spot-check would have passed it. find_rtos is the declared
read and nobody downstream rewrites it. The condition on which the six columns return: Z2b
re-measures after the §1.2 compaction; if it then exceeds 50 ms on the largest product, they come
back as a derived materialisation with its class written on it, never as verbatim columns.
Two indexes measured at 819 MB buying nothing. idx_scope_entry_comp_status and
idx_scope_entry_org_status cost 12% of the store and change the filter timing by 0.008 ms and
0.003 ms, because the organisation and component indexes already serve it. Dropped at Z2b. [2026-09-15, FIND-RTOS-01: not yet — Z2b's compaction has not run; idx_scope_entry_comp_status and idx_scope_entry_org_status are still in register row rto, and the planner serves the component filter from idx_scope_entry_component. A standing Z2b item.]
The 77 organisations named in scope with no organisation body are carried in the view
scope_organisation_unresolved, surface with NULL names in find_rtos as they should, and are on
the fetch list. [2026-09-14, NRT-APPEND-01: resolved — they were among C-SWEEP-05's 117, and the view reads 0 rows on register row rto (F-NA-1 below).] [2026-09-15, FIND-RTOS-01: find_rtos.legal_name is NULL only where a resolved organisation has no current legal name — 180 rows on 3 cancelled RTOs on 2026-09-15 (22606 · 22613 · 22721); legal_name_last carries the most recent name regardless of currency and is NULL on 0 rows (R-FR-1).]
C-SWEEP-05 — 117 RTOs TGA's search lists were absent from the organisation enumeration (LISTING-SWEEP-01 Gate 3, 2026-09-14;
cause named at NRT-APPEND-01 Gate 1, rows landed at its Gate 3). 111 of them carry a TGA updatedDate on or before the ed1 organisation fetch — absences, not new
registrations. Their bodies were held, and NRT-APPEND-01 landed them: ~/mirror/tga-fetch-20260914T001545Z/bodies/ (four organisation bodies each).
Their scope of record is scope-single/ for the 31 paged organisations and bodies/ for the 86 one-page ones —
bodies/' paged org/scope walk is the record of the paging defect, not of the scope (tga-ntr-api.md §5.8, C-SWEEP-06).
That scope is in the per-organisation org/scope shape and lands in scope_entry_org, not scope_entry (§5.1).
The cause (NRT-APPEND-01 Gate 1, NRT-APPEND-01-GATE-1-20260914T022751Z.md d40d8c71f3a1528c): both August captures of GET /api/search/organisation carried no query string, so no orderby, and each returned totalCount 13,146 rows — of only 12,496 distinct organisations on 2026-08-14 and 12,867 on 2026-08-15. Organisations served more than once stood in for organisations never served: C-SWEEP-06's shape on the search surface, a month earlier. None of the 117 is in either capture. The sorted listing of 2026-09-14 is the organisation population of record, and organisation now holds 13,151.
F-NA-1 — C-SWEEP-05 on the store (NRT-APPEND-01 Gate 4, outputs/nrt-append-01/NRT-APPEND-01-GATE-4-20260914T041918Z.md 05787d2ba59f82e1). The predecessor a75c7118659e125f held 20,677 delivery scope_entry rows on 77 organisations with no organisation row — the 77 of scope_organisation_unresolved above. Every one resolves under the 117 in 6956114f01106a89: scope_fact carries 0 unresolved rows, each figure read by an orphan query that saw a planted orphan first. The search defect's footprint in the store was those rows.
Scope is the largest table in the estate by rows and the reason it is its own database: it grows with the number of registrations, not with the number of products, and in the lifted store it and its six indexes are 96.3% of 6.708 GB.
5.1 · scope_entry_org — the same fact from TGA's second surface (NRT-APPEND-01, 2026-09-14; R-NA-2)¶
TGA serves a scope fact from two surfaces, and they are two shapes. A delivery body is per component and lists the
organisations; an org/scope body is per organisation and lists the components. scope_entry is the delivery shape and
scope_entry_org is the org/scope shape — two tables in register row rto, the same fact from two TGA surfaces, never one
table.
scope_entry |
scope_entry_org |
|
|---|---|---|
| source | per-component delivery body |
per-organisation org/scope body — one row per served value[] entry |
| the organisation | the served element's organisationId |
the ledger row that fetched the body — the body names none (R-SP-1) |
| status | status / status_label — the scope row's own |
component_status — the component's status, as org/scope serves it |
| columns it does not have | — | element_index · is_history · has_current_regulatory_decision · a scope-row status: org/scope serves none of them, so the columns do not exist — never NULL-filled |
| unique key | ux_scope_entry_fact (organisation_id_norm, component_nrt_id_norm, element_index, row_sha256) — the lift key, delivery's rule |
ux_scope_entry_org_fact (organisation_id_norm, component_nrt_id_norm, row_sha256), row_sha256 the digest of this served entry |
fetched_at · source_captured_at |
as lifted | the body's own ledger time |
status is a different attribute on each surface (NRT-APPEND-01 Gate 1 (d), measured on organisation 0022): of 3,414
facts matched across the two, org/scope's status equals scope_entry.status on 980 and equals the component's own status on
3,414. Mixing them in one column is forbidden: scope_fact carries scope_row_status (NULL on every org/scope row) and
component_status (NULL on every delivery row), and 0 rows cross (Gate 4).
The views. scope_fact = scope_entry UNION ALL scope_entry_org, with source_family ∈ delivery · org/scope;
scope_fact_current = its rows whose end date is NULL or after register_today; scope_organisation_unresolved is redefined over
scope_fact. scope_entry_current stays delivery-only, and find_rtos is engine code — its union with scope_entry_org is an
engine brief (Gate 2, choice 4). [2026-09-15, FIND-RTOS-01: done — find_rtos is a view of record, §5.2.] On 6956114f01106a89: scope_fact delivery 6,526,345 · org/scope 21,513;
scope_fact_current on 2026-09-14 delivery 1,344,274 · org/scope 766 (the current figures move with the date).
What scope_entry_org holds — 21,513 rows at ed5-2026-09-14-nrtappend01 (R-NA-3): the 117 organisations'
21,310 from their scope of record (31 from scope-single/, 86 from bodies/; 35 with none), and
Amendment 1's 203 facts on 31 organisations whose product the organisation's delivery rows lack, from their held
ed1 org/scope bodies. Not landed, by ruling: 761 body-only facts on a held product that differ by end date
(462 of them on 91371) — two captures two days apart disagreeing on currency, not a hole — and 102 not one-to-one
(C-SCOPE-02). Landing either would put two versions of one product's scope under one organisation.
DDL of record: outputs/nrt-append-01/NRT-APPEND-01-GATE-2-20260914T024718Z-scope_entry_org.sql 3b3b461d1663387b.
5.2 · find_rtos — the view of record (FIND-RTOS-01, 2026-09-15)¶
find_rtos is a VIEW of record over scope_fact (FIND-RTOS-01, 2026-09-15): LEFT JOIN, de-fanned, source_family carried; find_rtos_current beside it; measured through the view at 47.204 ms median and 51.068 ms max of 20 runs on the largest product, against the 50 ms wall — the figures of record, FIND-RTOS-01 Gate 4 (outputs/find-rtos-01/FIND-RTOS-01-GATE-4-20260915T034151Z.md 2f0415837da4be2f; 5,488 rows on 7e2b1297-0af7-4a65-9148-d7d0bed8e948, 5,461 delivery + 27 org/scope). Before the view existed, the same body as a query read 40.77 ms with legal_name_last (Gate 2) and 32.95 ms without it (Gate 1). No second write was made to the store: the index R-FR-2 ruled at Gate 4 was withdrawn at ATTRIB-01's close (outputs/find-rtos-01/VERDICT-FIND-RTOS-01-ATTRIB-01-CLOSE-2026-09-15.md da49488667b21ee1), and register row rto is the store as built. Its three scalar subqueries are side (c) of Lane Z-02 Gate 4 verbatim; the additions are FROM scope_fact, source_family, the scope columns, and legal_name_last (R-FR-1) — the most recent legal name regardless of currency, beside legal_name, which stays current-only. A consumer filters WHERE component_nrt_id_norm = ?; both halves of the union are served by an index (idx_scope_entry_component, idx_scope_entry_org_component).
What the union adds. scope_entry_org carries 21,513 rows for 113 organisations, 108 of which also have delivery rows. An organisation that holds a product only on its own scope page — organisation 40328 on ae939bb4-e061-46c8-8404-cb0de46c84fd, the Gate 1 control — is in find_rtos and was absent from side (c) over scope_entry.
NULL names. legal_name is NULL on 180 rows of 3 resolved, cancelled RTOs whose only legal name has ended; legal_name_last is NULL on 0 rows, equal to scope_organisation_unresolved.
Standing finding — the 50 ms wall (FIND-RTOS-01 ATTRIB-01, 2026-09-15; outputs/find-rtos-01/FIND-RTOS-01-ATTRIB-01-GATE-2-20260915T090916Z.md). The largest product reads at 47.204 ms median with a 51.068 ms max against a declared 50 ms wall (FIND-RTOS-01 Gate 4). One run in twenty crossed it. ATTRIB-01 (860af05c88d3eb37) measured the cause: the view's four correlated scalar subqueries each cost 6–10 ms per 5,488-row product, legal_name_last at −10.024 ms being the largest but not separable from rto_type or reg_start at a 2.938 ms floor. The temp B-tree on legal_name_last's ordering is worth at most the gap to its siblings, under that floor. No index or rewrite is warranted, and the wall is not a production figure: this path is served from a D1 promotion, not from the Mac, and the same file's own baseline drifted 9.7% between two sessions on one day. The wall is re-measured where it is served, at PROMOTE-01.
DDL of record: outputs/find-rtos-01/find-rtos-01-views-v1.sql c4a1a0f2503525c8.
6 · History — one rule, one derived database¶
Rule: no table in tga-nrt or tga-rto ever overwrites a dated row. Releases, mapping edges,
legal names, registrations, scope entries, decisions — all kept with their start and end dates.
"Show history off" is WHERE end_date IS NULL OR end_date > today. That is all the toggle is.
The register keeps history by superseded_at + edition, the content store's mechanism
(INGEST-APPEND-01 Phase 2, 2026-09-07). A body whose digest changed is re-parsed and its new rows
written beside the old ones under a new edition; the rows they replace are stamped
superseded_at, and the stamp is the fetched_at of the body that said otherwise about that
resource — not the build's clock, not a neighbour's. edition joins the key on release,
release_external_link and release_asset, because a column that says which edition a row belongs
to cannot distinguish two unless it is in the key. Array-valued children move as whole sets, not
deltas: a partial write would leave the old set half-current. An L1 body's releases[] is
itself an array-valued child and moves as a whole set — the whole array is lifted in the body's
own order, never element by element. release_index is array position within an edition, never
identity across editions: the same release can sit at a different index in the next body, and
INGEST-APPEND-02 measured exactly that (release_index +1 on 46 of 46 carried rows). A reader
that treats it as a key will silently pair the wrong two rows. Three _current views filter on
superseded_at IS NULL; release_current also keeps its older lane-of-record meaning — it joins
component_current on code and source lane.
A release's level-2 columns land as a new edition; the prior row is stamped by join to the body's ledger
fetched_at (CHILD-APPEND-01, R-CA1-2, 2026-09-14). The 175 releases NRT-APPEND-01 landed at ed5 with
specialisations_state no-body gained their level-2 bodies' columns as ed6 rows, and exactly those 175 ed5 rows
were stamped — the value read from the ledger row of the body that caused the new row, never from the clock; superseded
release rows 60 → 235, and release_current holds one ed6 row per release (outputs/child-append-01/CHILD-APPEND-01-GATE-4-20260914T072602Z.md afd36fcb96a68980).
release_current is current in the mirror's sense, never in TGA's: it never reads release.currency, and
16,242 of its 103,341 rows (15.7%) carry currency='replaced' (register of record
66ad976acfda6dc8; RELEASE-CURRENT-CONTRACT-01 close 8b28657a8dc56cfb, window/main-02 @ cb131f61).
A read that wants TGA's current adds currency = 'current': with release_current that is an exact
one-row key, 0 exceptions across 87,099 components (the same close). An earlier sentence here claimed
the view was current in TGA's sense as well; corrected 2026-09-11 (CANON-CORRECTIONS-02).
⚠ Two code sites read release_current by name — FROM release_current in
packages/engine-core/release.ts and content.ts — and cannot execute against their bound store.
apps/workspace/wrangler.jsonc's NRT_DB binding points at rto-nrt-db, which holds the old data
model and has no release_current: a generation mismatch
(CORRECTION-OF-RECORD-RELEASE-CURRENT-CONTRACT-01-THE-PITH-GENERATION 4ce793c0aa392fe1,
window/main-02 @ cb131f61). They read no contaminated row; a throw is the expected behaviour,
inferred and not observed. The fix is to unbind, handed to the prototype-disconnect brief.
⚠ Two things this mechanism does not yet cover, stated so they are not rediscovered.
release_link and release_content_bundle still key (release_id_norm, element_index) with
edition outside the key, and will meet the identical defect the first time an edition is appended
to either — owed to TGA-MODEL-03. And edition on a row is a label, not a foreign key: the ed1
rows read ed1-2026-08-28 while the ed1 row in the edition table has edition_id
2026-08-18T11:41:54+00:00. The two strings differ. The ed2 row's stamp and id DO match; the older
mismatch is recorded, not retro-fixed.
tga-history is the derived timeline across both sides, built locally by LANE-Z-03C (2026-09-04;
store-register.md row history) and recomputed from the databases after each pull. The Product
account's tga-history-01 (Edition 1, SEQ-MODEL-01) is an earlier build and is not this store;
promotion is PROMOTE-01's [2026-09-17, CANON-SWEEP-CLOSED-01]. It
answers "what changed between these two dates" and drives Compare. It is never the source of a fact.
7 · Size, against the ceilings¶
Every figure below carries its as-at and the artefact it was read from. The estimates this section carried until 2026-09-04 were wrong by 4.5× and 4.5× respectively, in the same direction, for the same reason — they priced the rows and forgot the indexes.
| database | measured / projected | as-at | source |
|---|---|---|---|
tga-nrt |
0.928 GB (928,423,936 B) measured, the register of record | 2026-09-12 | 8b84a2e98745b474 at the file — INGEST-02's successor (register row nrt); three tables added, see §1 |
tga-nrt |
0.873 GB (872,943,616 B) measured, of record at that date | 2026-09-11 | 4190424cbe933e5a at the file — VOCAB-BUILD-01 close ec22420785ac5f7b, re-read at the file at CANON-CORRECTIONS-04 Gate 1 |
tga-nrt |
0.145 GB measured | 2026-09-04 | e2-nrt-01.sqlite at the file (the superseded E2 store 2eff8cadfe218c21; row kept — this table carries an as-at per figure) |
tga-nrt-content |
3.980 GB (3,979,702,272 B) measured, the v3 store of record | 2026-09-05 | 2cbbd45df238a272 — CONTENT-V3-01 close 749bd8fba68caa98, re-read at the file at CANON-PASS-02 Gate 3 |
tga-nrt-content-legacy |
3.483 GB (3,483,017,216 B) measured, the legacy family | 2026-09-05 | 253c2bfd7b4769b4 — same close, same re-read |
tga-nrt-content |
3.806 GB at 16 of 36 types, measured | 2026-09-04 | 81af99aa1f184209 (superseded by the v3 pair above; row kept — this table carries an as-at per figure) |
tga-nrt-content |
3.328 GB at 16 of 36 types, measured | 2026-09-03 | 24dd0063bdf83a4f (superseded; row kept — this table carries an as-at per figure) |
tga-nrt-content |
6.720 GB projected at all 36 · 4.764 GB under §1.2 | 2026-09-03 | CANON-01 Gate 1 §3 |
tga-nrt-packaging |
0.868 GB measured, the database itself | 2026-09-04 | f04ff17c8e2bb640 |
tga-rto + tga-scope together |
6.708 GB measured, of which scope and its six indexes are 96.3% | 2026-09-03 | e93515bf7ed061d4 |
tga-history |
0.852 GB measured | 2026-08-19 | tga-history-01, Product |
Composed, against the 10 GB D1 ceiling and the 60% line — 0116 counted once, in
tga-nrt-packaging and not twice:
| without §1.2 | under §1.2 | |
|---|---|---|
tga-nrt-content (35 types) + tga-nrt-packaging, one database |
7.375 GB · 73.8% | 5.620 GB · 56.2% |
split — tga-nrt-content (35 types) |
6.029 GB · 60.3% | 4.274 GB · 42.7% |
split — tga-nrt-packaging |
1.346 GB · 13.5% | ≤1.346 GB (Q's own next DDL applies §1.2) |
How the projection was made, because a projection's method is its figure. The store of record
holds no body column, so an unparsed type's size cannot be read from it; it was read from the
bodies (§3). Two predictors were fitted over the 17 parsed types and compared out of sample by
leave-one-out: a simple bytes-per-body-byte ratio (median error 38.5%) and
bytes ≈ 1,471.7·sections + 3.353·body_bytes (median error 20.5%, max 107.6%). The fitted one
is carried. ⚠ A PROJECTION NAMES ITS ERROR. At +20% the unsplit, convention-applied figure is
6.744 GB and back over the 60% line — which is why the estate is six databases and not five.
Least sure on this section: 0000 (24,112 sections, 280.7 MB of body — 18% of all published
content bytes) and 0116 (201.6 MB) are 31% of the unparsed body bytes between them and are the
two types least like anything in the fit. If they parse denser than the fit expects, every figure
above is low.
Bodies in R2, uncapped. Growth under the diff model is only what changes per pull — megabytes, not gigabytes — provided scope inserts on change only (§1.1).
The ledger set is 20 ledgers (2026-09-12, BRIEF-LEDGER-V6): 15 _LEDGER.sqlite, the
pith-index-01 export, and 4 JSONL listings. It is enumerated by instrument, not by hand, and
pinned by digest at every emission — the completeness ledger's held figures are measured across
that whole set (§13), so the set's membership is part of what an edition asserts.
8 · What is signed here and what is not¶
Signed by this document: the estate's databases and their names — eight built (ADR-082, ADR-087 §A, ADR-091); one set, kept current by digest-diff, edition as a stamp; that the section index and the parsed rows share one database; that scope lives beside the RTO in its own database, and that packaging and the grid do too; that history is a rule plus one derived database; that every column is TGA's label or API field, and that our own columns follow §1.2's storage convention. Not signed: any brief, gate, or the order of the parsers. The addenda — one per NRT page type and one per RTO tab, each a labels-verbatim table from its screen cap — come next, one at a time.
9 · Least sure¶
- ~~The
scope_entrycolumns.~~ CLOSED 2026-09-03 — Lane Z read the whole DDL. All four site fields are columns; the one wrinkle isextent(§5). - ~~Usage recommendation and Classifications are not in the E2 register.~~ MEASURED 2026-09-04, and
the answer is in two halves. No column exists —
e2-nrt-01.componentcarriescode · code_norm · component_id · component_id_norm · type_code · type_label · title · source_lane · is_test_artefact · source_object_sha256and nothing else, and there is no taxonomy table. But both are held in bodies:usageRecommendationfor 87,099 distinct codes in the L1 register bodies, and classifications in 30,674classification/bodies. This is a build gap, not a fetch gap — LANDED 2026-09-05 at TGA-MODEL-03 (Gate 328c2f4b266bf871b, DDL of record37f735f5b941571f):component.usage_recommendationis populated on 87,107 component rows and NULL on the 38,741 accredited rows (§2.1), and classifications arecomponent_classificationand the two taxonomy tables, as §2.2 records. §13. - Whether
tga-historyshould be rebuilt locally during the stop. It exists only on the Product account today; the rule in §6 does not need it, but Compare does. Z3 owns this question and is drafted on Lane Z's close. [Answered 2026-09-04: LANE-Z-03C builttga-historylocally —store-register.mdrowhistory. Recorded 2026-09-17, CANON-SWEEP-CLOSED-01.]
10 · Speed is a UX imperative (Tim, 2026-09-02)¶
The prototype's type-ahead and in-page loading were unacceptably slow; that must not recur. Standing
constraints on every surface that reads these databases: type-ahead reads one small purpose-built
table (code · title · type · status) through the FTS5 prefix index, debounced, and never touches the
content or RTO databases; scope carries indexes on both join ends (RTO code, product id) and on the
current-status filter; every _norm join column is indexed; content search is FTS5, never LIKE;
the hot product header is one register query and is cached. Before 14 September: EXPLAIN QUERY
PLAN and a timing on the type-ahead and Find-RTOs queries at full population — a table scan on
either is a defect.
11 · Indexes are part of the schema, not an afterthought (Tim, 2026-09-02)¶
Indexing has been forgotten at every build because nothing required it. Three mechanisms make it structural:
- An index is declared in the table's DDL of record — the same file as its
CREATE TABLE. A DDL file that defines a table without its indexes is incomplete; a later "add index" migration is a defect report against that file, not a fix. - The schema-parity guard compares indexes as well as tables. A missing or extra index fails the
build the same way a missing column does. (
scripts/schema-parity-guard.mjs— its manifest and diff must covertype='index'insqlite_master; if it does not today, that is the first change.) -
Every brief that adds or alters a table names the queries that will read it, and Gate 4 runs
EXPLAIN QUERY PLANon each one at full population. A table scan on a named query fails the gate. The plan output is filed with the gate, not summarised. -
⚠ AN ESTIMATE PRICES INDEXES — on the disk and on the bill. Earned twice in two days and in two different currencies. On disk: §7's
tga-rtoestimate said 1–1.5 GB and the store came in at 6.708 GB, because the estimate priced 6.5 M rows and not the six indexes over them, which are 3.126 GB on their own. On the bill: a D1 job that inserted 200,000 rows was billed 0.60 M rows written, because D1 counts each index entry as a row written — a 60× miss on a stated estimate (CANON-01 Gate 1b §6, 2026-09-04). A size estimate that names only rows, and a cost estimate that names only inserts, are both incomplete by the same omission.
⚠ release_id_norm is a false friend (E13, Lane W Gate 1 §3; register row 106; runner R3b-18,
2026-09-06). In tga-nrt-content it is TEXT holding the dashed UUID, un-normalised, where
tga-nrt v1's is BLOB(16). A _norm name promises a normalisation the column must actually
carry; these two do not join on their names. A shape probe compares key representation, not column
names. (SUBSTRATE-NAME-MATCHES-SHAPE, standing-rules.md.)
Minimum index set for this model, to be declared in the DDL files: every _norm join column; on
scope_entry both join ends (RTO code, product id) and the current-status filter; the FTS5 tables
for the type-ahead (tga-nrt) and, when built, for content search (tga-nrt-content); on every
content row table its (release_id_norm, content_type_code) prefix; and on every table read
through a _current view, an index leading on that view's own correlation key. The last is not a
nicety: packaging_unit_current correlates on (bundle_id_norm, section_index), no index led on
it, and the view did not return in 60 seconds on a query the base table answers in 4.3 ms. One
added index — (bundle_id_norm, section_index, parser_version), 2.4 s to build — took it to
52.5 ms. A _current view without its correlation index is not slow; it is unusable, and it is
unusable silently (Lane Q Gate 4 §3, 2026-09-03).
12 · What was missing — found on a hard read, 2026-09-02, before the tap¶
Twelve things a data model has to state and this one did not. Each is small to add now and expensive to discover after a load. In order of consequence.
- The bundle is part of a content row's identity. A release can carry the same content type
twice: UEEAS0009 has Modification history
0012in both itsDefaultandAssesment Requirementsbundles. A key of(release, content type, position)collides. Every content row is keyed(release_id, bundle_id, content_type_code, item_index, …).unit_parseandevidence_itemtoday omit the bundle; that changes with the re-key under ALL-RELEASES-01. - A current-row resolution rule. Under the diff model and the parser-version rule, one section
can have rows from two parser versions and two body digests side by side. Readers need one
answer. Rule: the current rows for a section are those whose
body_sha256is the latest fetched body for that bundle and whoseparser_versionis the latest that parsed it; exposed as a view per content table (element_currentetc.) so no consumer writes the predicate itself. Older rows stay, readable by asking for a date or a version. - An RTO is identified by TGA's organisation id, never by its code — RULED AND BUILT
(RULING-RTO-KEY-01). The RTO page has a History → Code tab: codes change.
tga-rtois keyed onorganisation_id_norm; the code is a dated row inod_codes, exactly as TGA shows it. Tim's acceptance test is met on the store: all 90 superseded codes resolve throughorganisation_by_codeon an index lookup, an impossible code resolves 0 — where Z1 measured the same question as a JSON scan (Lane Z Gate 4 §1, 2026-09-03). Same principle for products:codeis stable in practice butidis the key TGA itself uses. - Every table is one of three classes, and the class is written on it: verbatim · derived ·
interpreted. Verbatim = TGA's words, keyed on position (elements, criteria, evidence, section
text). Derived = computed deterministically from verbatim (Number of required units, the
supersedes / superseded-by pair, the
unitgridcross-check,tga-history). Interpreted = a reading that could be wrong (choose-N packaging rules, licensing determinations) and carriesparser_version+confidence. A consumer may cite verbatim as TGA; it may not cite interpreted as TGA. - Inline markup inside text is content, not noise — and it is measured, not assumed.
section_text106,738 of 416,139 rows (25.6%) andmodification_history12,611 of 155,901 (8.1%) carry inline markup (2026-09-03). Stripping it would lose wording fidelity on a quarter of the prose corpus. "1420 weighting points" is bold in the source; unit codes inside prose are<ntr-tcref>links; there are<sub>,<sup>,<br>. Stripping to plain text loses wording fidelity and every link. Rule:textholds the plain string; a siblingspantable holds(row, start, end, kind, attrs)for emphasis and references, so the paragraph can be re-rendered with its bold and its links, and the plain text stays searchable. The round-trip test in §7 compares rendered output, so it catches loss here. - One universal reference table.
<ntr-tcref data-nrt-code data-nrt-type>occurs in prerequisites, entry requirements, packaging rules, skill-set requirements, modification histories and descriptions. Every occurrence becomes one row inreference: from(release, bundle, content type, row)to(code, type, title-as-published). "Which qualifications contain unit X" and "what does this unit's prerequisite chain look like" become one indexed query instead of a parse. - Vocabularies are tables with observed values, and an unknown value refuses. Content type
codes (36), bundle type codes (
0000001300120014), component types (6), release currency, usage recommendation, classification schemes, scope status. Each is avocab_*table loaded from what the corpus contains, and a parser meeting a value not in it stops with a reason — the same rule the layout parser already has for an unknown table shape. - Row digests for the non-body tables — BUILT, and their representation is now §1.2's. The diff
model (§1.1) compares bodies by digest. Register and RTO rows come from JSON, not bodies. Each
such row carries a canonical-JSON digest of its source element — in
tga-rtothese arerow_sha256(the element) andrecord_sha256(the body). They areBLOB(32)from each store's next DDL of record, per §1.2, which is where the 44% saving on a scope-shaped row comes from. - Absence is recorded, not inferred. A component, release, bundle or scope entry present in
the previous pull and missing from this one gets an
absent_fromdate on its row (never a delete). A section a bundle does not carry is ano sectionstate insection_parse; a section present but declaring "Not applicable" isdeclared-absentwith the declaration verbatim. Three states everywhere (§3), and the fourth — withdrawn by TGA — is a date, not a state.
⚠ section_state's remedy is EXISTS on the address (DDL delta row 34), and until that is
applied the view's raw declared-absent output is a BLOCK-ROW COUNT, never a section count (E18,
CONTENT-V3-01 G5-2, 2026-09-05). Any figure read off the un-remedied view is counting blocks and
must say so.
10. Images and non-text blocks. The evidence parser found 8 images across 3 documents and
recorded them as unrepresentable-block. Rule: the image bytes go to R2 keyed on their digest;
the block row keeps its position with block_kind='image' and the R2 key. Nothing that was in
the document is absent from the model, even when it is not text. Companion-volume figures
extend this (RULING-CVIG-FIGURES-01, Tim 2026-09-03): every figure is kept as an image plus a
written description of what it shows, produced by a model. The description is OURS, not TGA's —
it lives in the X-namespace with the other found-in-corpus items and is never presented as
something TGA said. A hand-read sample of the descriptions is the control in the brief that
produces them. F3.
11. Text normalisation is a comparison rule, not a storage rule. Stored text is the decoded
Unicode of the source, verbatim — entities decoded, nothing else touched (no whitespace
collapse, no quote straightening, no NFC). The round-trip diff normalises both sides the same
way (collapse whitespace, unify bullet glyphs and dashes) before comparing. This keeps the store
honest and the test fair.
12. The cold-rebuild statement. Every database must be reproducible from: the R2 bodies, the
DDL files of record, the parser scripts at their recorded versions, and the vocab tables. A
rebuild is run once, locally, before the first cloud load, and its counts reconciled to the
live stores — that is the proof that nothing in the databases exists only because someone
typed it. Retention for parser versions: current and previous live in D1; older to R2 as JSONL.
- The civil zone is
Australia/Sydney, stated once. The register publishes calendar dates with no zone; anything that needs an instant resolves them in the zone the national register runs on. The legislated DST rule is stated once, inmeta.dst_rule_source, and a zone table is a later brief — not a value copied into every row. - ⚠
NrtStatusIdis NOT the enum forscope_entry.status. The swagger declaresNrtStatusIdascurrent · pending · deleted · superseded · cancelled · nonCurrent(identical in TRAINING and ORGANISATION). Scope actually servescurrent4,037,551 ·nonCurrent893,805 ·cancelled837,281 ·withdrawn666,644 ·pending77,136 ·suspended14,032. Six declared, six observed, and they are different sixes:deletedandsupersededare declared and never occur;withdrawnandsuspendedoccur and are never declared. A validator built onNrtStatusIdrefuses 680,676 real rows and accepts two values that never appear. The vocabulary is loaded from what the corpus contains (item 7), never from the contract.
Also named, not elaborated: dates are ISO-8601 as TGA serves them, timezone as served; NULL means
absent-in-source, never default (ABSENT-NOT-DEFAULT-01); the fetch date on a row is the fetch
ledger's timestamp, not the load time; companion-volume listing rows (Title · Filetype · Status ·
Created date · Release, from the package page) join to Lane F's files on file digest; e2-nrt-01.release carries 103,335 rows for 103,327
distinct releases — eight release ids appear twice, once under training-L1 and once under
unit-L1, identical in every other column; and ten component code_norm values are non-unique for
the same reason. A count that joins release to component without collapsing them is inflated by
24 and says so nowhere (measured 2026-09-04) [2026-09-11, the register of record
4190424cbe933e5a, read-only: release 103,409 rows · 103,341 distinct ids · 103,349 not superseded ·
8 ids still appear twice among the non-superseded rows — the finding holds; the counts moved with §6's
editions (CANON-CORRECTIONS-04 Gate 1 d1f3348709322d10)]; a byte-pinned
control set is the acceptance fixture for every type — UEEAS0009 R2 · UEE63020 R4 · UEESS00210 R1 ·
10830NAT · RTO 0573 · ACM package — and a change to any parser re-runs all six.
Least sure, on this section: (5) — the span table is the right shape for fidelity but it is the one item here that adds real parser work; if the falsifier does not need bold and links rendered, it can be a later addition provided the text column is verbatim from day one, which it is.
13 · The tables the held bodies need (new, 2026-09-04)¶
Every family below is fetched and on disk today, with no rows. That state has a name —
held-bodies (§14) — and it is not owed: nothing has to be fetched for any of it. This is a
build gap, not a fetch gap, and BRIEF-TGA-MODEL-03 builds it.
One column decides the order of that work: does the body name its subject? A body that carries
its own resource id can become rows today. A body that lists only its contents — a delivery
body names the RTOs delivering a product without saying which product — cannot, until the fetch
ledger (pith-index-01, on the Product account) is exported and says which request produced it.
Fourteen of the families below are in the second class.
| table | parsed from | class | body names its subject? |
|---|---|---|---|
release.currency_change_date — a column, non-empty on 100.0% of 103,320 bodies |
releases |
verbatim | yes (id) |
release.approval_process · .isc_approval_date · .nqc_endorsement_date · .work_placement_hours · .ministerial_agreement_date — five nullable columns, NULL = absent-in-source |
releases |
verbatim | yes |
release_asset · release_external_link · release_specialisation (403 bodies) · release_packaging_information (2,214) |
releases |
verbatim | yes |
component.usage_recommendation — a column, for 87,099 codes |
the L1 register bodies | verbatim | yes (code) |
component_classification (isPrimary, purpose, purposeCode, scheme, schemeCode, schemeDescription, startDate, **endDate**, value, valueCode, valueDescription — eleven keys, not ten) |
classification — 30,674 |
verbatim | no |
taxonomy_industry_sector · taxonomy_occupation |
taxonomy-industry 1,244 · taxonomy-occupation 3,103 |
verbatim | no |
delivery_notification — in tga-scope, they are RTO × product facts (country, dateOfChange, nationalCode, notificationDate, notificationType, state) |
dnh 312,010 (+ delivery 89,738 folds in) |
verbatim | no — carries the product, never the RTO |
unitgrid_usage |
unitgridusage 27,156 |
verbatim | no |
completion_mapping and completion_usage — two tables, two shapes, two directions, as supersession_edge has one per direction |
completion 2,575 · completionusage 6,403 |
verbatim | no |
release_component |
release-components 929 |
verbatim | yes (releaseId per item) — a whole-register listing, 929 pages; not per-release |
accredited-course and accredited-unit releases, into release — the build gap below |
register-accredited-course-L1 19,439 · register-accredited-unit-L1 12,789 |
verbatim | yes (code, id) |
F-NA-2 — release_component is keyed by the MEMBER, so a component set aside leaves membership rows that no owner-keyed check sees (standing line; NRT-APPEND-01 Gate 4, outputs/nrt-append-01/NRT-APPEND-01-GATE-4-20260914T041918Z.md 05787d2ba59f82e1). release_component.component_code_norm names the component a release lists, not the release that lists it. When the 158 slash codes were set aside on 2026-08-17 (C-0914-05), 140 ed1 rows (lane release-components) went on naming them with a release_id that no release row carried: orphans by release_id 140 on 8b84a2e98745b474, 0 on bb3167d140d369c7, where all 140 resolve to the 175 new releases. Check release_component against release on release_id. A check keyed by the member reads these rows as level-2 children of the member's own release, which they are not (seat error 44, Gate 4 v2).
13a · Built by INGEST-02 (2026-09-12) — three of these families are now rows¶
| table | parsed from | class | body names its subject? |
|---|---|---|---|
component_recognition_manager (displayOrder, name, shortName, startDate, endDate, organisationWebAddress) — 123,537 rows over 125,838 codes; 3,938 codes served [] (every skill set and every training package), certified by two build_finding rows, never inferred |
recognitionmanager 125,838 |
verbatim | yes — via the ledger's subject_id |
release_training_package_usage (code, type, title, usageRecommendation, usageRecommendationLabel, releasesLabel) — 134,112 rows over 102,409 live releases of units, qualifications and skill sets; 932 training-package releases answer TGA's 400 and are build_finding rows, not table rows |
tpusage 102,409 |
verbatim | yes — via the ledger's subject_id, resolved to release_id_norm through release (component_code, release_number) WHERE superseded_at IS NULL, 102,409/102,409 with 0 ambiguous |
release_training_package_usage_release (package_release_number) — 500,453 rows from 500,454 array entries: the child is keyed as a SET, and TGA repeats 9.0 once for AURETR128\|1 → AUR (one fact served twice, a served-duplicate finding) |
tpusage |
verbatim | inherits its parent's subject |
A level-2 release body has FOUR children, not three (INGEST-02 Gate 2; the brief's own list named three): release_asset · release_content_bundle · release_external_link · release_link (href, rel). A served array dropped is a silent loss.
qualification_specialisation is body-driven from CHILD-APPEND-01 onward, lane qual-L2 (2026-09-14; R-CHILD-G3-5, R-CA1-3). Until then the table was carried from E2 row for row and no body-driven writer existed. outputs/child-append-01/child-append-01-rows-v1.py spec_rows writes one row per element of a release-detail body's specializations[], in array order (code, title, packagingInformation.core/elective/measure, workPlacementHours), source_lane qual-L2 (the DDL's CHECK set). Proven against a held E2 qual-L2 row set before it saw a CHILD body, and against the body element by element at Gate 4 (9 of 9, outputs/child-append-01/CHILD-APPEND-01-GATE-4-20260914T072602Z.md afd36fcb96a68980).
release_content_bundle — 35/140 is a shape, not a gap (standing line; CHILD-FETCH-01 Gate 3, CHILD-APPEND-01). Of the 175 releases whose level-2 bodies CHILD-FETCH-01 held, 35 list at least one content bundle (64 bundles, 377 items) and 140 list none — every one of the 140 serves all four child arrays empty. A release that lists no bundle has nothing to fetch: a count of releases without release_content_bundle rows is not a count of missing bundles.
taxonomy_industry_sector.description is nullable (R-CA1-7, CHILD-APPEND-01, 2026-09-14). TGA serves industry sector 1331 with no description key on RII30726 and RII41326; the absence is NULL, never '' (ABSENT-NOT-DEFAULT-01). The ed1 bodies never met such a sector, so the NOT NULL described ed1, not TGA's field. The DDL of record carries the amendment as a dated block (outputs/ingest-append-01/tga-nrt-v3.sql 213a21f7629eb8e4, the original line kept, struck by comment); applied in 767f89ac3afb6ffc only.
record_sha256 holds the SOURCE-OBJECT digest under that name (C-0912-04, logged not resolved). In scope_entry the column called record_sha256 carries the digest of the fetched body the row came from — the same thing source_object_sha256 names elsewhere in the estate. INGEST-02's 150 rows follow the existing convention rather than diverge from it; the naming is inconsistent across stores and is owed a ruling, not a silent fix.
⚠ Four corrections to this section (E1 + E17 folded; TGA-MODEL-03 Gates 1–4, 2026-09-05; the HELD-BODIES-RECON-01 close. Filed CANON-PASS-02-DELIBERATE).
- Packaging is three columns on
release, not a child table.release_packaging_informationabove is superseded by three columns carried onreleaseitself. A child table for a one-to-one fact is a join nobody needs. endDateiscomponent_classification's eleventh key — corrected in the row above. The ten-key list was short by one, and a key list that is short cannot be asserted complete.deliveryisscope_entry. Thedeliverybodies do not need a table of their own: they are RTO × product facts and that is whatscope_entryis.delivery_notificationkeeps its own row above because a notification is an event, not a scope fact.component_currency_periodis ONE table for every product type. Accredited products have currency periods, not releases — which is why 38,739 components carry no release row and why that is not a defect. One table, keyed on the component, covers accredited and non-accredited alike; a per-type table would model the register's shape rather than the fact.
⭐ The register gave 38,739 components no release row. 87,099 of 125,838 distinct codes carry a
release; the 38,739 that do not are the accredited families — accreditedCourse 19,439 and
accreditedUnit 19,302, zero releases between them, with a positive control run before the
zero was believed. Their L1 bodies are held and carry currencyPeriods and contentBundles.
A build gap, not a fetch gap.
Not tables — derived checks. scopesummary (3,674), org-training-packages (5,920) and
prerequisites (2,285) are TGA re-serving what the register or scope already holds. Each becomes a
round-trip check — compute from scope_entry / reference, compare to the body — and is
recorded render in the ledger with its check named. Same reasoning as the PDFs (§14).
A summary the source computes by an unpublished rule is checked against the source's own data, not against the summary (RULING-TGA-COMPUTED-SUMMARY-01, Tim 2026-09-11). Its rule is TGA's; a residual the source's own data explains is a standing finding about TGA, never an open against the mirror. (C-0911-22)
⭐ v6 — what held means, and the order in which the tests run (new, 2026-09-12; BRIEF-LEDGER-V6
Gates 1–5, on the advisory seat's four verdicts). Five rules, each law from this line on:
heldis a three-part test over the ledger set — never a directory listing. A body is held for a path when a row at that exactrequest_pathhasstatus = 200, itsstored_asexists on disk, its sha256 equals the row'sbody_sha256, and its length equals the row'sbytes. All three parts, or it is not held; andheldnever reaches across paths. C-0911-26 is law: a loader reads the index, never the directory. Aglob/listdir/scandirresult is not evidence that anything is held — it cannot distinguish a body a row names from one it does not.- A 200 whose body fails its family's reader is a broken row (C-0911-25). A path with no good
copy is
held-broken— neitherinnor held, owed a refetch or a ruling. A path with rows staysin; a path with one good copy and one broken copy is held on the good copy. - A file under a ledger's root that no row's
stored_asnames is counted and never loaded — it isunrequested-body, listed so the set is visible, and it enters no cell of the ledger. - Rows-based derivations run before bodies-based ones.
D6(for non-gap paths),D6aandD7aprecedeD10,D11andD12: a body is evidence only when nothing better exists. The exception is written intoD6itself — for agappath the partition IS its rows test, so that branch stays where LEDGER-01 put it; running a bare row count there would read another table's rows as evidence for content the mirror does not hold. - The register join key (C-0912-01): join on
code, or on the store's owncode_normrule — neverlower(code). A case-folding expression is not the store's key and will not stay true. The rule, measured (NRT-APPEND-01):code_norm= casefold · strip · collapse whitespace, and a/is kept (CEP/HCEG→cep/hceg). It holds on everycomponentrow ofbb3167d140d369c7— 0 disagree of 126,111; the 158 slash codes resolve by it to exactly themselves (158 of 158); and the served code resolves nothing where it differs from its norm (0 hits on 114) —outputs/nrt-append-01/NRT-APPEND-01-GATE-4-20260914T041918Z.md05787d2ba59f82e1.
⭐ v7 (2026-09-13; BRIEF-LEDGER-V7, Gates 1–5). The instruments read the store set from the store
register (docs/docs/ops/store-register.md), cited by §1; no instrument carries a mirror store path
or digest as a constant — repo artefacts (the swaggers, Lane Q's walk ledger) stay asserted by digest in
the instrument, the repo being their version control. A known-present row count is a fact about a store
and lives in the register with its citation (R-G2-V7-3). EVID gained two rows for the tables INGEST-02
built — not a rule surface, and measured both ways: 43/3 without them, 45/1 with. A digest cannot
protect a parse that picked the wrong row — the first parse rule selected a superseded content store and
every digest check still passed. A store digest is compared only under one root; cross-root equality is
asserted table by table over the measurement tables, provenance tables named and excluded (C-0913-01).
⭐ LIFT-AUDIT-01 (2026-09-13) — the lift key is now ENFORCED, and the six ed1 loaders are retired.
R-LA-2 asked for whichever section carries the "writer of record" language; canon carries none — that
phrase is in SEAT-CORRECTIONS-REGISTER-2026-09-04.md, not here — so the block sits in §13 beside the
tables, and this sentence says why.
- The lift key of record is
(organisation_id_norm, component_nrt_id_norm, element_index, row_sha256), andtga-rto.scope_entry.ux_scope_entry_factnow ENFORCES it (R-LA-4). Lane Z's lift keyed on a tuple that omittedelement_indexby design while the source store's ownUNIQUEincluded it, so 24 real rows for organisation 52583 were indistinguishable and dropped. A key is of record where it is enforced, not where it is written — the old index was the old key, and declaring a new one in prose while the DDL enforced the old one is how the two drifted apart. - The six ed1 loaders are retired as runnables and retained as cited specifications. Their files are byte-identical and stay so — three live instruments cite them by line and by digest, one inside a seat ruling, and a header would break every coordinate. This block is the register of retired runnables; there is no header to find in the files.
| script | functions | digest | defect | superseded by |
|---|---|---|---|---|
outputs/e2-build-2026-08-29/build.py |
w_nrt · w_spec · w_org · w_scope |
16a676780ac60198 |
silent skip (PARSE-SWEEP §(ii)) — w_scope drops an unparseable body without a word |
INGEST-APPEND-01/02, DNH-BUILD-01, INGEST-02 |
outputs/e2-build-2026-08-29/load-component-parent.py |
w |
3d1ad2a0220411b1 |
silent skip | ledger-driven loaders |
outputs/e2-build-2026-08-29/build-content-store.py |
w |
dcbfd05ccb21b4b2 |
silent skip | CONTENT-V3-01, ASSET-TEXT-BUILD-01 |
- The two
scopesummarymirror-side defects are closed, as measured. Restoring the 24 moves organisation 52583 from disagreeing to agreeing; superseding the orphan-sourced row changes no summary verdict (organisation 40328 is not among the disagreeing set). 65 organisations still disagree: 38date-windowand 6implicit-rowsit on TGA's side under RULING-TGA-COMPUTED-SUMMARY-01, and 21otherare unresolved and named, not attributed (C-0913-18, an open finding for a later recon). INGEST-02's 150 rows, across 146 organisations, moved no summary verdict at all.
14 · The completeness ledger — "whole" is emitted, never remembered (new, 2026-09-04)¶
Editions are per chain-date, and the manifest is the chain. PRIOR_EDITIONS.json names every
prior edition by filename and digest and is asserted whole before a byte is emitted, so the chain
survives a change of chain-date without a cross-date rule. (C-0912-03)
⚠ Open question, not law — whether the ledger should carry a check's RESULT for a render path.
Today D2's evidence is a fixed "owed check" string and the emitter has no field for a result, so a
check that has run cannot be said in the ledger to have run: LIFT-AUDIT-01 Gate 4 ran
scopesummary's round-trip and -ed7 still reads "owed check". Answering it is a rule-surface change
and belongs to a v8 of the ledger instruments. (C-0913-20)
Facet round-trip of 2026-09-14 — the third leg of done, measured after the append. 534 cells against TGA's own facets,
every non-match row named from TGA's listing. Pre-append (outputs/sweep-01/SWEEP-01-GATE-3-20260914T010043Z.md
5f675649134fa5d2): match 230 · currency 95 · tga-computed 154 · not-facetable 12 · gap 43, 37 of the gaps carrying the 158
fetch-excluded codes with bodies held and landing owed. Post-append, NRT-APPEND-01 landed (SWEEP-01-GATE-3-20260914T121552Z.md
f636d94b8f9a65af, instrument v4 f89ee2ec8a063546 — v3 plus a second body source, the 09-14 fetch ledger, for the 88 accredited
bodies not in ed1; seat verdict VERDICT-SWEEP-RECHECK-01-GATE-4-2026-09-14.md): match 299 · currency 62 · tga-computed 155 ·
not-facetable 12 · gap 6. Table (a) TGA-not-ours 0, (b) ours-not-TGA 0; listing 126,101 / 13,151 = ours. The 37 closed (29 match,
8 currency). The 6 are the rule-open cells and nothing else — our unproven derivations against a boolean TGA serves, oracle
in TGA's listing, not on the done path: C-SWEEP-02 HasLicensingInformation (2,121 components where our rule — a current
release carries a licensing assertion — is not TGA's flag; Δ ±1,041), C-SWEEP-03 IsCurrent (36 organisations; Δ ∓14),
C-SWEEP-04 HasCurrentRestrictionOrRegulatoryDecision (40; Δ ±10). No excluded, unexplained or ours-only row remains.
Together with TGA-MIRROR-COMPLETENESS-LEDGER-01-2026-09-14-ed9.md 39970370b57b098b (in 46 · held-bodies 0 · owed 0 ·
open_no_parse_row 0 on all 36 types) and LEDGER-ED9-EXCLUSIONS-2026-09-14.md 72787fe43d4a9cd7, all three legs of done are on
evidence as of main at this commit. Currency with TGA is not part of done. Government statistics is outside "everything TGA has" by RULING-GOVERNMENT-STATISTICS-01 (2026-09-16); it is a future Pith family, not a TGA source, and does not gate this certification. The TGA sweep is closed as at 2026-09-16: eight databases of record — tga-nrt · tga-nrt-content · tga-nrt-content-legacy · tga-nrt-packaging · tga-rto · tga-scope · tga-history · tga-cvig (store-register.md) — each certified at a Gate 4; the completeness ledger (-ed9, 39970370b57b098b) reads owed 0 · held-bodies 0 · open_no_parse_row 0 on all 36 types, with its three render paths and the documents named as exclusions, not closures, in LEDGER-ED9-EXCLUSIONS-2026-09-14.md 72787fe43d4a9cd7; currency with TGA is not part of done, and promotion to D1 is not part of done (Tim's definition, 2026-09-11: stored locally). Promotion is PROMOTE-01's. The instrument's "positive control: the join matched
125,838 both ways" and empty "By kind:" lines are emitter cosmetics (NAME-RETIRE-01).
TGA-MIRROR-COMPLETENESS-LEDGER-01 is the measure of whether this mirror is complete, and it is
a document derived from queries, regenerated by scripts/ledger-01-emit-v1.py — never hand-edited,
never typed. Its emission also replaces Part two of tga-source-atlas.md, by the same instrument.
It exists because the cure for "can anyone remember what we didn't do" is a filed ledger, not a
memory (PRIORITY-STATEMENT-FULL-MIRROR, Tim 2026-09-03).
Seven states, one per item, each assigned by a printed rule and never by hand:
| state | means |
|---|---|
| in | rows in a local store of record, keyed and digested, gate-certified |
| in flight | a brief is running or filed against it today |
| held-bodies | fetched, in the mirror, no rows and no parse — §13 |
| owed | TGA serves it, nothing holds it as rows or bodies, no brief filed |
| fetch-dependent | cannot close without contacting TGA |
| render | a TGA surface re-serving data held elsewhere; the round-trip test is the mirror |
| not TGA's / refused | outside the mirror by definition, or TGA answers 403 |
Over the 49 data paths TGA's own swagger declares: 6 in · 15 in flight · 14 held-bodies ·
13 owed · 1 fetch-dependent · 0 refused. The honest sentence is not that half the estate is
owed: 14 of 49 are already on disk and waiting for a table, ten of the thirteen owed are the
classification-scheme vocabulary tails, and recognitionmanager is the only path with nothing held
at all.
⚠ The fetch list is 7 releases, not 28,171 (corrected 2026-09-04; the seed and the first
emission both had this wrong). Of 103,327 releases, 75,156 have content and are held. Of the
remaining 28,171: 7 have no level-2 body at all, and those 7 are the fetch list. The other
28,164 carry a level-2 body in which TGA itself serves contentBundles: [] — a served
absence, already held and already correct, the release-grain case of item 9's declared-absent.
951 of those 28,164 also list a document asset, so for them the round-trip test has no rows to
run against; RULING-ASSET-ONLY-RELEASES-01 (Tim, 2026-09-04) holds those 951 documents as bodies in
R2, state held-bodies, no parse briefed. Every other release's document stays not held.
⚠ Served-absent is as-at the capture, not forever. All of it is measured on what TGA served on
28 August 2026; whether a fresh pull still returns [] is a question only the refresh pull answers.
The partition itself, its controls and the three refuted counter-hypotheses are not repeated here:
they are in the 2026-09-04 emission §5, and in outputs/canon-01/CANON-01-FETCHGAP-TEST.txt
7ff788ce77a1a864 and CANON-01-GATE2-CH3-USAGEREC.txt.
The fetch-ledger export — one read of pith-index-01 to a local file — is the second line after
the 7, because it is what turns every held-bodies row whose bodies do not self-identify into a
coverage figure.