TGA-SOURCE-ATLAS-01 — training.gov.au as a data source¶
Drafted 2026-08-20 (Architect seat, Cowork, Fable, session 33), on Tim's ask of the same day.
Status: APPROVED by Tim 2026-08-20 as a working reference — v1, from skeleton
95db533c47f3e763 / 12,180 B. A LIVING DOCUMENT, kept at hand: Part one is the stable
plain-language map, amended by dated addition; Part two is a named structure whose cells only
TGA-MIRROR-CORPUS-AUDIT-01 can fill, and it is generated by instrument, never hand-edited.
Home: docs/docs/ops/tga-source-atlas.md, beside substrate-findings.md, subordinate to it on
any figure conflict.
What this document is: the map of the source — what TGA looks like as a data pot, corner to corner, in language a person can hold. Written so that a future return to this territory (fault test, repair, addition) tools up to a known floor in one read.
What this document is never: a description of our programme. State, briefs, gates and priorities live in the anchor and the TMs. The moment this document grows a "what we're doing about it" section, it has become a second anchor, and both die of it.
Standing rule inherited from the anchor: a line here is not a measurement. Every figure carries its as-at and source. The coverage table (Part two) is the only part that claims currency, because an instrument regenerates it.
PART ONE — THE ATLAS (changes rarely; edit by dated addition)¶
1 · What TGA is¶
training.gov.au is the Australian Government's single national register of vocational education and training — every training product (qualifications, units, skill sets, accredited courses) and every registered training organisation, current and historical, for the whole country. There is no second register: even RTOs regulated by the two state regulators (VRQA in Victoria, TAC in WA) appear here. It is public, openly licensed, and served by an open REST API as well as the website.
One sentence to hold: TGA publishes facts; it does not publish judgements. What a fact means for compliance (may this RTO enrol a student in this product today?) is never in the data — that is what the rules engine (SEQ-MODEL-01) exists to compute.
2 · The surfaces¶
TGA offers its data through four distinct surfaces. A completeness claim must name which surface it ranges over.
- The REST API (
/api/…) — the primary machine surface. Returns JSON. Takes OData filters (server-side population selection). Publishes its own contract as a Swagger specification — the source's own enumeration of every endpoint it offers. The Swagger capture and diff-watch is TGA-UPSTREAM-INTELLIGENCE-CAPTURE-01; not yet built, as at 2026-08-20. - The website pages — human-rendered views over the same API. Matter to us twice: as the render target (Radar recreates the component page faithfully), and as the second enumeration instrument — the endpoints the site's own traffic calls is the usage-side denominator.
- The reports page — periodic published exports. A named tripwire surface; not a fetch family.
- The documents — the files behind each release of each component (the product as PDF,
DOCX and XML; assessment requirements; complete packages), plus companion volumes (CVIGs).
Reached via the release-detail endpoint, which is both the version record and the per-release
asset manifest (measured 2026-08-20; see Part two).
Added 2026-09-04: TGA's own documents are not held (Tim, 2026-09-03). The parsed rows plus
the round-trip test are the mirror of that content, so the files are a
render, not a gap. One exception, ruled 2026-09-04: 951 releases publish a document and no machine-readable content bundle, so for those the round trip has no rows to run against; those 951 documents are fetched and kept as bodies (RULING-ASSET-ONLY-RELEASES-01).
Adjacent sources that are NOT TGA and are treated as their own families with their own
captures: legislation.gov.au (the era instruments — the law the rules engine cites),
ASQA's site (the transition-extensions register, captured 2026-08-19), and the state regulator
portals (TAC, VRQA — named future families, parked). These live in regulatory-sources-01,
never in the TGA pith stores.
3 · The entity families, in plain language¶
- Training components — the products. 125,996 in the register of record (frozen population, 2026-08-15). Six types: qualification, unit, skill set, accredited course, accredited short course, module. Each carries its lifecycle status, classifications, and the supersession graph — which products replaced which, and whether the replacement was equivalent. Detail bodies exist per component and per release.
- Organisations — 13,146 records, of which 13,106 are RTOs (register body, 2026-08-15; the 40 non-RTOs are regulators, training-package developers and state authorities — extra data, not noise). The detail bodies carry the full corporate history inline: every address, contact, legal name, trading name, registration period and regulatory decision, each with start/end dates. The history is in the body — it does not need to be assembled from snapshots. (Parsed into 18 tables / 420,949 rows at A1-ORG-PARSE-01, 2026-08-20.)
- Scope — what each RTO may deliver/assess, per component, with start/end dates, an explicit/implicit flag and an extent ("Deliver and assess" vs "Assess only"). ~94% of scope is implicit — units flowing from on-scope qualifications — and TGA publishes the implicit lists rather than leaving them to be derived (sample measurement, n=200 bodies, 2026-08-20).
- Delivery notifications — where an RTO actually notified delivery of a component; bitemporal (the date notified and the date effective can differ by years). The largest family by volume (millions of pairs).
- Releases and documents — each component has numbered releases; each release lists its document assets with name, type, size and publish date. ~104,737 (component, release) pairs; the payload behind the manifests is ~225 GB (mean-basis, n=24, re-run owed — see Part two caveats).
- Classifications and taxonomies — ANZSCO, ASCED, industry/occupation taxonomies, and the scheme value dictionaries (7 endpoints, 0.40 MB — the full code lists including codes no component uses).
- Metadata — one endpoint publishing
dataExportSynchronisationDateTimeplus a build version: the source's only change signal, and it is global, not per-object.
Added 2026-09-04 — the mirror is two mirrors, and only one can answer a question about a
resource. pith-assets is keyed <family>/<sha256 of the body>.json: content-addressed,
carrying no resource id, verified on 36 objects across 12 families with 0 differing. It cannot be
asked "is there an object for release X". The content mirror's own listing
(tga-content-ed1-2026-08-28/_LISTING.jsonl, 105,911 rows over 75,156 distinct release_id) can.
Naming which mirror is being counted, before counting, is the whole of that figure's
reliability.**
4 · Temporal properties — the four facts that shape everything downstream¶
- No per-object change signal. No ETag, no Last-Modified. One global export-sync stamp. Re-sync designs poll the stamp; per-object hash comparison does not work (identical data re-serialises with different bytes — measured).
- History is embedded, not versioned. The detail bodies carry dated intervals inline; TGA does not publish "the body as it was last year". What we captured on a date is the only record of how it stood that date — which is why capture custody (the pith stores) exists.
- The register is bitemporal where it matters. Notification date ≠ effective date on delivery facts (observed lag up to five years). Rights are evaluated on effective dates; visibility questions use notification dates.
- All dates are register-published calendar dates. No timezones, no instants. Measured ISO-8601 across 624,491 values with zero exceptions (org family, 2026-08-20) — but that is one family's measurement, not a source-wide law; each parse re-measures.
5 · The honest limits¶
- TGA withholds by policy as well as by absence. One endpoint answered 403 "Completion
mapping is internal use only" on 2026-08-20. A refusal is a real answer and is recorded as one.
Corrected 2026-09-04: that refusal is falsified on the bodies. The 28 August build holds
2,575
completionand 6,403completionusagebodies — 8,978 in all, every one a JSON array of components, zero error bodies. Whatever answered 403 in August, what is on disk is data. A refusal is an observation with a date on it, never a property of an endpoint — and it is re-tested before it is carried forward. - Every claim is bounded by "what TGA publishes", never "what TGA knows."
- The source is not perfectly clean. Known warts, measured: inconsistent date formats in
ASQA's register (2/4-digit years, month-name variants, at least one malformed value); TGA
field-name traps (
registrationManageron organisations vsrecognitionManageron components — two fields, not a rename); status vocabularies that collapse distinct states when read by.nameinstead of.id. The full trap list lives insubstrate-findings.md; this atlas names that they exist so nobody assumes cleanliness. - No measured instance of TGA losing an artefact currently exists (the earlier sole-survivor claim was falsified 2026-08-19). Archive protection arguments rest on prudence, not on TGA fragility.
6 · How to read a "do we have X" answer¶
The three-states frame governs: A — in the corpus (ledgered, closure-proven; trust it) ·
B — only in the prototype's rto-nrt-db (no ledger, no closure; volume signal only, never
a figure) · C — never fetched. Part two ranges over the API surface and states A/C per
family; B exists only as history.
PART TWO — COVERAGE (generated by instrument; hand-editing is a defect)¶
Owner instrument: scripts/ledger-01-emit-v1.py e4fdf55fc452a74d, emission outputs/ledger/TGA-MIRROR-COMPLETENESS-LEDGER-01-2026-09-04.md 4bd17526896a44b4, as-at 2026-09-04T10:04:22.
1 · The seven states¶
The seed had six. held-bodies is the seventh, ruled at the Gate 1 verdict §2.1 because the stores needed it: a family whose objects are on disk and whose rows do not exist is not owed — nothing has to be fetched for it — and it is not in. The seed called those families owed, which is what a ledger without an instrument does.
| state | means |
|---|---|
| in | rows in a local store of record, keyed and digested, gate-certified |
| in flight | a brief is running or filed against it today |
| held-bodies | (new) fetched, in the mirror, no rows and no parse |
| owed | TGA serves it, nothing holds it as rows or bodies, no brief filed |
| fetch-dependent | cannot close without contacting TGA |
| render | a TGA surface re-serving data held elsewhere; the round-trip test is the mirror |
| not TGA's / refused | outside the mirror by definition, or TGA answers 403 |
The rule table, as applied — every state below comes from one of these and the rule id is printed with it:
| rule | state | assigned when |
|---|---|---|
S1 |
render | the path's class in the Gate 1 store is render |
S2 |
not TGA's | the path's class is not TGA's data |
S3 |
in flight | rows exist AND a brief is filed and running against it today |
S4 |
in | rows exist in a local store of record, keyed and digested |
S5 |
held-bodies | no rows, and >0 objects for it in a mirror listing (verdict §2.1) |
S6 |
fetch-dependent | no rows, no bodies, and it cannot close without contacting TGA |
S7 |
owed | no rows, no bodies, no brief — TGA serves it and nothing holds it |
⚠ refused has no members, and that is a measurement¶
The seed §8 counts refused 2 — /training/{code}/completion and /completionusage, on the atlas's 2026-08-20 record of a 403 "internal use only". Gate 2 was told to read one body of each. It read every body of both:
| family | bodies | error/dict bodies | empty arrays |
|---|---|---|---|
completion |
2,575 | 0 | 1 |
completionusage |
6,403 | 0 | 1 |
All 8,978 are JSON arrays of components. Zero are error bodies. Verbatim head of one of each:
completion/ [{"code":"VBQM230","endDate":"2012-12-31","hasPreRequisites":false,"isMandatory":false,"links":[{"rel":"training-component","href":"https://training.gov.au/api/training/vbqm230"}],"startDate":"2009-08
completionusage/ [{"code":"10183NAT","id":"e230ecba-e5c1-4a6a-87b6-991d24fff873","status":"nonCurrent","statusLabel":"Non-current","title":"Diploma of Education","type":"accreditedCourse"},{"code":"10326NAT","id":"2c6
Whatever answered 403 on 20 August, what is on disk from the 28 August build is data. The two paths are data paths in held-bodies; the refused state has no members today. This corrects the seed §8 and the Gate 1 verdict §2.2's "47 + 2 refused": the reconciliation stands, but the two are not refused.
2 · The denominator, and the seed's own arithmetic reconciled¶
88 paths across the eight captures, each at the digest the brief asserts. By class, from the Gate 1 store:
| class | paths |
|---|---|
| data path | 47 |
| render | 35 |
| not TGA's data | 4 |
| not TGA's / refused | 2 |
The seed §8 says "of the 88, 39 are renders or not TGA's data (12 EXPORT · 8 SEARCH · 3 REPORT · 1 FEEDBACK · 3 security-role paths · 10 export paths inside TRAINING/ORGANISATION · 2 files paths)". This instrument: render 35 + not TGA's data 4 = 39. The seed reconciles exactly, item by item.
and "of the remaining 49 data paths" — this instrument: 47 data paths + 2 classified refused = 49. The seed reconciles exactly. With the refusal falsified above, the 49 are data paths outright and the seed's total was right all along; only its last two labels were not.
3 · NRT — training products¶
| path | population | held | state | rule | source |
|---|---|---|---|---|---|
/training/{code} |
125,838 components (125,848 rows) | component · release · supersession_edge · component_parent in e2-nrt-01 |
in | S4 |
Gate 1 item 2 |
— its usageRecommendation and classifications |
125,838 | no column exists in e2-nrt-01.component, and no taxonomy table |
held-bodies | S5 |
Gate 1 item 2; bodies at classification/ 30,674 |
/training/{code}/releases |
125,838 components | 87,099 carry ≥1 release; releases-list/ holds 87,099 |
in | S4 |
Gate 1 item 3 |
/training/{code}/releases/{n} |
103,327 releases | 103,320 level-2 bodies (100.0%), 7 without | held-bodies | S5 |
Gate 1 item 6 |
…/document-bundle · /content/bundle/{id} · /content/item/{id} |
103,327 releases | 781,817 sections indexed; bodies in e2-content-01 |
in for fetched releases, fetch-dependent for the rest — §5 | S4/S6 |
Gate 1 items 4, 5 |
…/unitgrid |
12,902 walked | grid 10,269 · empty 2,633; join 12,902/12,902 | in flight — Q3 | S3 |
Gate 1 item 7 |
…/components · …/components/qualifications |
931 training packages | release-components/ 929 bodies |
held-bodies | S5 |
Gate 1 item 9 |
…/files · …/files/zip |
— | — | render — Tim ruled 2026-09-03 the PDFs are not held | S1 |
RULINGS-2026-09-03-TAPPED-THREE §4 |
/training/{code}/classification |
125,838 × n | classification/ 30,674 bodies, no rows |
held-bodies | S5 |
Gate 1 item 9 |
/taxonomyindustrysectors · /taxonomyoccupations |
2 vocabularies | taxonomy-industry 1,244 · taxonomy-occupation 3,103 |
held-bodies | S5 |
Gate 1 item 9 |
/training/{code}/prerequisites |
unmeasured | prerequisites/ 2,285 bodies; the content-side reference rows are in |
held-bodies | S5 |
Gate 1 item 9 |
/training/{code}/delivery |
unmeasured | delivery/ 89,738 bodies |
held-bodies | S5 |
Gate 1 item 9 |
/training/{code}/unitgridusage |
unmeasured | unitgridusage/ 27,156 bodies |
held-bodies | S5 |
Gate 1 item 9 |
/training/{code}/recognitionmanager · /restrictions |
unmeasured | comp-restrictions/ 3; no family for recognitionmanager |
held-bodies / owed | S5/S7 |
Gate 1 item 9 |
/training/{code}/completion · /completionusage |
unmeasured | completion/ 2,575 · completionusage/ 6,403 bodies, all data, none refused |
held-bodies | S5 |
Gate 2 §1 |
/releases/{releaseId}/training |
derivable | release.release_id_norm join |
in (derived) | S4 |
model §1 |
⭐ The register forgot to give 38,739 components a release. 87,099 of 125,838 distinct codes carry a release row; the 38,739 that do not are the accredited families — accreditedCourse 19,439 and accreditedUnit 19,302 rows, zero releases between them, with a positive control run before the zero was believed. Their L1 register bodies ARE held: register-accredited-course-L1 19,439 · register-accredited-unit-L1 12,789. This is a build gap, not a fetch gap.
4 · NRT content — the section types, parsed into rows¶
781,817 sections held over 36 content type codes. Held vs parse rows, per code, from the store of record:
| code | sections | parse rows | no-section rows |
open |
|---|---|---|---|---|
0012 |
104,754 | 104,754 | 0 | 0 |
0118 |
58,537 | 58,542 | 5 | 0 |
0103 |
58,494 | 58,494 | 0 | 0 |
0102 |
52,695 | 52,695 | 0 | 0 |
0121 |
46,470 | 46,470 | 0 | 0 |
0011 |
38,629 | 3,680 | 2 | 34,949 |
0200 |
30,985 | 30,985 | 0 | 0 |
0113 |
30,502 | 31,808 | 1,306 | 0 |
0120 |
30,500 | 31,806 | 1,306 | 0 |
0104 |
30,497 | 31,804 | 1,307 | 0 |
0112 |
30,476 | 0 | 0 | 30,476 |
0100 |
28,048 | 0 | 0 | 28,048 |
0111 |
28,043 | 0 | 0 | 28,043 |
0124 |
28,039 | 0 | 0 | 28,039 |
0125 |
27,995 | 0 | 0 | 27,995 |
0108 |
27,993 | 0 | 0 | 27,993 |
0119 |
26,645 | 0 | 0 | 26,645 |
0000 |
24,112 | 0 | 0 | 24,112 |
0001 |
15,745 | 19,425 | 90 | 0 |
0123 |
11,103 | 0 | 0 | 11,103 |
0116 |
10,077 | 0 | 0 | 10,077 |
0117 |
10,069 | 10,069 | 0 | 0 |
0110 |
9,604 | 9,604 | 0 | 0 |
0126 |
5,454 | 5,454 | 0 | 0 |
0127 |
5,452 | 5,462 | 10 | 0 |
0128 |
5,451 | 5,462 | 11 | 0 |
0109 |
4,620 | 0 | 0 | 4,620 |
0107 |
481 | 481 | 0 | 0 |
0010 |
74 | 0 | 0 | 74 |
0201 |
64 | 0 | 0 | 64 |
0002 |
48 | 0 | 0 | 48 |
0105 |
48 | 0 | 0 | 48 |
0122 |
47 | 0 | 0 | 47 |
0106 |
46 | 0 | 0 | 46 |
0202 |
10 | 0 | 0 | 10 |
0203 |
10 | 0 | 0 | 10 |
0011 open: 34,927 releases — the largest single open figure in the store, reproducing the BRIEF-TGA-MODEL-02 Gate 5 close.
5 · Fetch-dependent — and it is SEVEN releases, not twenty-eight thousand¶
28,171 releases of 103,327 have no row in the store of record's section. The first emission of this ledger called all 28,171 of them "never fetched". That was wrong in kind, and the partition below is why.
| cell | releases |
|---|---|
level-2 body held, contentBundles NON-EMPTY, and sections in the store |
75,156 |
level-2 body held, contentBundles non-empty, but no sections |
0 |
level-2 body held, contentBundles EMPTY, but sections in the store |
0 |
level-2 body held, contentBundles EMPTY, no sections — a SERVED ABSENCE |
28,164 |
| in the store with no level-2 body held | 0 |
| no level-2 body at all — ⭐ THE FETCH LIST | 7 |
A four-cell partition with zero exceptions. 28,164 releases carry a level-2 body in which TGA itself serves contentBundles: [] — the release-grain case of declared-absent. Nothing has to be fetched for them and nothing is missing. The fetch list is 7 releases.
Three counter-hypotheses were tested against the served-absent population and all three fail. They are not merely older releases of otherwise-bundled components: 28,163 of 28,164 belong to components where every release is empty. TGA does not hold the content behind a document instead: 100.0% of bundled releases list a document asset against 3.4% of these. And they are not current — measured against TGA's own usageRecommendation, held in the L1 register bodies for 87,099 codes, 0 of 28,164 belong to a component TGA marks current (61.6% deleted, 38.4% superseded), where the same predicate finds 22,715 current components on the bundled side. Cited, not recomputed here: outputs/canon-01/CANON-01-FETCHGAP-TEST.txt 7ff788ce77a1a864 and outputs/canon-01/CANON-01-GATE2-CH3-USAGEREC.txt.
⚠ 951 of the served-absent releases list a document asset — TGA publishes a downloadable document and no machine-readable bundle, so for those the round-trip test has no rows to run against. RULING-ASSET-ONLY-RELEASES-01 (Tim, 2026-09-04): those 951 documents are fetched and kept as bodies in R2, state held-bodies, no parse briefed. They are the second fetch-list line.
⚠ Served-absent is as-at the capture, not forever. Every figure above is measured on what TGA served on 28 August 2026. Whether a fresh pull still returns [] is a question only the refresh pull can answer.
The type table below ranges over the served-absent population and is retained for its shape; its own population is named in outputs/canon-01/CANON-01-GATE2-FANOUT.txt.
| type | current | replaced | total |
|---|---|---|---|
| qualification | 2,805 | 0 | 2,805 |
| skillSet | 0 | 2 | 2 |
| trainingPackage | 129 | 0 | 129 |
| unit | 25,234 | 1 | 25,235 |
| TOTAL | 28,168 | 3 | 28,171 |
The 25,235 is confirmed as a number and retired as a sentence. The Time Machine carried "25,235 sectionless current unit releases" and no filed artefact held it. Measured: unit releases with no section row = 25,235 exactly, but 25,234 current and 1 replaced. The figure was right; its population name was wrong by one row.
Controls, both reproducing: qualification releases with no 0116 2,835 (Q1's 2,835) · with no section at all 2,805 (-01's 2,805).
⚠ Which mirror answers this. pith-assets keys are <family>/<sha256 of the body>.json — content-addressed, carrying no resource id, verified on 36 objects across 12 families with 0 differing. It cannot be asked "is there an object for release X". tga-content-ed1-2026-08-28/_LISTING.jsonl can: 105,911 rows over 75,156 distinct release_id. Naming the path before counting from it is the whole of this row's reliability.
6 · RTO — organisations¶
⚠ Two databases answer to e2-rto-01 and only one is the register. The .local-stack store bound to that name holds 3 qualifications and 140 units — a prototype shape with no organisation table and no scope_entry. Every figure here is from ~/mirror/e2-stores-2026-09-01/e2-rto-01.sqlite.
| path | population | held | state | rule |
|---|---|---|---|---|
/organisation/{code} and its eleven tabs |
13,034 organisations | organisation + 17 od_* families; org-L1/ 13,034 bodies |
in flight — Lane Z | S3 |
/{code}/scope · /scopesummary |
6,526,449 scope rows | 6,526,449; scopesummary/ 3,674 bodies |
in flight — Lane Z | S3 |
/{code}/deliverynotificationhistory/{trainingcode} |
the largest family | dnh/ 312,010 bodies, no rows, and the model names no table |
held-bodies — was owed in the seed |
S5 |
/{code}/training-packages |
unmeasured | org-training-packages/ 5,920 bodies |
held-bodies | S5 |
/organisation-security-roles |
— | — | not TGA's | S2 |
/api/security-roles |
— | 1 body, metadata family |
not TGA's — declared by the METADATA swagger; TGA's own access-control vocabulary, not register data (C-0912-02) | S2 |
2 × scope/export/… |
— | — | render | S1 |
7 · held-bodies — what is on disk and has never become a row¶
This section did not exist in the seed, because the state did not. Every family below was fetched by the E2 build and holds objects today; none has rows.
| family | objects | coverage measurable locally? |
|---|---|---|
dnh |
312,010 | no — the body names its contents, never its subject |
releases |
103,320 | yes — the body names its subject (id) |
delivery |
89,738 | no — the body names its contents, never its subject |
releases-list |
87,099 | yes — the body names its subject (id (of each release)) |
register-unit-L1 |
56,725 | not examined at this gate |
classification |
30,674 | no — the body names its contents, never its subject |
unitgridusage |
27,156 | no — the body names its contents, never its subject |
unit-L1 |
24,913 | yes — the body names its subject (code + id) |
org |
24,110 | not examined at this gate |
register-accredited-course-L1 |
19,439 | not examined at this gate |
org-L1 |
13,034 | yes — the body names its subject (code + organisationId) |
register-accredited-unit-L1 |
12,789 | not examined at this gate |
register-qualification-L1 |
6,870 | not examined at this gate |
completionusage |
6,403 | no — the body names its contents, never its subject |
org-training-packages |
5,920 | not examined at this gate |
register-skill-set-L1 |
3,679 | not examined at this gate |
scopesummary |
3,674 | no — the body names its contents, never its subject |
taxonomy-occupation |
3,103 | no — the body names its contents, never its subject |
completion |
2,575 | no — the body names its contents, never its subject |
qual-L2 |
2,345 | not examined at this gate |
qual-L3 |
2,345 | not examined at this gate |
prerequisites |
2,285 | no — the body names its contents, never its subject |
unitgrid |
2,156 | not examined at this gate |
taxonomy-industry |
1,244 | no — the body names its contents, never its subject |
training-L1 |
1,174 | not examined at this gate |
release-components |
929 | yes — the body names its subject (releaseId) |
cv-list |
337 | not examined at this gate |
register-training-package-L1 |
259 | not examined at this gate |
metadata |
12 | not examined at this gate |
comp-restrictions |
3 | not examined at this gate |
cricos |
1 | not examined at this gate |
search-organisation |
1 | not examined at this gate |
search-training |
1 | not examined at this gate |
⚠ The Gate 1 verdict §5's least-sure, answered, and the answer is a split, not a yes or a no. Five families self-identify and their coverage is measurable locally today: releases, unit-L1, org-L1, release-components, releases-list. The rest name their contents and never their subject — a delivery body lists the RTOs delivering a product without saying which product; links is empty on every one read. For those, coverage needs the fetch ledger (pith-index-01, cloud), which is therefore a dependency for some families and a check for others, not a blanket dependency.
8 · The number, honestly — and the seed graded¶
The seed's §8 gave three counts over the 49 data paths, hand-derived: in or in flight 22 · owed 24 · fetch-dependent 1 (plus refused 2). Every one of the 49 is assigned a state below by the rule table, from the evidence named beside it, and the counts are SELECTed from that assignment — not typed.
| path | evidence | rule | state |
|---|---|---|---|
/api/content/bundle/{id} |
store of record.section | S4 |
in |
/api/content/item/{id} |
store of record.section | S4 |
in |
/api/classification-purpose |
— | S7 |
owed |
/api/classification-purpose/{code} |
— | S7 |
owed |
/api/metadata |
— | S7 |
owed |
/api/nrt-classification-scheme |
— | S7 |
owed |
/api/nrt-classification-scheme/{code} |
— | S7 |
owed |
/api/nrt-classification-scheme/{code}/values |
— | S7 |
owed |
/api/nrt-classification-scheme/{schemeCode}/values/{code} |
— | S7 |
owed |
/api/rto-classification-scheme |
— | S7 |
owed |
/api/rto-classification-scheme/{code} |
— | S7 |
owed |
/api/rto-classification-scheme/{code}/values |
— | S7 |
owed |
/api/rto-classification-scheme/{schemeCode}/values/{code} |
— | S7 |
owed |
/api/organisation/{code} |
e2-rto-01.organisation; Lane Z | S3 |
in flight |
/api/organisation/{code}/addresses |
od_addresses | S3 |
in flight |
/api/organisation/{code}/classification |
od_classifications | S3 |
in flight |
/api/organisation/{code}/contacts |
od_contacts | S3 |
in flight |
/api/organisation/{code}/cricoscode |
od_cricos_codes — 0 rows, served empty by TGA; the emptiness is the held fact | S4 |
in |
/api/organisation/{code}/deliverynotificationhistory/{trainingcode} |
dnh/ 312,010 bodies |
S5 |
held-bodies |
/api/organisation/{code}/legalname |
od_legal_names | S3 |
in flight |
/api/organisation/{code}/registration |
od_registrations | S3 |
in flight |
/api/organisation/{code}/registrationmanager |
od_registration_managers | S3 |
in flight |
/api/organisation/{code}/regulatorydecision |
od_regulatory_decisions | S3 |
in flight |
/api/organisation/{code}/restrictions |
od_restrictions | S3 |
in flight |
/api/organisation/{code}/role |
od_roles | S3 |
in flight |
/api/organisation/{code}/scope |
e2-rto-01.scope_entry | S3 |
in flight |
/api/organisation/{code}/scopesummary |
scopesummary bodies; Lane Z | S3 |
in flight |
/api/organisation/{code}/tradingname |
od_trading_names | S3 |
in flight |
/api/organisation/{code}/training-packages |
org-training-packages/ 5,920 bodies |
S5 |
held-bodies |
/api/organisation/{code}/webaddress |
od_web_addresses | S3 |
in flight |
/api/releases/{releaseId}/training |
e2-nrt-01.release (derived) | S4 |
in |
/api/training/{code} |
e2-nrt-01.component | S4 |
in |
/api/training/{code}/classification |
classification/ 30,674 bodies |
S5 |
held-bodies |
/api/training/{code}/completion |
completion/ 2,575 bodies |
S5 |
held-bodies |
/api/training/{code}/completionusage |
completionusage/ 6,403 bodies |
S5 |
held-bodies |
/api/training/{code}/delivery |
delivery/ 89,738 bodies |
S5 |
held-bodies |
/api/training/{code}/prerequisites |
prerequisites/ 2,285 bodies |
S5 |
held-bodies |
/api/training/{code}/recognitionmanager |
— | S7 |
owed |
/api/training/{code}/releases |
e2-nrt-01.release | S4 |
in |
/api/training/{code}/releases/{releaseNumber} |
releases/ 103,320 bodies |
S5 |
held-bodies |
/api/training/{code}/releases/{releaseNumber}/components |
release-components/ 929 bodies |
S5 |
held-bodies |
/api/training/{code}/releases/{releaseNumber}/components/qualifications |
release-components/ 929 bodies |
S5 |
held-bodies |
/api/training/{code}/releases/{releaseNumber}/document-bundle |
store of record.section for fetched releases; the rest is the fetch gap | S6 |
fetch-dependent |
/api/training/{code}/releases/{releaseNumber}/tpusage |
— | S7 |
owed |
/api/training/{code}/releases/{releaseNumber}/unitgrid |
Lane Q's walk; Q3 loads it | S3 |
in flight |
/api/training/{code}/restrictions |
comp-restrictions/ 3 bodies |
S5 |
held-bodies |
/api/training/{code}/taxonomyindustrysectors |
taxonomy-industry/ 1,244 bodies |
S5 |
held-bodies |
/api/training/{code}/taxonomyoccupations |
taxonomy-occupation/ 3,103 bodies |
S5 |
held-bodies |
/api/training/{code}/unitgridusage |
unitgridusage/ 27,156 bodies |
S5 |
held-bodies |
| the seed said | this emission says, derived | why it moved |
|---|---|---|
| in or in flight 22 | 21 (in 6 · in flight 15) | the register, content and RTO rows |
| owed 24 | owed 13 · held-bodies 14 | the seventh state: 14 of the seed's owed paths have their bodies on disk already and need a parser, not a fetch |
| fetch-dependent 1 | 1, and now with a size: 7 releases (was reported as 28,171 in the 2026-09-03 emission) | the seed named the row; the instrument counted it; the partition corrected it — 28,164 of the 28,171 are served absences, not a fetch |
| refused 2 | 0 | falsified on 8,978 bodies, §1 |
| — | total 49 of 49 data paths | every data path has a state |
The honest sentence the seed was reaching for: it is not that half the data paths are owed. It is that 14 of 49 (29%) are already on disk and waiting for a table — a parser job, not a fetch — one is a fetch of known size (28,171 releases), and only 13 have nothing held at all, of which 10 are the small classification-scheme and classification-purpose vocabularies.
The post-14-September list, in order¶
- The fetch list — 7 releases with no level-2 body held. Not 28,171: the other 28,164 are served absences (§5). 1b. The 951 asset-only releases — a document published with no content bundle (RULING-ASSET-ONLY-RELEASES-01). 1c. (retired figure, kept for the audit trail) the former "fetch gap" of 28,171 releases with no section row, 28,171 of them recorded as never fetched: unit 25,235 · qualification 2,805 · trainingPackage 129 · skillSet 2.
- The fetch-ledger export — one read of
pith-index-01to a local file under SPEND-ENVELOPE-01, which turns everyheld-bodiesrow whose bodies do not self-identify into a coverage figure. Ruled at the Gate 1 verdict §2.1. - The refresh pull, the
/metadatasentinel, and a refresh policy per family — seed §7, none built.
Before any of that, and needing no fetch at all¶
- Accredited-course and accredited-unit releases from their held L1 bodies — 38,739 components with no release row.
- The level-2 release fields —
currencyChangeDate·assets[]·externalLinks[]·specializations[]·packagingInformation— from 103,320 bodies already on disk, into tables the model must first name. - Classification and taxonomy per component from 30,674 + 4,347 bodies.
- Delivery notifications — 312,010
dnhand 89,738deliverybodies, into a table the model must first name.
Emitted 2026-09-04T10:04:22 by ledger-01-emit-v1.py e4fdf55fc452a74d. Every cell is a row of the Gate 1 store. To regenerate: re-run Gate 1's instrument, then this one.
Regeneration rule: the audit emits this table whole, with its own as-at, and the emission replaces the section between the PART TWO header and this line. A hand edit to any cell is a defect of the same class as back-filling provenance.
Status and open items¶
- ~~Tim reads Part one for domain wrongness.~~ DISCHARGED — approved 2026-08-20.
- OPEN: the audit brief (TGA-MIRROR-CORPUS-AUDIT-01) names this document's Part two as a deliverable when it is written. Until the audit runs, Part two is seeded knowledge with no completeness claim.
- ~~Home and hierarchy ruling.~~ DISCHARGED — ruled 2026-08-20: home as in the header; subordinate to substrate-findings on any figure conflict (findings are the measurement record; this is the map drawn from them).
Least sure¶
Whether Part one's plain-language register stays plain under maintenance. Every editor of this file will know more than its intended reader, and the drift pressure is one-directional — toward the jargon of whoever last touched it. The cure attempted here is structural (the volatile parts are machine-generated, so humans mostly never need to edit), but Part one has no such guard. What would make me wrong about the document's whole value: if in six months it reads like substrate-findings, it has failed at the only thing it exists for.
— Claude (Architect seat, Cowork, Fable, session 33), 2026-08-20. Approved by Tim the same day; a working reference, amended by dated addition.