Skip to content

TGA-SOURCE-ATLAS-01 — training.gov.au as a data source

Drafted 2026-08-20 (Architect seat, Cowork, Fable, session 33), on Tim's ask of the same day. Status: APPROVED by Tim 2026-08-20 as a working reference — v1, from skeleton 95db533c47f3e763 / 12,180 B. A LIVING DOCUMENT, kept at hand: Part one is the stable plain-language map, amended by dated addition; Part two is a named structure whose cells only TGA-MIRROR-CORPUS-AUDIT-01 can fill, and it is generated by instrument, never hand-edited. Home: docs/docs/ops/tga-source-atlas.md, beside substrate-findings.md, subordinate to it on any figure conflict.

What this document is: the map of the source — what TGA looks like as a data pot, corner to corner, in language a person can hold. Written so that a future return to this territory (fault test, repair, addition) tools up to a known floor in one read.

What this document is never: a description of our programme. State, briefs, gates and priorities live in the anchor and the TMs. The moment this document grows a "what we're doing about it" section, it has become a second anchor, and both die of it.

Standing rule inherited from the anchor: a line here is not a measurement. Every figure carries its as-at and source. The coverage table (Part two) is the only part that claims currency, because an instrument regenerates it.


PART ONE — THE ATLAS (changes rarely; edit by dated addition)

1 · What TGA is

training.gov.au is the Australian Government's single national register of vocational education and training — every training product (qualifications, units, skill sets, accredited courses) and every registered training organisation, current and historical, for the whole country. There is no second register: even RTOs regulated by the two state regulators (VRQA in Victoria, TAC in WA) appear here. It is public, openly licensed, and served by an open REST API as well as the website.

One sentence to hold: TGA publishes facts; it does not publish judgements. What a fact means for compliance (may this RTO enrol a student in this product today?) is never in the data — that is what the rules engine (SEQ-MODEL-01) exists to compute.

2 · The surfaces

TGA offers its data through four distinct surfaces. A completeness claim must name which surface it ranges over.

  1. The REST API (/api/…) — the primary machine surface. Returns JSON. Takes OData filters (server-side population selection). Publishes its own contract as a Swagger specification — the source's own enumeration of every endpoint it offers. The Swagger capture and diff-watch is TGA-UPSTREAM-INTELLIGENCE-CAPTURE-01; not yet built, as at 2026-08-20.
  2. The website pages — human-rendered views over the same API. Matter to us twice: as the render target (Radar recreates the component page faithfully), and as the second enumeration instrument — the endpoints the site's own traffic calls is the usage-side denominator.
  3. The reports page — periodic published exports. A named tripwire surface; not a fetch family.
  4. The documents — the files behind each release of each component (the product as PDF, DOCX and XML; assessment requirements; complete packages), plus companion volumes (CVIGs). Reached via the release-detail endpoint, which is both the version record and the per-release asset manifest (measured 2026-08-20; see Part two). Added 2026-09-04: TGA's own documents are not held (Tim, 2026-09-03). The parsed rows plus the round-trip test are the mirror of that content, so the files are a render, not a gap. One exception, ruled 2026-09-04: 951 releases publish a document and no machine-readable content bundle, so for those the round trip has no rows to run against; those 951 documents are fetched and kept as bodies (RULING-ASSET-ONLY-RELEASES-01).

Adjacent sources that are NOT TGA and are treated as their own families with their own captures: legislation.gov.au (the era instruments — the law the rules engine cites), ASQA's site (the transition-extensions register, captured 2026-08-19), and the state regulator portals (TAC, VRQA — named future families, parked). These live in regulatory-sources-01, never in the TGA pith stores.

3 · The entity families, in plain language

  • Training components — the products. 125,996 in the register of record (frozen population, 2026-08-15). Six types: qualification, unit, skill set, accredited course, accredited short course, module. Each carries its lifecycle status, classifications, and the supersession graph — which products replaced which, and whether the replacement was equivalent. Detail bodies exist per component and per release.
  • Organisations — 13,146 records, of which 13,106 are RTOs (register body, 2026-08-15; the 40 non-RTOs are regulators, training-package developers and state authorities — extra data, not noise). The detail bodies carry the full corporate history inline: every address, contact, legal name, trading name, registration period and regulatory decision, each with start/end dates. The history is in the body — it does not need to be assembled from snapshots. (Parsed into 18 tables / 420,949 rows at A1-ORG-PARSE-01, 2026-08-20.)
  • Scope — what each RTO may deliver/assess, per component, with start/end dates, an explicit/implicit flag and an extent ("Deliver and assess" vs "Assess only"). ~94% of scope is implicit — units flowing from on-scope qualifications — and TGA publishes the implicit lists rather than leaving them to be derived (sample measurement, n=200 bodies, 2026-08-20).
  • Delivery notifications — where an RTO actually notified delivery of a component; bitemporal (the date notified and the date effective can differ by years). The largest family by volume (millions of pairs).
  • Releases and documents — each component has numbered releases; each release lists its document assets with name, type, size and publish date. ~104,737 (component, release) pairs; the payload behind the manifests is ~225 GB (mean-basis, n=24, re-run owed — see Part two caveats).
  • Classifications and taxonomies — ANZSCO, ASCED, industry/occupation taxonomies, and the scheme value dictionaries (7 endpoints, 0.40 MB — the full code lists including codes no component uses).
  • Metadata — one endpoint publishing dataExportSynchronisationDateTime plus a build version: the source's only change signal, and it is global, not per-object.

Added 2026-09-04 — the mirror is two mirrors, and only one can answer a question about a resource. pith-assets is keyed <family>/<sha256 of the body>.json: content-addressed, carrying no resource id, verified on 36 objects across 12 families with 0 differing. It cannot be asked "is there an object for release X". The content mirror's own listing (tga-content-ed1-2026-08-28/_LISTING.jsonl, 105,911 rows over 75,156 distinct release_id) can. Naming which mirror is being counted, before counting, is the whole of that figure's reliability.**

4 · Temporal properties — the four facts that shape everything downstream

  1. No per-object change signal. No ETag, no Last-Modified. One global export-sync stamp. Re-sync designs poll the stamp; per-object hash comparison does not work (identical data re-serialises with different bytes — measured).
  2. History is embedded, not versioned. The detail bodies carry dated intervals inline; TGA does not publish "the body as it was last year". What we captured on a date is the only record of how it stood that date — which is why capture custody (the pith stores) exists.
  3. The register is bitemporal where it matters. Notification date ≠ effective date on delivery facts (observed lag up to five years). Rights are evaluated on effective dates; visibility questions use notification dates.
  4. All dates are register-published calendar dates. No timezones, no instants. Measured ISO-8601 across 624,491 values with zero exceptions (org family, 2026-08-20) — but that is one family's measurement, not a source-wide law; each parse re-measures.

5 · The honest limits

  • TGA withholds by policy as well as by absence. One endpoint answered 403 "Completion mapping is internal use only" on 2026-08-20. A refusal is a real answer and is recorded as one. Corrected 2026-09-04: that refusal is falsified on the bodies. The 28 August build holds 2,575 completion and 6,403 completionusage bodies — 8,978 in all, every one a JSON array of components, zero error bodies. Whatever answered 403 in August, what is on disk is data. A refusal is an observation with a date on it, never a property of an endpoint — and it is re-tested before it is carried forward.
  • Every claim is bounded by "what TGA publishes", never "what TGA knows."
  • The source is not perfectly clean. Known warts, measured: inconsistent date formats in ASQA's register (2/4-digit years, month-name variants, at least one malformed value); TGA field-name traps (registrationManager on organisations vs recognitionManager on components — two fields, not a rename); status vocabularies that collapse distinct states when read by .name instead of .id. The full trap list lives in substrate-findings.md; this atlas names that they exist so nobody assumes cleanliness.
  • No measured instance of TGA losing an artefact currently exists (the earlier sole-survivor claim was falsified 2026-08-19). Archive protection arguments rest on prudence, not on TGA fragility.

6 · How to read a "do we have X" answer

The three-states frame governs: A — in the corpus (ledgered, closure-proven; trust it) · B — only in the prototype's rto-nrt-db (no ledger, no closure; volume signal only, never a figure) · C — never fetched. Part two ranges over the API surface and states A/C per family; B exists only as history.


PART TWO — COVERAGE (generated by instrument; hand-editing is a defect)

Owner instrument: scripts/ledger-01-emit-v1.py e4fdf55fc452a74d, emission outputs/ledger/TGA-MIRROR-COMPLETENESS-LEDGER-01-2026-09-04.md 4bd17526896a44b4, as-at 2026-09-04T10:04:22.

1 · The seven states

The seed had six. held-bodies is the seventh, ruled at the Gate 1 verdict §2.1 because the stores needed it: a family whose objects are on disk and whose rows do not exist is not owed — nothing has to be fetched for it — and it is not in. The seed called those families owed, which is what a ledger without an instrument does.

state means
in rows in a local store of record, keyed and digested, gate-certified
in flight a brief is running or filed against it today
held-bodies (new) fetched, in the mirror, no rows and no parse
owed TGA serves it, nothing holds it as rows or bodies, no brief filed
fetch-dependent cannot close without contacting TGA
render a TGA surface re-serving data held elsewhere; the round-trip test is the mirror
not TGA's / refused outside the mirror by definition, or TGA answers 403

The rule table, as applied — every state below comes from one of these and the rule id is printed with it:

rule state assigned when
S1 render the path's class in the Gate 1 store is render
S2 not TGA's the path's class is not TGA's data
S3 in flight rows exist AND a brief is filed and running against it today
S4 in rows exist in a local store of record, keyed and digested
S5 held-bodies no rows, and >0 objects for it in a mirror listing (verdict §2.1)
S6 fetch-dependent no rows, no bodies, and it cannot close without contacting TGA
S7 owed no rows, no bodies, no brief — TGA serves it and nothing holds it

⚠ refused has no members, and that is a measurement

The seed §8 counts refused 2 — /training/{code}/completion and /completionusage, on the atlas's 2026-08-20 record of a 403 "internal use only". Gate 2 was told to read one body of each. It read every body of both:

family bodies error/dict bodies empty arrays
completion 2,575 0 1
completionusage 6,403 0 1

All 8,978 are JSON arrays of components. Zero are error bodies. Verbatim head of one of each:

completion/       [{"code":"VBQM230","endDate":"2012-12-31","hasPreRequisites":false,"isMandatory":false,"links":[{"rel":"training-component","href":"https://training.gov.au/api/training/vbqm230"}],"startDate":"2009-08
completionusage/  [{"code":"10183NAT","id":"e230ecba-e5c1-4a6a-87b6-991d24fff873","status":"nonCurrent","statusLabel":"Non-current","title":"Diploma of Education","type":"accreditedCourse"},{"code":"10326NAT","id":"2c6

Whatever answered 403 on 20 August, what is on disk from the 28 August build is data. The two paths are data paths in held-bodies; the refused state has no members today. This corrects the seed §8 and the Gate 1 verdict §2.2's "47 + 2 refused": the reconciliation stands, but the two are not refused.

2 · The denominator, and the seed's own arithmetic reconciled

88 paths across the eight captures, each at the digest the brief asserts. By class, from the Gate 1 store:

class paths
data path 47
render 35
not TGA's data 4
not TGA's / refused 2

The seed §8 says "of the 88, 39 are renders or not TGA's data (12 EXPORT · 8 SEARCH · 3 REPORT · 1 FEEDBACK · 3 security-role paths · 10 export paths inside TRAINING/ORGANISATION · 2 files paths)". This instrument: render 35 + not TGA's data 4 = 39. The seed reconciles exactly, item by item.

and "of the remaining 49 data paths" — this instrument: 47 data paths + 2 classified refused = 49. The seed reconciles exactly. With the refusal falsified above, the 49 are data paths outright and the seed's total was right all along; only its last two labels were not.

3 · NRT — training products

path population held state rule source
/training/{code} 125,838 components (125,848 rows) component · release · supersession_edge · component_parent in e2-nrt-01 in S4 Gate 1 item 2
— its usageRecommendation and classifications 125,838 no column exists in e2-nrt-01.component, and no taxonomy table held-bodies S5 Gate 1 item 2; bodies at classification/ 30,674
/training/{code}/releases 125,838 components 87,099 carry ≥1 release; releases-list/ holds 87,099 in S4 Gate 1 item 3
/training/{code}/releases/{n} 103,327 releases 103,320 level-2 bodies (100.0%), 7 without held-bodies S5 Gate 1 item 6
…/document-bundle · /content/bundle/{id} · /content/item/{id} 103,327 releases 781,817 sections indexed; bodies in e2-content-01 in for fetched releases, fetch-dependent for the rest — §5 S4/S6 Gate 1 items 4, 5
…/unitgrid 12,902 walked grid 10,269 · empty 2,633; join 12,902/12,902 in flight — Q3 S3 Gate 1 item 7
…/components · …/components/qualifications 931 training packages release-components/ 929 bodies held-bodies S5 Gate 1 item 9
…/files · …/files/zip — — render — Tim ruled 2026-09-03 the PDFs are not held S1 RULINGS-2026-09-03-TAPPED-THREE §4
/training/{code}/classification 125,838 × n classification/ 30,674 bodies, no rows held-bodies S5 Gate 1 item 9
/taxonomyindustrysectors · /taxonomyoccupations 2 vocabularies taxonomy-industry 1,244 · taxonomy-occupation 3,103 held-bodies S5 Gate 1 item 9
/training/{code}/prerequisites unmeasured prerequisites/ 2,285 bodies; the content-side reference rows are in held-bodies S5 Gate 1 item 9
/training/{code}/delivery unmeasured delivery/ 89,738 bodies held-bodies S5 Gate 1 item 9
/training/{code}/unitgridusage unmeasured unitgridusage/ 27,156 bodies held-bodies S5 Gate 1 item 9
/training/{code}/recognitionmanager · /restrictions unmeasured comp-restrictions/ 3; no family for recognitionmanager held-bodies / owed S5/S7 Gate 1 item 9
/training/{code}/completion · /completionusage unmeasured completion/ 2,575 · completionusage/ 6,403 bodies, all data, none refused held-bodies S5 Gate 2 §1
/releases/{releaseId}/training derivable release.release_id_norm join in (derived) S4 model §1

⭐ The register forgot to give 38,739 components a release. 87,099 of 125,838 distinct codes carry a release row; the 38,739 that do not are the accredited families — accreditedCourse 19,439 and accreditedUnit 19,302 rows, zero releases between them, with a positive control run before the zero was believed. Their L1 register bodies ARE held: register-accredited-course-L1 19,439 · register-accredited-unit-L1 12,789. This is a build gap, not a fetch gap.

4 · NRT content — the section types, parsed into rows

781,817 sections held over 36 content type codes. Held vs parse rows, per code, from the store of record:

code sections parse rows no-section rows open
0012 104,754 104,754 0 0
0118 58,537 58,542 5 0
0103 58,494 58,494 0 0
0102 52,695 52,695 0 0
0121 46,470 46,470 0 0
0011 38,629 3,680 2 34,949
0200 30,985 30,985 0 0
0113 30,502 31,808 1,306 0
0120 30,500 31,806 1,306 0
0104 30,497 31,804 1,307 0
0112 30,476 0 0 30,476
0100 28,048 0 0 28,048
0111 28,043 0 0 28,043
0124 28,039 0 0 28,039
0125 27,995 0 0 27,995
0108 27,993 0 0 27,993
0119 26,645 0 0 26,645
0000 24,112 0 0 24,112
0001 15,745 19,425 90 0
0123 11,103 0 0 11,103
0116 10,077 0 0 10,077
0117 10,069 10,069 0 0
0110 9,604 9,604 0 0
0126 5,454 5,454 0 0
0127 5,452 5,462 10 0
0128 5,451 5,462 11 0
0109 4,620 0 0 4,620
0107 481 481 0 0
0010 74 0 0 74
0201 64 0 0 64
0002 48 0 0 48
0105 48 0 0 48
0122 47 0 0 47
0106 46 0 0 46
0202 10 0 0 10
0203 10 0 0 10

0011 open: 34,927 releases — the largest single open figure in the store, reproducing the BRIEF-TGA-MODEL-02 Gate 5 close.

5 · Fetch-dependent — and it is SEVEN releases, not twenty-eight thousand

28,171 releases of 103,327 have no row in the store of record's section. The first emission of this ledger called all 28,171 of them "never fetched". That was wrong in kind, and the partition below is why.

cell releases
level-2 body held, contentBundles NON-EMPTY, and sections in the store 75,156
level-2 body held, contentBundles non-empty, but no sections 0
level-2 body held, contentBundles EMPTY, but sections in the store 0
level-2 body held, contentBundles EMPTY, no sections — a SERVED ABSENCE 28,164
in the store with no level-2 body held 0
no level-2 body at all — ⭐ THE FETCH LIST 7

A four-cell partition with zero exceptions. 28,164 releases carry a level-2 body in which TGA itself serves contentBundles: [] — the release-grain case of declared-absent. Nothing has to be fetched for them and nothing is missing. The fetch list is 7 releases.

Three counter-hypotheses were tested against the served-absent population and all three fail. They are not merely older releases of otherwise-bundled components: 28,163 of 28,164 belong to components where every release is empty. TGA does not hold the content behind a document instead: 100.0% of bundled releases list a document asset against 3.4% of these. And they are not current — measured against TGA's own usageRecommendation, held in the L1 register bodies for 87,099 codes, 0 of 28,164 belong to a component TGA marks current (61.6% deleted, 38.4% superseded), where the same predicate finds 22,715 current components on the bundled side. Cited, not recomputed here: outputs/canon-01/CANON-01-FETCHGAP-TEST.txt 7ff788ce77a1a864 and outputs/canon-01/CANON-01-GATE2-CH3-USAGEREC.txt.

⚠ 951 of the served-absent releases list a document asset — TGA publishes a downloadable document and no machine-readable bundle, so for those the round-trip test has no rows to run against. RULING-ASSET-ONLY-RELEASES-01 (Tim, 2026-09-04): those 951 documents are fetched and kept as bodies in R2, state held-bodies, no parse briefed. They are the second fetch-list line.

⚠ Served-absent is as-at the capture, not forever. Every figure above is measured on what TGA served on 28 August 2026. Whether a fresh pull still returns [] is a question only the refresh pull can answer.

The type table below ranges over the served-absent population and is retained for its shape; its own population is named in outputs/canon-01/CANON-01-GATE2-FANOUT.txt.

type current replaced total
qualification 2,805 0 2,805
skillSet 0 2 2
trainingPackage 129 0 129
unit 25,234 1 25,235
TOTAL 28,168 3 28,171

The 25,235 is confirmed as a number and retired as a sentence. The Time Machine carried "25,235 sectionless current unit releases" and no filed artefact held it. Measured: unit releases with no section row = 25,235 exactly, but 25,234 current and 1 replaced. The figure was right; its population name was wrong by one row.

Controls, both reproducing: qualification releases with no 0116 2,835 (Q1's 2,835) · with no section at all 2,805 (-01's 2,805).

⚠ Which mirror answers this. pith-assets keys are <family>/<sha256 of the body>.json — content-addressed, carrying no resource id, verified on 36 objects across 12 families with 0 differing. It cannot be asked "is there an object for release X". tga-content-ed1-2026-08-28/_LISTING.jsonl can: 105,911 rows over 75,156 distinct release_id. Naming the path before counting from it is the whole of this row's reliability.

6 · RTO — organisations

⚠ Two databases answer to e2-rto-01 and only one is the register. The .local-stack store bound to that name holds 3 qualifications and 140 units — a prototype shape with no organisation table and no scope_entry. Every figure here is from ~/mirror/e2-stores-2026-09-01/e2-rto-01.sqlite.

path population held state rule
/organisation/{code} and its eleven tabs 13,034 organisations organisation + 17 od_* families; org-L1/ 13,034 bodies in flight — Lane Z S3
/{code}/scope · /scopesummary 6,526,449 scope rows 6,526,449; scopesummary/ 3,674 bodies in flight — Lane Z S3
/{code}/deliverynotificationhistory/{trainingcode} the largest family dnh/ 312,010 bodies, no rows, and the model names no table held-bodies — was owed in the seed S5
/{code}/training-packages unmeasured org-training-packages/ 5,920 bodies held-bodies S5
/organisation-security-roles — — not TGA's S2
/api/security-roles — 1 body, metadata family not TGA's — declared by the METADATA swagger; TGA's own access-control vocabulary, not register data (C-0912-02) S2
2 × scope/export/… — — render S1

7 · held-bodies — what is on disk and has never become a row

This section did not exist in the seed, because the state did not. Every family below was fetched by the E2 build and holds objects today; none has rows.

family objects coverage measurable locally?
dnh 312,010 no — the body names its contents, never its subject
releases 103,320 yes — the body names its subject (id)
delivery 89,738 no — the body names its contents, never its subject
releases-list 87,099 yes — the body names its subject (id (of each release))
register-unit-L1 56,725 not examined at this gate
classification 30,674 no — the body names its contents, never its subject
unitgridusage 27,156 no — the body names its contents, never its subject
unit-L1 24,913 yes — the body names its subject (code + id)
org 24,110 not examined at this gate
register-accredited-course-L1 19,439 not examined at this gate
org-L1 13,034 yes — the body names its subject (code + organisationId)
register-accredited-unit-L1 12,789 not examined at this gate
register-qualification-L1 6,870 not examined at this gate
completionusage 6,403 no — the body names its contents, never its subject
org-training-packages 5,920 not examined at this gate
register-skill-set-L1 3,679 not examined at this gate
scopesummary 3,674 no — the body names its contents, never its subject
taxonomy-occupation 3,103 no — the body names its contents, never its subject
completion 2,575 no — the body names its contents, never its subject
qual-L2 2,345 not examined at this gate
qual-L3 2,345 not examined at this gate
prerequisites 2,285 no — the body names its contents, never its subject
unitgrid 2,156 not examined at this gate
taxonomy-industry 1,244 no — the body names its contents, never its subject
training-L1 1,174 not examined at this gate
release-components 929 yes — the body names its subject (releaseId)
cv-list 337 not examined at this gate
register-training-package-L1 259 not examined at this gate
metadata 12 not examined at this gate
comp-restrictions 3 not examined at this gate
cricos 1 not examined at this gate
search-organisation 1 not examined at this gate
search-training 1 not examined at this gate

⚠ The Gate 1 verdict §5's least-sure, answered, and the answer is a split, not a yes or a no. Five families self-identify and their coverage is measurable locally today: releases, unit-L1, org-L1, release-components, releases-list. The rest name their contents and never their subject — a delivery body lists the RTOs delivering a product without saying which product; links is empty on every one read. For those, coverage needs the fetch ledger (pith-index-01, cloud), which is therefore a dependency for some families and a check for others, not a blanket dependency.

8 · The number, honestly — and the seed graded

The seed's §8 gave three counts over the 49 data paths, hand-derived: in or in flight 22 · owed 24 · fetch-dependent 1 (plus refused 2). Every one of the 49 is assigned a state below by the rule table, from the evidence named beside it, and the counts are SELECTed from that assignment — not typed.

path evidence rule state
/api/content/bundle/{id} store of record.section S4 in
/api/content/item/{id} store of record.section S4 in
/api/classification-purpose — S7 owed
/api/classification-purpose/{code} — S7 owed
/api/metadata — S7 owed
/api/nrt-classification-scheme — S7 owed
/api/nrt-classification-scheme/{code} — S7 owed
/api/nrt-classification-scheme/{code}/values — S7 owed
/api/nrt-classification-scheme/{schemeCode}/values/{code} — S7 owed
/api/rto-classification-scheme — S7 owed
/api/rto-classification-scheme/{code} — S7 owed
/api/rto-classification-scheme/{code}/values — S7 owed
/api/rto-classification-scheme/{schemeCode}/values/{code} — S7 owed
/api/organisation/{code} e2-rto-01.organisation; Lane Z S3 in flight
/api/organisation/{code}/addresses od_addresses S3 in flight
/api/organisation/{code}/classification od_classifications S3 in flight
/api/organisation/{code}/contacts od_contacts S3 in flight
/api/organisation/{code}/cricoscode od_cricos_codes — 0 rows, served empty by TGA; the emptiness is the held fact S4 in
/api/organisation/{code}/deliverynotificationhistory/{trainingcode} dnh/ 312,010 bodies S5 held-bodies
/api/organisation/{code}/legalname od_legal_names S3 in flight
/api/organisation/{code}/registration od_registrations S3 in flight
/api/organisation/{code}/registrationmanager od_registration_managers S3 in flight
/api/organisation/{code}/regulatorydecision od_regulatory_decisions S3 in flight
/api/organisation/{code}/restrictions od_restrictions S3 in flight
/api/organisation/{code}/role od_roles S3 in flight
/api/organisation/{code}/scope e2-rto-01.scope_entry S3 in flight
/api/organisation/{code}/scopesummary scopesummary bodies; Lane Z S3 in flight
/api/organisation/{code}/tradingname od_trading_names S3 in flight
/api/organisation/{code}/training-packages org-training-packages/ 5,920 bodies S5 held-bodies
/api/organisation/{code}/webaddress od_web_addresses S3 in flight
/api/releases/{releaseId}/training e2-nrt-01.release (derived) S4 in
/api/training/{code} e2-nrt-01.component S4 in
/api/training/{code}/classification classification/ 30,674 bodies S5 held-bodies
/api/training/{code}/completion completion/ 2,575 bodies S5 held-bodies
/api/training/{code}/completionusage completionusage/ 6,403 bodies S5 held-bodies
/api/training/{code}/delivery delivery/ 89,738 bodies S5 held-bodies
/api/training/{code}/prerequisites prerequisites/ 2,285 bodies S5 held-bodies
/api/training/{code}/recognitionmanager — S7 owed
/api/training/{code}/releases e2-nrt-01.release S4 in
/api/training/{code}/releases/{releaseNumber} releases/ 103,320 bodies S5 held-bodies
/api/training/{code}/releases/{releaseNumber}/components release-components/ 929 bodies S5 held-bodies
/api/training/{code}/releases/{releaseNumber}/components/qualifications release-components/ 929 bodies S5 held-bodies
/api/training/{code}/releases/{releaseNumber}/document-bundle store of record.section for fetched releases; the rest is the fetch gap S6 fetch-dependent
/api/training/{code}/releases/{releaseNumber}/tpusage — S7 owed
/api/training/{code}/releases/{releaseNumber}/unitgrid Lane Q's walk; Q3 loads it S3 in flight
/api/training/{code}/restrictions comp-restrictions/ 3 bodies S5 held-bodies
/api/training/{code}/taxonomyindustrysectors taxonomy-industry/ 1,244 bodies S5 held-bodies
/api/training/{code}/taxonomyoccupations taxonomy-occupation/ 3,103 bodies S5 held-bodies
/api/training/{code}/unitgridusage unitgridusage/ 27,156 bodies S5 held-bodies
the seed said this emission says, derived why it moved
in or in flight 22 21 (in 6 · in flight 15) the register, content and RTO rows
owed 24 owed 13 · held-bodies 14 the seventh state: 14 of the seed's owed paths have their bodies on disk already and need a parser, not a fetch
fetch-dependent 1 1, and now with a size: 7 releases (was reported as 28,171 in the 2026-09-03 emission) the seed named the row; the instrument counted it; the partition corrected it — 28,164 of the 28,171 are served absences, not a fetch
refused 2 0 falsified on 8,978 bodies, §1
— total 49 of 49 data paths every data path has a state

The honest sentence the seed was reaching for: it is not that half the data paths are owed. It is that 14 of 49 (29%) are already on disk and waiting for a table — a parser job, not a fetch — one is a fetch of known size (28,171 releases), and only 13 have nothing held at all, of which 10 are the small classification-scheme and classification-purpose vocabularies.

The post-14-September list, in order

  1. The fetch list — 7 releases with no level-2 body held. Not 28,171: the other 28,164 are served absences (§5). 1b. The 951 asset-only releases — a document published with no content bundle (RULING-ASSET-ONLY-RELEASES-01). 1c. (retired figure, kept for the audit trail) the former "fetch gap" of 28,171 releases with no section row, 28,171 of them recorded as never fetched: unit 25,235 · qualification 2,805 · trainingPackage 129 · skillSet 2.
  2. The fetch-ledger export — one read of pith-index-01 to a local file under SPEND-ENVELOPE-01, which turns every held-bodies row whose bodies do not self-identify into a coverage figure. Ruled at the Gate 1 verdict §2.1.
  3. The refresh pull, the /metadata sentinel, and a refresh policy per family — seed §7, none built.

Before any of that, and needing no fetch at all

  1. Accredited-course and accredited-unit releases from their held L1 bodies — 38,739 components with no release row.
  2. The level-2 release fields — currencyChangeDate · assets[] · externalLinks[] · specializations[] · packagingInformation — from 103,320 bodies already on disk, into tables the model must first name.
  3. Classification and taxonomy per component from 30,674 + 4,347 bodies.
  4. Delivery notifications — 312,010 dnh and 89,738 delivery bodies, into a table the model must first name.

Emitted 2026-09-04T10:04:22 by ledger-01-emit-v1.py e4fdf55fc452a74d. Every cell is a row of the Gate 1 store. To regenerate: re-run Gate 1's instrument, then this one.

Regeneration rule: the audit emits this table whole, with its own as-at, and the emission replaces the section between the PART TWO header and this line. A hand edit to any cell is a defect of the same class as back-filling provenance.


Status and open items

  1. ~~Tim reads Part one for domain wrongness.~~ DISCHARGED — approved 2026-08-20.
  2. OPEN: the audit brief (TGA-MIRROR-CORPUS-AUDIT-01) names this document's Part two as a deliverable when it is written. Until the audit runs, Part two is seeded knowledge with no completeness claim.
  3. ~~Home and hierarchy ruling.~~ DISCHARGED — ruled 2026-08-20: home as in the header; subordinate to substrate-findings on any figure conflict (findings are the measurement record; this is the map drawn from them).

Least sure

Whether Part one's plain-language register stays plain under maintenance. Every editor of this file will know more than its intended reader, and the drift pressure is one-directional — toward the jargon of whoever last touched it. The cure attempted here is structural (the volatile parts are machine-generated, so humans mostly never need to edit), but Part one has no such guard. What would make me wrong about the document's whole value: if in six months it reads like substrate-findings, it has failed at the only thing it exists for.

— Claude (Architect seat, Cowork, Fable, session 33), 2026-08-20. Approved by Tim the same day; a working reference, amended by dated addition.