ANCHOR-ONE-GATHER-01 — INVENTORY¶
Seat: Driver (Opus). Date: 2026-08-14. Job: gathering only, per
outputs/ANCHOR-ONE-DRIVER-GATHER-RELAY-2026-08-14.md.
Posture: read-only. No account call, no live-estate query, no TGA request, no git, nothing
mutated. Every figure below is quoted from a repo document and carries its source; nothing was
re-measured against the estate. Two things were computed in this window, both offline on repo
bytes and both marked as such.
0 · POPULATION SEARCHED — so the absence claims mean something¶
The relay said search, don't assume the list is complete. What was searched, and how:
| instrument | population | result |
|---|---|---|
find over the whole repo (excluding node_modules, docs/site) for *tga* |
whole tree | 132 paths |
find for *ingest*·*parse*·*capture*·*mirror*·*cvig*·*companion*·*release*·*drain*·*pith*·*swagger* |
docs/docs, outputs, banked, trust-docs, evidence, sources, archive, crossings |
~380 paths |
grep -rn "Q-P2P-" |
outputs/, docs/ |
Q-P2P-001 … 008 exist; this report continues at 009 |
grep for 3,985 / 3887 / operating RTO |
outputs/, docs/docs/ |
one hit, in the North Star itself |
Instrument demonstrated before any zero was asserted. Each glob returned a known-present hit
before it was trusted (e.g. the *swagger* glob returned the delta-audit report; the *cvig* glob
returned the acquisition census).
One name-blindness caught and repaired mid-job. The first sweep's patterns
(tga|ingest|parse|capture|mirror) do not match the RTO/organisation strand, whose documents
are named RADAR-REGISTER-01-* and ORG-RAW-ARCHIVE-BACKFILL-01-*. Nine documents were added to
the population and read on that discovery. A filename is not a document — and in this repo, the
organisation half of the TGA ingest does not carry the letters T-G-A in its name.
Documents read in full: 51. Named in the source column throughout. Two computations performed
offline on repo bytes, marked [computed here]: the swagger operation count (§5) and arithmetic
cross-checks on the group tables.
1 · DATASET × SOURCE MAP¶
The API surface is a single unauthenticated JSON REST API at https://training.gov.au/api.
No SOAP, no WSDL, no key, no token — "Confirmed over 480,000+ unauthenticated calls"
(docs/docs/ops/tga-ntr-api.md:36). Eight spec groups. Two hard client facts that bound any
scripted ingestion: curl is refused at the TLS layer and returns HTTP 000 on every attempt
(tga-ntr-api.md:69–86) — a scheduled job built on plain curl or urllib fails 100% and
presents as an outage — while Node fetch, Playwright's context.request and Cloudflare Workers
all reach /api/* fine (tga-api-field-inventory.md:480;
outputs/TGA-SWAGGER-VS-REPORTS-DELTA-AUDIT-01-report.md:16–18).
| dataset | source operation(s) | arrival format | proven parser? |
|---|---|---|---|
| Units | SEARCH_V1 /api/search/training → TRAINING_V1 /api/training/{nrtId}?include=all → /releases/{n}?include=All → CONTENT_V1 /api/content/bundle/{bundleId} |
JSON envelope; the content bytes are HTML inside items[].content, content item code 0118 |
Yes, with a named unfixed defect |
| Quals | same chain, plus TRAINING_V1 /api/training/{code}/releases/{n}/unitgrid and content item 0116 PackagingRules |
JSON; unitgrid is a bare JSON array, no wrapper; packaging rules are prose HTML | Yes for cross-reference edges. No for packaging semantics. |
| RTOs | ORGANISATION_V1 /api/organisation/{code}?include=all (+ /registration, /regulatorydecision, /restrictions, /deliverynotificationhistory/{code}) |
JSON | NO |
| CVIGs | TRAINING_V1 /api/training/{code}/releases/{n}/files → second fetch to each row's uri |
JSON listing; bytes are PDF (509 of 562), xlsx (36), docx (16), zip (1) | Acquisition proven. Extraction: NO parser exists. |
| Release history | releases[] on component detail; TRAINING_V1 /api/training/{code}/releases/{n} |
JSON | Capture proven. Continuity/version record: NO — deferred by design. |
1.1 Units — proven, and the proof names its own defect¶
Capture and parse are separate and were deliberately kept so:
"Parsing is a separate, later, re-runnable pass against stored raw. If we parse on the way
in we bake today's mistake into tomorrow's database."
(outputs/TGA-CLEANROOM-INGEST-01-BRIEF-2026-08-04.md:32–41)
The proving document is outputs/TGA-CLEANROOM-INGEST-01-PARSE-REPORT-2026-08-04.md.
261,186 performance criteria extracted across 15,169 current units (:38); verdicts over the
whole population — parsed 14,934 / 98.5%, unparsed 153 / 1.0%, no 0118 82 / 0.5% (:32–36).
The question it existed to answer: "of the 6,505 current units with no performance criteria in
the Pith, 6,357 have them in content item 0118 and always did" (:3–4).
What "proven" meant here — four independent instruments, none of them a success flag:
an independent token-count tripwire that fired on 5 of 15,169 units "Every one carries
verdict: parsed. Without the tripwire all five ship as complete" (:64–72); a second-witness
cross-column comparison, 971 of 1,405 found verbatim, "73.5% vindicated by machine, with no human
reading anything" (VERIFICATION-03:16–21); a positional-signature prediction-then-measure test
(VERIFICATION-03:39–44); and a full-depth regression of stored Pith criteria against clean-room
extraction — 142,955 comparisons, 97.80% exact agreement, 2,618 disagreements over 304 units
(D8:14–19), with instrument comparability asserted explicitly (D8:23–24).
Three things it did not do, all stated by the document itself:
"no sample has been checked against the register by hand at this scale. The pilot verified 9
units against hand counts; this run verified none" (PARSE-REPORT:93–95);
"Nothing is written to R2 or D1. Nothing has been diffed." (:101); and the parser is
known defective — "The clean-room parser loses the first criterion after an element heading.
71.5 / 17.7 / 10.8 against a flat null is not ambiguous"
(VERIFICATION-03-REVIEW:33–34). PARSER-FIX was scoped at five gates and, on the record read,
never started (RETURNS-R1-R4:125).
1.2 Qualifications — the strongest proof in the estate, and its exact scope¶
QUAL-PARSE-01 is the only corpus-scale run in this strand that ended in persisted substrate.
626,914 validated edges over 50,808 distinct referenced codes, from 55,006 rows / 5,229 codes /
10,097 (code, release) pairs; 7 parks; every reconciliation count equal
(QUAL-PARSE-01-G4-VERDICT-2026-08-09.md:32–37). Gate 4 PASS, Gate 5 signed, arc closed
(G4-VERDICT:3; GATE5-CLOSE:148).
What "proven" meant here is the highest bar reached anywhere in the record: 13/13 artefact
digests independently re-hashed by the reviewing seat; the headroom arithmetic re-derived from its
own inputs "independently computed by this seat before the run and matching to the byte"
(G4-VERDICT:17–21); the database size read through a second channel (the Cloudflare account
API) and agreeing exactly (:21–24); guard polarity proven offline before call one — read guard
30/30 negatives refused and 12/12 positives allowed, write guard 28/28 refused and 16/16 allowed
(G4-REPORT:44); two independent instruments agreeing on the raw census, 627,389 = 627,389
(G4-REPORT:197–199); and a resume mechanism proven under real failure, twice
(G4-VERDICT:56–60).
Its scope, exactly. It is a reference graph — which unit codes a qualification's pages name —
not packaging semantics. No core/elective structure, no group membership, no select-rules are
claimed anywhere in the arc. And resolution was against our mirror only: "Against our mirror
only — TGA was never checked" (G4-REPORT:193–195).
The gap sitting next to it. Group A/B/C elective sub-structure — the thing a qualification tree
actually needs — is not published in any structured endpoint: "TGA does NOT publish this in any
structured endpoint — it only exists in the prose HTML of content bundle item 0116"
(tga-unitgrid-endpoint.md:96). tools/parse-packaging-rules.mjs exists and stamps
parser_version = 'unitgrid-v1' (:4–5), but no gate verdict for it appears in the population
searched.
1.3 RTOs — the dataset with no proven path. Say it plainly.¶
The endpoint is rich and documented — 20 ORGANISATION operations, per-code detail at
/api/organisation/{code}?include=all (TGA-SWAGGER-VS-REPORTS-DELTA-AUDIT-01-report.md:170;
docs/docs/infrastructure/tga-ingest.md:51). What is missing is any demonstration that we can
enumerate the register through it.
- The corpus was seeded once and never re-demonstrated: "12,413 of 12,506 rows created
2026-03-28 (one-time bulk-import spike at genesis)… The corpus was bulk-seeded once"
(
ORG-RAW-ARCHIVE-BACKFILL-01-part2-gate4-validation.md:8–10). The provenance of that seed is "some one-time load" (:47) — not named in any document read. - The weekly refresh path cannot enumerate, measured: "the Gate-2/3 design paginated
/search/organisationassuming it enumerates the corpus. It does not — it returns a ~240 randomly-shuffled pool (78% adjacent-page overlap). Caught empirically at Gate 4" (ORG-RAW-ARCHIVE-BACKFILL-01-part2-close.md:42–45). - One thing here is genuinely proven, and its reference is D1, not TGA: the R2 raw-archive
backfill walked all 12,506 codes in ~1.5 h, "zero 404s, zero fetch-failures, zero NULLs",
reconciled set-difference both directions = 0, coverage 0.9% → 100%
(
part2-close.md:3–15), with the failure path demonstrated rather than assumed (:29–31). It proves the archive matches the mirror. It says nothing about the mirror matching the register. - No organisation-side register reconciliation exists — the equivalent of
TGA-REGISTER-RECONCILIATION-2026-08-06.md, which reconciles training components against TGA's own live facets, has never been done for organisations. Silent across all nine RTO-strand documents read.
1.4 CVIGs — acquisition proven, extraction unbuilt, and a legal stop condition on the record¶
Acquisition is the cleanest single arc in the inventory. CVIG-ACQUIRE-01 Gate 4 certifies
562 unique files, 1,001,808,227 bytes (955.4 MiB), 51 of 54 current packages, every object
verified by read-back out of R2, bytes agree exactly (CVIG-ACQUIRE-01-GATE4-CERT:14–22, 236);
the three packages with none (CPC08, MEM05, MSA07) were swept across all 34 of their releases and
named rather than counted as gaps (CVIG-ACQUIRE-01-CENSUS:226–258). Coverage claim stated
exactly, and narrower than "complete": "complete with respect to TGA" (GATE4-CERT:241).
There is no CVIG parser. Zero occurrences of "CVIG", "companion volume" or "implementation
guide" across the entire 17-document pilot/parse/schema set (grep-verified). Extraction must be
planned against PDF — "a table-preserving extraction strategy built on DOCX would cover about
3% of the corpus" (tga-ntr-api.md:422).
And a filed stop condition contradicts the North Star's scoping. CVIG-LANDSCAPE-RECON-01
Finding 4: HLT R11.0's CVIG is CC Attribution-NonCommercial-ShareAlike verbatim, and "The NC
term restricts ingestion into a commercial product absent permission… no CVIG content enters
any RTOpacks store until a written legal position exists" (DOSSIER:131–147). The amendment
adds: "Site licence does not launder document licence" (AMENDMENT-01:44–48). See Q-P2P-010.
1.5 Release history — captured, never modelled¶
Release identifiers are captured and release-scoping is native: tga_training_releases holds
103,473 rows / 87,245 distinct component codes / 103,473 distinct (code, release) pairs — the
pair is unique (INGESTED-SUBSTRATE-INVENTORY-01-REPORT:214), and content is keyed by code and
release, 77,713 code-release pairs across 61,486 codes (nrt-atom-schema.md:13).
A version-continuity record does not exist and was deliberately deferred:
"Continuity is NOT identity. Cross-release links (identical / reworded / split-into /
merged-from / no-successor) are a separate assertion layer with provenance, designed in its own
brief." (nrt-atom-schema.md:28). That brief is not written.
The trap any fresh pull must respect: "RELEASE NUMBERS ARE STRINGS AND ARE NOT CONTIGUOUS.
Enumerate them; never generate them." — measured on 64 VET components, 53 agree and 11 disagree;
CPP07 advertises latest 15.0 but releases 1–6 do not exist and 14.1–14.7 do
(tga-ntr-api.md:270–286).
1.6 Datasets the North Star names with no proven path¶
Stated plainly, as the relay asked:
- RTOs — no demonstrated enumeration of the register, no register-side reconciliation, seed provenance unnamed. The richest endpoint in the API and the thinnest evidence in the repo.
- CVIGs beyond acquisition — no parser, no extraction gate, and a licence stop condition.
- Release history as a continuity record — captured as identifiers, never modelled; the assertion layer is explicitly a future brief.
2 · THE INGESTION DISCIPLINE, AS BUILT¶
Five mechanisms. One paragraph each — what it demonstrated, at what volume, with what verification, and what it did not cover.
TGA-INGEST-FINDING-AND-SPEC-01 (2026-08-04) — the founding finding, n=1. Proved on a single
unit (RIIBEF201E) against the live register that the missing criteria were ours, not the
register's: "The criteria have been published since 2020. We hold 97 characters." (:38).
It also established the structural fact the whole strand rests on — TGA publishes elements and
performance criteria as one content item, 0118, with no per-package variation at source
(:41–44). Verification was live read-only GETs plus spec reading. It covered one unit, and
says so: "the finding is proven on a single unit", RII being 680 of the 1,910 affected, "the
other five packages are untested" (:153–155).
TGA-INGEST-PILOT-01 (2026-08-04) — forty units chosen to break it. "forty units, ~160
requests", 39 real plus one deliberately absent code, run as a Worker against a dev D1, capturing
412 content items with sha256, source URL, HTTP status and duration on every row
(BRIEF-v6:50–56, 200; GATE2-VERDICT:75). Verification was hand-verified enumeration on the
regression specimens — "RIIBEF201E… 8 + 2 + 3 + 4 = 17. The parser reports 17… confirmed
against the hand-verified enumeration, not merely asserted" (GATE2-VERDICT-02:18–22) — plus
digest verification of the artefact at every gate, and a reconciliation control observed both
firing and passing. Gate 2 failed first time and cleared on re-submission at 33 of 39 parsed.
It did not cover: six named units still refused (CHCDIS022, CPCCBC4047A, PSPBDR016,
SHBBBOS007, SIRCCCS001, SITEEVT020); a third 0118 layout; the AMPR317 off-by-one; the
zero-release and missing-sequence sentinel paths, "both still unexercised and therefore unproven
per G7"; Cloudflare Queue retry/DLQ; and homes for content codes 0102 and 0121
(GATE2-VERDICT-02:43–45, 112–122). Its clearance was narrow and said so:
"Gate 2 cleared authorises the drain brief being written. It does not authorise firing at
15,169." (:30).
TGA-CLEANROOM-INGEST-01 (2026-08-04/05) — the corpus run. Full population: 15,169 current
units, 614 MB of raw captured, 261,186 criteria parsed offline from stored bytes, 98.5% parse
rate, layouts packed 10,286 / structured 4,645 / listed 3 (PARSE-REPORT:32–44;
VERIFICATION-03-REVIEW:92). Verification is described in §1.1. Its most valuable output is a
methodological one worth carrying into anchor one verbatim: "The listed layout — which cost a
whole remediation cycle in the pilot at 6 of 39 units — is 3 in 15,169. The pilot's 40-unit
sample over-represented it by two orders of magnitude. A sample chosen to break things does
exactly that, and its proportions are not the population's." (PARSE-REPORT:42–44).
It loaded nothing — that phrase appears in four separate documents. And every population label
in the strand turned out wrong once measured: "the 6,505 'destroyed' were 70.6% recoverable; the
1,311 'fabrications' were 73.5% my own under-extraction; and now the 8,693 'usable' carry 2,618
disagreements" (D8:72–74). Four numeric deltas remain published and unadjusted (RETURNS:141).
NRT-ATOM-SCHEMA-01 v1.2 (RATIFIED 2026-08-06) — the address grammar, not a run. Defines the
permanent citable address of a criterion —
{unit_code}:R{release}:{doc}:{section}:{slot_type}.{ordinal} — release-scoped and immutable, with
a provenance rule that is the content-level analogue of the hard separation rule: "the atoms table
holds register-published bytes only… Computed content never enters it — not as rows, not as
columns — regardless of quality or model" (:16–32). Its figures are provenance-tiered, [C]
measured this seat vs [A] Alex's census gate-reviewed on bytes but "not independently
re-derived" (:5). Its §6 build gates are the best specification of ingestion discipline in the
repo and should be lifted into the anchor-one brief: conservation both directions, instruments
proven on both polarities before trusted, every count names its population, structural invariants
checked independently of the parser, no derived column unmarked (:88–95).
The unit-content parser it governs has never run at corpus scale — "Gate 4 corpus launch only
after all three fix rounds close" (NRT-PARSE-02-V13-TAP-RECORD:38), and v1.3 rulings R26–R31 are
tapped and binding while the schema doc still reads v1.2 (:8–9).
CVIG-ACQUIRE-01 (2026-08-08) — the acquisition arc. 574 (package, release) pairs enumerated,
562 unique files pulled and read back, ~1,200 API calls, zero 429s, zero throttling, one transport
anomaly diagnosed as an instrument fault with the object proven intact (CENSUS:158, 502;
GATE4-CERT:25, 76–105). Ran at five gates, not two — the two-gate proposal was withdrawn
because exclusion 6 (external API) answered YES (AMENDMENT-01:23). It did not cover: extraction,
the 205 non-Current packages ("If any Current unit belongs to one of those 205 packages, its
companion volume is not in this corpus. I have not measured how many", CENSUS:400–405), and
TGA↔JSC version skew.
2.1 One correction to the North Star's own description of this discipline¶
§1 of the North Star describes the asset as "machine-readable corpus, byte-level diffing,
cross-model triangulation". Cross-model triangulation does not appear anywhere in the
ingestion record. Grep across the full 17-document pilot/parse/schema set returns zero hits for
cross-model, triangulat, multi-model or oracle. What the record actually contains, and what
is worth naming precisely because it is better than triangulation, is: digest verification at
every gate; dual-instrument agreement; second-witness cross-column comparison; prediction-then-
measure with frozen baselines; guard polarity proven offline before the first call; independent
tripwires that fire on parses wearing a success flag; and hand-verified enumeration at n=9.
Amending §1 is a ruling, not a Driver edit — flagged here, not made.
3 · CAPTURE-DEPTH STATUS (North Star gate 6)¶
The gate asks: whether our byte-capture history of TGA covers what the current mirror accumulated, or whether any part of the old mirror is quietly observed-truth.
Verdict from the record: PARTIALLY ANSWERED. One half is settled decisively; the other half has never been measured, and the instrument that would measure it is named, costs zero TGA requests, and has not been run.
3.1 What byte-capture exists¶
All in R2 bucket tga-content, plus one dev D1.
| prefix / store | what | count | window |
|---|---|---|---|
training/{code}/raw.json |
the raw /api/training/{code} level-1 body, stored unmodified |
84,728 — "fully enumerated, not truncated" | earliest object 2026-04-14T20:21:19Z, latest 2026-08-01T16:14:18Z |
versions/{collection}/{code}/{YYYY-MM-DD}.json |
dated snapshots of the same level-1 bodies, written per sync run whether or not content changed | ≥144,674 — a floor, from a truncated enumeration | earliest 2026-05-09T16:13:46Z |
organisations/{code}/raw.json |
org bodies | 12,537 (bucket recon) / 12,506 (reconciled to D1) | uninspected for date range |
dev D1 rto-nrt-tga-pilot-dev |
the only content-level (level-3) capture in the record | 412 items over 39 units | dates not stated |
Sources: TGA-CONTENT-BUCKET-RECON-2026-08-04.md:14–15, 27–38, 94–96;
TGA-DRAIN-01-VERSIONS-INSPECTION-2026-08-04.md:13–46, 110–118;
ORG-RAW-ARCHIVE-BACKFILL-01-part2-close.md:12–15;
TGA-PRE-DRAIN-RESOLUTION-01-BRIEF:7.
What is not captured at all, in production: "What is genuinely absent is level 2 (release)
and level 3 (bundle) raw, plus assets." (CONTENT-BUCKET-RECON:70–72). Content bytes for
15,169 units have no production capture — only 39 units in a dev D1 whose survival past 2026-08-04
is unstated anywhere.
3.2 The half that IS answered — and it answers YES¶
Nothing in the capture predates 2026-04-14. The load-bearing sentence:
"The bucket was created 2026-04-07 and its earliest object is 2026-04-14, so nothing in it
predates the cutoff. The question cannot discriminate on this bucket."
(TGA-CONTENT-BUCKET-RECON:94–95)
Three consequences follow, all established:
- The one enrichment run this corpus has ever had happened five weeks before capture began, and
its input no longer exists. 2026-03-09, 06:12:37 → 06:16:42, 15,200 rows
(
TGA-EXPORT-PROVENANCE-RECON-01-REPORT:9–10). After a 26,429-file search across repo, Downloads, Desktop and Documents: "the input to the only enrichment run this corpus has ever had is not recoverable from any surface we hold" (CORRECTION-01:44–46) — with one named boundary, 15 of 16 R2 buckets never enumerated at object level (:48–51). - Every vintage timestamp in the mirror is ours alone, and TGA has no field to regenerate them
from. "TGA publishes no corpus vintage. There is no field anywhere across the four specs
that says when the training component data was last changed"
(
TGA-MODIFIED-DATE-READING-01:90); "The only honest vintage available is ours:MAX(synced_at)… labelled as our sync vintage" (:100–102).dataExportSynchronisationDateTimeis environment provenance, not dataset vintage (:78–88). Sosynced_at, and the{YYYY-MM-DD}in everyversions/key, are irreducibly observed truth. - The whole RADAR observation strand is observed-truth and is already retired by ruling.
446,610 dated signals, 317,106 append-only observations carrying
change_count, 7,867 IP observations, 2,965 screenshots — "the entire spec-known intelligence layer… is frozen at 2026-04-03", refresh running at ~20 RTOs/month against 12,508, "a ~52-year full-cycle" (RADAR-SUBSTRATE-RECON-01-report.md:21–33, 85–97, 163–172). The all-data-dead ruling already covers this; recorded here only because gate 6 asks what cannot be regenerated, and this is the largest such holding in the estate.
3.3 The half that is NOT answered — and this is the finding¶
No document anywhere in the population read compares what the mirror holds against what TGA will
serve today. Not one. TGA-EXPORT-PROVENANCE-RECON-01 measures source depth by status — TGA
serves Deleted components at near-complete metadata depth and only 8/30 content depth
(:120–128, 178), and "Deletion removes content, not chronology" (:162–163) — but it never
asks what we hold. That 8/30 is a forward cost of a fresh pull, not a demonstrated loss in the
existing mirror.
Two instruments are named in the record, both cost zero TGA requests, and neither has been run:
- The
versions/size-diff census. "The fullversions/training/diff by size, grouped by code — the census that turns §3's n=3 into a number. Zero requests to TGA." (TGA-DRAIN-01-VERSIONS-INSPECTION:107–108). It was started and stopped at 20%:TGA-VERSIONS-CENSUS-01-REPORT-2026-08-04enumerated 29,028 keys of ≥144,674 (20.06%), compared 5,823 codes with ≥2 snapshots, and found zero differing series — but explicitly refuses to generalise: "exhausted: NO… What it does not support: any claim about the ~80% not reached, or about same-length edits" (:14–19, 25–38). The same-length sample is UNRUN and "the size-only error rate remains assumed" (AMENDMENT-01:19, 111–113). - A per-code join of
training/against the D1 code set. "A per-code join against all 15,169 was not run… The exact coverage number against the 15,169 remains uncomputed." (CONTENT-BUCKET-RECON:88, 117).
Per the relay, I have not investigated the live estate to close either. The gap is the report.
3.4 The carve-out, stated as the evidence supports it¶
On this record, the carve-out from "rebuild everything fresh" is not content — it is provenance.
The evidence that content is safely regenerable is strong and points one way: Class C is our own
defect in all six packages tested and "genuinely absent at source for none of them"
(TGA-CLASS-C-CHECK-01:63–66, n=1 per package over 6, population 1,910 unenumerated);
units.superseded_by_code is "demonstrably inverted" and re-capturing from source repairs it for
free (TGA-CAPTURE-ADDENDUM-01:52); and LGACOR001 reported three criteria against fourteen with
parse_verdict = 'parsed' and no anomaly row (PRE-DRAIN-RESOLUTION-01-VERDICT:16–17) — the
existing mirror's derived content is not authority.
What a fresh ingestion cannot regenerate, and what therefore needs a ruling rather than a design
decision: the pre-2026-04-14 accumulation has no byte backing at all; the 2026-03-09 enrichment
input is gone; and every synced_at and every dated versions/ key exists only because we wrote
it. See Q-P2P-009.
3.5 The standing storage ruling, for the design's inputs¶
TGA-CAPTURE-ADDENDUM-01 (binding on the drain brief) requires "Store the raw JSON body of
every response, at all three levels, with its sha256… Columns become derived", with the test
"could we answer a new question about a unit tomorrow without re-fetching it?" (:19–27).
The delegated storage question was ruled: raw bodies in D1, not R2 — ~0.28 GB at 15,169 units,
~3% of D1's 10 GB ceiling, with assets to R2 by hash-and-key, and a stated revisit trigger at 3 GB
(STORAGE-RULING:3, 15–44). Both documents are silent on retrospective retention — neither
rules on whether pre-existing captured bytes carry authority.
4 · VOLUMES, REPORTED-NOT-VERIFIED¶
Every figure below is as-written in its source. Nothing re-measured. Where two documents disagree, both are shown and neither is reconciled.
4.1 Units¶
| figure | population | source | date |
|---|---|---|---|
| 125,920 rows / 125,920 distinct codes | tga_training_components — whole register, all six types, all statuses |
INGESTED-SUBSTRATE-INVENTORY-01-REPORT:38 |
2026-08-06, measured |
| 75,267 | register codes of type unit, all statuses |
INVENTORY-REPORT:48; corroborated by live API CVIG-CENSUS:113 and TGA's own facet TGA-REGISTER-RECONCILIATION:26 (Δ 0) |
2026-08-06/08 |
| 15,169 | unit × status current, register mirror; content coverage 15,169 of 15,169, gap 0 |
INVENTORY-REPORT:83; re-verified VERDICT:35 |
2026-08-06 |
| 15,089 | units ⋈ tga_training_components, status_slug='current' — population 75,187 |
PITH-BASELINE:32; PITH-DELTA:113 |
2026-08-05, unchanged 08-08 |
| 15,165 | TGA's own live Current facet; ours 15,169; Δ +4 | TGA-REGISTER-RECONCILIATION:53–55 |
2026-08-06 |
| 15,200 | units.enriched_at IS NOT NULL — the KN corpus, sacred |
PITH-BASELINE:78; PITH-DELTA:123 |
Δ 0 across the cron cycle |
| 680,071 / 49,894 | unit content rows / unit codes with ≥1 content row | INVENTORY-REPORT:48 |
2026-08-06 |
| 25,373 / 33.7% | unit register codes with no content row | INVENTORY-REPORT:72 |
2026-08-06 |
| 261,186 | performance criteria extracted, clean-room, 15,169 current units | TGA-CLEANROOM-INGEST-01-PARSE-REPORT:38 |
2026-08-04 |
Three-way disagreement to carry into the design, unreconciled in the record: 15,169 · 15,089 ·
15,165. The arithmetic shape of the first is that units holds 75,187 rows against the register's
75,267 — 80 rows short, named in no document.
4.2 Qualifications¶
8,034 register codes all statuses (INVENTORY-REPORT:47; TGA live facet agrees, Δ 0) ·
1,164 current, coverage 1,164 of 1,164, gap 0 (:87) · 8,007 rows in the qualifications
table — 27 short of the register, unexplained (PITH-BASELINE:19) · 1,137 register-current
by the join (PITH-BASELINE:42) — again beside 1,164, unreconciled · 5,229 qual codes with
content over 55,006 rows (:47) · 12,902 release rows (:219) ·
626,914 validated cross-reference edges over 50,808 codes (QUAL-PARSE-01-G4-VERDICT:32).
4.3 RTOs — the thinnest dataset in the inventory, and the figures conflict¶
| figure | population | source | date |
|---|---|---|---|
| 12,515 | rtos — "all registered training organisations" |
docs/docs/data-sources.md:86 |
audited 2026-03-29 |
| 12,506 | tga_organisations rows = R2 archive codes, both directions Δ 0 |
ORG-RAW-ARCHIVE-BACKFILL-01-part2-close.md:12–15 |
2026-07-03 |
| 12,500 | tga_organisations |
data-sources.md:128 |
2026-03-29 |
| 12,508 | radar_dossiers; +2 over 12,506, "unreconciled" |
RADAR-SUBSTRATE-RECON-01-report.md:27, 39–41 |
2026-07-04 |
| 3,907 current / 4,016 with variants / 8,490 historical | tga_organisations.registration_status_label |
part2-gate4-validation.md:11–14 |
2026-07-03 |
| 3,916 current RTOs | re-ingested by the full-fleet enrich-sync run | ORG-RAW-ARCHIVE-BACKFILL-01-brief.md:19 |
2026-07-03 |
| 937 RTOs with regulatory history; 420 currently active with decisions against them | vw_rto_risk_profile |
data-sources.md:151–152 |
2026-03-29 |
| 1,293,177 rows / 3,920 orgs delivery-notification history | one-off March backfill, "nothing since" | RADAR-REGISTER-01-gate1-audit.md:120–138 |
frozen 2026-03-28 |
The ~3,985 operating RTOs figure could not be sourced. It appears once in the repo, in the North
Star itself (prototype-to-product-north-star.md:8), attributed to "the live TGA register".
3,887 returns zero hits repo-wide; the 76/22 breakdown appears nowhere; there is no
suspended-RTO status axis in any organisation document read. The nearest measured figures —
3,907 / 3,916 / 4,016, all dated 2026-07-03 — bracket 3,985 without any of them being it, and the
only instrument these documents describe for reading the live org register is proven to return a
~240 shuffled pool. See finding F-01.
4.4 CVIGs¶
562 unique files (501 Companion Volume + 61 Companion Volume Supporting Material) ·
1,001,808,227 bytes / 955.4 MiB, measured by HEAD on all 562 and confirmed byte-for-byte on
read-back · 51 of 54 current packages hold volumes; 3 have none in existence ·
574 (package, release) pairs queried · mean file 1,782,576 B, largest 25,260,989 B ·
.pdf 499 · .PDF 10 · .xlsx 36 · .docx 16 · .zip 1 · 562/562 HEAD 200, zero dead links.
Enumeration cost ~1,200 calls. Stage-B pull priced at 562 requests / 955.4 MiB / ~15–25
minutes at concurrency 2. (CVIG-ACQUIRE-01-CENSUS:15–18, 158–208, 446–448, 502;
GATE4-CERT:14–22.)
The trap, and it is load-bearing for any fresh pull: "Querying only each package's latest
release finds 106 of 562 volumes. 456 volumes — 81% of the corpus — are reachable ONLY through
historical release queries" (CENSUS:426–433). And /files/zip?includeHistory=true is not a
shortcut — ACM's current-release zip holds 9 files against a true total of 22 (:438–440).
4.5 Release history / versions¶
103,473 rows / 87,245 distinct component codes, the (code, release) pair unique
(INVENTORY-REPORT:214) — by type: unit 84,173 · qualification 12,902 · skillSet 5,461 ·
trainingPackage 931 (:218–219), and "It holds no row for any of the 38,681 accredited
components" (:220). Exactly one code in 125,920 has partial release coverage — UEESS00078
(:232–238). Release dates span 1997-09-23 → 2026-07-30 (:171); the latest release date on
any content-gap code is 2012-08-17, with zero gap codes dated 2015 or later (:167–168).
259 training packages, Status/Id facet {0: 54, −4: 14, −5: 20, −6: 171}, matching TGA's own UI
facet exactly (CVIG-CENSUS:115–120).
4.6 Cadence¶
TGA's own publication cadence is not recorded anywhere in the population read. No
releases-per-period figure attributed to TGA exists. What exists are proxies: register drift of
−44 rows over five days (2 accreditedCourse + 42 accreditedUnit new, plus 4 units transitioning
current → supersededEquivalent), with the composition explicitly labelled inference
(TGA-REGISTER-RECONCILIATION:32, 64–74); in-release amendment observed once, CHC R11.0 v1.0 → v1.1
in six weeks (DOSSIER:124, n=1); and modal snapshot series length 5 over the 20% of versions/
enumerated (TGA-VERSIONS-CENSUS-01-REPORT:31–33).
Our own cadence is what the documents actually quantify, and it is worse than it looks:
tga-sync runs 0 16 * * SAT; Saturday 2026-08-08 touched 453 rows of 125,920 — 0.36% of the
mirror — and every single one was already deleted (PITH-DELTA:150–157). 5,456 components
marked current have not been re-read from the register since June (:167). An off-schedule
writer exists — 7 rows carry synced_at on a Thursday, and neither cron fires on a Thursday
(:182).
4.7 Figures the record itself flags as wrong¶
Carried because a gather report that omits the withdrawn figures invites their reuse.
TGA-CORPUS-REINGEST-01-STUBis withdrawn at premise — "Do not execute, do not draft from, do not cite as premise." (:1–9). Its ~65,000 population is superseded by 84,728, which in turn contradicts the measured 125,920. All three stand in the record.- 87,024 — the census figure that was never reported, because four consecutive chunks returned
identical counts with a byte-identical cursor. "a wrong number wearing a success flag"
(
TGA-VERSIONS-CENSUS-01-REPORT:72–74). Candidate guard: "A paginated continuation proves progress by a monotonically advancing boundary key, never by an accumulating count." - Accredited-unit coverage by row presence is wrong by 1,189 — 1,189 NULL-content shells;
effective coverage zero of 19,247 (
INVENTORY-REPORT:339–341). data (5).xlsxis eliminated as the enrichment input; the original timestamp was an hour wrong becausestat -f%SBrenders in the current timezone and August is AEST while March was AEDT (TGA-EXPORT-PROVENANCE-RECON-01-CORRECTION-01:10–35).- The deletion-breaks-qualification-linkage conclusion is withdrawn — the measurement was real,
the relation was misread;
mappingInformationis unit-to-unit equivalence, not qualification membership (CORRECTION-01:91–97). - CVIG
titleandcreatedOnare advertised values, not truth — one file's are wrong by three years;releasesis the applicability truth (CENSUS:376–379). data-sources.mdis stale and wrong in places — last audited 2026-03-29; it carriestga_training_componentsat 84,728 against 125,920 measured, and claims a weekly incremental pipeline for delivery-notification history that grep proves does not exist (RADAR-REGISTER-01-gate1-audit.md:125–131).- A digest fabrication occurred and was self-caught — a commit message carried a full-length
sha256 whose tail was invented (
INGESTED-SUBSTRATE-INVENTORY-01-VERDICT:138–146).
5 · RATE AND SHAPE CONSTRAINTS¶
5.1 The 88-operation surface — verified in this window¶
[computed here] I counted the eight captured swagger specs at
outputs/tga-swagger-evidence/*_V1.swagger.json (retrieved 2026-07-04T19:41:10Z) with a JSON pass
over paths × HTTP methods:
| group | paths | operations |
|---|---|---|
| TRAINING | 29 | 29 |
| ORGANISATION | 20 | 20 |
| METADATA | 13 | 13 |
| EXPORT | 12 | 12 |
| SEARCH | 8 | 8 |
| REPORT | 3 | 3 |
| CONTENT | 2 | 2 |
| FEEDBACK | 1 | 1 |
| total | 88 | 88 |
88 paths and 88 operations — every path carries exactly one method, which resolves the
paths-vs-operations ambiguity between tga-ntr-api.md:101 ("88 documented paths") and the delta
audit's "88 operations". Both documents were right and were counting the same thing.
Five of eight groups have never been called in production — consumed: SEARCH, TRAINING,
ORGANISATION; not consumed: EXPORT, METADATA, REPORT, CONTENT, FEEDBACK
(TGA-SWAGGER-VS-REPORTS-DELTA-AUDIT-01-report.md:179–182). Any anchor-one plan that leans on
EXPORT or METADATA is operating on zero measured evidence.
5.2 Paging — the ceiling is real and already breached¶
offsetdocumented max 100,000;pageSize=100verified; upper bound onpageSizenot established (tga-ntr-api.md:193, 780).totalCountat offset 0 is 125,920. Headroom is negative. "25,920 components — 20.6% of the register — are unreachable byoffsetpaging at any page size… A paging fix built onoffsetalone will terminate 25,920 short and report complete" (TGA-RATE-CEILING-01-REPORT:92–100).- The workaround is to partition, not to page. Six
Type/Idbuckets sum to 125,920 against an unfiltered 125,920 — "Delta 0. Exact reconciliation." Largest single bucket 75,267; current units 15,169 (TGA-FILTER-AXIS-NUMBERS:26, 38, 44).RecognitionManager/Codeis not safe as a partition axis — its facet truncates at ten values and sums to 121,982 (:116). pageNumberis silently ignored. Pages 1, 200, 400, 600, 800 return the same first three records; 500 records yielded 102 distinct codes (TGA-SEARCH-FIELD-VALIDATION-01-REPORT:26).offsetworks — 500 returned, 500 distinct, 0 duplicates (TGA-RATE-CEILING-01-REPORT:38).orderbyis mandatory or rows vanish silently. An unordered walk of current units reached exhaustion 2,011 codes short of the register's owntotalCountand "re-running it returned the same short set";orderby=Code ascreturned all 15,169. "The omission is silent, stable and reproducible. It does not look like an error." (tga-ntr-api.md:198–205).- Read
totalCount, nevercount—countis the page count (:211–213). - No cursor, no ETag, no If-Modified-Since, no changed-since feed documented anywhere. Zero grep
hits for
cursoracross the API docs.
5.3 Throttling — no limit has ever been observed, and the operating rule is still open¶
| run | requests | concurrency | achieved | 429s |
|---|---|---|---|---|
| TGA-RATE-CEILING-01 ramp | 180 | 1 → 40 | 76.1 req/s at c40 | 0 |
| TGA-CLEANROOM-INGEST-01 | 98,142 | 20 held | ~216 req/s peak | 0 |
| NRT-CLONE-01 Stage C | 296,621 | 20 held | 163.1 req/s over 30.3 min | 0 |
"No HTTP 429 has ever been observed from this API, at any rate, across roughly 480,000
requests." No Retry-After, no X-RateLimit-*, no CAPTCHA (tga-ntr-api.md:711–730). Latency
did not degrade under sustained load — per-minute p50 stayed in a 76–93 ms band and the second half
of the 30-minute run was faster than the first (:740–749). "The ramp ended because it ran out of
levels, not because TGA pushed back" (TGA-RATE-CEILING-01-REPORT:123).
Three cautions that must ride with those numbers.
The ~15 RPS figure in circulation is ours, not TGA's — "BASE_DELAY_MS = 67… yields ~15
req/s and was never measured. It is our own throttle wearing the label of a measurement" (:770).
The operating limit is formally unresolved — "THE OPERATING LIMIT IS NOT SETTLED, AND IT IS
NOT SETTLED HERE"; the ruling of record is concurrency 20, and an earlier draft mis-transcribed
it as "20 req/s" — "different dials… they differ by ~8×" (:763–767). And the CVIG endpoint
family carries a tighter standing discipline: concurrency ≤2, ≥0.5 s delay, one backoff retry,
identifying user-agent (:488–491).
Longest continuous measurement is ~30 minutes — "Whether the absence of any rate limit persists
over a multi-hour run" is listed as not established (:778). A full-corpus pull is a multi-hour
run.
The real bottleneck is ours, not theirs: end-to-end training_refresh is ~1,250 ms per record
of which our own write path is ~1,168 ms — TGA is 6.6% of the time we spend per record
(TGA-RATE-CEILING-01-REPORT:128–134).
5.4 Bulk / export¶
Twelve EXPORT operations exist (bulk CSV/Excel: organisations, org scopes, training components,
qual unit-grids, skillset unit-grids, completion mappings). Not one has ever been called —
grep-confirmed zero /export|/csv|/excel calls across the codebase; fetch architecture is
"100% per-record or paginated search" (DELTA-AUDIT:129, 180, 187). Whether they are usable for
a full pull is not recorded anywhere; the strongest statement on file is a value judgement
made without fetching one ("efficiency, not new data"). Separately:
"There is no bulk path to unit content. Per-unit fetch is the only route."
(TGA-INGEST-FINDING-AND-SPEC-01:86).
5.5 Change detection — there is none, and C4/C5 say why¶
The relay asked after "the C4 delivery-notification and C5 no-change-endpoint findings". Both labels
exist in TGA-SWAGGER-VS-REPORTS-DELTA-AUDIT-01-report.md, and neither is about a notification
channel or a no-change endpoint — the phrase "delivery notification" there names a dataset
(per-RTO per-unit delivery events), not a push mechanism.
C4 (:105–135) is the moat list: nine Swagger-exposed, report-absent datasets ranked by product
value. Delivery-notification history is #1 and unhoovered as a live pipeline — 1,293,177 rows
from a one-off March backfill, no worker refreshes it. Also unhoovered: registration/recognition
managers, release document bundles + files, reverse-usage indexes, prerequisites.
C5 (:137–153) is the inverse, and it is the more useful finding for anchor one:
"No second upstream channel found" — the PowerBI reports are a strict derivation of the same
NTR substrate, "not evidence of a second data source". The API surface is the whole obtainable
set. And the gap it names is temporal: "there is no org-scope-change / org-structure-change /
accredited-course-change endpoint anywhere in the 88 ops" (:141). Training components do
carry real version history via releases; organisations and scope do not. A full fresh pull is a
snapshot; any change history is manufactured by diffing snapshots.
The only freshness instrument is GET /api/metadata's
dataExportSynchronisationDateTime — a single platform-wide scalar with no per-component
granularity. Standing discipline: "Stamp it at the start and end of every run. A change mid-run
means the source moved underneath the capture" (tga-ntr-api.md:355). Note that production is
currently serving 4.225.0-beta.1+46 — "a beta build in production… Pin nothing to it long-term"
(CVIG-CENSUS:508–513).
5.6 Licence¶
Site content is CC BY 4.0 with required attribution "© Commonwealth of Australia."
(tga-ntr-api.md:51–53) — but that does not extend to companion volumes, which are third-party
material under that licence's own exception; operating posture is ingest-and-derive with no
redistribution (:463–465). See Q-P2P-010.
6 · QUESTION BLOCKS¶
Three, each meeting the litmus in TWIN-WINDOW-PROTOCOL-01 §4. Everything else is in the findings list below.
Q-P2P-009 — Gate 6: what provenance, if any, crosses into the clean Pith
CONTEXT: Gate 6 asks whether any part of the old mirror is quietly observed-truth. Half is
now answered from the record. R2 `tga-content` was created 2026-04-07 and its earliest object
is 2026-04-14; the recon states "nothing in it predates the cutoff. The question cannot
discriminate on this bucket" (TGA-CONTENT-BUCKET-RECON-2026-08-04.md:94-95). TGA publishes no
per-component modified date on any of the four specs read
(TGA-MODIFIED-DATE-READING-01:32,40,90); the only vintage available is our own `synced_at`
(:100-102). So three classes of held data cannot be regenerated by any fresh pull, at any
cost: (a) `synced_at` on 125,920 component rows, spanning 2026-06-28 -> 2026-08-08; (b) the
dated `versions/` snapshot series, >=144,674 objects, earliest 2026-05-09, whose vintage is
in the key; (c) the 2026-03-09 enrichment run's provenance, whose input is gone
(TGA-EXPORT-PROVENANCE-RECON-01-CORRECTION-01:44-46). Class (c) is already dead by the KN
ruling. Classes (a) and (b) are not addressed by any ruling I can find.
The other half of gate 6 is NOT answered and I did not investigate the live estate to close
it: no document anywhere compares what our mirror holds against what TGA will serve today.
Two instruments are named in the record, both costing zero TGA requests, neither run - the
`versions/` size-diff census (started, stopped at 20.06% coverage, zero differing series over
5,823 codes, same-length sample UNRUN) and a per-code join of `training/` against the D1 code
set (TGA-DRAIN-01-VERSIONS-INSPECTION:107-108; TGA-CONTENT-BUCKET-RECON:88,117).
QUESTION: does anchor one carry any pre-existing provenance across into the clean Pith, or is
the new Pith's provenance stamped only from its own ingestion date forward?
OPTIONS:
(a) Nothing crosses. Cleanest, matches "rebuild over migrate" exactly; we permanently lose
any ability to say when a given row was first observed before the rebuild date.
(b) Carry the `versions/` date series as a separate observed-truth table in the backplane,
stamped as observed and never as register truth. Preserves a five-month observation
window at low cost; adds a table to the clean Pith that is not regenerable, which
complicates the "reconstructible at birth" property doctrine 2 wants.
(c) Deliberately abandon it in writing under doctrine 1's abandonment clause, as Radar's
history was. Same outcome as (a) but leaves a decision in the record rather than a
silence.
BLOCKED: the provenance columns in the clean-Pith schema, and whether the anchor-one brief
needs a "carry-across" stage at all. Also: whether the two unrun zero-cost censuses should be
run before the design freezes, or dropped.
Q-P2P-010 — CVIG licence stop condition versus anchor-one scope
CONTEXT: North Star §5 names CVIGs as one of the five anchor-one datasets. A filed recon
finding says they cannot be stored. CVIG-LANDSCAPE-RECON-01-DOSSIER-2026-08-08 §5 Finding 4:
HLT R11.0's companion volume carries "CC Attribution-NonCommercial-ShareAlike" verbatim; "The
NC term restricts ingestion into a commercial product absent permission"; and the stop
condition is stated as "no CVIG content enters any RTOpacks store until a written legal
position exists" (:131-147). CVIG-ACQUIRE-01-AMENDMENT-01:44-48 adds that training.gov.au's
site-wide CC BY 4.0 carries a third-party-material exception and "Site licence does not
launder document licence". Licence posture is measured to be inconsistent across documents
from the same publisher; CHC's variant is unspecified in the extracted text; UEE has no notice
found; the DEWR CC BY 4.0 default is not confirmed as applying.
Note the acquisition already happened under the ingest-and-derive posture: 562 files /
1,001,808,227 bytes are in R2 `companion-volumes-ref` in the Prototype account, Gate 4
certified. That corpus is regenerable in ~15-25 minutes at concurrency 2 (CENSUS:446-448), so
nothing is lost by leaving it where it is.
QUESTION: what is CVIGs' scope inside anchor one, given a filed stop condition on storage?
OPTIONS:
(a) Acquisition and metadata only - hold the file listing, uri, sha256, size, per-volume
licence field, and the release applicability range, with no document bytes in the
backplane. Anchor one ships four and a half datasets, and the fifth arrives when legal
does.
(b) Full ingest under the ingest-and-derive posture - bytes in, no verbatim customer-facing
reproduction, no redistribution. Faster, and it is what the acquisition arc already
assumed; but it proceeds past a stop condition a filed document set.
(c) Sequence the legal position before the anchor-one brief is cut. Cleanest, and it is a
calendar-waiting item that could run alongside an active brief per the drip amendment.
BLOCKED: whether the clean-Pith schema needs a CVIG document store at all, and whether the
anchor-one brief is written against four datasets or five.
Q-P2P-011 — "through the proven parsers": the three parsers are in three different states
CONTEXT: North Star §5 says anchor one "sits entirely in known territory - our upstream, our
parsers, demonstrated by triangulation - and touches nothing in the frozen estate: zero blast
radius." The zero-blast-radius half is correct and I found nothing against it. The parsers
half is not uniform, and one of the three is in worse shape than the sentence implies.
1. QUAL-PARSE-01 (qualification cross-reference edges) - genuinely proven. Corpus run,
626,914 validated edges, Gate 4 PASS on independently re-hashed digests and re-derived
arithmetic, Gate 5 signed, arc closed. No caveat except that resolution was against our
mirror, not TGA.
2. The clean-room unit parser - ran at 15,169 units and 98.5%, but is KNOWN DEFECTIVE and
the fix never ran: "The clean-room parser loses the first criterion after an element
heading" (TGA-CLEANROOM-INGEST-01-VERIFICATION-03-REVIEW:33-34); PARSER-FIX was scoped
at five gates and is recorded as not started (RETURNS-R1-R4:125). Its output was never
loaded anywhere. Full-depth regression against the existing Pith left 2,618
disagreements over 304 units with "which side is right NOT decided here" (D8:62-64).
3. NRT-PARSE-02, the parser the ratified atom schema governs - HAS NEVER RUN AT CORPUS
SCALE. "Gate 4 corpus launch only after all three fix rounds close"
(NRT-PARSE-02-V13-TAP-RECORD:38). Schema is ratified at v1.2 with rulings R26-R31 tapped
and binding while the document still reads v1.2 (:8-9).
Separately, and smaller: "cross-model triangulation" appears nowhere in the ingestion record -
zero hits across all 17 pilot/parse/schema documents. The verification that does exist is
stronger and differently named (dual-instrument agreement, guard polarity proven offline,
independent tripwires, prediction-then-measure, digest re-hashing).
QUESTION: does anchor one wait on PARSER-FIX and the NRT-PARSE-02 corpus launch, or does it
ingest raw-first and treat parsing as a later re-runnable pass against stored bytes?
OPTIONS:
(a) Raw-first, parse later - exactly the clean-room brief's own discipline: "Parsing is a
separate, later, re-runnable pass against stored raw. If we parse on the way in we bake
today's mistake into tomorrow's database" (TGA-CLEANROOM-INGEST-01-BRIEF:32-41). The
backplane's first edition would then be raw + provenance, with parsed atoms following
as a second edition. Anchor one starts now and is not blocked by either parser.
(b) Close PARSER-FIX and the NRT corpus launch first, then ingest once through fixed
parsers. Fewer editions, but anchor one waits on five gates plus three fix rounds.
(c) Split by dataset - quals through the proven parser now, units raw-first.
BLOCKED: the anchor-one sequencing, and whether North Star §5's "through the proven parsers"
wants amending to name which parser is proven for what. That amendment is a ruling, not a
Driver edit.
7 · FINDINGS — everything that did not meet the litmus¶
F-01 · The ~3,985 operating RTOs figure cannot be sourced, and it is load-bearing on the
canonical design ceiling. It appears once in the repo — in the North Star itself
(prototype-to-product-north-star.md:8), attributed to "the live TGA register". 3,887 returns
zero hits repo-wide; the 76 / 22 breakdown appears nowhere; no organisation document read has a
suspended-RTO status axis. The nearest measured figures are 3,907, 3,916 and 4,016 current, all
dated 2026-07-03, none reconciled to the others — and the only live-register instrument these
documents describe (/search/organisation) is measured to return a ~240 shuffled pool of 12,506
and cannot produce a register-wide count. This is doctrine 10's failure mode
("Figures carried between documents are re-derived or dated") three lines from where the figure
sits. The 4,000 / 20,000 ceiling itself is not threatened — 4,000 has headroom over every
candidate figure — so this is not an escalation. It wants a dated re-derivation before it is cited
again.
F-02 · The unit population has three answers and the design must pick one. 15,169 (register
mirror, current) · 15,089 (units ⋈ components) · 15,165 (TGA's own live facet, Δ +4 against
ours). The arithmetic shape is that units holds 75,187 rows against the register's 75,267 —
80 rows short, named in no document. Quals show the same pattern: 1,164 / 1,137, and 8,007 against
8,034. Whatever anchor one ingests, its population statement should be written before the first
call, not derived from what came back.
F-03 · Scope drives cost by roughly 5×, and the North Star does not state it. Current units 15,169; all-status units 75,267; the whole register 125,920. The clean-room brief chose everything ("All training components, every status… Every release of every component"), the finding-and-spec chose current-first ("the full 75,267 at the same rate is 6–16 days"). This is a design decision, named here, not taken.
F-04 · A full fresh pull must partition and must order. offset alone cannot reach 20.6% of
the register; the Type/Id × Status/Id partition reconciles to 125,920 exactly; omitting
orderby=Code asc loses rows silently, stably and reproducibly. Either omission produces a short
corpus that reports success.
F-05 · 81% of companion volumes are reachable only through historical releases. Latest-release enumeration finds 106 of 562. The zip endpoint is not a shortcut.
F-06 · Delivery-notification history is a moat asset with no pipeline. 1,293,177 rows over
3,920 orgs, hoovered once on 2026-03-28 by a Python script that is its only writer, nothing since.
data-sources.md claims a live weekly incremental path; grep proves no TS/JS worker writes the
table. It is regenerable from TGA (the endpoint exists), so it is disposable under doctrine 1 —
but the volume precedent is useful: a per-RTO-per-unit event stream at that scale is tractable.
F-07 · docs/docs/data-sources.md is four months stale and wrong in named places. Last audited
2026-03-29. It carries tga_training_components at 84,728 against 125,920 measured; claims "Weekly"
cadence for paths its own gaps table describes as "one-time seeds from March 2026"; and states the
delivery-notification cadence refuted in F-06. It should not be used as a source for anchor one.
F-08 · Two zero-cost censuses are owed and unrun. The versions/ size-diff (stopped at 20.06%)
and the per-code join of training/ against the D1 code set. Both cost zero TGA requests. Both
would convert gate 6's unanswered half into a number. Named in Q-P2P-009 because whether to run
them before the design freezes is the Architect's call.
F-09 · curl cannot reach training.gov.au. TLS-layer refusal, HTTP 000 after full timeout, 100%
failure. Node fetch, Playwright and Cloudflare Workers all work. Any anchor-one runner built on
shell curl or Python urllib will present as an outage rather than a bug. Workers are proven
against /api/* by the running fleet.
F-10 · Our own write path, not TGA, is the throughput ceiling. TGA is 6.6% of per-record time. Sizing the ingestion against TGA's rate is sizing against the wrong dial.
F-11 · The record's own verification vocabulary is worth lifting into the brief verbatim.
nrt-atom-schema.md:88–95 §6 is the best statement of ingestion gates in the repo. Add to it the
three guards the strand paid for: G21 ("an ignored API parameter returns success… The test is not
'did it return 200'. It is 'did the number move'"), G23 ("a parser is checked against an
independent count of what it should have found, never only against its own success flag"), and the
candidate minted at the versions census ("a paginated continuation proves progress by a
monotonically advancing boundary key, never by an accumulating count").
F-12 · North Star §1's "cross-model triangulation" is not in the record. See §2.1. Amending is a ruling; flagged, not made.
Least sure, and what would make me wrong: the RTO conclusion. I am asserting there is no proven
path to the organisation register, and that assertion rests on nine documents I found only after my
first sweep's name patterns missed the whole strand — which is exactly the instrument failure this
report warns about elsewhere. If there is an organisation-side reconciliation filed under a name
that contains neither "tga", "org", "rto" nor "radar", I did not see it, and §1.3 overstates the
gap. The cheapest check is to ask Alex for any document that reconciles tga_organisations against
a live TGA organisation facet the way TGA-REGISTER-RECONCILIATION-2026-08-06 does for training
components. Everything else here is quoted from bytes I read.
— Driver seat (Opus), 2026-08-14. Read-only. Nothing mutated.