TGA API — observed reference¶
Working notes on ws.training.gov.au as it actually behaves, recorded where it diverges
from its published contract. This file exists because a divergence has already cost measurement
time, and the next instrument built against the swagger should not have to rediscover it.
Standing: observations, each with the measurement that produced it. Not canon; not an ADR. Where an entry contradicts the swagger, the entry is what the substrate did.
Divergence 1 — UnitGridUsage.type is declared but absent from every live row¶
TRAINING_V1 declares UnitGridUsage.type. It is carried by none of the rows we have seen.
| Measured | 1,425 of 1,425 unitgrid rows sampled across 40 bodies carry no type field |
| Source | A1-CORPUS-01 Phase C, phase-c.mjs (1de16d40…), T0 freeze 2026-08-15 |
| Filed | outputs/A1-CORPUS-01-PHASE-C-CONFIRMBACK-2026-08-15.md §3 |
A live row is shaped:
{"code":"ACMGEN101","hasPreRequisites":false,"isEssential":true,
"isEssentialLabel":"Core","links":[…],"title":"…"}
Consequence, and the reason this is written down: any instrument that filters unitgrid rows on
type returns zero, silently. It will not error. It will produce an empty result that looks
like a measurement.
In Phase C the typeIdAllowlist filter excluded nothing precisely because nothing had a type to
be filtered on — the fail-closed branch kept every row and flagged 2,331 anomalies, exactly one
per non-empty body (2,345 bodies − 14 empty grids = 2,331, exact). Had the filter been
fail-open, the extractor would have produced 0 and the run would have halted on its
known-present control. Both behaviours are correct; the point is that a fail-open filter here
produces a plausible zero.
If you are building against unitgrid: treat type as absent, derive unit-vs-accreditedUnit
type from the register rather than the grid, and give any type-conditional filter a
known-present control that fires on zero.
Divergence 2 — completed is returned where the contract says complete¶
The published contract names complete; the API returns completed. Recorded during the PROBE
arc against the UCCA engine's TGA-facing jobs; it broke the published contract for consumers
matching on the exact string.
Consequence: an equality match on complete never fires. Match on both, or normalise.
Divergence 3 — NrtStatusId is not the enum for scope status¶
The swagger declares one set of six values and the scope data serves a different six. Declared
identically in TRAINING.swagger.json and ORGANISATION.swagger.json:
| the six | |
|---|---|
declared NrtStatusId |
current · pending · deleted · superseded · cancelled · nonCurrent |
observed in scope_entry |
current 4,037,551 · nonCurrent 893,805 · cancelled 837,281 · withdrawn 666,644 · pending 77,136 · suspended 14,032 |
Four are shared. deleted and superseded are declared and never occur; withdrawn and
suspended occur and are never declared.
| Measured | all 6,526,172 rows of the lifted tga-rto.scope_entry, via vocab_scope_status |
| Source | CANON-01 Gate 1 §2.6, 2026-09-04; swagger 26d6271a4ec43942 / b1e31a01dec591ce |
Consequence: a validator or enum type generated from NrtStatusId refuses 680,676 real rows
and accepts two values that never appear. Load the vocabulary from what the corpus contains, and let
an unknown value refuse — never from the contract.
Divergence 4 — a 403 that was not a property of the endpoint¶
/api/training/{code}/completion and /completionusage answered 403 "Completion mapping is
internal use only" on 2026-08-20. On 2026-08-28 the same paths served data.
| Measured | 8,978 bodies on disk from the 28 August build — completion 2,575, completionusage 6,403 — every one a JSON array of components; zero error bodies; two empty arrays |
| Source | TGA-MIRROR-LEDGER-01 Gate 2 §1, 2026-09-03 |
Consequence, and it is the general one: a refusal is an observation with a date on it, not a property of a path. Recording "403" once and carrying it forward as a fact about the endpoint cost this estate two families it already held. Re-test a refusal before citing it.
Divergence 5 — recognitionmanager serves an ARRAY, not the single object the swagger declares¶
TRAINING_V1.swagger.json declares GET /api/training/{code}/recognitionmanager → 200 as one
RecognitionManagerAssignment object. Every live body is a JSON array, and a code can carry two
assignments (a national council and a state authority), which the declared shape cannot express.
| Measured | 125,838 bodies, one per training code — every one a JSON array, zero objects. Lengths: 0 → 3,938 · 1 → 120,263 · 2 → 1,637 |
| Empty is an answer | The 3,938 empty arrays are exactly every skill set (3,679) and every training package (259). No unit, qualification, accredited course or accredited unit is empty |
| Source | FETCH-01 Gate 3 pull 20260911T095213Z (ledger 1315b3e605c05b70); census at Gate 4, outputs/fetch-01/FETCH-01-GATE-4-20260911T095213Z.md |
An instrument built from the swagger would read body.name and find nothing. Read body[0], and
expect 0, 1 or 2 entries.
Observed — tpusage is per RELEASE, and TGA scopes it in the 400¶
GET /api/training/{code}/releases/{releaseNumber}/tpusage returns the training packages a component
sits in. The swagger declares a 400 but not what it means; TGA's body says it:
{"status":400,"detail":"Method only available for units, skill sets and qualifications"}
| Measured | 102,409 bodies over the live releases of every unit, qualification and skill set — all JSON arrays, all parse. 932 training-package releases answer 400 and are ledgered as served absences |
| Per release, not per code | Of 11,247 products with more than one release, 0 serve the same body across their releases. Release membership is real structure — a store keyed per code would lose it |
| Bodies repeat anyway | 102,409 bodies carry 9,505 distinct digests, and the same body is served for different products (e.g. AHC30116\|1 and AHC50116\|1). Dedupe by digest is safe; the subject→body mapping is not |
| Source | FETCH-01 Gate 2 probe, Gate 3 pull, Gate 4 census (as above) |
Observed throughput and its shape¶
Not a contract question, but the number every ingest design needs.
| observation | figure | source |
|---|---|---|
| Sustained end-to-end fetch rate, A1-CORPUS-01 Phase E | ~3.1 fetches/s, flat from 4 to 20 concurrent lanes | PHASE-E-LADDER-SNAPSHOT-2026-08-16.json (ddd94082…) |
| 429s observed across 8,000 fetches, tiers 4→20 | zero | same |
| Per-lane latency at 4 lanes → 20 lanes | 1,276 ms → 6,271 ms (derived, not measured) | same |
| Sustained ONE-LANE unpaced rate, FETCH-01 Gate 3 | 6.77 fetches/s over 228,262 requests in 9 h 22 m; fastest 1,000-request window 7.35/s | outputs/fetch-01/FETCH-01-GATE-4-20260911T095213Z.md, from the pull's own ledger timestamps |
| 429s observed across 228,262 fetches at one lane | zero | same |
The one-lane figure is 2.2× the A1-CORPUS-01 rate above, on a single lane — which supports the suspicion already recorded here that ~3.1/s was our pipeline and not TGA. This row is a measurement, not a ruling: RULING-LOCAL-FETCH-01 still names ~3.1/s as its envelope, and whether that is restated is Tim's (FETCH-01 Gate 3 verdict, 2026-09-12).
Read carefully: the flat throughput is a property of our pipeline, not a proven property of TGA. Each lane touches TGA, then D1 three times, then R2, in series — so the ~3.1/s ceiling is attributable to any of them from this data alone. The prototype's 151.5 rec/s against the same service is the reason to suspect the ceiling is ours rather than theirs, and that inference depends on the prototype figure being end-to-end rather than a local processing rate — unconfirmed at the time of writing.
Zero 429s across 8,000 fetches is the solid part: TGA did not rate-limit us at any concurrency we tried.
Maintained as observations accrue. Entries carry their measurement or they do not belong here.