Skip to content

TGA API — observed reference

Working notes on ws.training.gov.au as it actually behaves, recorded where it diverges from its published contract. This file exists because a divergence has already cost measurement time, and the next instrument built against the swagger should not have to rediscover it.

Standing: observations, each with the measurement that produced it. Not canon; not an ADR. Where an entry contradicts the swagger, the entry is what the substrate did.


Divergence 1 — UnitGridUsage.type is declared but absent from every live row

TRAINING_V1 declares UnitGridUsage.type. It is carried by none of the rows we have seen.

Measured 1,425 of 1,425 unitgrid rows sampled across 40 bodies carry no type field
Source A1-CORPUS-01 Phase C, phase-c.mjs (1de16d40…), T0 freeze 2026-08-15
Filed outputs/A1-CORPUS-01-PHASE-C-CONFIRMBACK-2026-08-15.md §3

A live row is shaped:

{"code":"ACMGEN101","hasPreRequisites":false,"isEssential":true,
 "isEssentialLabel":"Core","links":[…],"title":"…"}

Consequence, and the reason this is written down: any instrument that filters unitgrid rows on type returns zero, silently. It will not error. It will produce an empty result that looks like a measurement.

In Phase C the typeIdAllowlist filter excluded nothing precisely because nothing had a type to be filtered on — the fail-closed branch kept every row and flagged 2,331 anomalies, exactly one per non-empty body (2,345 bodies − 14 empty grids = 2,331, exact). Had the filter been fail-open, the extractor would have produced 0 and the run would have halted on its known-present control. Both behaviours are correct; the point is that a fail-open filter here produces a plausible zero.

If you are building against unitgrid: treat type as absent, derive unit-vs-accreditedUnit type from the register rather than the grid, and give any type-conditional filter a known-present control that fires on zero.


Divergence 2 — completed is returned where the contract says complete

The published contract names complete; the API returns completed. Recorded during the PROBE arc against the UCCA engine's TGA-facing jobs; it broke the published contract for consumers matching on the exact string.

Consequence: an equality match on complete never fires. Match on both, or normalise.


Divergence 3 — NrtStatusId is not the enum for scope status

The swagger declares one set of six values and the scope data serves a different six. Declared identically in TRAINING.swagger.json and ORGANISATION.swagger.json:

the six
declared NrtStatusId current · pending · deleted · superseded · cancelled · nonCurrent
observed in scope_entry current 4,037,551 · nonCurrent 893,805 · cancelled 837,281 · withdrawn 666,644 · pending 77,136 · suspended 14,032

Four are shared. deleted and superseded are declared and never occur; withdrawn and suspended occur and are never declared.

Measured all 6,526,172 rows of the lifted tga-rto.scope_entry, via vocab_scope_status
Source CANON-01 Gate 1 §2.6, 2026-09-04; swagger 26d6271a4ec43942 / b1e31a01dec591ce

Consequence: a validator or enum type generated from NrtStatusId refuses 680,676 real rows and accepts two values that never appear. Load the vocabulary from what the corpus contains, and let an unknown value refuse — never from the contract.


Divergence 4 — a 403 that was not a property of the endpoint

/api/training/{code}/completion and /completionusage answered 403 "Completion mapping is internal use only" on 2026-08-20. On 2026-08-28 the same paths served data.

Measured 8,978 bodies on disk from the 28 August build — completion 2,575, completionusage 6,403 — every one a JSON array of components; zero error bodies; two empty arrays
Source TGA-MIRROR-LEDGER-01 Gate 2 §1, 2026-09-03

Consequence, and it is the general one: a refusal is an observation with a date on it, not a property of a path. Recording "403" once and carrying it forward as a fact about the endpoint cost this estate two families it already held. Re-test a refusal before citing it.


Divergence 5 — recognitionmanager serves an ARRAY, not the single object the swagger declares

TRAINING_V1.swagger.json declares GET /api/training/{code}/recognitionmanager → 200 as one RecognitionManagerAssignment object. Every live body is a JSON array, and a code can carry two assignments (a national council and a state authority), which the declared shape cannot express.

Measured 125,838 bodies, one per training code — every one a JSON array, zero objects. Lengths: 0 → 3,938 · 1 → 120,263 · 2 → 1,637
Empty is an answer The 3,938 empty arrays are exactly every skill set (3,679) and every training package (259). No unit, qualification, accredited course or accredited unit is empty
Source FETCH-01 Gate 3 pull 20260911T095213Z (ledger 1315b3e605c05b70); census at Gate 4, outputs/fetch-01/FETCH-01-GATE-4-20260911T095213Z.md

An instrument built from the swagger would read body.name and find nothing. Read body[0], and expect 0, 1 or 2 entries.


Observed — tpusage is per RELEASE, and TGA scopes it in the 400

GET /api/training/{code}/releases/{releaseNumber}/tpusage returns the training packages a component sits in. The swagger declares a 400 but not what it means; TGA's body says it:

{"status":400,"detail":"Method only available for units, skill sets and qualifications"}

Measured 102,409 bodies over the live releases of every unit, qualification and skill set — all JSON arrays, all parse. 932 training-package releases answer 400 and are ledgered as served absences
Per release, not per code Of 11,247 products with more than one release, 0 serve the same body across their releases. Release membership is real structure — a store keyed per code would lose it
Bodies repeat anyway 102,409 bodies carry 9,505 distinct digests, and the same body is served for different products (e.g. AHC30116\|1 and AHC50116\|1). Dedupe by digest is safe; the subject→body mapping is not
Source FETCH-01 Gate 2 probe, Gate 3 pull, Gate 4 census (as above)

Observed throughput and its shape

Not a contract question, but the number every ingest design needs.

observation figure source
Sustained end-to-end fetch rate, A1-CORPUS-01 Phase E ~3.1 fetches/s, flat from 4 to 20 concurrent lanes PHASE-E-LADDER-SNAPSHOT-2026-08-16.json (ddd94082…)
429s observed across 8,000 fetches, tiers 4→20 zero same
Per-lane latency at 4 lanes → 20 lanes 1,276 ms → 6,271 ms (derived, not measured) same
Sustained ONE-LANE unpaced rate, FETCH-01 Gate 3 6.77 fetches/s over 228,262 requests in 9 h 22 m; fastest 1,000-request window 7.35/s outputs/fetch-01/FETCH-01-GATE-4-20260911T095213Z.md, from the pull's own ledger timestamps
429s observed across 228,262 fetches at one lane zero same

The one-lane figure is 2.2× the A1-CORPUS-01 rate above, on a single lane — which supports the suspicion already recorded here that ~3.1/s was our pipeline and not TGA. This row is a measurement, not a ruling: RULING-LOCAL-FETCH-01 still names ~3.1/s as its envelope, and whether that is restated is Tim's (FETCH-01 Gate 3 verdict, 2026-09-12).

Read carefully: the flat throughput is a property of our pipeline, not a proven property of TGA. Each lane touches TGA, then D1 three times, then R2, in series — so the ~3.1/s ceiling is attributable to any of them from this data alone. The prototype's 151.5 rec/s against the same service is the reason to suspect the ceiling is ours rather than theirs, and that inference depends on the prototype figure being end-to-end rather than a local processing rate — unconfirmed at the time of writing.

Zero 429s across 8,000 fetches is the solid part: TGA did not rate-limit us at any concurrency we tried.


Maintained as observations accrue. Entries carry their measurement or they do not belong here.