Skip to content

TGA National Training Register Web API — reference

Status: Reference doc filed under the EXT-API RULE (standing-rules.md) — any external API needs a reference doc in docs/docs/ops/ before deploy. This doc is a precondition of NRT-CLONE-01 Stage D.

Last verified: 2026-08-05. Site version 4.225.0.1. Register data export synchronisation at time of writing: 2026-08-05T04:37:39+00:00.

Amended 2026-08-08 by CVIG-ACQUIRE-01 Gate 5. That brief profiled the release-files endpoint in full and censused the national companion-volume corpus. Three changes, each marked in place: §1 the OpenAPI coverage claim is corrected (the spec documents far more than was thought); §1.2 the client requirement is new; §3.10 the files endpoint is promoted out of §3.7's "not fully profiled" list. Site version at amendment 4.225.0.1, register synchronisation 2026-08-08T06:40:50+00:00.

Provenance of every figure below: measured live against the API — by Claude at gate (browser, same-origin fetch), or by Alex in NRT-CLONE-01 Stages A, B and C. Nothing here is inferred from documentation. Where something is not measured it says so.

Where a figure comes from a full-population sweep rather than a sample, it says so and gives the denominator. Stage B swept all 42 type × status cells (125,920 components, 1,291 requests) and Stage C swept all 87,239 VET components; those two runs replace several sampled figures that appeared in the first draft of this doc, and the corrections are marked [was sampled].


1 · Identity

Service NTR Web API (NTR Web API, per its own OpenAPI info.title)
Owner Department of Employment and Workplace Relations (Commonwealth), operator of training.gov.au
Base https://training.gov.au
Secondary host https://content.training.gov.au — Drupal JSON:API for site chrome (menus, banners, contextual help). Not part of the register data surface; do not build on it.
Auth None. No key, no token, no cookie required for any endpoint below. Confirmed over 480,000+ unauthenticated calls.
Versioning api-version=1.0 query parameter. Optional on every endpoint measured; see §3.5.
OpenAPI https://training.gov.au/swagger/index.html. ~~Documents only the two content endpoints.~~ CORRECTED 2026-08-08 — the spec documents 8 groups and 88 paths, including search, training, releases and files. See §1.3.

1.1 Terms of use, licence and robots — read 2026-08-05, and there is no automation clause

Terms of use — https://training.gov.au/terms-use, last updated December 2017. Read in full. It is a disclaimer of warranty and a security-of-information notice, and nothing else.

It contains NO clause about automated access, scraping, crawling, robots, bulk download, API use, request rate, or acceptable use of any kind. The only obligation it places on a user is to take their own precautions against malicious code. There is no rate term to breach.

Copyright — https://training.gov.au/copyright-information. © Commonwealth of Australia 2024.

The content is licensed CC BY 4.0 — Creative Commons Attribution 4.0 International — excepting the Department's logo, the Commonwealth Coat of Arms, trademarked material and third-party material. Required attribution string, verbatim: © Commonwealth of Australia.

That is a mirror-and-redistribute licence with attribution. It is the licence basis for holding the clone and for surfacing register content to customers, and it must be carried on any surface that shows this content.

robots.txt does not exist. https://training.gov.au/robots.txt returns HTTP 200 with the Nuxt application shell — the SPA catch-all serves HTML for every unmatched path. A robots parser reading that response finds no directives and concludes "no restrictions", which is a correct conclusion reached by an accident: there is no robots.txt to disallow anything. Do not cite the 200 as permission; cite its absence as absence.

These pages are rendered client-side. curl on any of them returns the same 3,325-byte shell. The Drupal content API at content.training.gov.au/jsonapi returns 403 at its root. They must be read in a browser, and a curl that "returns 200" on a terms page has read nothing.

1.2 Client requirement — curl is refused at the TLS layer (new 2026-08-08)

curl cannot reach training.gov.au at all. TCP 443 connects — nc -vz training.gov.au 443 succeeds — and then the TLS/HTTP exchange stalls until timeout. Every curl attempt returns exit 28 after the full timeout with HTTP 000, on both the apex and www, with or without a browser user-agent string.

The site fronts with Dynatrace RUM instrumentation carrying an owasp=1 flag. The refusal is client-fingerprint shaped, not IP-shaped and not rate-shaped — it is not a ban, and waiting does not clear it. A real browser engine reaches every endpoint immediately and without special measures; headless Chromium driven by Playwright is the repo-standard instrument, and Playwright's context.request is sufficient for pure JSON and file fetches (no page needed).

This is distinct from the client-side-rendering note above. That note says a curl of a terms page returns the SPA shell instead of the prose. This says a curl of anything on this host, including the JSON API, returns nothing at all.

Operational consequence: a scheduled job built on plain curl or urllib will fail 100% and present as an outage or as "the API is down". It is neither. Before diagnosing a TGA outage, confirm the instrument is a browser engine. This is also why the OpenAPI spec was long thought unreachable — see §1.3.

A custom identifying user-agent is fine. RTOpacks-CVIG-Acquire/1.0 (+contact: …) was tested against the WAF before use and passed on both JSON and file endpoints; UA is not what is being screened.

1.3 OpenAPI coverage — corrected 2026-08-08

The earlier reading — "documents only the two content endpoints, the remaining eleven were captured from the site's own network requests" — was wrong, and the cause was the instrument, not the spec. The spec was never sparse; it was simply unreachable by curl (§1.2). Read with a browser engine, https://training.gov.au/swagger/index.html (title: NTR Web API) declares 8 spec groups totalling 88 documented paths:

Group Paths Spec URL
Content 2 /swagger/Content - v1/swagger.json
Export 12 /swagger/Export - v1/swagger.json
Feedback 1 /swagger/Feedback - v1/swagger.json
Metadata 13 /swagger/Metadata - v1/swagger.json
Organisation 20 /swagger/Organisation - v1/swagger.json
Report 3 /swagger/Report - v1/swagger.json
Search 8 /swagger/Search - v1/swagger.json
Training 29 /swagger/Training - v1/swagger.json

Note the literal spaces in each group name; they must be URL-encoded. The group list itself is declared in /swagger/index.js, not in index.html.

Every one of the 8 specs declares zero security schemes, which corroborates the "no auth" finding in §1 from the spec side as well as from measurement. The SPA's own runtime config (window.__NUXT__.config) independently carries authentication: { enabled: false }.

Endpoints previously listed as undocumented are in fact documented, including /api/search/training, /api/training/{code}, /api/training/{code}/releases, /api/training/{code}/releases/{n}/files and /api/training/{code}/releases/{n}/files/zip. The spec is now the better source for parameter names than network capture.


2 · The addressing rule — READ THIS FIRST

nrtId is the address. code is a display and join attribute.

Every component record carries nrtId, a GUID. Every path that takes a component identifier accepts it.

Component codes are register-authored strings and may contain the URL path separator. Slash-bearing codes exist in production — 004/01, 023/02, 127/20, 158/14 and others — and no encoding makes them addressable:

shape result
/api/training/004%2F01 404
/api/training/004%252F01 404
/api/training/004/01 404
/api/training/cf1d32fa-025e-499a-9df0-c400761e6de0 200, returns code: "004/01"

Measured on 4 shapes × 3 codes at gate, end-to-end on 10 consecutive slash-bearing components at gate, and again by Alex on 30 consecutive slash-bearing components — 30/30 detail 200, releases 1.00 each, and the returned code checked against the requested component on every one.

Check the returned code against the component you asked for. A GUID lookup that silently returned the wrong component is indistinguishable from success at the count level.

2.1 Prevalence — now measured over the whole register [was sampled]

158 of 125,920 components carry a slash — 0.13%, confined to three cells:

cell slash codes of rate
Unit of competency / Deleted 137 21,077 0.65%
Accredited unit-module / Non-current 18 12,777 0.14%
Unit of competency / Superseded (Non-Equivalent) 3 9,425 0.03%
all other 21 populated cells 0 92,641 0

The first draft carried 2.2%, from the first 2,000 Deleted units by Code asc. The true rate for that cell is 0.65% — the sample overstated it 3.4×, because Code asc clusters every slash code at the front of the page. Extrapolated, that sample would have predicted ~460 affected components against a true 158. A sample drawn under Code asc is not representative of code shape.

Consequence: a loader keying on code reports these as not found rather than not addressable. Thirty failed calls in the Stage C-prereq run are verified as this defect. Thirty more in Stage A are inferred from a matching count and cell, and the log that would prove it was never persisted — so the honest tally is 30 verified, 30 inferred, the evidence unrecoverable. The Stage C-prereqs ruling's "all 60 failed calls across both runs are this one defect" overstates by exactly half.

code is nevertheless globally unique — 125,920 rows, 125,920 distinct codes, zero collisions across all six types and all seven statuses. It is a sound join key and an unsound path segment, both at once.


3 · Endpoints

3.1 Search — GET /api/search/training

/api/search/training?api-version=1.0&pageSize=100&offset=0&includeTotalCount=true
  &orderby=Code%20asc&filter=Type/Id%20eq%204%20and%20Status/Id%20eq%20-4
param notes
searchText may be omitted entirely; an empty string also returns everything
offset / pageSize pageSize=100 verified. pageSize=1000 honoured (LISTING-SWEEP-01 Gate 2, 2026-09-14: a 200 with 1,000 rows, 1,000 distinct codes, totalCount 18,620 on Type/Id eq 1 and Status/Id eq -1; the whole 126,101-row listing then walked at 1,000, rows = totalCount = distinct on all 24 cells). offset documented max 100,000. Upper bound above 1,000 not established.
orderby lower-case orderby is what the API accepts (the swagger spells it orderBy). Code asc verified — and verified on /api/search/organisation too (LISTING-SWEEP-01 Gate 2: 13,151 rows = totalCount = distinct codes; the field sourced from the organisation record's code).
filter OData subset. Type/Id eq N, Status/Id eq N, and verified. Parentheses optional. contains(Code,'/') returns HTTP 500 — string functions are not supported.
includeTotalCount populates totalCount — the population total for the filter, not the page count.

3.1.1 orderby IS MANDATORY FOR ANY PAGED SWEEP.

Without an explicit order, offset paging permanently omits records. Measured during TGA-CLEANROOM-INGEST-01: an unordered walk of current units of competency reached exhaustion — pages stopped returning new codes — 2,011 codes short of the register's own totalCount, and re-running it returned the same short set. Adding orderby=Code asc returned all 15,169.

The omission is silent, stable and reproducible. It does not look like an error. A run that pages without an order will report success and a short corpus, every time, with the same shortfall.

Stage B swept all 42 cells with orderby=Code asc and reproduced the register's own totals cell for cell, 125,920 of 125,920, zero disagreements.

3.1.2 Read totalCount, not count.

count is the page count. totalCount, populated by includeTotalCount=true, is the population total and it agrees with the facets endpoint on all 24 populated cells. The first draft of this doc said the response total was a page count and directed loaders to the facets endpoint for totals; that is wrong on the field name, and a loader that follows it makes six extra calls for a number it already had.

3.1.3 A meaningless-but-well-formed filter returns HTTP 200 and nothing.

Status/Id eq 5 — 5 is not a status — returns HTTP 200, totalCount: 0, empty data. No error, no warning. A harness that does not print the population total before it fetches will record a clean "0 of 0 passed" and call it a result. Verify the number moved.

Record shape — VET example:

searchScore · code · title · titleUpper · nrtId · type{id,name,sortOrder}
status{id,isCurrent,name,sortOrder} · statusTypeRanking · usageRecommendation{name,startDate}
currencyPeriod{startDate,endDate} · latestRelease{date,number}
supersedes[{isEquivalent,code,title}] · supersededBy[{isEquivalent,code,title}]
preRequisiteUnits[] · taxonomyIndustry[] · taxonomyOccupation[] · trainingPackage{code,title}
trainingPackageDeveloper{legalName,organisationId} · recognitionManager{code,name,shortName}
hasWorkPlacementHours · isConfidential · asced6{code,description,name}

type.id and status.id in the record are STRING SLUGS, not the numeric ids the filter takes — "unit", "supersededEquivalent". See §4.1/§4.2 for both vocabularies. The number you filter with is never the value you get back.

supersedes[].isEquivalent / supersededBy[].isEquivalent is the per-relationship equivalence flag. See §5.2 — load-bearing, and it has no equivalent in our current substrate.

3.2 Facets — GET /api/search/training/facets

/api/search/training/facets?api-version=1.0&searchText=&filter=&facetProperty=Type/Id

Returns {count, data:[], facets:[{propertyName, counts:{<id>:<n>}}]}. facetProperty verified for Type/Id and Status/Id. Six calls (one per type, faceting on Status/Id) yield the complete type × status matrix.

Facets and totalCount agreed on all 24 populated cells. Either is a valid denominator; facets is cheaper for the matrix, totalCount is free on a sweep you are running anyway.

Measured 2026-09-13/14 (FACET-ROUNDTRIP-01, LISTING-SWEEP-01):

  • Verified facetProperty values: training Type/Id · Status/Id · QualificationLevel/Code · RecognitionManager/Code · TrainingPackage/Code · UsageRecommendation/Name · HasWorkPlacementHours · HasLicensingInformation · Asced4/Code · Asced6/Code · Anzsco/Code · Asco/Code · TaxonomyIndustry/Id · TaxonomyOccupation/Id; organisation Registration/Status · RegistrationManager/Code · RtoType/Code · IsRto · IsCurrent · HasCurrentRestrictionOrRegulatoryDecision · Roles/Id.
  • Not facetable — HTTP 500 on two attempts, each a well-formed property of the search record: IsConfidential, TrainingPackageDeveloper/OrganisationId, HeadOfficeAddress/State/Code, and all nine Delivery/{Act,Nsw,Nt,Qld,Sa,Tas,Vic,Wa,International}. The swagger sources no second spelling for any of them. Same class as componentType (CVIG-ACQUIRE-01).
  • A facet block returns at most 50 values. TrainingPackage/Code, Asced4/Code, Asced6/Code, Anzsco/Code, Asco/Code, TaxonomyIndustry/Id and TaxonomyOccupation/Id each returned exactly 50. A sum over a capped block is not a population; compare value by value over the 50 TGA shows.
  • On multi-valued and classification axes the facet counts fewer than TGA's own records carry — 1 to 22 short on 135 cells where the mirror equals TGA's listing exactly (TaxonomyOccupation/Id, Asco/Code, TaxonomyIndustry/Id, Anzsco/Code, Asced6/Code). The facet's arithmetic is TGA's; check against the records (RULING-TGA-COMPUTED-SUMMARY-01), never against the facet.
  • Facet keys are not always the record's value: TrainingPackage/Code keys come back case-mangled (hlT07, zzZ00), and Registration/Status keys are numbers while the record carries slugs — 0 current · 1 suspended · 2 pending · -1 nonCurrent · -2 cancelled · -3 withdrawn, verified by filtered reads.
  • File the listing beside the facets (standing-rules.md LISTING-NAMES-THE-ROWS-01 clause 3).

3.3 Component detail — GET /api/training/{nrtId}?api-version=1.0&include=all

Returns, for VET components:

code · id · title · type · releases[] · currencyPeriods · usageRecommendation
usageRecommendationLabel · mappingInformation · parent · preRequisites
developmentStandard · trainingPackageDeveloper

releases[] entries carry id and releaseNumber but contentBundles is EMPTY at this level. Bundles are populated only by §3.4. A loader reading bundles here gets zero for every component.

3.3.1 RELEASE NUMBERS ARE STRINGS AND ARE NOT CONTIGUOUS. Enumerate them; never generate them.

releaseNumber is a string that may carry a decimal part, and a component's release set is neither 1..N nor guaranteed to start at 1.

Measured on 64 VET components (4 from each populated cell) against latestRelease.number: 53 agree, 11 disagree. Every disagreement was a training package.

CPP07   latestRelease.number "15.0"  actual [7,8,9,10,11,12,13,14,14.1,…,14.7,15]   ← 1–6 DO NOT EXIST
AHC     latestRelease.number "11.0"  actual [1,1.1,2,3,4,5,6,6.1,6.2,7,7.1,7.2,8,9,10,11]
AMP     latestRelease.number  "9.1"  actual [1,2,2.1,2.2,3,4,5,5.1,6,7,7.1,8,8.1,9,9.1]
AHC10   latestRelease.number  "8.0"  actual [2,2.1,3,4,5,6,7,8]
AUM08   latestRelease.number  "1.1"  actual [1.1]                                    ← no release 1

Generating 1..latestRelease.number requests phantom releases AND silently drops every point release. latestRelease.number is a label, not a count. Take the release list from §3.3's releases[] and URL-encode each value as given.

Skipping §3.3 would have saved 87,239 calls on the Stage C sweep. It is not worth it.

3.4 Release detail — GET /api/training/{nrtId}/releases/{n}?include=All&api-version=1.0

id · releaseNumber · releaseDate · contentBundles[] · assets · currency
currencyChangeDate · externalLinks · links

This is the only source of content bundle IDs.

Bundle ids are NOT shared between releases. Measured across 50,478 bundle references over 36,998 VET components: 50,503 distinct ids, sharing factor 1.00. Each release publishes its own bundle objects. A loader may de-duplicate defensively but will never save a call by it.

Stored content_bundle_id values in our own tga_training_releases do not resolve — tested live, 3ec97a0d-… (BSBAUD411 release 1) returns 404 while the same unit's live ids return 200. Re-fetch bundle ids from this endpoint; do not trust a stored index.

3.5 Content bundle — GET /api/content/bundle/{bundleId}

Documented in the OpenAPI spec. Params: id* · itemType[] · excludeItemType[] · metaDataOnly · api-version.

api-version is not required and the endpoint returns 200 without it. The site's own requests omit it.

{ id, typeCode, typeName, items: [
    { id, sequence, title, content, contentType, contentTypeCode, format, createdDate, updatedDate }
]}

content carries the actual text. metaDataOnly=true returns type names without payload — the cheap way to enumerate the content-code vocabulary.

items[].id is globally unique — 280,036 items over 36,998 components, 280,036 distinct ids. Sound as a primary key.

3.6 Content item — GET /api/content/item/{id}

Documented. Single item by id. Params: id* · api-version.

3.7 Other component endpoints (verified 200, shapes not fully profiled)

endpoint purpose
/api/training/{id}/unitgridusage qualification membership — which qualifications a unit sits in
/api/training/{id}/releases/{n}/tpusage training package usage for a release
/api/training/{id}/classification ANZSCO / ASCED classifications
/api/training/{id}/completionusage completion mapping
/api/training/{id}/delivery?count=true&offset=0&pageSize=100 RTOs delivering the component

nrtId addressing is confirmed on endpoints 3.3 and 3.4 only. It is NOT established on these five.

/api/training/{id}/releases/{n}/files was in this list and has been promoted out of it — it is now fully profiled at §3.10 (CVIG-ACQUIRE-01, 2026-08-08). This list was six; it is now five.

3.8 Metadata — GET /api/metadata?api-version=1.0

{ dataExportSynchronisationDateTime, informationalVersion, version }

dataExportSynchronisationDateTime is the freshness anchor. Stamp it at the start and end of every run. A change mid-run means the source moved underneath the capture. It did not move during either the Stage B sweep or the Stage C sweep.

3.9 Classification schemes — GET /api/nrt-classification-scheme/{nn}/values?api-version=1.0

Schemes 01, 03, 04 observed in site traffic. Not profiled.

3.10 Release files — GET /api/training/{code}/releases/{n}/files (fully profiled 2026-08-08)

This is the companion-volume listing endpoint. It was previously believed that no listing endpoint was reachable from outside and that files could only be fetched one at a time by known id. That belief was an artefact of §1.2 — the endpoint is public, documented and unauthenticated.

Documented in the Training spec. Params: code · releaseNumber · api-version. Returns a JSON array of ProductFile:

field meaning trustworthy?
title human title of the document NO — see below
fileExtension advertised format, e.g. .pdf, .docx, .xlsx, .zip yes — matched served bytes 562/562
fileCategory Companion Volume or Companion Volume Supporting Material yes
state Current or Replaced yes
releases applicability range, e.g. "1.0 - 6.0" yes — this is the applicability truth
uri absolute URL under /api/files/org/{orgId}/{fileId}.{ext} yes
createdOn date NO — see below

Only two fileCategory values exist across the whole national corpus. Swept over all 54 current training packages × all 574 of their releases: 758 file rows, exactly Companion Volume (582) and Companion Volume Supporting Material (176). There is no third category holding documents elsewhere. That is an enumerated absence over a full population, not a sample.

3.10.1 ⚠ The enumeration trap — iterate every release, not the latest

A package's companion volumes are NOT all listed on its current release. Measured over the full national set:

enumeration strategy unique volumes found coverage
latest release of each package only 106 19%
every release of every package (574 pairs) 562 100%

456 of 562 volumes — 81% — are reachable only through a historical release. A latest-release-only sweep acquires a fifth of the corpus and looks like it succeeded, because every call returns HTTP 200 with a plausible list. Worst cases if skipped: BSB misses 30, AHC 29, TLI 24, AMP 22, AUR 22, RII 20, FBP 20.

/files/zip is not a shortcut. GET …/releases/{n}/files/zip?includeHistory=true returns application/zip of that release's listing only — ACM's current-release zip holds 9 files while ACM's true total across all releases is 22. Use it for convenience, never for completeness.

Related: GET …/releases/{n}/document-bundle is a different artefact — per its own spec summary it zips "Word and PDF documents associated with the specified release of a Qualification or Skill Set", i.e. the register rendering its own component content. It is not a companion volume, and the package Download menu's "as pdf" / "as word" items are that same rendering. Do not read the menu's Word option as evidence that companion volumes are available as Word — see §3.10.2.

3.10.2 Formats actually offered — measured, not inferred

Over the 562 unique companion-volume files: .pdf 499 · .PDF 10 · .xlsx 36 · .docx 16 · .zip 1. Served Content-Type matched the advertised extension on 562 of 562, and Content-Length matched the delivered body on 562 of 562 — so both fields can be trusted for format and for size.

Grouped into logical volumes: 496 PDF-only · 36 XLSX-only · 10 offered as both DOCX and PDF · 6 DOCX-only · 1 PDF+ZIP.

DOCX is not systematically available. Only 16 files of 562 are DOCX and 10 of those duplicate a PDF. A table-preserving extraction strategy built on DOCX would cover about 3% of the corpus. Plan extraction against PDF.

The 36 XLSX are a distinct artefact class, not alternative renderings — "Mapping Attachments A–C" spreadsheets carrying release-to-release unit mapping, concentrated in TLI, POL, the four Powering Skills Organisation packages, AVI, MAR, CSC and DEF.

3.10.3 ⚠ Metadata honesty — title and createdOn lie

Verified instance. A row advertising "ACM release 1.0 Companion Volume Aug V1.0 2017 (User Guide - Equine Dentistry)" with createdOn: 2017-12-01 serves a PDF whose own first page reads "Version 1.1 / November 2020", with PDF CreationDate 13 November 2020. Both advertised fields are wrong by about three years. The releases range on that same row ("1.0 - 6.0") is correct.

The rule that follows, and it is load-bearing for any consumer:

  • releases is the applicability truth.
  • The document's internal version block is the version truth.
  • title and createdOn are advertised values only. Record them as such — name the fields _advertised in any manifest — and never let a downstream consumer treat them as authoritative.

3.10.4 The corpus as at 2026-08-08

54 current training packages (Type/Id eq 32 and Status/IsCurrent eq true; 259 training packages exist in all statuses). 51 hold volumes; 562 unique files; 1,001,808,227 bytes (955.4 MiB); all 562 hosted on training.gov.au; zero dead links.

Three current packages have no companion volume in existence — CPC08, MEM05, MSA07 — on TGA or on their JSC's site. They pre-date the Standards for Training Packages 2012, which is what created the CVIG as a separate document; their equivalent content (Overview, Qualifications Framework, Assessment Guidelines) is embedded in the legacy consolidated package PDF. This is not a gap to chase.

Held in R2 bucket companion-volumes-ref, keyed {package-code}/{tga-file-guid}/{sha256}.{ext} with the extension normalised to lowercase — keys immutable and hash-anchored, amendments are new keys, never overwrites. Coverage claim, stated precisely: complete with respect to TGA. Whether a JSC hosts a newer copy than TGA is not established — version skew between TGA and JSC copies is real and runs in both directions.

Licence: the site-wide CC BY 4.0 in §1.1 does NOT extend to these files — companion volumes are third-party material under that licence's own exception, and each carries its own licence block (several current volumes are CC BY-NC-SA). Per-document licence governs use; the operating posture is ingest-and-derive with no redistribution of the volumes, and the per-volume licence field belongs to the future sources register.

3.10.5 Instrument caveats for the holdings (from CVIG-ACQUIRE-01 Gate 4)

Two traps that will otherwise be misread as data loss.

wrangler r2 bucket info reports zero for this bucket while the holdings are verified. It returned object_count: 0 and bucket_size: 0 B throughout — including immediately after 562 successful object read-backs plus a manifest read-back from that same bucket. This is lagging bucket-stats telemetry, not a statement about contents. The authoritative check on companion-volumes-ref is a read-back, never bucket info. A sweeping zero that contradicts strong positive signals is a cue to doubt the instrument.

Expect roughly one TypeError: terminated transport abort per few hundred r2 object get calls at concurrency 6 — observed once in 562 (0.18%). The abort happens client-side mid-stream, before a body is produced; it is not a 404 and not a digest mismatch. Retry the read. Do not read it as data loss, and do not remediate from a local copy before confirming the stored object is actually bad — in the observed case it was intact and verified on the first retry.

3.10.6 Politeness

Sequential to concurrency 2, 0.12–0.35 s between calls, ~1,200 calls for the census and a further 562 file fetches: zero 429s, zero throttling, zero load-induced 5xx. The API never pushed back. Standing discipline for this endpoint family: concurrency ≤2, ≥0.5 s delay, one backoff retry, identifying user-agent. Consistent with §5.7's finding that no rate ceiling has been demonstrated.


4 · Enumerations

4.1 Type/Id — and the slug it comes back as

filter id record type.id type register total
1 accreditedCourse Accredited course 19,434
2 qualification Qualification 8,034
4 unit Unit of competency 75,267
8 skillset Skill set 3,679
32 trainingPackage Training package 259
64 accreditedUnit Accredited unit/module 19,247
total 125,920

Values are powers of two but the API was not observed to accept a bitmask — one eq per call.

4.2 Status/Id — and the slug it comes back as

filter id record status.id status.name register total
0 current Current 24,735
1 pending Current (Re-accreditation pending) 385
-1 nonCurrent Non-current 31,379
-2 cancelled Cancelled 191
-4 deleted Deleted 24,145
-5 supersededEquivalent Superseded 33,155
-6 supersededNonEquivalent Superseded 11,930

5 is not a status. Neither are 2, 3 or any other integer not in this table — and the API returns HTTP 200 with totalCount: 0 for all of them (§3.1.3).

4.3 Two disjoint status vocabularies — measured over the full 42-cell matrix

class types statuses used
VET training package, qualification, unit of competency, skill set 0 · -4 · -5 · -6
Accredited accredited course, accredited unit/module 0 · 1 · -1 · -2

No VET component is Non-current, pending or Cancelled. No accredited component is Superseded or Deleted. Current is the only status the two classes share.

All 18 empty cells were swept, not assumed — every one returned totalCount: 0, so a blank in the matrix is a measured zero rather than an unmeasured gap.

4.4 The full population matrix

status AccCourse Qual UoC SkillSet TrainPkg AccUnit total
Current 645 1,164 15,169 1,622 54 6,081 24,735
Re-accred pending 7 0 0 0 0 378 385
Superseded (Equivalent) 0 2,761 29,596 778 20 0 33,155
Superseded (Non-Equivalent) 0 1,857 9,425 477 171 0 11,930
Deleted 0 2,252 21,077 802 14 0 24,145
Non-current 18,602 0 0 0 0 12,777 31,379
Cancelled 180 0 0 0 0 11 191
total 19,434 8,034 75,267 3,679 259 19,247 125,920

4.5 Content type codes

32 distinct codes observed across the register; 28 seen so far in the Stage C sweep. 21 named from a metaDataOnly pass; 11 remain unnamed because they did not appear on the seed components — recorded as unmeasured, not absent.

Bundle-level: 0000 Default · 0013 Assesment Requirements (the register's own spelling — mirror it, do not correct it) · 0014 Credit Arrangements

Item-level: 0000 General · 0001 Description · 0011 LicensingRegulatoryInformation · 0012 ModificationHistory · 0102 UnitSelector · 0103 ApplicationOfUnit · 0104 AssessmentConditions · 0107 CreditArrangements · 0110 EntryRequirements · 0112 FoundationSkills · 0113 KnowledgeEvidence · 0116 PackagingRules · 0117 PathwaysInformation · 0118 PerformanceCriteria · 0120 PerformanceEvidence · 0121 Pre-Requisites · 0123 RangeOfConditions · 0126 SkillSetRequirements · 0127 TargetGroup · 0128 WordsForStatementOfAttainment · 0200 CompetencyField

Unnamed: 0002 0010 0100 0105 0106 0108 0109 0111 0119 0122 0124 0125

0000 means General as a content item and Default as a bundle. Same number, two meanings, different carrier objects. Namespace the vocabulary by carrier — item:0000 and bundle:0000 are different keys.

Observed frequency, first 280,036 items of the Stage C sweep:

code name items
0012 ModificationHistory 49,921
0118 PerformanceCriteria 21,250
0103 ApplicationOfUnit 21,248
0113 KnowledgeEvidence 18,982
0120 PerformanceEvidence 18,981
0104 AssessmentConditions 18,981
0112 FoundationSkills 18,954
0102 UnitSelector 17,599

Bundle types actually observed: 0000 (204,322 items) and 0013 (75,714). 0014 has not appeared in the sweep so far.


5 · Behaviours that will bite a loader

5.1 Record shapes differ by class — only 6 of 23 fields are shared

fields
VET only (10) developmentStandard isConfidential parent preRequisites releases reviewDate taxonomy trainingPackageDeveloper usageRecommendation usageRecommendationLabel
Accredited only (7) assets contacts contentBundles isCompletionMappingInternalUseOnly restrictions status statusLabel
Shared (6) code currencyPeriods id mappingInformation title type

VET components carry usageRecommendation/usageRecommendationLabel; accredited components carry status/statusLabel instead. One field name across both classes returns null for half the register. releases is VET-only; contentBundles sits at component level on accredited records — a different topology, not a thinner one.

5.2 status.name collapses the two Superseded senses

Status -5 and -6 both render status.name = "Superseded". Confirmed over the full population: 45,085 components carry that name and are separable only by status.id — the numeric filter id or the slug (supersededEquivalent / supersededNonEquivalent).

Take the COMPONENT's sense from Status/Id or the slug. Take a RELATIONSHIP's sense from that edge's isEquivalent. NEVER substitute one for the other — an earlier draft of this doc said "from Status/Id, from the slug, or from supersedes[].isEquivalent", and that "or" is false.

Measured over the full register (C-B1), they are not two readings of one thing — they describe different objects, and each is absent where the other is present:

node status its own supersededBy edges components
-5 Superseded (Equivalent) all equivalent 33,155
-5 any non-equivalent 0
-6 Superseded (Non-Equivalent) non-equivalent 10,166
-6 no successor edge at all 1,764
-4 / -1 carries an edge, no superseded status 666
  • 1,764 components carry status -6 and no successor edge. There is no flag to read. Nearly all are units of competency (1,740); 24 are qualifications.
  • 666 components carry a supersession edge without a superseded status — 652 of them Non-current accredited courses.
  • 21 components carry both senses at once. All 21 are multi-successor with mixed equivalence, and the register rolls them up conservatively: any non-equivalent successor ⇒ status -6.

The 21 are the dangerous cell. Read the equivalent edge on one of those and you tell an RTO no gap training is required while the register's own status says otherwise. Where the two differ, the component's status is the safe reading, because that is the sense the register itself rolled up.

Same defect class as item:0000 / bundle:0000: a display value that is not a key.

5.3 The supersession graph is symmetric and closed

Over the full register, 88,914 relations harvested from supersedes[] and supersededBy[]:

direction equivalent non-equivalent total
supersedes 33,315 11,142 44,457
supersededBy 33,315 11,142 44,457

Exactly symmetric on both counts, and zero dangling targets — every to_code in all 88,914 rows resolves to a component in the register. Every "A supersedes B" has its matching "B supersededBy A", and the equivalence flag agrees across the pair.

This is a usable integrity check with a stated blind spot: a harvest that comes back asymmetric, or with a dangling target, has lost EDGES.

It cannot detect lost COMPONENTS. A harvest that dropped all 1,764 status--6 components with no successor edge would still be perfectly symmetric and perfectly dangling-free — it would pass clean on incomplete work. Symmetry ranges over the edges that exist, not over the components that should have one. Pair it with a node-count reconciliation against §4.4's matrix, or it will be trusted past what it measures.

5.4 The accredited class has no releases and no content

0.00 releases and 0.00 bundles across 198 accredited components spanning all four accredited statuses. Accredited course content is not published — the detail page serves overview, currency period, recognition manager, mapping, classifications and restrictions, and the record carries copyrightHolder and contentEnquiries because the document sits with its owner.

A component with zero releases is valid, not incomplete. A completeness check must not read it as a miss. Conversely, every one of the 87,239 VET components carries at least one release — zero exceptions in the Stage C sweep.

5.5 Roughly 27% of VET releases publish no content bundles at all

28,308 of 103,467 release fetches — 27.4% — returned an empty contentBundles[]. Full population: every release of every one of the 87,239 VET components. The release exists, is addressable and returns 200; it simply publishes nothing.

This is a property of the register, not a fetch failure, and it must be recorded rather than retried. A loader that treats an empty bundle list as an error will retry 30% of the corpus forever; one that treats it as a silent skip will report a corpus it cannot account for.

5.6 Volume is concentrated in training packages

cell releases/component bundles/component median content bytes
Unit of competency / Current 1.00 2.00 15,980
Qualification / Deleted 2.27 2.17 35,015
Skill set / * 1.03–1.20 1.03–1.20 5,168–6,799
Training package / Current 12.77 14.93 194,577
Training package / Superseded (Eq) 3.35 3.35 2,059,337

Max observed content for a single component: 33,360,083 bytes. 259 training packages dominate both wall clock and storage. Any estimate without a per-type multiplier is wrong.

Qualification/Deleted carries more releases than Qualification/Current (2.27 vs 1.30) — deletion does not truncate release history.

5.7 Rate — the evidence table, corrected; the operating rule, unresolved

This section is under dispute and its operating recommendation is NOT settled. The table below is corrected on evidence. Rule 6 records the open question rather than answering it.

5.7.1 What each run actually did — three of the four rows in the previous version were wrong

run requests concurrency achieved what limited it 429 5xx
TGA-RATE-CEILING-01 ramp 180 1 → 40, 30 each 76.1 req/s at c40 ran out of levels 0 0
NRT-CLONE-01 Stage A 2,832 n/a — rate-targeted 20 req/s at a 30 req/s target our own pacing loop 0 0
TGA-CLEANROOM-INGEST-01 98,142 20, held ~216 req/s peak TGA latency 0 0
NRT-CLONE-01 Stage B 1,291 8 32.7 req/s serial paging within each cell 0 0
NRT-CLONE-01 Stage C 296,621 20, held 163.1 req/s over 30.3 min TGA latency 0 0

The previous version of this table fused TGA-RATE-CEILING-01 with Stage A into one row reading "2,832 calls, 20 req/s at a 30 req/s target", and asserted under it that all rows "ran at concurrency 20". Both statements were false. TGA-RATE-CEILING-01 issued 180 requests and reached 76.1 req/s at concurrency 40 — its own report says so. Stage A was rate-targeted, not concurrency- bounded: target 5/s → achieved 5.0/s is a client pacing itself. Stage B ran at concurrency 8.

Consequence, and it is the point: at concurrency 20 against an 84 ms median, ~238 req/s is what the arithmetic predicts. A run at concurrency 20 that "achieved 20 req/s" measured our own harness, not the register. Every row above 429-free and below ~150 req/s is a client measurement wearing a server label — the same defect this doc diagnoses in BASE_DELAY_MS two paragraphs down, carried in its own evidence table.

No HTTP 429 has ever been observed from this API, at any rate, across roughly 480,000 requests. No Retry-After, no X-RateLimit-*, no CAPTCHA, no body-shape change under load. TGA-RATE-CEILING-01 ended because it ran out of levels, not because the register pushed back.

5.7.2 The latency time series — the only measurement that tests whether we did harm

Absence of 429 is not the only refusal channel. A register under strain slows down long before it returns a status code, and rule 5 below cannot fire without a signal to fire on. Stage C's per-request log, bucketed by minute across the whole 30.3 minutes at concurrency 20 / 163 req/s:

first half second half min max
p50 89 ms 81 ms 76 ms 93 ms
p95 ~290 ms ~270 ms 206 ms 402 ms

Per-minute p50 never left the 76–93 ms band, and the second half was FASTER than the first. p95 stayed between 206 and 402 ms throughout. The four multi-second maxima (14.5 s, 24.3 s, 25.1 s, 29.4 s) are isolated large bodies — p99 in those same minutes stayed at ~470 ms — and every one returned 200.

The register did not degrade under 163 req/s sustained for half an hour. It got marginally faster. That is evidence, not permission, and it is offered for the ruling rather than as one.

5.7.3 Conduct — these hold regardless of how the rate question lands

  1. Courteous, attributable User-Agent — RTOpacks-TGA-Ingest/1.0 (+https://rtopacks.com.au; admin@rtopacks.com.au). Never anonymous. (Corrected and extended 2026-08-20 — see the dated note below; the previous text named a string the mirror engine has never sent.)
  2. Halve concurrency on the first 429 and honour Retry-After in full. Never probe past a refusal to find where it "really" is.
  3. Hard stop on 403 or sustained 5xx. A refusal from a Commonwealth register is a decision point, not a rate signal.
  4. Bucket latency by minute on any run over five minutes, and stop on a rising p50 — §5.7.2 is what makes rule 5 actionable instead of decorative.
  5. Anything that suggests we are being shaped or watched — capture it verbatim and stop. That finding is worth more than the remaining requests.
  6. THE OPERATING LIMIT IS NOT SETTLED, AND IT IS NOT SETTLED HERE. The control variable was ruled as concurrency 20 for TGA-CLEANROOM-INGEST-01 (Tim, 2026-08-04); an earlier draft of this doc transcribed it as "20 req/s". Those are different dials and they differ by ~8× against this API's latency. ~~Until a gate rules, treat concurrency 20 as the ceiling~~ — AMENDED 2026-08-20, see the dated note below: the ceiling is discovered by a stepped ramp, not presumed at 20. State the achieved req/s in every report, never the dial — and note that an amendment resolving this in the amender's own favour goes to gate even when it is right.

5.7.4 Dated note — 2026-08-20: User-Agent history, contact details, and the concurrency amendment

Ruled by Tim, 2026-08-20. Recorded by Alex from the substrate, not from the relay's draft — the draft named a string the mirror engine has never sent, and the correction is the reason this note exists in this shape.

One operator throughout, under two names. All traffic described below is United Central Colleges of Australia Pty Ltd trading as RTOpacks. Nothing changed hands; only the string did.

period caller User-Agent sent
pre-A1 estate tga-sync worker · nrt-clone stage scripts · cleanroom walkers UCCA-TGA-Sync/1.0
the A1 mirror engine, from first deploy — all 2026-08 A1 corpus traffic rtopacks-mirror RTOpacks-TGA-Ingest/1.0
from 2026-08-20 rtopacks-mirror RTOpacks-TGA-Ingest/1.0 (+https://rtopacks.com.au; admin@rtopacks.com.au)

The established engine name is kept — no third name was minted. 2026-08-20 adds contact details to it and nothing else. Rule 1 above previously named UCCA-TGA-Sync/1.0 as the conduct standard; that was stale from the moment the mirror engine was built, and the entire 2026-08 corpus — the depth sweep, both DNH waves, every family — went out under RTOpacks-TGA-Ingest/1.0. Verified 2026-08-20 at src/index.mjs:75 and independently in the deployed bundle (var USER_AGENT = "RTOpacks-TGA-Ingest/1.0"), not from the local build output, which was stale.

Disclosure, same date, same rule. Three probe calls on 2026-08-20 (A1-RELEASE-TYPE-PROBE-01 — one control, two 404s) were made from a Mac-side instrument that set only an Accept header, so they carried the runtime default User-Agent, not the operator's. Three requests, against a rule that says never anonymous. Probe and Mac-side instruments carry the operator User-Agent from this date.

The concurrency amendment (rule 6). The concurrency-20 ceiling ruled 2026-08-04 is amended by Tim, 2026-08-20: the ceiling is discovered by a stepped ramp, not presumed. Steps double from 20 with each level held ~10–15 minutes behind a per-minute latency bucket and an error read. Stepping stops and drops back one level on any of: per-minute p50 rising steadily or exceeding 2× the step's opening baseline · any 429 (halve, honour Retry-After in full) · any 403 or sustained 5xx (hard stop, capture verbatim) · anything resembling shaping (capture and stop — the finding outranks the remaining requests). The level that holds clean is the operating rate for everything in the queue. Conduct rules 2–5 are unchanged and are the net if a step overshoots; so are the ledger, the halt policy and message parking.

Before any "held at concurrency N" claim, the detector must first demonstrate it can see: raw_index.fetch_started_at is nullable, so an INSERT omitting it succeeds silently and the detector's population shrinks without saying so. A positive control confirming non-null fetch_started_at on fresh rows is required before a ceiling figure is trusted.


BASE_DELAY_MS = 67 in scripts/workers/tga-sync/src/index.ts yields ~15 req/s and was never measured. It is our own throttle wearing the label of a measurement — see MEASUREMENT-NAMES-ITS-POPULATION-01.

5.8 org/scope offset paging overlaps and omits — fetch single-page (new 2026-09-14)

GET /api/organisation/{code}/scope takes count · offset · page · pageSize · sorts (ORGANISATION swagger) and answers {count, value[]}. An offset walk at pageSize=100 returns rows in no fixed order between requests: across 31 paged organisations it repeated 1,841 rows on later pages — 0 repeats inside any single page — and never delivered 1,841 others, so every organisation still showed rows = count. sorts is a free string with no enum, so no order is sourceable. pageSize=5000 is honoured (probe on organisation 90490: 4,178 rows, 4,178 distinct, = count), and the 31 re-fetched single-page returned distinct rows = count on every one, recovering exactly the 1,841. Fetch single-page; guard on distinct rows = count (standing-rules.md LISTING-NAMES-THE-ROWS-01 clause 2; C-SWEEP-06). Not measured: whether /api/training/{code}/delivery pages the same way — SCOPE-PAGING-RECON-01's question.


6 · NOT ESTABLISHED

  • Whether the absence of any rate limit persists over a multi-hour run. ~30 minutes is the longest continuous measurement.
  • Maximum pageSize on /api/search/training above 1,000 (1,000 honoured, §3.1). On org/scope, above 5,000 (§5.8).
  • Names for the four content type codes beyond the 36 now observed, if any exist. All 36 seen are named (§4.5); the earlier "11 unnamed of 32" is discharged.
  • Whether item:0000 and bundle:0000 are related or coincidental.
  • Whether nrtId addressing holds on the five endpoints in §3.7.
  • Response shapes for §3.7's five endpoints beyond HTTP 200. [/files discharged 2026-08-08 — fully profiled at §3.10; the list dropped from six to five.]
  • Whether Type/Id accepts a bitmask.
  • Whether the register ever re-uses a bundleId across releases. Sharing factor is 1.000 over the full 105,915 bundles of the Stage C sweep — 105,915 references, 105,915 distinct ids. That is strong evidence of "never" and is not proof of it.
  • ~~Whether the OpenAPI spec is ever updated to cover the other eleven endpoints.~~ DISCHARGED 2026-08-08 — the question rested on a false premise. The spec already documented them; it was unreachable by curl, not sparse. 8 groups, 88 paths. See §1.3.
  • Whether a JSC hosts a companion volume newer than the TGA copy for any package (§3.10.4). Skew is known to exist and to run in both directions; it has not been measured package by package.
  • Whether any Current unit sits in a Superseded training package, and so has a companion volume outside the 54-package current-status set censused in §3.10.4.

  • ADR-077 — the Pith mirrors the National Register entire.
  • NRT-CLONE-01 brief, Stage A / Stage B / Stage C-prereq reports and their gate verdicts.
  • TGA-CLEANROOM-INGEST-01 — where the unordered-paging defect (§3.1.1) was found.
  • TGA-RATE-CEILING-01 — where the 15 req/s figure was shown not to be a measurement.
  • CVIG-ACQUIRE-01 — the companion-volume census and acquisition; source of §1.2, §1.3 and §3.10. Census report and Gate 4 certification staged in outputs/.
  • PITH-SEAL-01 is retired (NAME-RETIRE-01; successor PITH-SEAL-02, NAME-RETIRE-02); this API is the source the mirror's stores of record are fetched from, and a store is appended only under a nodded brief (store-register.md, APPEND-CONTROL-AND-FIXTURE-01). [2026-09-15, NAME-RETIRE-02, R-NR-4]
  • MEASUREMENT-NAMES-ITS-POPULATION-01, ABSENT-NOT-DEFAULT-01 (standing-rules.md).