FLEET-STUDY-CLOSE-01 — the three open gaps, closed — 2026-08-14¶
Written by: Claude (Driver seat, Opus) · Closes: §3 of time-machine-2026-08-14-DRIVER-v1.0
Status: findings, not canon. Nothing here amends the North Star until the Architect seat rules on it.
Register: every figure below is quoted from Cloudflare's own published documentation, read this
session and linked at the foot. Nothing has been run. No account was mutated.
0 · THE ONE-PARAGRAPH VERSION¶
All three gaps are closed enough to design against. Fleet migration is not a Cloudflare feature — it is a thing we build, and the documented ceiling that bounds it is the same 1,200-requests-per- five-minutes API limit that bounds provisioning. R2 has no object versioning at all, so retention and per-tenant restore must be built from immutable keys, lifecycle rules and bucket locks rather than assumed. And the study turned up two errors in our own canon about residency, one of which overstates a risk and one of which understates the option set.
1 · GAP 1 — FLEET MIGRATION TOOLING¶
1.1 What Cloudflare gives us¶
| Fact | Value |
|---|---|
| Tracking table | d1_migrations — "Creating migrations will keep a record of applied migrations in the d1_migrations table found in your database." Renameable via migrations_table. |
| Migration directory | migrations_dir, default migrations/ |
| Glob discovery | migrations_pattern, default migrations/*.sql. "When migrations_pattern is set, migrations_dir must also be set, and migrations_pattern must start with whatever migrations_dir is set to." |
| Nested layouts | Supported, for ORM output such as Drizzle: "migrations_pattern": "migrations/*/migration.sql". Each migration is recorded "as a path relative to migrations_dir". |
| Commands | wrangler d1 migrations create / list / apply |
| Foreign keys | PRAGMA defer_foreign_keys = true before changes that would violate them |
1.2 The structural fact — there is no fleet path¶
Every documented migration mechanism is per-database and wrangler-bound. wrangler d1 migrations
apply resolves its target through a Wrangler configuration file, by binding name or database name.
The documentation offers no bulk, fan-out or namespace-wide migration command, and no REST endpoint
for applying migrations at all.
The only programmatic route is the D1 REST query endpoint, and Cloudflare says plainly what it is for: "D1's built-in REST API is best suited for administrative use as the global Cloudflare API rate limit applies."
So: a fleet migration runner is ours to build. Minimum shape the evidence supports —
- Enumerate tenant databases from the control plane's own registry, not from a
d1 listcall. - Apply ordered
.sqlper tenant overPOST /accounts/{id}/d1/database/{db_id}/query. - Keep our own applied-migrations row per tenant database —
d1_migrationsis only written by wrangler, so a REST-driven runner must maintain the ledger itself or it has no idea where it got to. - Be resumable and idempotent, because a 3,000-tenant wave will be interrupted by the rate limit.
1.3 The ceiling that bounds it¶
| Limit | Value |
|---|---|
| Client API, per user / account token | 1,200 per 5 minutes |
| Client API, per IP | 200 per second |
| On exceeding | HTTP 429, and "blocks all API calls for the next five minutes" |
| Headers returned | Ratelimit, Ratelimit-Policy, retry-after |
| Max SQL statement length | 100,000 bytes |
| Max SQL query duration | 30 seconds |
Arithmetic, ours not Cloudflare's: at 3,000 tenants and one batched call per tenant, a fleet-wide migration wave is ≥12.5 minutes of pure rate-limit budget. Provisioning from cold — D1 create plus script upload, two calls per tenant — is ≥25 minutes for 6,000 calls. Both are acceptable. Neither is instant, and both must be built as resumable jobs rather than a loop that assumes it will finish.
Batching is what keeps this cheap. A migration wave should send one multi-statement call per tenant, not one call per statement — the 100 KB statement ceiling is generous enough that this is almost always possible.
1.4 The ambiguity I could not resolve — and it matters¶
The prose says the limit is per user: "1,200 requests per five minute period per user, and applies cumulatively regardless of whether the request is made via the dashboard, API key, or API token." The limits table says per user/account token. These are not the same claim.
If it is per user, an automated fleet migration and a human clicking around the dashboard draw on one shared bucket, and the control plane can be starved by a person looking at a graph. If it is per token, a dedicated account-owned token has its own budget and the two never collide.
This is not resolvable from the documentation. It is cheap to measure and the spike should measure it.
1.5 Two constraints that shape the design¶
- No gradual deployments for user Workers. Confirmed again: "Changes made to user Workers create a new version that deployed all-at-once to 100% of traffic." Cohort rollout is ours to build.
- Time Travel is not a fleet instrument. 30-day retention, always on, but "Restoring a database to a specific point-in-time is a destructive operation, and overwrites the database in place", capped at 10 restores per 10 minutes per database, and documented only through wrangler — no REST restore. A fleet-wide rollback of 3,000 databases has no documented mechanism.
2 · GAP 2 — R2 AND BACKUP AT FLEET SCALE¶
2.1 Fleet limits — not the constraint¶
| Limit | Value |
|---|---|
| Buckets per account | 1,000,000 |
| Objects per bucket · data per bucket | Unlimited |
| Object size | 5 TiB |
| Bucket management operations | 50 per second per bucket |
| Concurrent writes to the same key | 1 per second |
| Cloudflare REST API (R2) | 1,200 per 5 minutes — object operations belong on the S3 API, not here |
2.2 The finding that changes the backup design¶
R2 does not implement object versioning. On the S3 compatibility page, both
PutBucketVersioning and GetBucketVersioning are marked ❌ unimplemented.
Consequence: an overwrite is final and a delete is final. There is no "restore the previous version" primitive to fall back on. Every retention guarantee we make to an RTO — and evidence retention is the point of Record — has to be built out of:
- Immutable keys. Never overwrite; write a new key per version and resolve current in the index.
- Lifecycle rules — max 1,000 per bucket; expire by age or by date; transition Standard → Infrequent Access; abort incomplete multipart (a default rule already expires these at 7 days). Deletion is not instant: "Objects will typically be removed from a bucket within 24 hours", and existing objects may lag when a rule changes.
- Bucket locks — max 1,000 rules; time-based, date-based or indefinite; "Bucket lock rules take precedence over lifecycle rules"; and "A bucket cannot be emptied while any bucket lock rules are configured."
That last sentence is load-bearing in two directions. It is exactly the primitive that makes a compliance retention promise real. It is also the reason "delete this tenant" stops being one call — a locked bucket cannot simply be emptied, so offboarding becomes a defined multi-step process that must be designed, not discovered at the first cancellation.
2.3 D1 backup mechanics¶
| Fact | Value |
|---|---|
| Time Travel | Always on, no opt-in. 30 days on Workers Paid. Bookmarks created automatically. |
| Restore semantics | Destructive, overwrites in place |
| Restore rate | 10 restores per 10 minutes per database |
| Export | wrangler d1 export <db> --remote --output=./database.sql — SQL dump; --table, --no-data, --no-schema |
| Import | wrangler d1 execute <db> --remote --file=<f.sql>, "limited to 5GiB files" |
| REST import/export | Not documented. wrangler only. |
Per-tenant backup is ours to build too. Time Travel is a 30-day in-place undo, not a backup: it overwrites rather than producing an artefact, and it cannot restore into a different database. (That last clause is inference from the documented restore semantics, not a quoted statement — the docs do not address it directly.) A real per-tenant restore path means periodic export to R2 under immutable keys, with lifecycle for retention and locks where an obligation requires them.
3 · GAP 3 — PROVISIONING RATE LIMITS — CLOSED¶
The section that truncated twice is the API Rate Limits section at the foot of the Workers for Platforms limits page. It carries the standard account table, reproduced at §1.3 above. There is no special provisioning throughput allowance for Workers for Platforms: user-Worker uploads draw on the same 1,200 / 5 minutes as everything else.
One figure in that table is a design constraint in its own right:
| Quota | Value |
|---|---|
| User API token quota | 50 |
| Account API token quota | 500 |
A token per tenant is impossible at 3,000 tenants. The control plane runs on a small number of account-scoped tokens, and per-tenant isolation is carried entirely by bindings — which is what the reference architecture already told us: "Each User Worker can only access the bindings that are explicitly attached to it."
Still unmeasured: dispatch namespaces per account. Not stated anywhere I could find. We need one, possibly a second for staging, so it is very unlikely to bind — but it is unmeasured, not absent, and should not be written down as "unlimited".
4 · TWO CORRECTIONS TO OUR OWN CANON¶
Both are in where-we-are.md and in DRIVER v1.0 §2.1. Neither changes the residency conclusion.
4.1 Jurisdictions are three, not two¶
We wrote "eu and fedramp only". The D1 create endpoint now documents eu, fedramp, us.
R2's jurisdiction header remains default | eu | fedramp.
There is still no Australian or Oceania jurisdiction, so the conclusion stands — oc is a hint,
not residency, and must never be described to an RTO as a guarantee. But the sentence as written is
factually wrong and will be caught by anyone who checks.
4.2 We assert an immutability the docs do not state¶
We wrote of the location hint: "This cannot be corrected after a resource is created."
What the documentation actually supports is narrower:
- Jurisdiction — explicitly create-time only: "can only be set at database creation time via wrangler, REST API or the UI and cannot be added/updated after the database already exists."
primary_location_hint— set at create (POST /accounts/{id}/d1/database, allowed valueswnam enam weur eeur apac oc). The PATCH endpoint on an existing database exposes onlyread_replication. So it is not changeable through any documented API — which supports the practical conclusion — but Cloudflare nowhere says it is immutable.
Recommended rewording: "there is no documented means of changing it after creation". Same operational consequence, and it survives contact with someone who reads the page.
4.3 One thing that is reversible, and we should say so¶
read_replication.mode is settable at create and editable via PATCH (auto | disabled).
D1 replicates to every supported region — ENAM, WNAM, WEUR, EEUR, APAC, OC. So enabling
replication remains a residency decision, exactly as canon says, but unlike the location hint it is
a decision we can take back.
5 · THE API SHAPES, FOR THE SPIKE AND FOR THE CONTROL PLANE¶
# create dispatch namespace
POST /accounts/{account_id}/workers/dispatch/namespaces
{"name": "<namespace>"}
# create tenant database, Oceania hint, no jurisdiction
POST /accounts/{account_id}/d1/database
{"name": "<tenant-db>", "primary_location_hint": "oc"}
# upload user Worker with the tenant's D1 bound, multipart
PUT /accounts/{account_id}/workers/dispatch/namespaces/{ns}/scripts/{tenant}
metadata={"main_module":"worker.mjs",
"bindings":[{"type":"d1","name":"DB","id":"<tenant-db-id>"}],
"tags":["tenant-<id>"],
"compatibility_date":"<date>"}
# administrative SQL against a tenant database
POST /accounts/{account_id}/d1/database/{database_id}/query
# list / delete by tag
GET /accounts/{account_id}/workers/dispatch/namespaces/{ns}/scripts
DELETE /accounts/{account_id}/workers/dispatch/namespaces/{ns}/scripts?tags=<tag>:yes
# create tenant bucket, Oceania hint
POST /accounts/{account_id}/r2/buckets
{"name":"<tenant-bucket>","locationHint":"oc"}
Note the R2 field is locationHint (camelCase) while D1's is primary_location_hint (snake). They
are different spellings of the same idea and a provisioning routine that assumes one shape will
silently create a bucket in the wrong hemisphere.
6 · WHAT THIS LEAVES OPEN¶
- Cost at populated tenant scale — still blocked on a schema, therefore on the empty-tile spec. Unchanged from DRIVER v1.0.
- Dispatch namespaces per account — undocumented (§3).
- Whether the 1,200/5min bucket is per user or per token (§1.4) — measurable, and the single most consequential unknown left in the study, because it decides whether the control plane can be starved by a human on the dashboard.
- Everything in §1 and §2 remains documentation, not demonstration.
SPIKE-P2P-01exists to fix that for the provisioning path.
Least sure, and what would make me wrong. The arithmetic in §1.3 is mine, not Cloudflare's, and it assumes one API call per tenant per wave — if a migration needs a read-back or a verification call per tenant, every figure doubles. And §1.2's central claim is an absence claim: I am asserting that no fleet migration mechanism exists because I could not find one across the D1 migrations page, the wrangler configuration reference and two documentation searches. That instrument has returned plenty of known-present hits, so the absence is worth something — but a Cloudflare-side feature I did not search for by the right name would refute it, and the cheapest refutation is Alex finding one in a changelog while building the spike.
Sources: D1 migrations · Wrangler configuration · D1 limits · D1 import/export · D1 Time Travel · D1 read replication · D1 create API · D1 edit API · D1 jurisdiction changelog · Build an API to access D1 · R2 limits · R2 object lifecycles · R2 bucket locks · R2 S3 API compatibility · R2 create bucket API · WfP limits · WfP platform examples · Infrastructure as code · API rate limits