Skip to content

FLEET-STUDY-CLOSE-01 — the three open gaps, closed — 2026-08-14

Written by: Claude (Driver seat, Opus) · Closes: §3 of time-machine-2026-08-14-DRIVER-v1.0 Status: findings, not canon. Nothing here amends the North Star until the Architect seat rules on it. Register: every figure below is quoted from Cloudflare's own published documentation, read this session and linked at the foot. Nothing has been run. No account was mutated.


0 · THE ONE-PARAGRAPH VERSION

All three gaps are closed enough to design against. Fleet migration is not a Cloudflare feature — it is a thing we build, and the documented ceiling that bounds it is the same 1,200-requests-per- five-minutes API limit that bounds provisioning. R2 has no object versioning at all, so retention and per-tenant restore must be built from immutable keys, lifecycle rules and bucket locks rather than assumed. And the study turned up two errors in our own canon about residency, one of which overstates a risk and one of which understates the option set.


1 · GAP 1 — FLEET MIGRATION TOOLING

1.1 What Cloudflare gives us

Fact Value
Tracking table d1_migrations — "Creating migrations will keep a record of applied migrations in the d1_migrations table found in your database." Renameable via migrations_table.
Migration directory migrations_dir, default migrations/
Glob discovery migrations_pattern, default migrations/*.sql. "When migrations_pattern is set, migrations_dir must also be set, and migrations_pattern must start with whatever migrations_dir is set to."
Nested layouts Supported, for ORM output such as Drizzle: "migrations_pattern": "migrations/*/migration.sql". Each migration is recorded "as a path relative to migrations_dir".
Commands wrangler d1 migrations create / list / apply
Foreign keys PRAGMA defer_foreign_keys = true before changes that would violate them

1.2 The structural fact — there is no fleet path

Every documented migration mechanism is per-database and wrangler-bound. wrangler d1 migrations apply resolves its target through a Wrangler configuration file, by binding name or database name. The documentation offers no bulk, fan-out or namespace-wide migration command, and no REST endpoint for applying migrations at all.

The only programmatic route is the D1 REST query endpoint, and Cloudflare says plainly what it is for: "D1's built-in REST API is best suited for administrative use as the global Cloudflare API rate limit applies."

So: a fleet migration runner is ours to build. Minimum shape the evidence supports —

  1. Enumerate tenant databases from the control plane's own registry, not from a d1 list call.
  2. Apply ordered .sql per tenant over POST /accounts/{id}/d1/database/{db_id}/query.
  3. Keep our own applied-migrations row per tenant database — d1_migrations is only written by wrangler, so a REST-driven runner must maintain the ledger itself or it has no idea where it got to.
  4. Be resumable and idempotent, because a 3,000-tenant wave will be interrupted by the rate limit.

1.3 The ceiling that bounds it

Limit Value
Client API, per user / account token 1,200 per 5 minutes
Client API, per IP 200 per second
On exceeding HTTP 429, and "blocks all API calls for the next five minutes"
Headers returned Ratelimit, Ratelimit-Policy, retry-after
Max SQL statement length 100,000 bytes
Max SQL query duration 30 seconds

Arithmetic, ours not Cloudflare's: at 3,000 tenants and one batched call per tenant, a fleet-wide migration wave is ≥12.5 minutes of pure rate-limit budget. Provisioning from cold — D1 create plus script upload, two calls per tenant — is ≥25 minutes for 6,000 calls. Both are acceptable. Neither is instant, and both must be built as resumable jobs rather than a loop that assumes it will finish.

Batching is what keeps this cheap. A migration wave should send one multi-statement call per tenant, not one call per statement — the 100 KB statement ceiling is generous enough that this is almost always possible.

1.4 The ambiguity I could not resolve — and it matters

The prose says the limit is per user: "1,200 requests per five minute period per user, and applies cumulatively regardless of whether the request is made via the dashboard, API key, or API token." The limits table says per user/account token. These are not the same claim.

If it is per user, an automated fleet migration and a human clicking around the dashboard draw on one shared bucket, and the control plane can be starved by a person looking at a graph. If it is per token, a dedicated account-owned token has its own budget and the two never collide.

This is not resolvable from the documentation. It is cheap to measure and the spike should measure it.

1.5 Two constraints that shape the design

  • No gradual deployments for user Workers. Confirmed again: "Changes made to user Workers create a new version that deployed all-at-once to 100% of traffic." Cohort rollout is ours to build.
  • Time Travel is not a fleet instrument. 30-day retention, always on, but "Restoring a database to a specific point-in-time is a destructive operation, and overwrites the database in place", capped at 10 restores per 10 minutes per database, and documented only through wrangler — no REST restore. A fleet-wide rollback of 3,000 databases has no documented mechanism.

2 · GAP 2 — R2 AND BACKUP AT FLEET SCALE

2.1 Fleet limits — not the constraint

Limit Value
Buckets per account 1,000,000
Objects per bucket · data per bucket Unlimited
Object size 5 TiB
Bucket management operations 50 per second per bucket
Concurrent writes to the same key 1 per second
Cloudflare REST API (R2) 1,200 per 5 minutes — object operations belong on the S3 API, not here

2.2 The finding that changes the backup design

R2 does not implement object versioning. On the S3 compatibility page, both PutBucketVersioning and GetBucketVersioning are marked ❌ unimplemented.

Consequence: an overwrite is final and a delete is final. There is no "restore the previous version" primitive to fall back on. Every retention guarantee we make to an RTO — and evidence retention is the point of Record — has to be built out of:

  • Immutable keys. Never overwrite; write a new key per version and resolve current in the index.
  • Lifecycle rules — max 1,000 per bucket; expire by age or by date; transition Standard → Infrequent Access; abort incomplete multipart (a default rule already expires these at 7 days). Deletion is not instant: "Objects will typically be removed from a bucket within 24 hours", and existing objects may lag when a rule changes.
  • Bucket locks — max 1,000 rules; time-based, date-based or indefinite; "Bucket lock rules take precedence over lifecycle rules"; and "A bucket cannot be emptied while any bucket lock rules are configured."

That last sentence is load-bearing in two directions. It is exactly the primitive that makes a compliance retention promise real. It is also the reason "delete this tenant" stops being one call — a locked bucket cannot simply be emptied, so offboarding becomes a defined multi-step process that must be designed, not discovered at the first cancellation.

2.3 D1 backup mechanics

Fact Value
Time Travel Always on, no opt-in. 30 days on Workers Paid. Bookmarks created automatically.
Restore semantics Destructive, overwrites in place
Restore rate 10 restores per 10 minutes per database
Export wrangler d1 export <db> --remote --output=./database.sql — SQL dump; --table, --no-data, --no-schema
Import wrangler d1 execute <db> --remote --file=<f.sql>, "limited to 5GiB files"
REST import/export Not documented. wrangler only.

Per-tenant backup is ours to build too. Time Travel is a 30-day in-place undo, not a backup: it overwrites rather than producing an artefact, and it cannot restore into a different database. (That last clause is inference from the documented restore semantics, not a quoted statement — the docs do not address it directly.) A real per-tenant restore path means periodic export to R2 under immutable keys, with lifecycle for retention and locks where an obligation requires them.


3 · GAP 3 — PROVISIONING RATE LIMITS — CLOSED

The section that truncated twice is the API Rate Limits section at the foot of the Workers for Platforms limits page. It carries the standard account table, reproduced at §1.3 above. There is no special provisioning throughput allowance for Workers for Platforms: user-Worker uploads draw on the same 1,200 / 5 minutes as everything else.

One figure in that table is a design constraint in its own right:

Quota Value
User API token quota 50
Account API token quota 500

A token per tenant is impossible at 3,000 tenants. The control plane runs on a small number of account-scoped tokens, and per-tenant isolation is carried entirely by bindings — which is what the reference architecture already told us: "Each User Worker can only access the bindings that are explicitly attached to it."

Still unmeasured: dispatch namespaces per account. Not stated anywhere I could find. We need one, possibly a second for staging, so it is very unlikely to bind — but it is unmeasured, not absent, and should not be written down as "unlimited".


4 · TWO CORRECTIONS TO OUR OWN CANON

Both are in where-we-are.md and in DRIVER v1.0 §2.1. Neither changes the residency conclusion.

4.1 Jurisdictions are three, not two

We wrote "eu and fedramp only". The D1 create endpoint now documents eu, fedramp, us. R2's jurisdiction header remains default | eu | fedramp.

There is still no Australian or Oceania jurisdiction, so the conclusion stands — oc is a hint, not residency, and must never be described to an RTO as a guarantee. But the sentence as written is factually wrong and will be caught by anyone who checks.

4.2 We assert an immutability the docs do not state

We wrote of the location hint: "This cannot be corrected after a resource is created."

What the documentation actually supports is narrower:

  • Jurisdiction — explicitly create-time only: "can only be set at database creation time via wrangler, REST API or the UI and cannot be added/updated after the database already exists."
  • primary_location_hint — set at create (POST /accounts/{id}/d1/database, allowed values wnam enam weur eeur apac oc). The PATCH endpoint on an existing database exposes only read_replication. So it is not changeable through any documented API — which supports the practical conclusion — but Cloudflare nowhere says it is immutable.

Recommended rewording: "there is no documented means of changing it after creation". Same operational consequence, and it survives contact with someone who reads the page.

4.3 One thing that is reversible, and we should say so

read_replication.mode is settable at create and editable via PATCH (auto | disabled). D1 replicates to every supported region — ENAM, WNAM, WEUR, EEUR, APAC, OC. So enabling replication remains a residency decision, exactly as canon says, but unlike the location hint it is a decision we can take back.


5 · THE API SHAPES, FOR THE SPIKE AND FOR THE CONTROL PLANE

# create dispatch namespace
POST /accounts/{account_id}/workers/dispatch/namespaces
     {"name": "<namespace>"}

# create tenant database, Oceania hint, no jurisdiction
POST /accounts/{account_id}/d1/database
     {"name": "<tenant-db>", "primary_location_hint": "oc"}

# upload user Worker with the tenant's D1 bound, multipart
PUT  /accounts/{account_id}/workers/dispatch/namespaces/{ns}/scripts/{tenant}
     metadata={"main_module":"worker.mjs",
               "bindings":[{"type":"d1","name":"DB","id":"<tenant-db-id>"}],
               "tags":["tenant-<id>"],
               "compatibility_date":"<date>"}

# administrative SQL against a tenant database
POST /accounts/{account_id}/d1/database/{database_id}/query

# list / delete by tag
GET    /accounts/{account_id}/workers/dispatch/namespaces/{ns}/scripts
DELETE /accounts/{account_id}/workers/dispatch/namespaces/{ns}/scripts?tags=<tag>:yes

# create tenant bucket, Oceania hint
POST /accounts/{account_id}/r2/buckets
     {"name":"<tenant-bucket>","locationHint":"oc"}

Note the R2 field is locationHint (camelCase) while D1's is primary_location_hint (snake). They are different spellings of the same idea and a provisioning routine that assumes one shape will silently create a bucket in the wrong hemisphere.


6 · WHAT THIS LEAVES OPEN

  1. Cost at populated tenant scale — still blocked on a schema, therefore on the empty-tile spec. Unchanged from DRIVER v1.0.
  2. Dispatch namespaces per account — undocumented (§3).
  3. Whether the 1,200/5min bucket is per user or per token (§1.4) — measurable, and the single most consequential unknown left in the study, because it decides whether the control plane can be starved by a human on the dashboard.
  4. Everything in §1 and §2 remains documentation, not demonstration. SPIKE-P2P-01 exists to fix that for the provisioning path.

Least sure, and what would make me wrong. The arithmetic in §1.3 is mine, not Cloudflare's, and it assumes one API call per tenant per wave — if a migration needs a read-back or a verification call per tenant, every figure doubles. And §1.2's central claim is an absence claim: I am asserting that no fleet migration mechanism exists because I could not find one across the D1 migrations page, the wrangler configuration reference and two documentation searches. That instrument has returned plenty of known-present hits, so the absence is worth something — but a Cloudflare-side feature I did not search for by the right name would refute it, and the cheapest refutation is Alex finding one in a changelog while building the spike.

Sources: D1 migrations · Wrangler configuration · D1 limits · D1 import/export · D1 Time Travel · D1 read replication · D1 create API · D1 edit API · D1 jurisdiction changelog · Build an API to access D1 · R2 limits · R2 object lifecycles · R2 bucket locks · R2 S3 API compatibility · R2 create bucket API · WfP limits · WfP platform examples · Infrastructure as code · API rate limits