Skip to content

SPIKE-P2P-03 — FINDINGS — fleet migration runner

Ran: 2026-08-14, 07:2xZ–07:3xZ · Seat: Alex (bridge) · Cleared by: Tim, direct Status: findings, not canon. Nothing here amends the North Star. The Architect seat holds the design decisions this informs. Brief: outputs/SPIKE-P2P-03-BRIEF-DRAFT-2026-08-14.md (drafted by this seat; run on Tim's direct clearance rather than Architect review — noted so the provenance is not overstated). Filed because: these findings existed only in a session transcript and a scratch directory. ARCHITECTURE-FOUNDATIONS-01 §0 — a partial document in the repository beats a complete document in a session.

Account: RTOpacks Product (c93403113c55c1df41d43e4c878d3054), scratch lane, zero gates. Token: product-scratch-01, sha256[:8] 9a8c0bd2, read via rtopacks-product-sa. Totals: 5.5 min wall-clock · 522 API calls · zero HTTP 429 · account returned to baseline.


0 · THE ONE-PARAGRAPH VERSION

A fleet migration runner of the simplest possible shape works: it skips tenants already at the target version, resumes cleanly after an abort with no special arguments, and quarantines a failing tenant rather than halting the fleet. Provisioning is ~1 s per database and every database honoured the oc hint. The rate ceiling does not bind a serial 4,000-tenant migration, but it binds any concurrent one, hard. The most valuable finding was not designed: a ledger read failed once in twenty, the runner read that failure as "not yet migrated", and re-applied a migration to an already-migrated tenant. It was survivable only because the statement happened to fail loudly.


1 · MEASURED RESULTS

Provisioning (step 1)

20/20 databases created, every one created_in_region: OC against a primary_location_hint of oc. Fleet total 20.2 s. Slowest 1.17 s · fastest 0.90 s · mean 1.01 s.

Consequence: a full tenant stand-up (database + schema + Worker) lands in the 5–10 s band, so signup can complete synchronously rather than needing a queued job and a holding screen. This retires the provisioning-throughput question without a throughput spike.

Clean run (step 4) — reported and verified separately

Runner's own report applied 20 · skipped 0 · failed 0 · 17.0 s · 60 API calls
Independent verification usi column present 20/20 · schema_version = 2 20/20 · 0 disagreements

Stated as two claims because they are two claims. The runner saying it worked is not evidence that it worked; the independent read is.

Resumability (step 5)

Aborted after exactly 10 tenants. State at abort, independently read: 10 at v3, 10 at v2. Re-run from the top with no special arguments: applied 10 · skipped 9 · failed 1.

Final state: 20/20 at v3, 20/20 carrying the column. Run 2 cost 41 calls against 30 for run 1's ten tenants — it did not re-do completed work, apart from the one tenant in §2.

Resumability of the shape is demonstrated. A runner that reads a per-tenant version and skips ahead needs no external state, no cursor, and no special resume mode.

Partial failure (step 6)

mig-tenant-07 seeded with the target column, then migration 004 run fleet-wide.

applied 19 · failed 1 · read_failures 0 · 19.3 s · 59 calls

{"tenant": "mig-tenant-07", "http": 400,
 "errors": ["duplicate column name: status: SQLITE_ERROR"]}

It quarantines. The other 19 completed. Independent read: 19 at v4, 1 stranded at v3.

The stranded tenant is identifiable from the ledger alone — but finding it costs one read per tenant. At 4,000 that is 4,000 reads to answer "who is behind?". This is the measured argument for a central index alongside the per-tenant ledger; see §3.


2 · THE FINDING NOBODY ORDERED

On the resume run, mig-tenant-06 — successfully migrated in run 1 — was not skipped. The runner attempted to re-apply and got duplicate column name: enrolled_at: SQLITE_ERROR.

Its ledger, read afterwards, was entirely correct: version 3 recorded at 07:28:49, column present.

So the ledger read returned no usable value for a row that was already committed, and the runner treated an unreadable ledger as an unmigrated tenant.

What was ruled out. A bounded read-after-write probe — 5 tenants × 8 write-then-immediately-read cycles, 40 cycles total — returned 40/40 consistent, zero stale reads. This is not a D1 read-after-write consistency problem at this volume. A transient API error is the likely cause.

What could not be established, and why. The first runner's current_version() collapsed every failure mode — HTTP error, malformed body, parse failure — into None. The cause is therefore unrecoverable after the fact. The instrument reported its own control flow rather than its effect on the world, which is precisely the defect class RTOPACKS-CONDITION-REVIEW-2026-08-11 §5 names. It was rebuilt to record read failures explicitly before step 6, which then logged read_failures: 0.

The design lesson, which stands regardless of cause:

An unreadable ledger must fail closed. It must never be read as "not yet migrated."

Here it was survivable only because ALTER TABLE ADD COLUMN fails loudly on a repeat. Had migration 003 been an INSERT or an UPDATE, the runner would have silently applied it twice to a live tenant, and nothing in the system would have reported a problem. At 4,000 tenants a 1-in-20 transient read failure is ~200 tenants per migration taking a second application.

This is also the only realistic failure the spike observed. Step 6's failure was hand-picked, as the brief's own least-sure section admitted. The accidental one behaved worse than the designed one.


3 · THE 4,000-TENANT ARITHMETIC

Extrapolated from 20. Not measured at 4,000.

Measured: 3.0 API calls per tenant per migration · 0.85 s per tenant (step 4).

4,000 × 3 calls 12,000 API calls
Serial execution 4,000 × 0.85 s ≈ 57 minutes
Rate-limit floor 12,000 ÷ 1,200 × 5 = 50 minutes

The ceiling does not bind a serial run — but only just. Execution time (57 min) marginally exceeds the budget floor (50 min).

Any concurrency inverts this immediately. Parallelise and the 1,200-per-5-minutes ceiling becomes the sole constraint, hard-flooring a full-fleet migration at ~50 minutes regardless of worker count. And that is a one-statement migration: a three-statement migration triples the call count and pushes the floor past 2.5 hours.

Observed rate during this job: 522 calls over 5.5 min = 95 calls/min, against a 240/min sustained budget. No 429 encountered. The per-user-vs-per-token question remains open and now matters more — if the budget is per user, a human on the dashboard competes with a running fleet migration.


4 · WHAT THIS DOES AND DOES NOT SETTLE

Settles: the runner shape works; resumability needs no special mode; failure quarantines rather than halting; provisioning is fast enough to be synchronous; oc placement is honoured at fleet scale; the rate ceiling's real shape.

Does not settle: ledger location. This tested shape (a) — a per-tenant table, matching D1's own d1_migrations convention. It says nothing about shape (b), a central ledger in the control plane. Step 6's "20 reads to find one stranded tenant" is the strongest indirect evidence for (b) that (a) can produce. The likely answer is both — (a) as truth, (b) as an index — but that is the Architect's decision, not this spike's.

Does not settle: realistic failure modes. Network timeout mid-write and a briefly-unavailable tenant database are the failures that will actually occur. Neither was modelled. §2 is an accidental sample of one.


5 · TEARDOWN

All 20 databases deleted, zero failures. Post-teardown listings against the pre-flight baseline:

After Baseline
D1 databases 0 0
Dispatch namespaces 0 0
Workers 0 —

No residue. Counted against a measured starting state, not asserted.


6 · GUARDS

Zero calls of any kind against the RTOpacks Prototype account · only product-scratch-01 used · no wrangler at any point, REST only · no DNS, route, custom domain or zone · no Workers, R2, KV, Durable Objects, queues or cron — D1 only · all data synthetic and self-labelling · scratch directory outside the repo · no repo commits during the run · fleet capped at 20 as instructed, extrapolation by arithmetic rather than by building.


Least sure, and what would make this wrong. §2's cause is unestablished and unreproduced — if it were in fact a D1 consistency behaviour rather than a transient error, the mitigation is different (read-your-writes handling rather than fail-closed) and the 40-cycle probe was simply too small to catch it. Second, every figure in §3 extrapolates linearly from 20 tenants; nothing here tests whether per-call latency or error rate degrades as the fleet grows, and a non-linear term would change the arithmetic entirely. Third, mig-tenant-07's quarantine behaviour is a property of this runner, which I wrote to continue on failure — it is evidence that quarantine is achievable, not that it is what any runner does by default.