Architecture

Status: living · 13.06.2026 11:57

6. Runtime view

Living document. The scenarios that most define the system's behavior. See diagrams/ for sequence diagrams.

6.1 Write path (the commit point)

The headline scenario; see diagrams/write-path-sequence.mermaid. Detailed step-by-step in section 5 (L4 write protocol) and normatively in the commit-protocol spec.

Step 0 — resolve the home zone (L2). Before any data moves, the access layer resolves the target against the global namespace (NamespaceStore, ADR-0020): an existing object returns its home zone (cached after first use); a new object is assigned a home zone and replica set by the placement service from policy — residency, replication factor, capacity. Steps 1–4 then run entirely within that home zone.

Steps 1–4 — stage and commit (home zone). As detailed in section 5: the client registers chunk IDs in the pending ledger, erasure-codes and writes fragments directly to the home zone's D servers — bulk data crosses no shared component, the basis of the throughput-scaling claim — then issues the single atomic metadata mutation that is the commit point. The file does not exist until that mutation and fully exists after it.

Two linearization points, not one. Introducing or moving a name (create, rename) is linearized globally by L2, so it is visible in every zone and never lost or resurrected (section 8.1); a file's content version is linearized at its home zone by the commit point. For a new object the global name is published only once its first content version has committed, so no reader anywhere sees a name without content. The exact interleaving is owned by the commit-protocol spec.

6.2 Read path

See diagrams/read-path-sequence.mermaid.

  1. Resolve the home zone (L2). Resolve the path/key against the global namespace → file identity, home zone, and an ACL check (cached after first use). Consistency-sensitive reads route to the home zone (section 6.6); stale-tolerant reads may target a nearer replica.

  2. Fetch the chunk map (L4). From the home zone's metadata store — or a version-fenced replica — get the chunk list, fragment locations, and EC scheme (cached after first read). In steady state both metadata hops are cached, so a hot read goes straight to the data path. One metadata value has a size ceiling, so a large object's map is segmented: the root carries the group identity and a segment table, and the chunks live in that group's own records. Resolving one is the root plus one bounded range read of that group's records — scoped to that group and that epoch, never a global scan — assembled in parsed-index order. The work that read may cost is the reader's number, not the record's: a table naming more segments than the reader's ceiling allows is refused unread, and each page asks for a fixed number of rows the reader chooses. (How many bytes a value costs to receive is the metadata store's bound, not the reader's — the store seam inherits each backend's native value limit.) Resolution is single-sourced: one call turns a committed root into that ordered list, so every consumer that can resolve — the whole-object, streaming and ranged read paths, and the maintenance passes that build the protected reference set (section 6.3) — learns which chunks an object owns the same way, and a consumer that has not yet adopted it refuses a segmented map outright rather than answering from the root alone. Three arms close the race with retirement, told apart by what a single re-read of the root sees: a generation superseded mid-resolve drops the stale resolution and restarts against the live root — never a torn or empty answer; a root found gone is a deletion, answered as "no such object" where it is seen rather than retried, since no later generation exists to restart onto; and an anomaly on a generation the root still names (an absent, undecodable, or mismatched segment — or a root record one of those re-reads cannot parse at all) fails closed with a typed error, scoped to that one object. Every fault a resolve raises about an object's own bytes is typed that way, so a maintenance consumer can tell "this object is unreadable" (contain it, keep walking) from "the store is failing" (end the pass). Restarts are budgeted, and running out is a typed error about that object — never "it owns no bytes".

    A maintenance pass resolves every committed object this way to build the set of fragments it must protect, and one object it cannot read leaves that set incomplete — which every pass reading it must then say out loud, because reclamation-safety and certification-honesty are the same property read twice. No pass that can reclaim bytes acts on an incomplete reference set: it deletes nothing, fleet-wide, since an unreadable map hides which chunks its object owns and no fragment can be shown not to be one of them. A pass that only verifies does not certify one either — it verifies every other object, names the one it could not, and reports the store not certified rather than clean. The operator-run pass that reconciles the fragment tier after a metadata restore answers the same way: it still reports every object it could read — what it marked collectable, and which chunks are unreadable, misplaced or under-replicated — names the ones it could not, and refuses to call the run clean, because a report drawn over part of a store is not a clean bill for it, and one damaged record must not leave an operator with no report at all at the moment they most need one. Its marks and its report rest on one reading: a record either read of the committed namespace could not read withholds every mark in the fleet, and a fragment either read finds referenced is never marked — two readings that can disagree are two conclusions, and the operator is shown one of them. A report-only surface reading that set (a drain's reconciliation status) is held to the same rule and never answers "satisfied" over an incomplete one: telling an operator a server is safe to decommission is the reclamation decision in another form. That surface also names the blocking record, which is what distinguishes it from the ordinary "not yet" of a server still holding referenced fragments — one says an evacuation is running and will finish, the other that nothing will converge until a named record is repaired, and an operator who cannot tell them apart is left waiting on a stall with no way out. The containment is per object throughout — the damaged record is attributed and the walk continues, so one unreadable object costs the store one object's availability, not every healthy object's protection.

  3. Read fragments directly from D servers. Holding the EC logic, the client needs only k of n fragments and reconstructs from whichever k arrive first — turning erasure coding into a tail-latency advantage: a slow or dead disk does not slow the read.

  4. Verify checksums on every fragment, catching bit rot at read time; re-read elsewhere on failure.

6.3 Repair (self-healing, intra-zone)

Driven by custodians (L4):

  1. Scrub — continuously read fragments, verify checksums against metadata, detect bit rot before the data is needed.
  2. Reconstruct — on D-server loss or failed checksum, read any k surviving fragments, recompute the missing ones, place them on healthy servers in correct failure domains, and update the chunk's location via a single atomic metadata mutation — the same commit-point pattern as a write. That mutation rewrites whichever record holds the chunk's reference: the inode record of a flat map, or the one seg: record of a segmented map (section 6.2), whose root is pinned and never rewritten. It is a compare-and-swap on the root generation, and, for a segmented map, on the segment record as the move itself re-reads it, and on the chunk's own reference — so a concurrent move of a sibling chunk in the same segment record is merged, while a superseded generation or a chunk moved under the plan makes the repair lose and retry next pass — and it also drains the repair obligation and marks each displaced fragment an orphan. An object the move cannot rewrite (a segment record that is absent, undecodable or would cross the value ceiling, or a segmented root stored under a key that is not the canonical inode:<id> spelling) is named for the operator, nothing is written for it, its obligation stays queued, and the pass does not certify.
  3. Rebalance — proactively move data off draining/hot servers, preserving failure-domain invariants.

Every recovery action is itself commit-point-atomic. A crashed repair job leaves garbage (collected later), never corruption.

The metric that matters: time-to-repair vs. failure rate. Durability ("the nines") is essentially the probability that more than m fragments fail within one repair window. Fast, parallel repair matters more than wide encoding. This is why repair-queue depth and time-to-repair are first-class telemetry (ADR-0011), not vanity metrics.

Repair vs. serve. A D-server loss makes both clients (read reconstruction, section 6.2) and custodians (repair reconstruction) read the surviving fragments, so they contend. Repair reads are throttled below foreground reads to protect read latency — but repair priority rises as redundancy falls, so a chunk near its durability floor (close to losing its m-th fragment) preempts foreground work. Durability is gate-zero (goal 1); latency yields to it only when redundancy is genuinely threatened. This dynamic priority is part of the admission/backpressure model (section 8.9).

6.4 Cross-zone replication and zone-loss recovery

  • Replication (L3): after a home-zone commit, replication workers copy chunks to other zones per policy. The remote replica becomes readable only when its record commits in L2 — never mid-copy.
  • Zone loss (L3 global custodians): detect the dead zone, find every file now below its policy replica count, re-replicate from survivors. Effective geographic durability is bounded by cross-region bandwidth and rebuild parallelism — the same repair-time principle at planetary scale.

6.5 Disaster recovery ordering

After a real disaster, restore in dependency order:

  1. Trust plane (provider CA) first of all — re-establish the CA (step-ca now, SPIRE reserved — ADR-0036) and distribute the trust bundle, because every step below dials over fail-closed mTLS (ADR-0005, ADR-0025) and cannot even complete a handshake without it. (In the single-binary profile the built-in dev-CA makes this a no-op; a real fleet must stand up the CA before anything else.)
  2. L5 coordination next — so components can find each other and elect leaders.
  3. L2 / L4 metadata — so the map to the bytes exists. Restored from out-of-band backups that do not depend on the system being restored.
  4. L3 replica verification and re-replication — so the bytes are protected again.

Restoring bytes before the map is useless; restoring the map before coordination cannot even start; and coordination itself cannot complete a single mutually-authenticated dial before the trust plane exists. This ordering is a runbook section and must be written and drilled before it is needed.

A metadata restore also brings multipart uploads back: an upload torn down after the restore point returns open while its staged bytes may already be reclaimed, and a retried Complete would publish over them (proposal 0016, D-B). So the post-restore reconciliation, run with writers stopped and before any gateway serves multipart requests, fences every open upload, and every one caught mid-Complete, last, after every chunk verdict is drawn: one metadata commit per upload moves the session to aborting and installs the retirement obligation owing all its staged records and parts, and it lands only if the session is still exactly as read and no obligation already holds a key it installs. One caught mid-Complete also gets a second obligation, the only deleter of the segment records its attempt wrote (0016 X57); its parts are owed as "all", since none can join once a Complete starts and a sparse list may not fit one metadata value. An upload the fence cannot settle — its record will not decode, it is at its last epoch, it changed mid-pass, or an obligation key is taken — is left as it was and named for an operator, and the command exits non-zero. Every run also names for an operator a fenced attempt whose segment records include one naming a chunk no part holds, not decoding or under a stray key, or whose second obligation will not decode, owes anything but that range, or is missing while such records remain: a drain deletes those records without marking any bytes, or nothing ever does. If a fence commit fails, the summary of every verdict is still written to the audit trail, marked incomplete and never clean, before the error; a re-run finishes the work and fences nothing twice.

Each run of that pass also records its progress as the restore-fence generation (section 5's mpufence record). Before its first write it records a new generation, not complete; after its last, it records that generation complete only if every write it made landed and no upload is left for an operator. Every run re-reads each fenced upload's records instead of trusting an earlier run's verdict, so neither a run that was cut short nor one over residue nobody repaired completes the generation. A restore rewinds this record too: an image taken after an earlier run completed reads complete until the next run starts. So until gateways check the record beside a restore-scoped signal (#508), multipart is kept safe by order alone: the pass runs before any gateway is re-enabled.

6.6 Consistency at runtime

See section 8 (the consistency contract). In brief: the namespace is globally strongly consistent; a file's writes are linearizable at its home zone; per-session read-your-writes and monotonic reads are provided (v1: by routing consistency-sensitive reads to the home zone); stale-tolerant reads may be served from a lagging replica; cross-session visibility of a remote write is eventually consistent, bounded by replication lag.

6.7 Delete and space reclamation

Delete is a commit-point operation too, and its cost is paid lazily, off the critical path.

  1. Unlink (the commit point). Removing the dirent — and tombstoning the inode/version — is a single atomic metadata mutation, linearized globally by L2 for the name so no zone ever resurrects a deleted file (section 8.1). The delete acknowledges here; the object is now invisible to every consistency-sensitive reader and its chunks are unreferenced.

  2. Reclaim (custodian GC, L4). The orphaned fragments are reclaimed by a GC custodian after a grace period — long enough that an in-flight reader holding the old version is never torn — exactly the pending-ledger sweep pattern (section 5). Reclamation is background work, not on the delete's latency path.

    The ledger of these grace records can outgrow what one metadata listing may return — a single delete of the largest segmented object writes more records than that — so GC walks it in bounded windows rather than reading it whole. Each pass reads at most a fixed number of records, in pages, starting where the previous pass stopped, and returns to the start of the ledger once it reaches the end; where to resume is kept in one small record of its own in the metadata store, so the walk carries over from pass to pass and across a change of leader. A pass that stops short of the ledger's end reports the walk as partial rather than the store as converged, so a caller that drives reconciliation until it is satisfied keeps going; and a resume record that no longer spells a key the ledger could hold (torn, oversized, damaged) is named to the operator and the walk restarts from the head, rewriting the record, rather than failing every pass on it. A pass concludes only what its window shows: a fragment whose record lies outside the window is left for a later pass, never reclaimed meanwhile on other evidence such as an expired write lease; a record whose value cannot be read still counts as a record — its fragment is kept and the record is named to the operator — and only a record spelled exactly as the delete path writes it can license a reclaim. The chunk-wide lease entry of an interrupted write is retired only once no fragment it accounts for is left, since other windows may still hold some. GC's own clean-up writes go out in commits of a fixed size, never one commit sized by the pass. The post-restore pass reads the same ledger page by page too, to its end, before deciding which fragments already carry a record.

    Reclaiming a marked fragment is recorded before its bytes are destroyed. GC first moves the fragment's record to a terminal reclaiming state — a compare-and-swap against the exact value it read, many records to one commit of the same fixed size — and deletes the fragment, and after it the record, only once that commit has landed. A mover that pre-marked a position and later adopts it preconditions its adoption on the record's earlier value, so from the instant the reclamation is recorded the adoption fails, instead of landing between the fragment's deletion and the record's and publishing a placement over bytes that are gone. A record that changed after GC read it loses its swap and keeps its fragment without costing any other record in the commit anything, and a pass that lost one and reclaimed nothing does not report the store converged. A record left reclaiming over a fragment still on disk — a pass that stopped between the swap and its deletes — is finished by the next pass with no second grace wait; one left over bytes already deleted is safe, since nothing can adopt it, and though a walk driven by the fragments on disk never revisits it, the sweep of records with no fragment beneath them, below, removes it. A fault that ends a pass after it has deleted fragments still commits their record deletes first, as far as it can. The bytes of an interrupted write carry no record and are reclaimed on their expired lease as before. A multipart retirement that is still draining never has its fragments reclaimed: its drain is what marks them, each record naming the retirement, and GC looks that retirement up by key — one read per fragment it would reclaim, never a listing of the retirement namespace — so a fragment it marked is reclaimed only once the retirement has finished draining and the record's own grace has run.

    GC also removes records with no fragment beneath them. A walk driven by the fragments on disk never reaches such a record — the old position of a repaired missing fragment, a record left reclaiming after its fragment was deleted, and later the planned positions a failed upload marks or a move marks before it writes — so without a sweep of its own it would stay in the ledger for good. A pass deletes a record of its window whose position none of that pass's own listings reported, once the record is older than the late-write deadline: the longest a mover may take between marking a position and authorizing the write into it, plus the longest an authorized write may take to land, plus the most two hosts' clocks may disagree. That rule is applied only to records written after their fragment — the plain grace record a delete or a vacated source leaves, and GC's own reclaiming record: a record that names the event which wrote it may have been written ahead of its fragment, and a write whose effect the store could not verify in time may still land beneath it, so such a record is left for the pass that owns its writer. The sweep's deletes, like GC's reclaim intents, go out in commits bounded by the transaction's operation budget rather than its byte budget, and a commit whose result the store could not report is judged as the one all-or-nothing commit it was, from a read of every record in it, once the store says it is out of flight: a record still holding what the pass read shows the commit did not land, so a record gone was another writer's delete; with no such survivor a record gone is recorded as gone but claimed as nobody's delete, since the pass cannot tell its own from another writer's; and a commit that may still land is named to the operator and left to the next pass. Every writer that marks a position before its fragment lands must hold to the first two, and the deadline sits strictly inside the grace window, which the build checks. The pass's clock is read before any of its listings is taken, so a listing that shows the position empty past that deadline shows it will stay empty; a listing an earlier pass took never counts, and a record on a server the pass did not list is left alone. The sweep answers to the reclaim's own protections — a record at a position a committed object or a staged upload names is kept, and while the set of committed objects is incomplete nothing is swept — and a record whose value cannot be read is kept, as is any key not spelled exactly as the delete path writes it. Each delete is conditioned on the exact value GC read, in commits of the same fixed size, so a record rewritten in the meantime survives; a delete that loses is judged again on a fresh read of the record, never on the lost condition alone; and each delete is audited and counted only after its commit has landed.

    Bytes a multipart upload has staged but not yet published — the fragments its committed parts and its in-flight staging entries name — are protected as a class of their own, apart from committed references: GC never reclaims a staged fragment, and the post-restore pass never marks one, whatever state the upload is in. Both passes read each upload's staging entries, then its committed parts, and only then the committed objects, so a part commit or a publication that lands while they read leaves every chunk visible in at least one of the three. The upload records are read a page at a time — the list of uploads, then each upload's own records — never in one listing of a whole namespace. A staged record a pass cannot read at all stops it from reclaiming or marking anything until the record is repaired, and the post-restore pass names that record in its report; a staged record a pass can read but not trust holds only that chunk's fragments; and a store fault while reading staged records fails the pass. The post-restore pass protects the upload records already written when it reads them; it is run with writers stopped. The drain-status query reads these records too, and counts a staged fragment as held: a drain is reported finished only once no byte that can still become referenced is left on the server — neither one a committed object places there nor one an upload has staged there — because telling an operator a server is safe to wipe while a live upload's bytes sit on it is the same permanent loss as deleting them. That answer covers the records the query read; keeping an upload from staging a new byte onto a server after its drain was requested is a precondition on the upload's own staging write, which no writer carries yet because none exists. A staged record that query cannot read, or can read but not trust, blocks every drain the same way an unreadable or untrustworthy committed record does, and is named in the answer. The evacuation loop is the other side of that: it moves committed content only and never rewrites an upload's own records, so a server holding only staged bytes has nothing for it to move and waits instead for those uploads to finish, abort or be reaped. Scrub reads committed references and, of an upload's own records, its committed parts alone — never its in-flight staging entries, since checking a fragment needs the committed scheme a part record carries; it too reads the committed parts before the committed objects, so a chunk a publication hands over mid-pass is still checked instead of certified unread. Once a committed object names a chunk, scrub checks the chunk where that object places it, not where the upload's part record does: the part record stays until the upload's records are retired and is not updated when the repair loop moves a fragment, so the position it names may hold nothing, and checking it there would have scrub queue, every pass, a repair the repair loop finds nothing to do for and drains. A part record scrub can read but whose placement it cannot trust is named, and for a chunk no committed object names it stops scrub from reporting the store checked, since that chunk was not. The repair loop settles an obligation for a chunk a committed object names against that object alone, as it always has, whatever a leftover part record says. For a chunk no committed object names, it reads the staged classes first, so a publication landing mid-pass is seen in at least one of them. While the upload is still open, a chunk its committed part names is repaired like a committed one — its obligation drained as a duplicate finding if nothing is missing — by a move that strands nothing it writes. The pass rebuilds the missing fragment and chooses a destination in a free failure domain, on a server of its fleet that carries no drain or decommission record of any value; it marks the destination position before writing there, with a fresh mark stamped from the reconstruction clock whatever mark was there before; it authorizes the writes only while that mark is younger than the longest a mover may take between marking a position and authorizing the write into it, and sends all of one move's writes together, each with a deadline the mark's stamp fixes, which the D server refuses to accept at or after — so one slow write never uses up the time another has, and a chunk missing several fragments is repaired in one pass when every write lands in time; and it then adopts the fragment in one commit, conditioned on the upload's record, the part record, the destination's mark and the absence of a drain record on the destination, which repoints the part record, removes the destination's mark, marks the position it vacates and removes the obligation. If anything changed in between — the upload fenced by a complete, an abort or a reaper, the part record rewritten, a drain recorded, GC deciding to reclaim the mark — or the D server refused the write or could not confirm it landed in time, nothing is adopted: the part record is left as it was, the destination's mark stays over whatever was written, for GC to reclaim after its grace, and the obligation stays queued. A position whose mark GC is already reclaiming, or whose mark cannot be read, is never written, and rules out only that position, not its server; a vacated position whose mark cannot be read, or an upload record that cannot be read, withholds the repair before anything is written and is named for an operator. The pass keeps the obligation, and does not report the pass finished, while the chunk is named only by an in-flight staging entry, while its upload has left the open state — the repair is retried once the chunk is published — or while an upload record holds it with a placement that cannot be used. A destination mark a stopped move leaves with no fragment beneath it names the move that wrote it, so GC's sweep of such records leaves it in place, as it leaves every record that names its event. One case is bounded rather than prevented: if the published object is deleted or overwritten before the upload's records are retired, after the repair loop moved one of its fragments, the leftover part record again names the chunk with no committed object to overrule it, at a position that now holds nothing, so scrub queues the chunk and the repair loop keeps it, every pass, until the drain that retires the upload's records deletes the part record.

  3. Propagate across zones (L3). If the object had replicas, the deletion propagates through the replica catalog: each holding zone drops its catalog record and its GC reclaims the local fragments. Until propagation completes, a stale-tolerant read in a lagging zone may still see the old object, bounded by replication lag (section 8.1); consistency-sensitive reads route to the home zone and never do.

  4. Prove it (audit + crypto-erase). The deletion is recorded in the append-only audit/event log — the operator's compliance and GDPR-deletion story (section 8.3). Where per-tenant envelope encryption is enabled, dropping the object/tenant key crypto-shreds the data immediately, making it unrecoverable independent of when GC runs (section 8.5).