8. Cross-cutting concepts
Living document.
8.1 The consistency contract
The most-cited paragraph of the commit-protocol spec. v1 adopts the simplest strong contract (Option A, home-zone authority) with the mechanism for the locality optimization (Option C, version-fenced reads) reserved so it can be added later without weakening any guarantee. See ADR-0015.
- Namespace operations are linearizable globally. Once acknowledged, visible in all zones. Create / rename / delete / share / ACL change never produce a "lost just-created file" or "resurrected deleted file".
- A file's writes are linearizable at its home zone. The commit point totally orders its versions.
- Per-session read-your-writes and monotonic reads. v1: by routing consistency-sensitive reads to the home zone. Forward-compatible path: version-fencing replica reads against the strongly-consistent namespace version (the existing
meta:versioncounter — reserved, costs nothing now). - Stale-tolerant reads may be served from a replica and may lag the latest by up to the replication lag. Cross-session visibility of a remote write is eventually consistent, bounded by replication lag.
- Synchronous N-zone replication is a per-tenant opt-in: no data-loss window, near-zero staleness, at the cost of cross-region write latency.
The version high-water mark also handles failover: if the home zone dies mid-session and the client is routed to a replica, the fence refuses any replica behind the last version seen — monotonic reads survive disaster.
The collaborative-editor future is served by this, not despite it: an OT/CRDT engine needs a single serialization point per document, and "the document's home zone linearizes its operation log" (point 2) is exactly that. The storage layer provides the strong primitive; the merge logic stays in the application.
8.2 Redundancy and recovery model
Durability is handled independently per layer, each using the mechanism suited to its data shape. This asymmetry is deliberate.
| Data | Mechanism | Why |
|---|---|---|
| Chunk data (D servers) | Reed-Solomon erasure coding, e.g. RS(6,3) at ~1.5× overhead | Large volume; EC is far cheaper than replication for the same durability |
| Chunk data, across zones | Whole-copy async replication (L3) | WAN latency forbids cross-zone EC; replicate copies, EC within each |
| Zonal metadata (L4) | Consensus replication (Raft), never EC | Small but precious; losing it orphans all chunks |
| Global namespace (L2) | Synchronous consensus replication, geo-distributed | Must never lag or diverge; can afford it because it is tiny |
| Everything | Out-of-band backup to independent storage | Replication faithfully replicates logical disasters (bad migration, errant delete) |
The bootstrap rule: backups must not depend on the system they back up. The namespace cannot be backed up into the file system whose namespace it is; the recursion bottoms out in boring, independent storage.
Configurable durability per zone (ADR-0008): none (dev), replication(n) (small), rs(k,m) (production). The chunk format records the scheme per chunk, so a zone that grows from replication into EC carries mixed-era data correctly.
8.3 Observability — three planes
ADR-0011. Instrument with OpenTelemetry (ADR-0012); the planes are what to instrument:
- Request plane — RED metrics (rate, errors, duration) per layer; traces following a write through chunk → EC → commit. The easy, standard part.
- Durability plane — the part that must be designed in. Under-replicated chunk count, repair-queue depth, time-to-repair distribution, replication lag per zone pair, scrub coverage and scrub-detected corruption rate. These come from instrumenting the custodians, which must emit them as a first-class output. This plane is the differentiator: a storage system silently below its redundancy floor reports all-green on the request plane.
- Capacity plane — per-server, per-zone, per-failure-domain utilization and growth rate, as a leading indicator.
Plus an append-only audit/event log of significant state transitions (placement, repairs, admissions, policy changes, deletions) — operational debugging and the provider's compliance story (GDPR deletion proof).
8.4 Manageability — declarative reconciliation
ADR-0011, ADR-0013. The management model is declarative, not imperative: operators change desired state (zone draining, tenant replication factor, server decommissioned) and the custodians continuously reconcile reality toward it — the same control-loop pattern as Kubernetes, on the substrate already present (L2 holds desired state, L3/L4 custodians reconcile). The management API is therefore mostly "read/write desired state + observe reconciliation progress": a small, safe, auditable surface.
The operations that must be first-class, safe, resumable, and observable: adding and draining capacity (the operation that separates real storage systems from toys), rolling upgrades (version skew is the normal state), policy changes as managed rollouts ("changed" in L2 vs. "satisfied" by actual replicas are different moments), and backup/restore.
8.5 Security and trust
The trust model follows from single-provider, closed federation (ADR-0005): one operator, an internal PKI, no cross-org or untrusted-operator trust to negotiate. That shrinks the security surface to four questions — who is calling, what they may touch, who can read the bytes at rest, and how tenants are kept apart. The adversarial view — assets, trust boundaries, the D-server-compromise blast radius, and the storage attack catalog with each attack mapped to its mitigating decision — is the threat model (section 14).
Authentication — two planes. Users and applications authenticate at the access layer (L1): OIDC / OAuth2 bearer tokens for the Drive / WebDAV / SDK surfaces, S3 Signature V4 (HMAC over per-tenant access keys) for the S3 surface, and OIDC + mTLS for the management API. The gateway is the authentication boundary; nothing below L1 re-authenticates the external principal. Services authenticate to each other with mTLS under the provider CA (ADR-0005) — zones, D servers, custodians, coordination; identity is the certificate, and there are no shared service secrets on the wire. The CA technology is step-ca now, with SPIRE reserved for the fleet scale where secret-less workload attestation earns its operational weight (ADR-0036); the certificate is consumed as an abstract PeerIdentity behind a seam so that future switch is a composition change, not a rewrite. etcd's own auth is defense-in-depth, never the primary boundary — coordination is network-isolated to the zone's control components and fronted by the mTLS fabric.
Authorization — enforced at the gateway, recorded in L2. The authoritative ACLs and sharing grants live in the global namespace (L2), globally linearizable (section 8.1), so a decision is never made against a stale grant. Enforcement is at the access layer: the gateway resolves the caller, fetches the cached, version-fenced ACL, and authorizes before issuing any storage operation. The metadata store and D servers are dumb about identity — they trust that a request which reached them was already authorized (defense in depth: they are network-isolated to the zone, and with encryption on they hold only ciphertext anyway). Fine-grained, relationship-based authorization (Zanzibar-class) is not v1; it is a reserved future consumer of the system, with its consistency-token hooks already noted (ADR-0018). v1 is POSIX-ish ACLs plus bucket / prefix policy.
Encryption. In transit, TLS everywhere — the external APIs and the internal mTLS fabric. At rest, optional per-tenant envelope encryption (ADR-0021): the client library encrypts before erasure coding and decrypts after reconstruction, so the metadata store and D servers hold only ciphertext; a provider KMS holds the keys behind a KeyService trait, and dropping a key is a crypto-erase (the section 6.7 delete fast-path).
Tenant boundary. The namespace, quotas, and placement policy are per-tenant (section 5, L2); envelope encryption makes the boundary cryptographic, not merely logical — one tenant's ciphertext is unreadable without that tenant's key. The fuller multi-tenancy model (isolation classes, noisy-neighbour control, quota enforcement points) is scoped separately.
8.6 The thick client and the conformance suite
The client library embeds chunking, EC, the commit protocol, and failover — so a second-language client (e.g. a Python SDK for an application team) cannot be a thin shim; it must re-implement all of it identically or it will write data the reference client cannot read, or commit non-atomically. This is why the on-disk format is a real spec with conformance vectors (specs/), and why a longer-term option is to expose the thick logic via a Rust core with FFI bindings rather than inviting risky reimplementation.
8.7 Compatibility and version skew
A half-upgraded fleet is the normal state during a rolling upgrade (constraint 2.1, scenario Q8), so version skew is designed for, not treated as an incident.
Two compatibility axes. Wire: every inter-component contract is versioned protobuf (proto), evolved by addition — neighbours interoperate across at least a one-version gap, so fields are never repurposed and removals lag deprecation by a release. On-disk: the chunk/fragment format carries its own version and EC-scheme id (ADR-0002, ADR-0019); a reader accepts every format version it claims to support, because data outlives the software that wrote it — old-format data is read, never rejected.
Tested, not hoped. Skew is exercised under the deterministic-simulation harness (ADR-0009): a simulated zone runs mixed-version nodes from a seed and asserts the commit protocol and read path stay correct across the gap. On-disk compatibility is pinned by conformance vectors per format version (specs/conformance/) that every reader must accept, and a mixed-version matrix in CI gates the one-version-gap guarantee (Q8) — the structural complement to the load/fault scenarios in section 10.
A metadata record's shape evolves the same way (proposal 0016 decision 7(a)). An inode's chunk map is either flat (the chunk list inline, a JSON array — the shape every record written before segmentation carries, needed because one metadata value has a size ceiling far below the launch object-size target) or segmented (a JSON object: the segment-group identity {nonce, epoch}, a segment count, and the segment table, with the chunks themselves in per-segment seg:<nonce>:<epoch>:<index> records). The two are told apart by JSON type, so a pre-existing record decodes and re-encodes byte-identically — a correctness requirement, not hygiene: every metadata compare-and-swap in the system is require(key, encode(prior)) compared byte-for-byte against the stored value, so an encoding that gained a tag or a wrapper would turn every overwrite, backfill, reconstruction and rebalance of every pre-existing object into a permanent conflict. seg:/seggrp: are new key prefixes over which nothing older reads, so their introduction needs no migration and rewrites no existing record. Structural invariants of the segmented shape — the segment count agreeing with the table, a gap-free ascending index sequence, every index inside the fixed-width key space the seg: grammar can address (which is therefore also the format's maximum segment count), a contiguous byte tiling that spans exactly the object's recorded size (and one whose arithmetic stays inside u64, since a wrapped span would agree with a forged size), a well-formed nonce, and one segment record's chunk lengths summing (checked, never wrapped) to the span it declares — are enforced at decode: a malformed record is an error, never a value a consumer could half-resolve. Capacity is deliberately not one of those invariants: the ceiling on how many segments a root may name is a number a deployment picks, so it is enforced only where a table becomes work — a publication that writes one, and the range read that would spend one — and a limit that moves later can never make an already-published object undecodable. That ceiling is itself sized with 2× headroom: a worst-case segment table encodes to at most half the tightest backend's value limit and the other half is a reserve for the rest of the record — the object metadata a client supplies, whose width is the caller's, and whatever field a later revision adds — because a root that only just fits is one a later field addition would make unwritable, and an object whose root can no longer be re-written is an object whose placement can never be repaired. The nonce is a validated type rather than a string for the same class of reason: the seg: range derived from it is what a cleanup pass deletes, so a nonce carrying the key grammar's separator would name another generation's live range. The key space is the opposite case and that is why it is checked at decode: a segment past it has no canonical key at any setting, so no reader, GC pass or repair could ever address its record. Symmetrically, a shape this build cannot yet resolve or publish is refused on the way in as well: until the staged-publication committer exists, creating an inode with a segmented map, or superseding one, is a typed error rather than a root written over segments that do not exist (or one whose seg: records nothing would ever reclaim). Every existing consumer of a chunk map obeys the same rule from the other side — the read path, the delete that must orphan an object's fragments, the maintenance loops that reclaim or move them, the id recovery that must not re-mint them all treat a shape they cannot resolve as a typed error for that object rather than as an empty chunk list, because "this object owns no chunks" is indistinguishable from a zero-length object and is precisely how a live object's fragments would become unreferenced and be collected. Landing the shape and its codec ahead of any producer or resolver (this slice) keeps the byte-identical, always-flat behaviour of every object in the fleet unchanged until a later slice actually publishes a segmented map.
An orphan mark's value evolves the same way (proposal 0016 decision 4.2). The grace record a delete or a repair leaves for a fragment it strands (orphan:<dserver>:<chunk>:<index>) has always been a bare decimal — the instant the fragment became unreferenced — and every mark already stored still is. The value now has three shapes: that bare decimal (legacy, with no event identity); a JSON object carrying the instant and the identity of the unreference event that wrote the mark (structured); and the same object with "reclaiming":true added (reclaiming: the garbage collector's recorded decision to reclaim the fragment, carrying an event only if the mark it replaced had one). They are told apart by form — digits or an object — so no stored mark needs migrating, and one codec, kept beside the key's own definition, is the only reader and writer of all three. Decoding accepts exactly what encoding writes: a value is re-encoded after it is read and refused unless it comes back byte for byte, so a mark that is only read is never rewritten and never has its grace clock restarted, a compare-and-swap built from a decoded mark matches the stored bytes, and a value in any other spelling — a leading zero, reordered, unknown or null fields, a default written out, an event identity outside its bounded grammar — is not a mark any maintenance pass acts on: the fragment is kept and the record named for repair. reclaiming is terminal, and no writer replaces it: the collector deletes the key once the fragment is gone, and would take any value written over it with it. A downgraded custodian reads a structured or reclaiming value as a mark it cannot read and keeps the fragment — the safe direction.
The segmented shape has a write side too (proposal 0016 decision 7(f)). A maintenance pass that moves one chunk's placement goes through one primitive (repoint_chunk), the write counterpart of the single resolver: it rewrites whichever record holds the chunk's reference — the flat inode record, or one seg: record, leaving the segmented root byte-identical — as a compare-and-swap on the root generation, on that record's bytes as the move itself re-reads them, and on the chunk's reference. The custodian's reconstruction pass repairs a chunk held in a seg: record this way (section 6.3), adding the obligation's delete and the orphan marks to the same batch. The re-encoded record is weighed against the backend value ceiling before anything is written, and a segment record the root still names but the move cannot read is the same typed, per-object error the read side raises: the pass contains it, and never rewrites it or reads it as a lost race.
8.8 Multi-tenancy
The primary deployment is a provider serving many tenants on shared infrastructure (section 1.4), so a tenant is a first-class unit — a namespace partition plus an identity domain, carrying its own policy (residency, replication factor, encryption, quotas, rate limits). See ADR-0022.
Multi-tenancy is logical, not physical (single-provider, ADR-0005): tenants share zones, D servers, and the metadata tiers. Isolation is the composition of four boundaries, each enforced at a definite point:
- Namespace — the L2 global namespace is partitioned per tenant; cross-tenant naming or traversal is impossible by construction (ADR-0020).
- Data — per-tenant envelope encryption makes the boundary cryptographic; D servers hold opaque, mixed-tenant ciphertext (ADR-0021).
- Capacity — per-tenant quotas checked at admission in the gateway against the tenant's L2 record (section 8.9).
- Performance — per-tenant rate limits at the access layer, and placement spreads a tenant across failure domains, containing noisy neighbours.
Enforcement lives at L1 (authentication, rate) and L2 (authorization, quota), never below — D servers and the metadata store stay tenant-oblivious, trusting an admitted request. The storage tier stays dumb; the policy tier stays centralized.
8.9 Admission control and backpressure
The system fails closed under pressure: a write is admitted only when identity, quota, capacity, and failure-domain room all allow it, and is refused with a clear, retryable signal otherwise — never a silent half-write or a durability corner cut.
- Quota / rate — the gateway checks the tenant's quota and rate limit (section 8.8) at admission; over a hard limit it rejects (429-style), backpressuring the client.
- Zone full — per-failure-domain utilization is the binding capacity signal (section 7.3): when a domain has no room for the configured EC scheme, placement (L2) redirects new writes to a zone with capacity, or rejects if residency policy forbids the redirect. Running out of room in one domain blocks EC writes before total capacity is exhausted.
- Metadata tier saturated — backpressure propagates to clients; the metadata tier is shardable (goal 2), so sustained pressure is a scaling signal surfaced by telemetry, not a failure.
- Repair vs. serve — see section 6.3: repair reads are throttled below foreground reads, but their priority rises as redundancy falls, so a chunk near its durability floor preempts. Durability (goal 1) outranks latency.
- Fragment-write authorization deadline (
W_write) — a caller-side timeout bounds only how long a writer waits, never when an already-accepted fragment write takes effect. So theChunkStore::put_fragmentseam and thePutFragmentRPC (FragmentPutRequest.deadline_millis) carry an optional authorization deadline in epoch milliseconds, and the D server enforces it against its own clock at its publication point — the last instant before the single step that makes the fragment visible, with the accept queue, the thread-pool queue and the data write already behind it, so a write parked anywhere upstream is refused rather than queued. The judgment deliberately precedes publication rather than following it: a refused write was never published, so it leaves nothing stored even if the process dies — whereas publishing first and retracting afterwards has a readable, crash-exposed window in which the bytes exist and can also destroy a concurrent same-id writer's acknowledged fragment. A refusal restores the store's pre-write state exactly (no fragment, no scratch); a store that cannot restore it reports a backend fault rather than the clean verdict, so a definite "nothing landed" is never returned over residue. A refusal touches only state private to its own write — never container state shared with concurrent writers, such as the chunk directory a filesystem store would be left holding empty — because a cleanup that writes a shared path can strip it from under a live writer, and bounding that with retries only moves the failure to one more racing refusal. Shared containers left empty are reclaimed instead where no write of that store is in flight by construction (for the filesystem store: at open, atomically in their emptiness), so cleanup is never a hazard to a live write. The publishing step itself is then verified, not assumed: it takes non-zero time and can straddle the deadline on a slow device, so the D server re-reads its clock immediately after it and, unless that reading is still inside the window, declines to acknowledge the write —Oktherefore means "published strictly before the deadline", checked at both ends. It reports that case as undetermined, not as a late landing: a clock read dates the read, not the syscall before it, so a timely publication followed by a descheduled thread is indistinguishable from one that overran, and asserting the worse reading would label in-window writes as leaks. The outcome is a typed, operator-classifiable one (WriteDeadlineExpired, distinct from a backend fault) carrying which of the two happened, because they say different things about the bytes: refused-and-nothing-landed (FAILED_PRECONDITIONon the wire, terminal — re-authorize, do not retry) or publication-unverified (ABORTED, indeterminate — that D server may hold the bytes, in window or not, so the caller must not count the write and must re-read rather than assume; bytes that did land late are garbage the position's evidence covers, and are deliberately not unlinked because retracting is neither atomic with publication nor safe against a concurrent same-id writer). The field is additive: an absent deadline is the pre-existing unbounded behaviour, so every current writer is unchanged — and, until a capability exchange exists, an old D server silently ignores it, so a mixed-version fleet does not get the guarantee. This is the mechanism only; the deadline's value and its strict margin against the reclaimer's orphan grace (G_orphan > W_write + δ_clock, whereδ_clockabsorbs the skew between the authorizing and applying clocks) belong to the multipart-commit protocol (proposal 0016 decision 5) and the window-sizing slice.
The principle throughout: shed or slow load predictably, surface it on the capacity and durability planes (section 8.3), and never trade correctness for admission.
8.10 Build, test, and CI
CI is the cross-cutting enforcement arm: one gate applies uniformly to every building block, run identically on a laptop and in CI because the logic lives in cargo xtask ci, not in workflow YAML (ADR-0009, ADR-0016). The workflow files under .github/workflows/ are thin callers; that — together with branch protection and CONTRIBUTING.md (the contributor-facing how-to) — is the authoritative pipeline definition, deliberately not restated here so it cannot drift.
- The merge gate (every PR).
fmt --check,clippy -D warnings, build, and test — where the deterministic-simulation property suite runs, the correctness authority (ADR-0009; section 10–13) — plus format conformance vectors (ADR-0002). Policy gates ride alongside: DCO sign-off and thecargo-denylicense/advisory wall (ADR-0003),require-issue, and ADR-immutability. A contribution that breaks one fails the build. - Repo-hygiene guards (same gate, before the build). Sub-second checks that turn a convention into a rule, listed as data in
xtask::repo_guard::CI_GUARDSso a test can see every guard the gate runs: no stray gitlinks in the index,#![forbid(unsafe_code)]in every crate root, and the blackbox property of proposal 0017 §9 — nothing in the normal dependency closure ofwyrd-validatemay be awyrd-*crate, or the validator would check Wyrd against Wyrd's own types. That guard reads the resolve graph fromcargo metadata --locked --all-features, so a Wyrd crate behind an off-by-default feature is still found, and it never rewritesCargo.lock. It follows only normal edges, transitively, and separately checks the declared dependency list, optional entries included. Dev-dependencies stay unconstrained, because the validator's test fixtures need them. Metadata it cannot fully read is a failure, never a pass.cargo xtask blackbox-guardruns it alone (--metadata FILEscans a savedcargo metadatadocument);cargo xtask ci-dry-run [--workspace DIR]runs the gate's guards for real on a workspace and only prints the cargo steps that would follow. - Heavier tiers (nightly / on-demand, from M2). The container integration suite and the real-hardware fault tiers — Tier-1 disk-faults, Tier-2 kill-and-reconstruct (proposal 0004's test taxonomy; section 10–13) — run off the PR path on slower runners, so the merge gate stays fast.
- The compounding loop. A DST seed that finds a bug is committed as a permanent regression test (ADR-0009); a recurring defect class or process friction becomes a systemic guardrail via the Act beat (ADR-0023) — so the same lesson is never re-learned instance by instance.
- Outbound integrity. Releases are signed (Sigstore/cosign keyless) with SLSA build provenance, and Actions are pinned by commit SHA (ADR-0030) — the outbound complement to the inbound advisory wall.
The principle: the rules that protect correctness — atomicity, the license wall, format compatibility, ADR immutability, issue traceability — are machine-enforced in CI, not honour-system, and single-sourced so the gate a contributor runs locally is exactly the gate that blocks the merge.