Skip to content

Normative · Part

Scope, backfill, eviction and deletion

Not everything a node stores may leave it. SSP/1 bounds replication in two tiers: a hard floor that policy cannot override, and a policy layer above it.

A conforming mesh MUST define a floor: a partition of its tables into three categories, fixed by the implementation rather than by configuration.

Category Meaning Policy may reach it
Syncable May replicate, subject to policy Yes
Never syncable Structurally excluded No — unrepresentable in policy
Node-local derived Reconstructible locally; replicating it is waste No

The floor is a mechanism, not a table of names. SSP/1 does not enumerate which schemas fall where, because that is a property of the application built on it rather than of the protocol. What SSP/1 requires is that the partition exists, that it is identical on every node of the mesh, and that the never-syncable category is structurally unrepresentable in policy rather than merely denied by default.

The distinction matters. A deny-by-default rule can be overridden by a misconfiguration or a widening operation. Structural exclusion cannot, and the categories that belong there are the ones where a single mistake is unrecoverable: credential and key material, transport state, and anything whose replication would leak the mesh’s own trust configuration.

Enforcement is dual-sided. A producer applies the floor in its scan predicate and treats a local violation as an error, because a producer emitting out-of-floor data has a bug that silence would hide. A receiver drops and logs, because raising there lets a hostile peer wedge the apply loop.

Each node holds a local policy store that is never replicated, keyed (peer_site_id, table_schema, table_name, layer) with effect allow or deny.

* is a valid wildcard for peer and for table. It is never valid for schema: a schema-level wildcard is how an accidental widening becomes a total one.

Layers are ordered system < role < user. Resolution for a given (peer, schema, table):

  1. Outside the floor, deny unconditionally.
  2. The highest layer containing any matching row wins.
  3. Within that layer, the most specific undominated row wins: an exact peer beats *, and an exact table beats *.
  4. Two incomparable maximal rows: deny and log. Never guess open.
  5. No matching row: deny. Fail closed.

Policy is a sender-side decision and never appears on the wire. A receiver therefore cannot widen its own feed by lying about its policy, because it is not the one consulting it. This is why scope is enforced where the data lives rather than where it is wanted.

Narrowing is always permitted. Widening is floor-checked and triggers backfill.

The role layer needs one thing the other layers do not: the assignment itself must travel between nodes, or two nodes resolve policy against different role views with no signal that they disagree.

A role assignment is a chained COSE_Sign1 record — the same chained-record shape the object plane uses for authorization records and sequencing receipts, one mechanism instantiated per authority:

RoleAssignment = COSE_Sign1(
protected = {
alg : ES256,
content type : "sovm/ssp-role-v1",
kid : issuer site_id (16 bytes)
},
payload = deterministic CBOR {
subject_site_id : bstr 16,
role : tstr,
generation : unsigned,
previous_record_hash : bstr 32 / null
},
external_aad = mesh_id (16 raw bytes)
)

Rules:

  • The issuer MUST be authorized by the mesh’s trust policy to assign roles. A record from an unauthorized issuer is dropped and surfaced.
  • Records chain per subject: generation is dense for that subject_site_id — exactly one greater than the generation of the record it supersedes, per substrate section 11 — previous_record_hash is the BLAKE3 hash of the exact bytes of the record it supersedes, and the genesis record carries generation 0 with a null predecessor.
  • The current role is the head of the subject’s chain. A verifier holding two records that claim the same predecessor MUST surface the conflict and MUST resolve the subject’s role layer as absent until it is resolved — which is deny-safe, since resolution fails closed on a missing layer. Because generations are dense, two records claiming one predecessor necessarily share a generation, so this is exactly the fork section 11 already requires be surfaced; the instantiation adds no same-predecessor index of its own.
  • Records travel in the sync exchange’s records member or through the mailbox.

The chain is why two nodes cannot silently disagree: either they hold the same head, or one holds a strict prefix of the other and converges on delivery, or the chain has forked — and a fork is detectable evidence of a misbehaving issuer, not a quiet divergence.

Where historical access is deployed, it is granted per domain to a named subject. That grant is a (subject_site_id, opaque_domain_id) fact and it must travel: SOP/1 §22.10 step 3 requires a member holding retained application KEK history to redistribute it after a genesis reset, and the member that survives to do so is not in general the node that issued the grant. A grant recorded only in the never-replicated §15.2 policy store is therefore unavailable exactly when it is needed.

A custodian designation is a chained COSE_Sign1 record, the same mechanism §15.3 instantiates, differing only in its authority:

CustodianDesignation = COSE_Sign1(
protected = {
alg : ES256,
content type : "sovm/ssp-custodian-v1",
kid : issuer site_id (16 bytes)
},
payload = deterministic CBOR {
subject_site_id : bstr 16,
opaque_domain_id : bstr,
entitlement_class : tstr,
generation : unsigned,
previous_record_hash : bstr 32 / null
},
external_aad = mesh_id (16 raw bytes)
)

Rules:

  • Records chain per (subject_site_id, opaque_domain_id), not per subject. A subject may be entitled in one domain and not in another, which a per-subject chain resolving to one head cannot express.
  • Issuer authorization is evaluated for the domain designated in. An issuer authorized to designate custodians in one domain is not thereby authorized in another. This narrows §15.3’s authorization question, which is asked of an issuer alone, to one asked of an (issuer, domain) pair. It is a narrowing in the sense chained records §12 permits — stated here because that section requires a narrowing to be stated on the instantiation’s page — and it weakens no verification rule.
  • Every other rule of §15.3 applies unchanged: generation dense per (subject, domain) pair — exactly one greater than the generation of the record it supersedes, as chained records §11 requires — with a genesis at 0 and a null predecessor, previous_record_hash over the exact bytes of the record superseded, the head as the current value, and a pair in conflict resolved as absent rather than as a branch.
  • entitlement_class is a tstr rather than an enumeration, so BOUNDED_HISTORY can be carried later without this record’s schema having to anticipate its parameters.
  • The record names who is entitled. It carries no key material, so it is a syncable record and not a §15.1 never-syncable one; it travels in the sync exchange’s records member or through the mailbox, the same path role assignments and transport bindings take.

The custodian count of a domain is the number of distinct subject_site_id whose (subject, domain) chain resolves to a head carrying FULL_HISTORY and whose subject is a current member. It is node-local derived state: reconstructible from the replicated records, computable offline by every member, and MUST NOT itself replicate.

A pair whose chain is in conflict contributes zero custodians, never one. The direction is deliberate. Ambiguity drives the count down, so a domain that may have one custodian is reported as having one rather than two, and the recovery risk a count of 1 exists to surface shows instead of hiding behind an unresolvable branch.

A mesh may require tier-3 modules of its members. Where SOP/1 is deployed, the capability declaration carries that requirement in MLS group state. A mesh running SSP/1 without an MLS control group has no group state to put it in, and a requirement nobody can carry is a requirement nobody enforces. This section is the second carrier, and it is the one that does not assume a group.

A capability requirement is a chained COSE_Sign1 record, the mechanism §15.3 instantiates, differing in its authority: it chains per mesh, not per subject. There is one requirement chain, as there is one membership-authorization chain in SOP/1 §5.7.

CapabilityRequirement = COSE_Sign1(
protected = {
alg : ES256,
content type : "sovm/ssp-capability-v1",
kid : issuer site_id (16 bytes)
},
payload = deterministic CBOR {
required_modules : [* tstr],
generation : unsigned,
previous_record_hash : bstr 32 / null
},
external_aad = mesh_id (16 raw bytes)
)

Rules:

  • required_modules carries conformance-module identifiers — the SOVM-* strings the profiles index defines, the same ones §55.1 item 8 obliges an implementation to state. They are not MLS extension type values. One namespace, two encodings: a mesh with an MLS control group lists the extension type in its declaration, a mesh carrying this chain lists the identifier here. An identifier the verifier does not recognize is not silently ignored — it is a requirement it cannot establish it meets, so it fails closed.
  • Only tier-3 modules may appear. A tier-1 module is required of everybody already, and listing a tier-2 capability here is the prohibition §55.3 states, not a stricter mesh.
  • The issuer MUST be authorized by the mesh’s trust policy to set the mesh’s requirement, exactly as §15.3 requires of a role issuer. A record from an unauthorized issuer is dropped and surfaced.
  • Records chain per mesh: generation is dense, previous_record_hash is the BLAKE3 hash of the exact bytes of the record superseded, and the genesis record carries generation 0 with a null predecessor — substrate §11 unchanged and unnarrowed.
  • Records travel in the sync exchange’s records member or through the mailbox. The payload is module identifiers only; no personal data may enter it, and the chain’s history is retained as evidence and is therefore not erasable.

A forked requirement chain resolves to the union of the branches. §15.3 resolves a forked role chain as absent, which is deny-safe there because resolution fails closed on a missing layer. Absent is not deny-safe here: it would read as “require nothing”, and a fork would lower the mesh’s floor to the ground. A verifier holding two records that claim the same predecessor MUST surface the fork, exactly as substrate §11 requires, and MUST resolve required_modules as the union of every branch — the strictest reading — until the fork is resolved.

Raising the floor is a deliberate eviction, and the honest rule says so. §6.2.1 requires a proposer tightening an MLS requirement to verify every current member first, on the principle that a mesh MUST NOT be able to strand its own membership. A chain mesh has no atomic view of its membership, so that rule cannot be met here and MUST NOT be stated as if it could. Instead: minting a record that adds a module evicts every member that does not cover it, the eviction takes effect at each peer’s next record delivery rather than atomically, and an implementation MUST surface the set of members it computes as evicted to the operator before minting. Loud, not prevented.

Three places consult the head:

  1. Join acceptance — a joiner whose module-set head does not cover the requirement head is refused, not admitted degraded.
  2. Sender-side policy — a peer that does not cover the requirement is denied, which is where revocation step 2 already puts “stop offering it data”.
  3. The apply chokepoint, per event. Every event already resolves its signer through the keyring; the requirement head and that signer’s module-set head are two further lookups on a path already performing one.

The third is what this carrier buys. The requirement is current state, not an admission snapshot: a mesh that raises its floor at generation N+1 stops non-covering members at the next record delivery, with no admission event to wait for and no Commit to coordinate.

A requirement is only checkable against a claim. §55.1 item 8 already obliges every implementation to state its module set, but it states it in documentation, where no verifier can reach it. This record puts the same statement on the wire, signed.

ModuleSet = COSE_Sign1(
protected = {
alg : ES256,
content type : "sovm/ssp-modules-v1",
kid : site_id (16 bytes)
},
payload = deterministic CBOR {
modules : [* tstr],
generation : unsigned
},
external_aad = mesh_id (16 raw bytes)
)

Rules:

  • The authority is the node itself: a module-set record is signed by the node’s root identity key, or by its event-signing subkey, and speaks only for that site_id.
  • The record is a chained record, narrowed to generation-only — the narrowing transport §28 already takes, for the same reason. The payload above is exhaustive and carries no previous_record_hash: a module set is wholly superseded, never appended to, so every record is a genesis record and there is no link for the chain to carry. That is the narrowing chained records §12 permits, stated here on the instantiation’s page as that section requires, and it weakens no verification rule — substrate §11’s density rule is stated relative to a predecessor and this chain has none.
  • generation is monotonic per site_id; a record supersedes all lower generations for that node, and a successor need only exceed its predecessor.
  • modules carries the same identifiers §15.5 uses. A node covers a requirement when every identifier in the requirement head appears in that node’s module-set head.
  • Records travel in the sync exchange’s records member or through the mailbox. The payload is module identifiers and the site_id in kid, both already public mesh facts; no personal data may enter it.

This record is the direct analogue of the MLS LeafNode capabilities list, and §55.2’s rule applies to it unchanged: an implementation MUST NOT claim a module it does not implement. Signing the claim does not make it true. What signing changes is that a node whose implementation has degraded must keep re-signing a statement it no longer honours in order to stay reachable, which makes the non-compliance attributable rather than silent — the open gate on the profiles index is narrowed by that, not closed.

When policy widens, history that was previously withheld must flow, or the peer’s view stays permanently incomplete in a way no live tail repairs.

The sender enqueues a durable job per (peer, schema, table) that re-sends each row’s covering set — the original events that between them determine the row’s current materialized state — ordered by row_key, resuming from a durable resume_after_row_key cursor.

The unit of backfill is a row; the unit of merge is a column. Those differ, and the difference is the whole of this rule. An event carries only the columns it changed, so the latest event for a row does not carry the winning value of every column of that row: a row whose column a was last written at HLC 5 and whose column b was last written at HLC 9, by two separate events, has no single event that reconstructs it. Re-sending only the HLC 9 event backfills b and silently drops a. Nor may the sender synthesize one event carrying both, because that event would have to be signed now, under a new HLC, which section 9.1 forbids.

What has to cross is therefore not the row’s values but the clocks that produce them. Every rule in materialization is decided by comparing an incoming HLC against a stored clock — a column write must beat the column clock and the tombstone clock, existence is the highest-HLC existence event — so two nodes agree on a row’s future only once they hold the same clocks for it, and a clock is carried only by the event that owns it.

A row’s covering set is therefore the events owning the clocks the row carries, deduplicated by event_id:

  • for each column carrying a clock, the event that owns it: the highest-HLC applied upsert carrying that column;
  • the event that currently wins existence;
  • the highest-HLC delete ever applied to the row, whether or not it won existence — the event the tombstone clock records.

For the ordinary row written once by one event, all three resolve to that event and the covering set is a single message. The last two entries are the ones an implementation will be tempted to drop, because omitting them still reproduces the row’s current state; what they carry is its future behaviour. Without the losing delete, the receiver’s tombstone clock sits below the sender’s, so a late write the sender rejects the receiver accepts — a divergence that first appears long after the backfill finished.

The first entry selects on the clock and not on the value, and that distinction is load-bearing rather than pedantic. A column whose current value is null still carries a clock, and dropping it reproduces the same divergence one level down. An upsert may carry an explicit null for a column (13.1), so a null cell is not always a delete’s work; omit its supplier and the receiver finishes the backfill holding no clock for that column while the sender holds one, and a later write below that clock is then rejected by the sender and accepted by the receiver. Where the null was a delete’s work, the delete that wrote it is already the second or third entry and event_id deduplication collapses the two, so selecting on the clock costs an implementation nothing in the case the value test was reaching for.

The covering set is also sufficient: backfill is not obliged to re-send the whole log. Resurrection rebuild consults only upserts that beat the tombstone clock, the tombstone clock is monotonic non-decreasing, and an upsert that lost a column to a higher-HLC upsert cannot win it back under a higher tombstone. No event outside the covering set can ever reach the receiver’s materialized state.

Requirements:

  • A backfilled event is the original COSE_Sign1 message, byte for byte. It is a re-delivery, not a new event: the derived event_id is therefore unchanged, and a receiver that already has it converges to a no-op via idempotent apply. Re-encoding or re-signing is forbidden — it would mint a new identity for an old mutation and attribute it to a new moment.
  • The sender MUST re-check policy at emit time and drop rows a standing narrow has since excluded. Policy can change while a backfill is in flight.
  • Live tail SHOULD be sent before backfill within a cycle, so current activity is not starved by history.
  • Backfill progress MUST commit only after a successful push.
  • A row’s covering set MUST be pushed within one cycle, and resume_after_row_key MUST NOT advance past a row until all of it has been pushed. A cursor that can certify half a row is a cursor that certifies a row missing columns.
  • Implementations SHOULD bound rows per cycle. The bound is a resource decision, not a protocol invariant. It is a bound on rows, not on events: a row’s covering set is indivisible.

A storage-constrained node MAY evict cold history below a per-table HLC horizon. Eviction is strictly local and MUST NOT be confused with deletion.

Requirements:

  • Eviction MUST NOT emit any event. It is a local storage operation, and emitting one would replicate a local capacity decision as a semantic one.
  • The horizon MUST be monotonic non-decreasing.
  • Events below the horizon count as applied for watermark purposes but write nothing. A peer must not be asked to re-send data the receiver deliberately discarded.
  • Producers MUST suppress deletes inferred from the absence of evicted rows. This is the trap eviction sets: a naive change-detector sees a missing row and concludes it was deleted, then replicates that conclusion to peers that still have it.
  • Only categories the implementation declares evictable may be evicted, and the declaration MUST be documented.
  • The declaration MUST state that resurrection on the category may come back partial. A category whose rows must resurrect exactly MUST NOT be declared evictable — that is how a table opts into full retention, per table rather than mesh-wide.

Eviction is local in its mechanism and not in its effect. Two nodes holding the same events under different horizons materialize different rows; convergence states which rows, and what does and does not repair them. Declaring a category evictable makes that trade on the application’s behalf, which is why the declaration is where it has to be visible.

Rehydration zeroes the horizon and relies on backfill to re-deliver the covering set. Eviction is where the per-column shape of that set stops being a corner case: the columns a rehydrating node has lost were last written at different HLCs, so the events that restore them are different events, and a latest-event-per-row backfill would leave the node permanently short of exactly those columns whose last write was not the row’s last write.

A sender cannot re-send what it has itself evicted, so rehydration reaches only as far below the horizon as some peer still retains. Where every node has evicted below the same horizon, that history is gone — a consequence of the mesh’s retention choices rather than a protocol failure, and the reason the evictable declaration above is a retention policy and not a storage detail.

Ambiguous eviction markers MUST fail open — keep the row. This is the one place SSP/1 prefers availability over strictness, and the reason is that the cost of being wrong is asymmetric: failing open wastes local storage, failing closed destroys data.

Two modes, one of which is not available.

An operation: "delete" event with a null payload. It replicates and merges like any other event, participates in causal-length existence, advances the tombstone clock, and is resurrectable by a later higher-HLC upsert.

A soft delete hides a row and clears its column values. It does not erase history: the retained log is what makes resurrection rebuild possible.

Reserved and not available. The "purge" token is registered so that no future extension reuses it, but no purge message format is specified. It is unrelated to the object plane’s physical purge, which is a specified storage operation on PME objects rather than a sync-layer message — the reservation here is what keeps the two apart on the wire.

An implementation MUST fail a purge request closed — an explicit error, never a silent downgrade to soft delete. A caller that asked for permanent erasure and received a tombstone has been given a false assurance, and that is worse than a rejection.

Where permanent erasure is required, the mechanism is destruction of the keys protecting the data at a layer above this one. See limitations.