Normative · Part
Scope, backfill, eviction and deletion
15. The scope mechanism
Section titled “15. The scope mechanism”Not everything a node stores may leave it. SSP/1 bounds replication in two tiers: a hard floor that policy cannot override, and a policy layer above it.
15.1. The floor
Section titled “15.1. The floor”A conforming mesh MUST define a floor: a partition of its tables into three categories, fixed by the implementation rather than by configuration.
| Category | Meaning | Policy may reach it |
|---|---|---|
| Syncable | May replicate, subject to policy | Yes |
| Never syncable | Structurally excluded | No — unrepresentable in policy |
| Node-local derived | Reconstructible locally; replicating it is waste | No |
The floor is a mechanism, not a table of names. SSP/1 does not enumerate which schemas fall where, because that is a property of the application built on it rather than of the protocol. What SSP/1 requires is that the partition exists, that it is identical on every node of the mesh, and that the never-syncable category is structurally unrepresentable in policy rather than merely denied by default.
The distinction matters. A deny-by-default rule can be overridden by a misconfiguration or a widening operation. Structural exclusion cannot, and the categories that belong there are the ones where a single mistake is unrecoverable: credential and key material, transport state, and anything whose replication would leak the mesh’s own trust configuration.
Enforcement is dual-sided. A producer applies the floor in its scan predicate and treats a local violation as an error, because a producer emitting out-of-floor data has a bug that silence would hide. A receiver drops and logs, because raising there lets a hostile peer wedge the apply loop.
15.2. Policy above the floor
Section titled “15.2. Policy above the floor”Each node holds a local policy store that is never replicated, keyed
(peer_site_id, table_schema, table_name, layer) with effect allow or deny.
* is a valid wildcard for peer and for table. It is never valid for schema:
a schema-level wildcard is how an accidental widening becomes a total one.
Layers are ordered system < role < user. Resolution for a given
(peer, schema, table):
- Outside the floor, deny unconditionally.
- The highest layer containing any matching row wins.
- Within that layer, the most specific undominated row wins: an exact peer beats
*, and an exact table beats*. - Two incomparable maximal rows: deny and log. Never guess open.
- No matching row: deny. Fail closed.
Policy is a sender-side decision and never appears on the wire. A receiver therefore cannot widen its own feed by lying about its policy, because it is not the one consulting it. This is why scope is enforced where the data lives rather than where it is wanted.
Narrowing is always permitted. Widening is floor-checked and triggers backfill.
15.3. Role assignment records
Section titled “15.3. Role assignment records”The role layer needs one thing the other layers do not: the assignment itself
must travel between nodes, or two nodes resolve policy against different role
views with no signal that they disagree.
A role assignment is a chained COSE_Sign1 record — the same chained-record shape
the object plane uses for authorization records and sequencing receipts, one
mechanism instantiated per authority:
RoleAssignment = COSE_Sign1( protected = { alg : ES256, content type : "sovm/ssp-role-v1", kid : issuer site_id (16 bytes) }, payload = deterministic CBOR { subject_site_id : bstr 16, role : tstr, generation : unsigned, previous_record_hash : bstr 32 / null }, external_aad = mesh_id (16 raw bytes))Rules:
- The issuer MUST be authorized by the mesh’s trust policy to assign roles. A record from an unauthorized issuer is dropped and surfaced.
- Records chain per subject:
generationis dense for thatsubject_site_id— exactly one greater than the generation of the record it supersedes, per substrate section 11 —previous_record_hashis the BLAKE3 hash of the exact bytes of the record it supersedes, and the genesis record carries generation0with a null predecessor. - The current role is the head of the subject’s chain. A verifier holding two
records that claim the same predecessor MUST surface the conflict and MUST
resolve the subject’s
rolelayer as absent until it is resolved — which is deny-safe, since resolution fails closed on a missing layer. Because generations are dense, two records claiming one predecessor necessarily share a generation, so this is exactly the fork section 11 already requires be surfaced; the instantiation adds no same-predecessor index of its own. - Records travel in the sync exchange’s
recordsmember or through the mailbox.
The chain is why two nodes cannot silently disagree: either they hold the same head, or one holds a strict prefix of the other and converges on delivery, or the chain has forked — and a fork is detectable evidence of a misbehaving issuer, not a quiet divergence.
15.4. Custodian designation records
Section titled “15.4. Custodian designation records”Where historical access is deployed, it is
granted per domain to a named subject. That grant is a
(subject_site_id, opaque_domain_id) fact and it must travel:
SOP/1 §22.10 step 3
requires a member holding retained application KEK history to redistribute it
after a genesis reset, and the member that survives to do so is not in general
the node that issued the grant. A grant recorded only in the never-replicated
§15.2 policy store is therefore unavailable
exactly when it is needed.
A custodian designation is a chained COSE_Sign1 record, the same mechanism
§15.3 instantiates, differing only in its
authority:
CustodianDesignation = COSE_Sign1( protected = { alg : ES256, content type : "sovm/ssp-custodian-v1", kid : issuer site_id (16 bytes) }, payload = deterministic CBOR { subject_site_id : bstr 16, opaque_domain_id : bstr, entitlement_class : tstr, generation : unsigned, previous_record_hash : bstr 32 / null }, external_aad = mesh_id (16 raw bytes))Rules:
- Records chain per
(subject_site_id, opaque_domain_id), not per subject. A subject may be entitled in one domain and not in another, which a per-subject chain resolving to one head cannot express. - Issuer authorization is evaluated for the domain designated in. An issuer
authorized to designate custodians in one domain is not thereby authorized in
another. This narrows §15.3’s authorization
question, which is asked of an issuer alone, to one asked of an
(issuer, domain)pair. It is a narrowing in the sense chained records §12 permits — stated here because that section requires a narrowing to be stated on the instantiation’s page — and it weakens no verification rule. - Every other rule of §15.3 applies unchanged:
generationdense per(subject, domain)pair — exactly one greater than the generation of the record it supersedes, as chained records §11 requires — with a genesis at0and a null predecessor,previous_record_hashover the exact bytes of the record superseded, the head as the current value, and a pair in conflict resolved as absent rather than as a branch. entitlement_classis atstrrather than an enumeration, soBOUNDED_HISTORYcan be carried later without this record’s schema having to anticipate its parameters.- The record names who is entitled. It carries no key material, so it is a
syncable record and not a §15.1 never-syncable one; it
travels in the sync exchange’s
recordsmember or through the mailbox, the same path role assignments and transport bindings take.
The custodian count of a domain is the number of distinct subject_site_id
whose (subject, domain) chain resolves to a head carrying
FULL_HISTORY and whose subject is a current
member. It is node-local derived state: reconstructible from
the replicated records, computable offline by every member, and MUST NOT itself
replicate.
A pair whose chain is in conflict contributes zero custodians, never one.
The direction is deliberate. Ambiguity drives the count down, so a domain that
may have one custodian is reported as having one rather than two, and the
recovery risk a count of 1 exists to surface shows instead of hiding behind an
unresolvable branch.
15.5. Capability requirement records
Section titled “15.5. Capability requirement records”A mesh may require tier-3 modules of its members. Where SOP/1 is deployed, the capability declaration carries that requirement in MLS group state. A mesh running SSP/1 without an MLS control group has no group state to put it in, and a requirement nobody can carry is a requirement nobody enforces. This section is the second carrier, and it is the one that does not assume a group.
A capability requirement is a chained COSE_Sign1 record, the mechanism
§15.3 instantiates, differing in its authority: it
chains per mesh, not per subject. There is one requirement chain, as there is
one membership-authorization chain in
SOP/1 §5.7.
CapabilityRequirement = COSE_Sign1( protected = { alg : ES256, content type : "sovm/ssp-capability-v1", kid : issuer site_id (16 bytes) }, payload = deterministic CBOR { required_modules : [* tstr], generation : unsigned, previous_record_hash : bstr 32 / null }, external_aad = mesh_id (16 raw bytes))Rules:
required_modulescarries conformance-module identifiers — theSOVM-*strings the profiles index defines, the same ones §55.1 item 8 obliges an implementation to state. They are not MLS extension type values. One namespace, two encodings: a mesh with an MLS control group lists the extension type in its declaration, a mesh carrying this chain lists the identifier here. An identifier the verifier does not recognize is not silently ignored — it is a requirement it cannot establish it meets, so it fails closed.- Only tier-3 modules may appear. A tier-1 module is required of everybody already, and listing a tier-2 capability here is the prohibition §55.3 states, not a stricter mesh.
- The issuer MUST be authorized by the mesh’s trust policy to set the mesh’s requirement, exactly as §15.3 requires of a role issuer. A record from an unauthorized issuer is dropped and surfaced.
- Records chain per mesh:
generationis dense,previous_record_hashis the BLAKE3 hash of the exact bytes of the record superseded, and the genesis record carries generation0with a null predecessor — substrate §11 unchanged and unnarrowed. - Records travel in the sync exchange’s
recordsmember or through the mailbox. The payload is module identifiers only; no personal data may enter it, and the chain’s history is retained as evidence and is therefore not erasable.
A forked requirement chain resolves to the union of the branches.
§15.3 resolves a forked role chain as absent,
which is deny-safe there because
resolution fails closed on a missing layer. Absent
is not deny-safe here: it would read as “require nothing”, and a fork would
lower the mesh’s floor to the ground. A verifier holding two records that claim the
same predecessor MUST surface the fork, exactly as
substrate §11 requires, and MUST resolve
required_modules as the union of every branch — the strictest reading — until the
fork is resolved.
Raising the floor is a deliberate eviction, and the honest rule says so. §6.2.1 requires a proposer tightening an MLS requirement to verify every current member first, on the principle that a mesh MUST NOT be able to strand its own membership. A chain mesh has no atomic view of its membership, so that rule cannot be met here and MUST NOT be stated as if it could. Instead: minting a record that adds a module evicts every member that does not cover it, the eviction takes effect at each peer’s next record delivery rather than atomically, and an implementation MUST surface the set of members it computes as evicted to the operator before minting. Loud, not prevented.
Three places consult the head:
- Join acceptance — a joiner whose module-set head does not cover the requirement head is refused, not admitted degraded.
- Sender-side policy — a peer that does not cover the requirement is denied, which is where revocation step 2 already puts “stop offering it data”.
- The apply chokepoint, per event. Every event already resolves its signer through the keyring; the requirement head and that signer’s module-set head are two further lookups on a path already performing one.
The third is what this carrier buys. The requirement is current state, not an admission snapshot: a mesh that raises its floor at generation N+1 stops non-covering members at the next record delivery, with no admission event to wait for and no Commit to coordinate.
15.6. Per-node module-set records
Section titled “15.6. Per-node module-set records”A requirement is only checkable against a claim. §55.1 item 8 already obliges every implementation to state its module set, but it states it in documentation, where no verifier can reach it. This record puts the same statement on the wire, signed.
ModuleSet = COSE_Sign1( protected = { alg : ES256, content type : "sovm/ssp-modules-v1", kid : site_id (16 bytes) }, payload = deterministic CBOR { modules : [* tstr], generation : unsigned }, external_aad = mesh_id (16 raw bytes))Rules:
- The authority is the node itself: a module-set record is signed by the node’s
root identity key, or by its event-signing
subkey, and speaks only for that
site_id. - The record is a chained record,
narrowed to generation-only — the narrowing transport
§28 already takes, for the same
reason. The payload above is exhaustive and carries no
previous_record_hash: a module set is wholly superseded, never appended to, so every record is a genesis record and there is no link for the chain to carry. That is the narrowing chained records §12 permits, stated here on the instantiation’s page as that section requires, and it weakens no verification rule — substrate §11’s density rule is stated relative to a predecessor and this chain has none. generationis monotonic persite_id; a record supersedes all lower generations for that node, and a successor need only exceed its predecessor.modulescarries the same identifiers §15.5 uses. A node covers a requirement when every identifier in the requirement head appears in that node’s module-set head.- Records travel in the sync exchange’s
recordsmember or through the mailbox. The payload is module identifiers and thesite_idinkid, both already public mesh facts; no personal data may enter it.
This record is the direct analogue of the MLS LeafNode capabilities list, and
§55.2’s rule applies to it unchanged: an
implementation MUST NOT claim a module it does not implement. Signing the claim
does not make it true. What signing changes is that a node whose implementation has
degraded must keep re-signing a statement it no longer honours in order to stay
reachable, which makes the non-compliance attributable rather than silent — the
open gate on the profiles index is narrowed by that, not
closed.
16. Widening backfill
Section titled “16. Widening backfill”When policy widens, history that was previously withheld must flow, or the peer’s view stays permanently incomplete in a way no live tail repairs.
The sender enqueues a durable job per (peer, schema, table) that re-sends each row’s
covering set — the original events that between them determine the row’s current
materialized state — ordered by row_key, resuming from a durable
resume_after_row_key cursor.
The unit of backfill is a row; the unit of merge is a
column.
Those differ, and the difference is the whole of this rule. An event carries only the
columns it changed, so the latest event for a row does not carry the winning value of
every column of that row: a row whose column a was last written at HLC 5 and whose
column b was last written at HLC 9, by two separate events, has no single event that
reconstructs it. Re-sending only the HLC 9 event backfills b and silently drops a.
Nor may the sender synthesize one event carrying both, because that event would have
to be signed now, under a new HLC, which
section 9.1
forbids.
What has to cross is therefore not the row’s values but the clocks that produce them. Every rule in materialization is decided by comparing an incoming HLC against a stored clock — a column write must beat the column clock and the tombstone clock, existence is the highest-HLC existence event — so two nodes agree on a row’s future only once they hold the same clocks for it, and a clock is carried only by the event that owns it.
A row’s covering set is therefore the events owning the clocks the row carries,
deduplicated by event_id:
- for each column carrying a clock, the event that owns it: the highest-HLC applied upsert carrying that column;
- the event that currently wins existence;
- the highest-HLC
deleteever applied to the row, whether or not it won existence — the event the tombstone clock records.
For the ordinary row written once by one event, all three resolve to that event and the covering set is a single message. The last two entries are the ones an implementation will be tempted to drop, because omitting them still reproduces the row’s current state; what they carry is its future behaviour. Without the losing delete, the receiver’s tombstone clock sits below the sender’s, so a late write the sender rejects the receiver accepts — a divergence that first appears long after the backfill finished.
The first entry selects on the clock and not on the value, and that distinction is
load-bearing rather than pedantic. A column whose current value is null still
carries a clock, and dropping it reproduces the same divergence one level down. An
upsert may carry an explicit null for a column
(13.1),
so a null cell is not always a delete’s work; omit its supplier and the receiver
finishes the backfill holding no clock for that column while the sender holds one, and
a later write below that clock is then rejected by the sender and accepted by the
receiver. Where the null was a delete’s work, the delete that wrote it is already
the second or third entry and event_id deduplication collapses the two, so selecting
on the clock costs an implementation nothing in the case the value test was reaching
for.
The covering set is also sufficient: backfill is not obliged to re-send the whole log. Resurrection rebuild consults only upserts that beat the tombstone clock, the tombstone clock is monotonic non-decreasing, and an upsert that lost a column to a higher-HLC upsert cannot win it back under a higher tombstone. No event outside the covering set can ever reach the receiver’s materialized state.
Requirements:
- A backfilled event is the original
COSE_Sign1message, byte for byte. It is a re-delivery, not a new event: the derivedevent_idis therefore unchanged, and a receiver that already has it converges to a no-op via idempotent apply. Re-encoding or re-signing is forbidden — it would mint a new identity for an old mutation and attribute it to a new moment. - The sender MUST re-check policy at emit time and drop rows a standing narrow has since excluded. Policy can change while a backfill is in flight.
- Live tail SHOULD be sent before backfill within a cycle, so current activity is not starved by history.
- Backfill progress MUST commit only after a successful push.
- A row’s covering set MUST be pushed within one cycle, and
resume_after_row_keyMUST NOT advance past a row until all of it has been pushed. A cursor that can certify half a row is a cursor that certifies a row missing columns. - Implementations SHOULD bound rows per cycle. The bound is a resource decision, not a protocol invariant. It is a bound on rows, not on events: a row’s covering set is indivisible.
17. Eviction is not deletion
Section titled “17. Eviction is not deletion”A storage-constrained node MAY evict cold history below a per-table HLC horizon. Eviction is strictly local and MUST NOT be confused with deletion.
Requirements:
- Eviction MUST NOT emit any event. It is a local storage operation, and emitting one would replicate a local capacity decision as a semantic one.
- The horizon MUST be monotonic non-decreasing.
- Events below the horizon count as applied for watermark purposes but write nothing. A peer must not be asked to re-send data the receiver deliberately discarded.
- Producers MUST suppress deletes inferred from the absence of evicted rows. This is the trap eviction sets: a naive change-detector sees a missing row and concludes it was deleted, then replicates that conclusion to peers that still have it.
- Only categories the implementation declares evictable may be evicted, and the declaration MUST be documented.
- The declaration MUST state that resurrection on the category may come back partial. A category whose rows must resurrect exactly MUST NOT be declared evictable — that is how a table opts into full retention, per table rather than mesh-wide.
Eviction is local in its mechanism and not in its effect. Two nodes holding the same events under different horizons materialize different rows; convergence states which rows, and what does and does not repair them. Declaring a category evictable makes that trade on the application’s behalf, which is why the declaration is where it has to be visible.
Rehydration zeroes the horizon and relies on backfill to re-deliver the covering set. Eviction is where the per-column shape of that set stops being a corner case: the columns a rehydrating node has lost were last written at different HLCs, so the events that restore them are different events, and a latest-event-per-row backfill would leave the node permanently short of exactly those columns whose last write was not the row’s last write.
A sender cannot re-send what it has itself evicted, so rehydration reaches only as far below the horizon as some peer still retains. Where every node has evicted below the same horizon, that history is gone — a consequence of the mesh’s retention choices rather than a protocol failure, and the reason the evictable declaration above is a retention policy and not a storage detail.
Ambiguous eviction markers MUST fail open — keep the row. This is the one place SSP/1 prefers availability over strictness, and the reason is that the cost of being wrong is asymmetric: failing open wastes local storage, failing closed destroys data.
18. Deletion
Section titled “18. Deletion”Two modes, one of which is not available.
18.1. Soft delete
Section titled “18.1. Soft delete”An operation: "delete" event with a null payload. It replicates and merges like any
other event, participates in
causal-length existence, advances the
tombstone clock, and is resurrectable by a
later higher-HLC upsert.
A soft delete hides a row and clears its column values. It does not erase history: the retained log is what makes resurrection rebuild possible.
18.2. Purge
Section titled “18.2. Purge”Reserved and not available. The "purge" token is registered so that no future
extension reuses it, but no purge message format is specified. It is unrelated
to the object plane’s physical purge,
which is a specified storage operation on PME objects rather than a sync-layer
message — the reservation here is what keeps the two apart on the wire.
An implementation MUST fail a purge request closed — an explicit error, never a silent downgrade to soft delete. A caller that asked for permanent erasure and received a tombstone has been given a false assurance, and that is worse than a rejection.
Where permanent erasure is required, the mechanism is destruction of the keys protecting the data at a layer above this one. See limitations.