Normative · Part
Consistency, sequencing and coordination
22. Consistency, sequencing, and coordination
Section titled “22. Consistency, sequencing, and coordination”22.1. Data plane
Section titled “22.1. Data plane”APPEND data converges by set union.
APPEND_DELETE data converges from the set of data objects plus delete objects.
Neither requires a leader.
22.2. Mutable state
Section titled “22.2. Mutable state”MUTABLE state converges according to SSP/1.
22.3. Membership plane
Section titled “22.3. Membership plane”MLS groups maintain ordered epochs.
That linearity is accepted because group membership is security control state, not ordinary user-data mutation.
22.4. Membership changes are rare
Section titled “22.4. Membership changes are rare”The architecture MUST NOT route ordinary append operations through MLS group Commits.
Expected MLS Commit triggers include:
- node add;
- node remove;
- domain-scope membership change;
- security-motivated group update;
- explicit group reinitialization.
22.5. Blind Commit sequencer
Section titled “22.5. Blind Commit sequencer”RFC 9420 requires applications to have an established mechanism for resolving conflicting Commits for the same epoch. SOVM prevents ordinary honest-client forks by serializing Commits through a blind per-group sequencer at the SSP Delivery Service.
MLS PrivateMessage framing exposes group_id, epoch, and content_type
without exposing the encrypted Commit body. SOVM MLS group identifiers MUST
therefore be random opaque values suitable for routing.
For each opaque group_id, the sequencer maintains the currently accepted MLS
epoch and applies an atomic first-Commit-wins rule:
accept commit iff: content_type == commit AND commit.epoch == sequencer.current_epoch(group_id) AND no commit has already been accepted for that epoch
on acceptance: durably store/forward opaque MLS bytes advance sequencer.current_epoch(group_id) by 1A second Commit based on the same epoch is rejected as stale and MUST be rebuilt by its sender after processing the canonical Commit.
The sequencer reads only standard MLS framing fields that are already unencrypted; it MUST NOT decrypt the Commit body or decide whether the membership change is semantically authorized.
Clients MUST still perform full MLS validation and SSP-policy validation. An authorized malicious or buggy member can submit a syntactically routed but semantically invalid Commit and cause an availability failure; the sequencer does not convert an untrusted Delivery Service into a cryptographic validator. Recovery from that exceptional case remains part of G4.
The CAS rule prevents honest races; it does not by itself constrain a malicious sequencer, and the threat model (Section 23.1) includes hostile relay infrastructure. A malicious sequencer cannot forge a valid Commit, but it can withhold messages, misreport its epoch state, or equivocate by presenting different valid competing Commits as the winner to different members. Therefore each accepted Commit MUST be acknowledged with a signed hash-chained sequencing receipt:
receipt_n = COSE_Sign1_sequencer( { group_id, epoch, commit_hash, previous_receipt_hash })Clients retain receipts for their groups and gossip current receipt heads through the control group. Two valid receipts for the same (group_id, epoch) with different commit_hash values, or divergent chains, are proof of sequencer equivocation.
Detected equivocation is handled as sequencer compromise: the fleet migrates the group to a new sequencer (Section 22.6), audits group state before proceeding, and surfaces the event to the user. Receipts do not remedy withholding, which remains an availability attack handled by Commit queueing and sequencer migration.
Acceptance imposes requirements on the receipt itself, not only on the opaque
MLS bytes and the epoch. Receipt signing under ES256 is randomised wherever the
receipt-signing key is held in a protected keystore: the substrate profile
requires deterministic RFC 6979 nonces only of software signers
(signature profile section 6),
and a protected keystore generates its nonce internally. Two signings of one
receipt payload therefore produce two byte-distinct receipts with two distinct
previous_receipt_hash successors.
This section reads divergent receipt chains as proof of sequencer equivocation, and Section 22.6 step 4 fails that closed into audit, migration and — where no evidence survives — the genesis reset of Section 22.10. A sequencer that re-signs after a crash therefore manufactures cryptographic proof of its own equivocation, and directs the fleet’s most destructive remedy at a group that was never attacked.
The crash is not required to be accidental. An attacker who can induce a restart in that window — by ordinary resource pressure, without any access to key material and without compromising the sequencer — can drive that outcome deliberately. This is an externally triggerable attack on the fleet’s most destructive remedy, not a reliability hazard. Therefore:
- A sequencer MUST durably persist the exact encoded
COSE_Sign1receipt bytes before those bytes are released to any member or written to any forwarding path, and MUST thereafter re-transmit exactly those bytes rather than re-encoding or re-signing. - A sequencer MUST NOT sign a second receipt for a
(group_id, epoch)for which a receipt has already been accepted. Where it cannot establish that a receipt for that pair was never released, it MUST treat the pair as already signed. - The epoch advance and the receipt persistence MUST be a single atomic
commitment. No implementation state may be observable in which
current_epoch(group_id)reflects an accepted Commit whose receipt is not durable — including as observed by the sequencer itself on resume after a restart, and not only as observed by a peer.
Clause 3 is stated rather than left to the implementer because the ordering between the two writes is otherwise unconstrained, and an implementation that performs them independently has a window it did not choose and cannot see.
Persist-before-release is the only available fix here, not the preferred one.
The synchronization layer meets the same defect one layer up and resolves it
differently: a producer that cannot establish whether a signed event was
released ticks its hybrid logical clock and signs a new event rather than
re-signing the old one (SSP/1 data model). That resolution
is unavailable here, and the reason is structural rather than incidental. It
works because a producer owns its own clock. A sequencer does not. A
receipt’s identity is (group_id, epoch), and the epoch is dictated by the CAS
rule from the Commit being acknowledged, not chosen by the signer. There is no
counter to tick and no second identity to sign under. RFC 6979 is likewise
unavailable: it is a rule for software signers and cannot be imposed on a
protected keystore, so requiring it would require the receipt-signing key to be
in software. A protocol whose correctness improved as that key moved out of
hardware would have exactly the wrong gradient.
The sequencer’s receipt-signing identity is registered with the fleet and is
rotatable; its lifecycle is part of the SOVM control-plane profile. Its public
key is registered and compared only in the 33-byte compressed encoding the
substrate signature profile
requires, because the check that makes equivocation actionable is a public-key
comparison over exactly that key.
Sequencing receipts and membership authorization records (Section 5.7) share
one chained-record schema in the SOVM COSE profile: one mechanism,
instantiated per authority. For receipts that authority is the group — the
chain is keyed by group_id and its generation is the MLS epoch — and the
sequencer bound to the group signs its links. Two valid receipts at one
(group_id, epoch) are therefore a fork of one chain under
/substrate/chained-records/ section 11
whichever sequencer identity signed them, which is what makes step 4 of
Section 22.6’s
recovery procedure checkable.
22.6. Sequencer state, bootstrap, singularity, and recovery
Section titled “22.6. Sequencer state, bootstrap, singularity, and recovery”Sequencer correctness depends on durable per-group epoch state and on singularity. Four cases are normative.
Bootstrap: the first Commit-bearing message observed for an unknown opaque
group_id initializes the sequencer’s epoch tracking for that group. A group
creator SHOULD explicitly register a new group_id with its initial epoch
before inviting members, so the initialization point is unambiguous.
Singularity: an MLS group MUST be bound to exactly one sequencer at a time. Two sequencers with independent state cannot jointly enforce first-Commit-wins. Migrating a group between sequencers is an explicit control-plane operation that transfers the current accepted epoch value and the head of the receipt chain (Section 22.5), so the new sequencer’s first receipt chains to the old sequencer’s last; the old and new sequencer MUST NOT operate concurrently for the same group.
Because the receipt chain’s authority is the group and not the signing sequencer (Section 22.5), the hand-off is an ordinary link in one chain rather than the start of a second one, and the identity permitted to sign a link MUST itself be resolvable at that link’s generation:
- The control-plane binding of a sequencer to a group MUST record the generation — the MLS epoch — from which the binding takes effect.
- A receipt whose signer was not the bound sequencer at that link’s generation MUST be rejected and retained for attribution, not applied. A rotated-out sequencer that keeps signing is the violation of this section’s concurrency prohibition, and this is where it becomes visible: an attributable rejection rather than a fork that freezes an honest group.
- A verifier that cannot resolve the binding in force at a link’s generation
MUST hold the link rather than apply or reject it, mirroring
/substrate/chained-records/section 11’s posture for an unknown predecessor. This fails closed without manufacturing an equivocation verdict. - Two bindings for one group at one generation are a fork of the binding record and MUST be surfaced there. Contested hand-offs are resolved where hand-offs are recorded.
Whether that binding is a signed record or control-plane state is left to the SOVM control-plane profile, on the same reasoning this section gives for the discontinuity record below: the requirement is satisfiable either way and introduces no registered value, and making it a signed record is a wire change belonging to that profile.
State loss: a sequencer that loses its epoch state MUST NOT resume from a naked member epoch assertion.
Recovery uses the signed receipt history introduced in Section 22.5:
- reachable current members submit their latest retained sequencing receipt chain (or a chain suffix anchored at a previously registered receipt head) together with an authenticated attestation of their current MLS epoch;
- the recovering sequencer verifies receipt signatures, hash-chain continuity, and consistency between the receipt head and attested epoch;
- if all available valid evidence converges on one receipt head/epoch, the sequencer resumes CAS enforcement from that state and chains its next receipt to the recovered head;
- two valid divergent receipt chains, or two receipts for the same
(group_id, epoch)with different Commit hashes, are proof of prior sequencer equivocation and MUST fail closed; the group is audited and migrated to a new sequencer; and - if no usable receipt evidence survives, recovery MUST use the domain genesis-reset procedure in Section 22.10 rather than inventing history.
Receipts are deliberately retained client-side because they are small security-audit records. A storage compaction policy MAY checkpoint old receipt prefixes, but MUST preserve a verifiable chain to the currently registered head.
Epoch advanced, receipt absent: the atomicity requirement of
Section 22.5 bounds how
often a sequencer reaches a state in which current_epoch has advanced past
a Commit whose receipt is not durable. It cannot bound it to zero. Post-commit
durability loss — an fsync that did not reach the medium, a promoted replica
behind the committed write, a restore from backup — reaches that state
through a correctly atomic implementation. A window that cannot be closed
needs a path, not a smaller window.
The three preceding cases all describe a sequencer that lost state and asks the fleet to reconstruct it. This state is the opposite shape: the epoch survived and the receipt did not. Without a stated path, a serving path finds stored MLS bytes, finds no persisted receipt, and re-signs — which Section 22.5 forbids and which step 4 of the recovery procedure then reads as equivocation.
A sequencer MUST:
- Detect it. On resume, and before serving any receipt for a group, verify
that its durable receipt-chain head for that group is consistent with
current_epoch(group_id). An epoch its own state records as accepted and for which no durable receipt exists is the condition this case names. - Never re-sign it, under Section 22.5’s non-repetition rule.
- Refuse to serve it. A request for that epoch’s receipt MUST be answered with a stated absence, and MUST NOT be answered with a receipt.
- Record the gap as a receipt-chain discontinuity — stated, not inferred — and surface it to the fleet with the loudness Section 22.5 gives a detected equivocation. Surfacing it to some members and not others does not satisfy this clause: a discontinuity disclosed selectively is a discontinuity used to partition the fleet’s view of the chain. The discontinuity is part of the sequencer’s registered state and MUST transfer on migration alongside the accepted epoch value and the receipt-chain head.
- Chain forward across the recorded gap, so the chain remains verifiable on both sides of it and the gap remains visible.
A member validating a chain across a recorded discontinuity MUST treat it as a durability gap rather than as equivocation evidence, and MUST do so only where the gap is one the sequencer recorded and the chain is continuous on both sides of it. Anything else remains equivocation.
A member who holds a valid sequencer-signed receipt for an epoch the sequencer has declared a discontinuity MUST NOT accept the discontinuity claim for that epoch, and MUST treat the discrepancy as equivocation evidence.
That second rule is what stops a malicious sequencer turning this case into a way to excuse equivocation. Without it: a sequencer releases receipt A to one member set and receipt B to another for the same epoch, then declares that epoch a discontinuity and chains forward from A. The chain it presents is continuous on both sides of the gap and the gap is one it recorded, so the first rule is satisfied — and the members holding B are left resolving a contradiction the text does not tell them how to resolve. A conforming implementation could discard its own receipt as stale. The sequencer’s claim that no durable receipt exists for that epoch is falsified by a member holding one, and accepting the claim anyway is accepting an assertion over cryptographic proof.
Members who never held a receipt for that epoch cannot check the claim, and must rely on it. That blind spot is not created here: Section 22.5 already needs two divergent chains to detect equivocation, and a member who witnessed neither has neither. This case neither widens it nor closes it.
The distinction step 4 of the recovery procedure has to draw is between two divergent chains, which is equivocation, and one chain with a stated hole, which is durability loss. Without this case it would have no way to draw that distinction; this case is that way. A discontinuity is a worse outcome than an unbroken chain and a far better one than a false verdict of equivocation against an operator who did nothing wrong.
Whether the discontinuity record is itself a signed record carrying a registry row, or sequencer state surfaced through the control plane, is deliberately not decided here. The requirement above is satisfiable either way and introduces no registered value; making it a signed record is a wire change and belongs to the specification that defines the sequencer’s control-plane profile, not to this section.
22.7. Relay unavailability
Section titled “22.7. Relay unavailability”If the blind sequencer is unavailable, membership-changing Commits queue.
Ordinary SOP data creation and existing-authority data synchronization MAY continue.
Clients MUST NOT create divergent offline membership histories and later expect automatic merge.
22.8. No server authority over data
Section titled “22.8. No server authority over data”Sequencing MLS Commits does not make the relay authoritative over trust or application data.
The relay can order or reject stale opaque Commit envelopes but cannot decide:
- whether a node is trusted;
- whether an Add/Remove is authorized by SSP policy;
- which application rows are valid;
- which object state wins.
Clients still perform all MLS and SSP validation.
22.9. Schema
Section titled “22.9. Schema”Schema evolution remains a control-plane problem.
SOP/1 does not define a schema CRDT.
22.10. Group recovery: domain genesis reset
Section titled “22.10. Group recovery: domain genesis reset”MLS protocol state is reformable, not load-bearing for data. Recovery
requires three components — SOP ciphertext, its ObjectManifest/AccessRecord
catalog state, and applicable KEK history (Section 12.7) — none of which is
MLS protocol state: KEKs and object DEKs are application state (Section 7.8),
and manifest/access metadata is durably replicated through catalog checkpoints
named by RecoveryRoot records. Loss of usable MLS group state
therefore never implies loss of data access while one recoverable copy of
each component exists (Appendix B, invariant 22).
The recovery procedure when no surviving node has usable current MLS state for a domain group is the domain genesis reset:
- surviving nodes re-establish mutual trust through the standard SSP pairing ceremony where required;
- a new MLS group is created under a fresh opaque
group_id, registered with the blind sequencer (Section 22.6 bootstrap), starting a fresh receipt chain; - a member holding retained application KEK history redistributes it through the standard COSE key-history package (Section 8.5) inside the new group;
- where local catalogs were lost, the domain catalog is restored from the
latest
RecoveryRootand catalog checkpoint (Section 12.7); - the old
group_idis marked dead in the catalog and its sequencer binding is retired.
The degenerate single-node case requires no MLS at all: the node resolves
its objects locally from its retained key store and catalog, or from a
RecoveryRoot/catalog checkpoint plus retained KEK history.
The genesis reset MUST be exercised as a recurring drill, and full disaster recovery from a blind archive plus a bare key store is gated by G10; recovery procedures that are never rehearsed do not exist.