Skip to content

Normative · Part

Consistency, sequencing and coordination

22. Consistency, sequencing, and coordination

Section titled “22. Consistency, sequencing, and coordination”

APPEND data converges by set union.

APPEND_DELETE data converges from the set of data objects plus delete objects.

Neither requires a leader.

MUTABLE state converges according to SSP/1.

MLS groups maintain ordered epochs.

That linearity is accepted because group membership is security control state, not ordinary user-data mutation.

The architecture MUST NOT route ordinary append operations through MLS group Commits.

Expected MLS Commit triggers include:

  • node add;
  • node remove;
  • domain-scope membership change;
  • security-motivated group update;
  • explicit group reinitialization.

RFC 9420 requires applications to have an established mechanism for resolving conflicting Commits for the same epoch. SOVM prevents ordinary honest-client forks by serializing Commits through a blind per-group sequencer at the SSP Delivery Service.

MLS PrivateMessage framing exposes group_id, epoch, and content_type without exposing the encrypted Commit body. SOVM MLS group identifiers MUST therefore be random opaque values suitable for routing.

For each opaque group_id, the sequencer maintains the currently accepted MLS epoch and applies an atomic first-Commit-wins rule:

accept commit iff:
content_type == commit
AND commit.epoch == sequencer.current_epoch(group_id)
AND no commit has already been accepted for that epoch
on acceptance:
durably store/forward opaque MLS bytes
advance sequencer.current_epoch(group_id) by 1

A second Commit based on the same epoch is rejected as stale and MUST be rebuilt by its sender after processing the canonical Commit.

The sequencer reads only standard MLS framing fields that are already unencrypted; it MUST NOT decrypt the Commit body or decide whether the membership change is semantically authorized.

Clients MUST still perform full MLS validation and SSP-policy validation. An authorized malicious or buggy member can submit a syntactically routed but semantically invalid Commit and cause an availability failure; the sequencer does not convert an untrusted Delivery Service into a cryptographic validator. Recovery from that exceptional case remains part of G4.

The CAS rule prevents honest races; it does not by itself constrain a malicious sequencer, and the threat model (Section 23.1) includes hostile relay infrastructure. A malicious sequencer cannot forge a valid Commit, but it can withhold messages, misreport its epoch state, or equivocate by presenting different valid competing Commits as the winner to different members. Therefore each accepted Commit MUST be acknowledged with a signed hash-chained sequencing receipt:

receipt_n = COSE_Sign1_sequencer(
{
group_id,
epoch,
commit_hash,
previous_receipt_hash
}
)

Clients retain receipts for their groups and gossip current receipt heads through the control group. Two valid receipts for the same (group_id, epoch) with different commit_hash values, or divergent chains, are proof of sequencer equivocation.

Detected equivocation is handled as sequencer compromise: the fleet migrates the group to a new sequencer (Section 22.6), audits group state before proceeding, and surfaces the event to the user. Receipts do not remedy withholding, which remains an availability attack handled by Commit queueing and sequencer migration.

Acceptance imposes requirements on the receipt itself, not only on the opaque MLS bytes and the epoch. Receipt signing under ES256 is randomised wherever the receipt-signing key is held in a protected keystore: the substrate profile requires deterministic RFC 6979 nonces only of software signers (signature profile section 6), and a protected keystore generates its nonce internally. Two signings of one receipt payload therefore produce two byte-distinct receipts with two distinct previous_receipt_hash successors.

This section reads divergent receipt chains as proof of sequencer equivocation, and Section 22.6 step 4 fails that closed into audit, migration and — where no evidence survives — the genesis reset of Section 22.10. A sequencer that re-signs after a crash therefore manufactures cryptographic proof of its own equivocation, and directs the fleet’s most destructive remedy at a group that was never attacked.

The crash is not required to be accidental. An attacker who can induce a restart in that window — by ordinary resource pressure, without any access to key material and without compromising the sequencer — can drive that outcome deliberately. This is an externally triggerable attack on the fleet’s most destructive remedy, not a reliability hazard. Therefore:

  1. A sequencer MUST durably persist the exact encoded COSE_Sign1 receipt bytes before those bytes are released to any member or written to any forwarding path, and MUST thereafter re-transmit exactly those bytes rather than re-encoding or re-signing.
  2. A sequencer MUST NOT sign a second receipt for a (group_id, epoch) for which a receipt has already been accepted. Where it cannot establish that a receipt for that pair was never released, it MUST treat the pair as already signed.
  3. The epoch advance and the receipt persistence MUST be a single atomic commitment. No implementation state may be observable in which current_epoch(group_id) reflects an accepted Commit whose receipt is not durable — including as observed by the sequencer itself on resume after a restart, and not only as observed by a peer.

Clause 3 is stated rather than left to the implementer because the ordering between the two writes is otherwise unconstrained, and an implementation that performs them independently has a window it did not choose and cannot see.

Persist-before-release is the only available fix here, not the preferred one. The synchronization layer meets the same defect one layer up and resolves it differently: a producer that cannot establish whether a signed event was released ticks its hybrid logical clock and signs a new event rather than re-signing the old one (SSP/1 data model). That resolution is unavailable here, and the reason is structural rather than incidental. It works because a producer owns its own clock. A sequencer does not. A receipt’s identity is (group_id, epoch), and the epoch is dictated by the CAS rule from the Commit being acknowledged, not chosen by the signer. There is no counter to tick and no second identity to sign under. RFC 6979 is likewise unavailable: it is a rule for software signers and cannot be imposed on a protected keystore, so requiring it would require the receipt-signing key to be in software. A protocol whose correctness improved as that key moved out of hardware would have exactly the wrong gradient.

The sequencer’s receipt-signing identity is registered with the fleet and is rotatable; its lifecycle is part of the SOVM control-plane profile. Its public key is registered and compared only in the 33-byte compressed encoding the substrate signature profile requires, because the check that makes equivocation actionable is a public-key comparison over exactly that key. Sequencing receipts and membership authorization records (Section 5.7) share one chained-record schema in the SOVM COSE profile: one mechanism, instantiated per authority. For receipts that authority is the group — the chain is keyed by group_id and its generation is the MLS epoch — and the sequencer bound to the group signs its links. Two valid receipts at one (group_id, epoch) are therefore a fork of one chain under /substrate/chained-records/ section 11 whichever sequencer identity signed them, which is what makes step 4 of Section 22.6’s recovery procedure checkable.

22.6. Sequencer state, bootstrap, singularity, and recovery

Section titled “22.6. Sequencer state, bootstrap, singularity, and recovery”

Sequencer correctness depends on durable per-group epoch state and on singularity. Four cases are normative.

Bootstrap: the first Commit-bearing message observed for an unknown opaque group_id initializes the sequencer’s epoch tracking for that group. A group creator SHOULD explicitly register a new group_id with its initial epoch before inviting members, so the initialization point is unambiguous.

Singularity: an MLS group MUST be bound to exactly one sequencer at a time. Two sequencers with independent state cannot jointly enforce first-Commit-wins. Migrating a group between sequencers is an explicit control-plane operation that transfers the current accepted epoch value and the head of the receipt chain (Section 22.5), so the new sequencer’s first receipt chains to the old sequencer’s last; the old and new sequencer MUST NOT operate concurrently for the same group.

Because the receipt chain’s authority is the group and not the signing sequencer (Section 22.5), the hand-off is an ordinary link in one chain rather than the start of a second one, and the identity permitted to sign a link MUST itself be resolvable at that link’s generation:

  • The control-plane binding of a sequencer to a group MUST record the generation — the MLS epoch — from which the binding takes effect.
  • A receipt whose signer was not the bound sequencer at that link’s generation MUST be rejected and retained for attribution, not applied. A rotated-out sequencer that keeps signing is the violation of this section’s concurrency prohibition, and this is where it becomes visible: an attributable rejection rather than a fork that freezes an honest group.
  • A verifier that cannot resolve the binding in force at a link’s generation MUST hold the link rather than apply or reject it, mirroring /substrate/chained-records/ section 11’s posture for an unknown predecessor. This fails closed without manufacturing an equivocation verdict.
  • Two bindings for one group at one generation are a fork of the binding record and MUST be surfaced there. Contested hand-offs are resolved where hand-offs are recorded.

Whether that binding is a signed record or control-plane state is left to the SOVM control-plane profile, on the same reasoning this section gives for the discontinuity record below: the requirement is satisfiable either way and introduces no registered value, and making it a signed record is a wire change belonging to that profile.

State loss: a sequencer that loses its epoch state MUST NOT resume from a naked member epoch assertion.

Recovery uses the signed receipt history introduced in Section 22.5:

  1. reachable current members submit their latest retained sequencing receipt chain (or a chain suffix anchored at a previously registered receipt head) together with an authenticated attestation of their current MLS epoch;
  2. the recovering sequencer verifies receipt signatures, hash-chain continuity, and consistency between the receipt head and attested epoch;
  3. if all available valid evidence converges on one receipt head/epoch, the sequencer resumes CAS enforcement from that state and chains its next receipt to the recovered head;
  4. two valid divergent receipt chains, or two receipts for the same (group_id, epoch) with different Commit hashes, are proof of prior sequencer equivocation and MUST fail closed; the group is audited and migrated to a new sequencer; and
  5. if no usable receipt evidence survives, recovery MUST use the domain genesis-reset procedure in Section 22.10 rather than inventing history.

Receipts are deliberately retained client-side because they are small security-audit records. A storage compaction policy MAY checkpoint old receipt prefixes, but MUST preserve a verifiable chain to the currently registered head.

Epoch advanced, receipt absent: the atomicity requirement of Section 22.5 bounds how often a sequencer reaches a state in which current_epoch has advanced past a Commit whose receipt is not durable. It cannot bound it to zero. Post-commit durability loss — an fsync that did not reach the medium, a promoted replica behind the committed write, a restore from backup — reaches that state through a correctly atomic implementation. A window that cannot be closed needs a path, not a smaller window.

The three preceding cases all describe a sequencer that lost state and asks the fleet to reconstruct it. This state is the opposite shape: the epoch survived and the receipt did not. Without a stated path, a serving path finds stored MLS bytes, finds no persisted receipt, and re-signs — which Section 22.5 forbids and which step 4 of the recovery procedure then reads as equivocation.

A sequencer MUST:

  1. Detect it. On resume, and before serving any receipt for a group, verify that its durable receipt-chain head for that group is consistent with current_epoch(group_id). An epoch its own state records as accepted and for which no durable receipt exists is the condition this case names.
  2. Never re-sign it, under Section 22.5’s non-repetition rule.
  3. Refuse to serve it. A request for that epoch’s receipt MUST be answered with a stated absence, and MUST NOT be answered with a receipt.
  4. Record the gap as a receipt-chain discontinuity — stated, not inferred — and surface it to the fleet with the loudness Section 22.5 gives a detected equivocation. Surfacing it to some members and not others does not satisfy this clause: a discontinuity disclosed selectively is a discontinuity used to partition the fleet’s view of the chain. The discontinuity is part of the sequencer’s registered state and MUST transfer on migration alongside the accepted epoch value and the receipt-chain head.
  5. Chain forward across the recorded gap, so the chain remains verifiable on both sides of it and the gap remains visible.

A member validating a chain across a recorded discontinuity MUST treat it as a durability gap rather than as equivocation evidence, and MUST do so only where the gap is one the sequencer recorded and the chain is continuous on both sides of it. Anything else remains equivocation.

A member who holds a valid sequencer-signed receipt for an epoch the sequencer has declared a discontinuity MUST NOT accept the discontinuity claim for that epoch, and MUST treat the discrepancy as equivocation evidence.

That second rule is what stops a malicious sequencer turning this case into a way to excuse equivocation. Without it: a sequencer releases receipt A to one member set and receipt B to another for the same epoch, then declares that epoch a discontinuity and chains forward from A. The chain it presents is continuous on both sides of the gap and the gap is one it recorded, so the first rule is satisfied — and the members holding B are left resolving a contradiction the text does not tell them how to resolve. A conforming implementation could discard its own receipt as stale. The sequencer’s claim that no durable receipt exists for that epoch is falsified by a member holding one, and accepting the claim anyway is accepting an assertion over cryptographic proof.

Members who never held a receipt for that epoch cannot check the claim, and must rely on it. That blind spot is not created here: Section 22.5 already needs two divergent chains to detect equivocation, and a member who witnessed neither has neither. This case neither widens it nor closes it.

The distinction step 4 of the recovery procedure has to draw is between two divergent chains, which is equivocation, and one chain with a stated hole, which is durability loss. Without this case it would have no way to draw that distinction; this case is that way. A discontinuity is a worse outcome than an unbroken chain and a far better one than a false verdict of equivocation against an operator who did nothing wrong.

Whether the discontinuity record is itself a signed record carrying a registry row, or sequencer state surfaced through the control plane, is deliberately not decided here. The requirement above is satisfiable either way and introduces no registered value; making it a signed record is a wire change and belongs to the specification that defines the sequencer’s control-plane profile, not to this section.

If the blind sequencer is unavailable, membership-changing Commits queue.

Ordinary SOP data creation and existing-authority data synchronization MAY continue.

Clients MUST NOT create divergent offline membership histories and later expect automatic merge.

Sequencing MLS Commits does not make the relay authoritative over trust or application data.

The relay can order or reject stale opaque Commit envelopes but cannot decide:

  • whether a node is trusted;
  • whether an Add/Remove is authorized by SSP policy;
  • which application rows are valid;
  • which object state wins.

Clients still perform all MLS and SSP validation.

Schema evolution remains a control-plane problem.

SOP/1 does not define a schema CRDT.

22.10. Group recovery: domain genesis reset

Section titled “22.10. Group recovery: domain genesis reset”

MLS protocol state is reformable, not load-bearing for data. Recovery requires three components — SOP ciphertext, its ObjectManifest/AccessRecord catalog state, and applicable KEK history (Section 12.7) — none of which is MLS protocol state: KEKs and object DEKs are application state (Section 7.8), and manifest/access metadata is durably replicated through catalog checkpoints named by RecoveryRoot records. Loss of usable MLS group state therefore never implies loss of data access while one recoverable copy of each component exists (Appendix B, invariant 22).

The recovery procedure when no surviving node has usable current MLS state for a domain group is the domain genesis reset:

  1. surviving nodes re-establish mutual trust through the standard SSP pairing ceremony where required;
  2. a new MLS group is created under a fresh opaque group_id, registered with the blind sequencer (Section 22.6 bootstrap), starting a fresh receipt chain;
  3. a member holding retained application KEK history redistributes it through the standard COSE key-history package (Section 8.5) inside the new group;
  4. where local catalogs were lost, the domain catalog is restored from the latest RecoveryRoot and catalog checkpoint (Section 12.7);
  5. the old group_id is marked dead in the catalog and its sequencer binding is retired.

The degenerate single-node case requires no MLS at all: the node resolves its objects locally from its retained key store and catalog, or from a RecoveryRoot/catalog checkpoint plus retained KEK history.

The genesis reset MUST be exercised as a recurring drill, and full disaster recovery from a blind archive plus a bare key store is gated by G10; recovery procedures that are never rehearsed do not exist.