Normative · Part
Protocol semantics
This page defines how a receiver turns a set of delivered events into materialized state, and why the result does not depend on the order they arrived in.
12. The apply chokepoint
Section titled “12. The apply chokepoint”Every transport MUST deliver received events through one apply path performing these checks in this order:
- Signature verification. COSE_Sign1 verification per
authentication: resolve the protected
kidin the trusted-peer keyring and verify withexternal_aadset to the localmesh_id. An event that fails — unknownkid, invalid signature, or an HLC above the signer’s retirement cutoff — is dropped and surfaced, never held. - Clock-skew bound. Reject
hlc_physical_msmore than 300 000 ms in the receiver’s future. - Scope check. Events outside the receiver’s syncable scope are dropped and logged, not raised as errors.
- Idempotent insert keyed on the derived
event_id— BLAKE3 of the received bytes. Duplicates are no-ops. - HLC observe. Merge the event’s HLC into the local clock.
Order matters at two points. Signature verification is first so that no later check consumes attacker-controlled input it has not authenticated. And step 3 drops rather than raises specifically so that a hostile peer cannot wedge the apply loop by sending one out-of-scope event: an exception there would halt processing of everything behind it, which converts a filtered event into a denial of service.
A single apply path is itself normative. Two paths mean two places for these checks to drift apart, and the checks are the security boundary.
13. Materialization and conflict resolution
Section titled “13. Materialization and conflict resolution”Events are materialized in HLC order, grouped per (site_id, schema, table).
13.1. Per-column last-writer-wins, bounded by the tombstone clock
Section titled “13.1. Per-column last-writer-wins, bounded by the tombstone clock”Each (schema, table, row_key, column) carries a column clock. A column value
is written only if its event’s HLC — including the site_id tiebreak — beats
both the stored column clock and the row’s tombstone clock.
Column clocks advance only for writes that are actually applied. A rejected write leaves the clock untouched, so a losing event cannot raise the bar for a later legitimate one.
Resolution is per column, not per row. Two nodes editing different columns of the same row while partitioned both keep their edit; whole-row LWW would discard one arbitrarily.
A column value MAY be null, and an upsert carrying an explicit null for a column is
an ordinary write of that column: it is subject to the same two clock
comparisons, and on winning it advances that column’s clock like any other value.
The null form of payload in section
7 is the tombstone encoding and constrains the
payload as a whole; it says nothing about a cell inside a payload map. The two are
not interchangeable. Clearing one column is a write bounded by that column’s clock,
where a delete tombstones the row, moves the tombstone clock, and clears every
column the tombstone clock beats — so a mesh with no way to write a null cell would
have to delete the row to clear a field, discarding concurrent writes to its other
columns and reintroducing exactly the whole-row loss this section exists to prevent.
A null therefore carries no information about which event wrote it, which is why
backfill selects a row’s covering set on the
column clock rather than on the column value.
13.2. Causal-length existence
Section titled “13.2. Causal-length existence”Row existence is tracked separately from column values, and the highest-HLC
existence event wins. A delete tombstones the row; a later, higher-HLC upsert
resurrects it.
This row tombstone is the sync layer’s construct and resurrection is its defined behaviour. It is not the object plane’s permanent identity tombstone, which no later write may resurrect — the two share a word because both mark a deletion, and nothing else (terminology).
A resurrection SHOULD be surfaced to the application rather than applied silently, because a row reappearing is usually either a genuine re-creation or a sign of a clock problem, and the application is the only layer that can tell which.
13.3. The tombstone clock
Section titled “13.3. The tombstone clock”Each row additionally tracks the highest-HLC delete ever applied to it,
independent of whether that delete won existence.
A delete — winning or losing — MUST clear (set to null) every column whose clock it beats, and advances the tombstone clock if it exceeds it.
This is what makes a delete’s column-clearing effect independent of arrival order. A column write below the tombstone is rejected whether it arrives before the delete or after it. Without a separate tombstone clock, “delete then late-arriving write” and “late-arriving write then delete” produce different states, and the protocol loses convergence.
13.4. Resurrection rebuild
Section titled “13.4. Resurrection rebuild”When an upsert resurrects a tombstoned row, the receiver MUST rebuild the row’s columns from its retained event log. For each column, the value is supplied by the highest-HLC upsert that:
- carries that column, and
- beats the tombstone clock, and
- passed the apply checks, and
- is not below the eviction horizon.
This restores concurrent post-tombstone writes that arrived while the row was absent. Without the rebuild, a write that arrived during the tombstoned interval is lost even though it causally succeeds the delete — a cross-site delivery-order defect that is invisible in single-node testing and only appears when two peers interleave a delete and a write.
The last bullet is a retention bound, and it is where two nodes holding the same events stop agreeing. A rebuild is complete when the row’s tombstone clock is at or above the table’s eviction horizon: every candidate then lies above the horizon and is still retained. When the tombstone clock is below the horizon, the candidate window between the two has been discarded, the rebuild is partial, and a column whose only surviving evidence lay in that window comes back null.
A receiver MUST NOT present a partial rebuild as a complete one. It MUST surface the resurrection as partial, naming the row and the horizon that truncated it — the signal 13.2 asks for, raised to a MUST here because the application is not being told that a value is surprising but that a value is missing.
That bound is stated against the horizon, not against which bytes happen to survive. Eviction is lazy in every practical implementation, so a receiver will often still hold events below its own horizon; it MUST NOT feed them to a rebuild. Admitting them would make the result depend on local reclamation timing, which no peer shares, in place of the declared horizon, which is at least stated.
13.5. Convergence
Section titled “13.5. Convergence”With the preceding rules, materialized state is a pure function of the delivered event set and the receiver’s eviction horizons. On a node that evicts nothing the horizons drop out and it is a pure function of the delivered event set alone. Any delivery order permitted by delivery — per-peer in-order, arbitrary cross-peer interleaving, at-least-once duplication — converges to the same state.
“The delivered event set” is the whole input only for the columns the receiver’s schema actually has; a receiver that lacks a column applies the rest of the event and rebuilds the column when it gains one (13.6).
Agreement between two nodes therefore rests on two preconditions, not one.
A node MUST NOT reuse an HLC across distinct events. Two rules in the data model discharge it, and both are required: the tick discipline while a node runs, and clock durability across a restart. The second exists because the first is a rule about a value in memory: a clock re-initialized from the wall clock, or restored from a stale backup, satisfies the tick discipline exactly and still re-issues HLCs the node has already released. An implementation that reuses an HLC breaks convergence, and no other rule on this page compensates.
Both hold the same horizons. Eviction is the one permitted departure from the pure-function property, and it is a deliberate trade rather than an oversight: a storage-constrained node gives up exactness on old rows to hold the table at all. Two nodes with the same delivered event set and different horizons differ in exactly two ways, both confined to rows whose history reaches below the higher of the two horizons. An event arriving below a node’s horizon counts as applied and writes nothing, where a node with the lower horizon applies it. And a resurrection rebuild whose tombstone clock sits below the horizon comes back partial, where the node with the lower horizon rebuilds the column. Rows whose history lies entirely above both horizons are unaffected, which is the bound worth stating: divergence is confined to evicted history, not spread across the table.
The divergence is accepted rather than automatically repaired, and it is worth being exact about why, because the obvious candidate does not do it. Widening backfill fires on a policy widening and not on a difference between two nodes, so nothing triggers it here; and a re-sent event landing below the receiver’s horizon is counted applied and written nowhere, so while the horizon stands the receiver defeats the re-delivery whatever the sender sends. The repair that does exist is deliberate and operator-driven: rehydration drops the horizon back to zero and takes the history from a peer that still holds it. Where no peer still holds it, nothing restores it.
SSP/1 defines no state digest, so a difference is not detectable on the wire either. The partial-rebuild signal in 13.4 is the only notice an application gets, which is why it is a MUST there rather than a SHOULD.
Implementations SHOULD verify convergence by property-based testing over randomized delivery orders rather than by example, because the failure modes live in interleavings a hand-written test is unlikely to pick. Such a test MUST fix the horizons equal across the nodes it compares, or it is exercising the horizons rather than the merge.
13.6. Held and dropped groups
Section titled “13.6. Held and dropped groups”| Condition | Behaviour | Watermark |
|---|---|---|
| Destination table missing | Group is held and retried in a later cycle | Does not advance past it |
hlc_physical_ms beyond the skew bound |
Event is rejected and surfaced, and not retained | Does not advance past it |
| Signature verification fails | Event is dropped and surfaced | Advances past it |
row_key is null or an empty map |
Event is dropped and surfaced | Advances past it |
row_key names any column the destination table does not have |
Event is dropped | Advances past it |
payload names columns the destination table does not have |
The columns it does have are applied; the rest are ignored and their column clocks left untouched | Advances past it |
The signature row deserves its rationale, because holding would seem safer and is
not. Events are attributed by their signed kid, not by the connection that
delivered them — so if verification failures held the attributed site’s group, a
forger could stall the victim’s legitimate stream by sending garbage carrying
the victim’s kid. A forgery must cost the forger, not the node it impersonates.
row_key is the only thing that names a row, and every rule in this section is
keyed on (schema, table, row_key, column). A null or empty key therefore addresses
nothing: no column clock, no existence bit and no tombstone clock can be located for
it, so refusing it is the only behaviour left to define. The data
model types row_key as a non-empty map for
both operations for that reason.
The alternative — reading a null key as an inert marker event that materializes
nothing — is rejected deliberately. Such a marker is a signed, in-scope, watermarked
event carrying an arbitrary payload that no receiver rule ever inspects, which is
a covert channel through the sync plane with nothing to weigh against it: SSP/1 has
exactly two operations and both mutate a row.
The key-resolution row is an adversarial-key defense: a peer that sends events keyed on non-existent columns must not be able to stall the watermark forever. No destination mutation runs, and the watermark advances. The condition is any missing column rather than all of them, because the partially-resolvable composite key is the dangerous case — a receiver that discarded the unresolvable components and matched on the remainder would address every row sharing the surviving prefix, turning one event into a mass update. Either the whole key resolves or the event does not apply.
Unknown payload columns are applied in part, and the reason is that the benign
case is the common one: two nodes at different schema versions. A node that has not
yet run the migration adding a column still receives events carrying it, and both
dropping the event and holding the group would make an ordinary rolling upgrade lose
or stall data the receiver could have applied.
Holding is specifically unavailable here in a way it is not for a missing table.
Groups are per (site_id, schema, table), so a group held for a table that does not
exist contains only events for that table — a peer naming a fictional table stalls a
fictional group. A column lives inside a real table’s group, so holding on an
unknown column would stall a legitimate stream, which is precisely the stall a
hostile peer is after.
What makes partial application safe is the rule already stated in 13.1: a column clock advances only for a write actually applied. An ignored column leaves its clock untouched, so ignoring it never raises the bar for the same value arriving later.
An upsert asserts row existence independently of its columns (13.2). An upsert whose payload names no destination column at all therefore still applies: the row exists and its key columns are written, and no value columns are. That is not a special case — it is what partial application says when the intersection is empty.
Such a row is existent and empty, and empty is not the same claim as cleared. 13.3 makes a null a real, clocked assertion — a delete sets it — whereas a column the local schema never materialized carries no clock at all. Both surface to a reader as an absent value, so a receiver MUST NOT present an unmaterialized column to the application as a default, a zero, or a cleared value. The distinction lives in the column clock and the receiver is the only layer that can see it, which is the same reason a resurrection is surfaced rather than applied silently.
A receiver SHOULD also meter the events whose payload columns it ignored. Ignoring is expected and transient while a mesh rolls through a schema upgrade, and it is a schema-drift fault when it persists; the individual event looks identical in both cases and only the rate over time separates them.
That leaves the ignored value itself. The watermark advanced, so the sender will not re-offer the event and the receiver’s own retained log is the only remaining copy. A receiver that gains a column by migration MUST rebuild it from that log by the resurrection rebuild procedure — the highest-HLC upsert that carries the column, beats the tombstone clock, passed the apply checks and is not below the eviction horizon. Values whose events fall below that horizon are not recoverable, which is eviction working as specified rather than a defect of this rule. Retention has a cost on the other side of the ledger: the log now demonstrably holds column values that appear in no local table, which is why security §43.2 scopes at-rest protection to the log and not to the materialized store alone.
Convergence is therefore a pure function of the delivered event set and the receiver’s schema. Two nodes holding the same events under the same schema agree. Two nodes holding the same events under different schemas agree on the columns they share, and agree completely once the narrower one migrates and rebuilds.
The skew row is the one condition that may resolve and is nonetheless not held. The receiver keeps nothing — a retained far-future event is storage a hostile peer chooses the size of — so the retry unit is the sender’s next sync cycle rather than a receiver-side queue, and not advancing the watermark is what leaves the event in the set the sender still re-offers. The rationale in full, including why advancing would acknowledge the site’s whole backlog, is section 11.1.
Held-versus-dropped is the distinction to get right. Holding is for conditions that may resolve and that a hostile peer cannot induce against a legitimate stream — a table that has not been created yet. Dropping is for conditions that will not resolve, and for anything a hostile peer could otherwise weaponize into a stall. Rejecting is the third verb (section 38): like dropping it retains nothing, and like holding it leaves the watermark where it is.
14. Delivery
Section titled “14. Delivery”14.1. Cursors
Section titled “14.1. Cursors”A sender maintains one durable cursor per peer per transport path, and that cursor —
not anything the peer advertises — is what
selects the events the sender offers. A cursor for
a peer the sender has never sent to starts at genesis, so a node joining an
established mesh receives that sender’s history from the beginning rather than only
what was produced after it arrived. A cursor MUST NOT be initialised or advanced from
a peer’s advertised have_max_hlc.
Cursors advance only after a successful push. A failed push leaves the cursor untouched, so the same batch is re-offered on the next cycle.
Delivery is therefore at-least-once, and correctness rests on receiver idempotency rather than sender exactly-once. That is the cheaper and more robust half of the trade: idempotency is a local property a receiver can guarantee alone, whereas exactly-once delivery requires agreement the transport cannot provide.
There is no retry-backoff protocol. The retry unit is the next sync cycle.
14.2. Applied watermarks
Section titled “14.2. Applied watermarks”A receiver advances a local applied watermark per contiguous applied chunk, and
reports it to peers. A watermark is an HLC value per (site_id, schema, table)
meaning “everything from this site for this table at or below this HLC has been
applied”.
Watermarks MUST be monotonic and MUST fail closed: the absence of a watermark means not-applied, never applied. A receiver MUST NOT infer progress it cannot demonstrate.
Watermarks are an acknowledgement optimization, not a correctness mechanism. A sender that ignores them entirely and re-offers from its cursor is still correct, because apply is idempotent. Their purpose is to let a sender stop re-offering data the peer demonstrably has, and to let an operator see how far behind a peer is.