One sovereign data lake, meshed across devices you trust and storage you don't.
One trust authority holds the keys; every machine under it can hold the data — analytical tables, live mutable state, and large binary assets. Machines keep producing while offline, with no central node to reach.
Authorized machines recover ordinary Parquet that can be queried with DuckDB, Arrow, Polars, Spark, or similar tools.
Experimental protocol · Specification only · No released implementation
1. The architecture in one figure
A dataset can live in many places and still belong to exactly one authority.
2. Three invariants
Storage ≠ access
A replica can hold every byte of a dataset without possessing the cryptographic capability to read it.
Writes ≠ coordination
A node should be able to produce data while disconnected, without contacting a master, a consensus group, or a global sequence allocator.
Encryption ≠ proprietary analytics
Authorized nodes recover ordinary analytical objects rather than a special database representation.
3. How an object moves through sovm
3.1
Produce
A node batches observations into a Parquet object, encrypts it, signs it, and names it by the hash of its ciphertext. Nothing on the network has to be reachable for this to happen.
3.2
Replicate
The object is copied outward through whatever storage and transport is available. Every hop holds bytes it has no capability to decrypt, so replicas can be cheap and untrusted.
3.3
Authorize
Membership, identity and key state travel on the control plane rather than with the object. An authorized node verifies the producer signature and decrypts.
3.4
Query
What comes out is ordinary Parquet. DuckDB, Arrow, Polars or Spark read it directly — there is no proprietary representation to convert first.
Transport independence test — the property the design is aiming at
An SOP object can be copied with cp, stored in S3, fetched using iroh, pinned in IPFS, carried on a USB drive, or sent through another transport while retaining its identity and security properties. An authorized node can receive a directory of SOP objects, identify them, verify integrity, determine dataset membership, decrypt the objects it is authorized for, and query the result — without contacting a sovm-operated central service.
This is the specification's target, not a demonstrated capability: there is no released implementation to test it against yet.
4. Two protocol planes
Two kinds of state need two different replication models.
Membership, authorization, identities, trust state and key evolution are mutable and small, and every node has to agree on what they currently are. That state needs deterministic convergence.
Observations, telemetry, agent traces, event batches and analytical data are immutable and large. Nothing concurrently mutates a batch that was written once, and nothing needs a global order over batches produced independently. Forcing that traffic through a convergence protocol buys coordination nobody asked for.
So sovm is two protocols over one substrate, and the split is the reason it works.
SSP/1
The control plane. Synchronizes the small mutable state — membership, identity, policy, catalog metadata — and supplies trust to everything above it. Group membership and key distribution use MLS (RFC 9420).
SOP/1
The object plane. Carries high-volume immutable data as encrypted Parquet objects named by the hash of their ciphertext. No agreement between peers is required to write one.
5. The architectural stack
The transport and storage layers are meant to be replaceable.
iroh is the preferred live transport because it dials by public-key identity and traverses NATs, but it is a substrate choice, not the innovation. Any of the bottom two rows can be swapped without the rows above noticing, and that is the whole point of naming an object by the hash of its ciphertext.
- DuckDB / Polars / Arrow / Sparkcompute
- Parquetdata format
- SOPobject semantics
- SSP / MLStrust / membership
- iroh / HTTP / S3 / IPFS / local FStransport / storage
- disk / NAS / cloud / edge / USBphysical custody
specified by sovminterchangeable
6. Why not X?
These are not competitors. Most of them compose with sovm.
IPFS
Content addressing, discovery, transport and distributed persistence are solved there, and solved well. sovm adds the layer above: authorization, identity, membership, key evolution, dataset and object semantics, and an analytical representation.
SOP objects can be replicated over IPFS.
Databricks / Delta Lake
ACID analytical tables, a central catalog, governance and large-scale compute are what a lakehouse is for. sovm addresses the stage before or outside centralization: disconnected producers, blind replicas, no global write coordination, trust that does not depend on the storage layer.
Edge nodes → sovm → ingestion → Delta Lake or Databricks is a coherent architecture.
Blockchain and Web3
Blockchains solve consensus between mutually distrustful parties. sovm assumes the dataset belongs to one authority even where the storage infrastructure does not. There is no reason to pay for consensus between devices that already answer to the same owner.
No token, no chain and no global ledger appear anywhere in the specification.
Syncthing
Syncthing already does encrypted replication through untrusted devices, and does it well. It is organized around files, folders and device synchronization. sovm is organized around immutable analytical objects, producer identity, dataset membership, explicit authorization and independently generated streams.
Either can move bytes; they answer different questions about what those bytes are.
S3 + KMS + Parquet + DuckDB + Tailscale + glue code
For a centrally managed organization with reliable connectivity, this is often simpler and better, and we would rather say so than pretend otherwise. sovm earns its complexity when the topology includes disconnected producers, heterogeneous networks, storage you want to use but not trust, several interchangeable replication paths, or datasets that should outlive the current vendor and application.
The same object stores and the same query engines sit underneath sovm.
7. Where sovm fits
7.1 Near-term wedge
Security telemetry and agent telemetry.
Thousands of laptops, servers and gateways produce highly sensitive event streams: process execution, authentication, network events, endpoint observations, forensic logs. The traffic is append-heavy, the producers go offline, several backup copies are desirable, and a compromise of the storage layer must not expose the telemetry.
A security dataset can exist in ten places without creating ten readable copies.
AI agents have the same shape. They produce conversations, tool calls, observations, decisions, actions and outputs continuously, and that record is the audit trail for everything the agent did on someone's behalf.
Models are replaceable. A decade of personal or organizational context is not.
7.2 Architectural showcase
Industrial, remote and fleet telemetry.
Mines, factories, ships, farms, oil platforms and scientific stations share one topology: intermittent connectivity, local autonomy, append-only sensor data, opportunistic replication, cheap local caches, and central analytics that happens later rather than live.
A vehicle fleet draws the same picture in miniature — vehicle, then phone, then depot Wi-Fi, then regional cache, then cloud. Every intermediate system in that chain holds encrypted objects without read authority, which is what makes the chain safe to build out of whatever hardware is already there.
7.3 Long-term vision
Sovereign human and organizational memory.
A personal data lake whose operator is the individual: devices, messages, activity, files, wearables, financial records and the actions their agents took for them. The same structure describes an organization that wants its own record rather than a vendor's copy of it.
A person's data should be able to outlive the current cloud provider, the current model, the current application, the current laptop and the current agent implementation.
The honest boundary
One authority holds the keys, and that authority can be a person or a household. Lose the key material and the backups and the lake is gone — see section 10.
8. Additional use cases
8.1
Retail and branch analytics
Sites generate their own data and stay useful when the link to head office is down. DuckDB owns the computation, Parquet owns the representation, sovm owns custody and replication.
8.2
Research data collection
Irreplaceable observations gathered in poor-connectivity environments, replicated opportunistically into several encrypted copies without breaking confidentiality obligations.
8.3
Archival to untrusted storage
Including rented capacity, with cryptographic rather than contractual assurance that the host cannot read what it holds.
8.4
Small-team operational data
Stays inside one trust boundary and survives any single machine failing, without a central service to run.
9. Vector databases and private AI
A vector index is a materialization, in exactly the way DuckDB is an execution engine.
Neither of them is the identity of the data. An embedding index over sovm objects is derived state: it carries provenance back to the objects it was built from, and it can be thrown away and rebuilt.
The vector database is never the source of truth. It is a rebuildable, provenance-bearing index over sovm data.
Private AI falls out of that rather than motivating it. When valuable historical data stays inside a sovereign trust boundary, compute can be brought to the authorized copy instead of exporting another plaintext dataset to another provider.
Rebuildable
The index is derived from objects named by the hash of their ciphertext, so rebuilding it from the same input set is a defined operation rather than a migration.
Provenance-bearing
Every object carries an immutable producer signature, so a row in a derived index can be traced to the node that produced it.
No export step
SOP objects are encrypted Parquet. Once an authorized node decrypts, Arrow, Polars, DuckDB and standard loaders read them directly, column pruning intact.
Selective by construction
The encrypted catalog carries time ranges and value bounds, so a working set is selected on those and only the matching objects are decrypted.
10. What sovm deliberately does not solve
Being explicit about this matters more than the feature list.
sovm owns custody, identity and replication. It does not own computation, and it makes no claim about what happens after an authorized node decrypts. Anything in the list on the right is somebody else's layer, and the specification says so rather than leaving a reader to discover it.
Note — No models and no inference
Note — No statistical privacy
Note — No consensus between strangers
Note — Revocation does not reach back
Note — Metadata still leaks
Note — No recovery key
11. Project status
Project status · Experimental
Specification published. Implementation not yet released.
SOVM/1 is at Revision 1, published 2026-09-03. A revision versions this document, not the protocol: nothing normative moved, and none of the ten decision gates that would promote SOP/1 to a normative production protocol has passed. There is no released implementation. The changelog has what a revision is and does; the status page has where every part of the specification stands today, gate by gate.
See open technical gates →12. Specifications
Go deeper
For architects
Topology, integration patterns, and protocol boundaries.
- Gateway nodes collect while disconnected; no master node, no global sequence
- One explicit, auditable export boundary — never transparent background replication
- Revocation is an MLS membership operation with a defined cryptographic effect
For security reviewers
Threat model, key custody, and revocation semantics.
- Per-object encryption; encrypted footers hide schema, statistics and row counts
- Opaque identifiers on the wire; object identity is the hash of its ciphertext
- One authority holds the keys — and it can be a person or a household
Thesis / Why now
Why increasingly valuable private datasets need a different custody model.
- Training corpus exhaustion is driving demand for private data access
- Centralization-only architectures cannot decouple placement from access
- The architectural bet: compute moves to data, custody is provable not contractual