One sovereign data lake, meshed across devices you trust and storage you don't.

One trust authority holds the keys; every machine under it can hold the data — analytical tables, live mutable state, and large binary assets. Machines keep producing while offline, with no central node to reach.

Authorized machines recover ordinary Parquet that can be queried with DuckDB, Arrow, Polars, Spark, or similar tools.

Experimental protocol · Specification only · No released implementation

A dataset can live in many places and still belong to exactly one authority.

Figure 1 · the trust boundary
TRUST AUTHORITYidentity · membership · keysnot a master data node — it controls trust and key membershipbut is not required to coordinate every writeencrypted objecttrust / key membershipOFFLINEno connectivity requiredmembershipread authorityINDEPENDENT PRODUCERSproducer / laptopOFFLINEwrites to local disk, no networkproducer / phoneONLINEsyncs when a link is availableproducer / field sensorOFFLINEbatch upload on physical mediaencrypted objectimmutable · id = hash(ciphertext)AUTHORIZED COMPUTEholds the read keysverifies · decryptsplaintext exists only hereencrypted objectfetched from any of themplaintextSTANDARD TOOLSDuckDBArrowPolarsSparkBLIND STORAGEoperated by anyone · trusted by no onelaptopNASS3relayIPFScacheUSBobject storeVPSstores and transports bytes — does not possess read authority
Figure 1: Holding the bytes conveys no ability to read them. Producers need no coordination and no connectivity to write; storage needs no trust to carry.

Storage ≠ access

A replica can hold every byte of a dataset without possessing the cryptographic capability to read it.

Writes ≠ coordination

A node should be able to produce data while disconnected, without contacting a master, a consensus group, or a global sequence allocator.

Encryption ≠ proprietary analytics

Authorized nodes recover ordinary analytical objects rather than a special database representation.

3.1

Produce

A node batches observations into a Parquet object, encrypts it, signs it, and names it by the hash of its ciphertext. Nothing on the network has to be reachable for this to happen.

3.2

Replicate

The object is copied outward through whatever storage and transport is available. Every hop holds bytes it has no capability to decrypt, so replicas can be cheap and untrusted.

3.3

Authorize

Membership, identity and key state travel on the control plane rather than with the object. An authorized node verifies the producer signature and decrypts.

3.4

Query

What comes out is ordinary Parquet. DuckDB, Arrow, Polars or Spark read it directly — there is no proprietary representation to convert first.

Transport independence test — the property the design is aiming at

An SOP object can be copied with cp, stored in S3, fetched using iroh, pinned in IPFS, carried on a USB drive, or sent through another transport while retaining its identity and security properties. An authorized node can receive a directory of SOP objects, identify them, verify integrity, determine dataset membership, decrypt the objects it is authorized for, and query the result — without contacting a sovm-operated central service.

This is the specification's target, not a demonstrated capability: there is no released implementation to test it against yet.

Two kinds of state need two different replication models.

Membership, authorization, identities, trust state and key evolution are mutable and small, and every node has to agree on what they currently are. That state needs deterministic convergence.

Observations, telemetry, agent traces, event batches and analytical data are immutable and large. Nothing concurrently mutates a batch that was written once, and nothing needs a global order over batches produced independently. Forcing that traffic through a convergence protocol buys coordination nobody asked for.

So sovm is two protocols over one substrate, and the split is the reason it works.

SSP/1

The control plane. Synchronizes the small mutable state — membership, identity, policy, catalog metadata — and supplies trust to everything above it. Group membership and key distribution use MLS (RFC 9420).

SOP/1

The object plane. Carries high-volume immutable data as encrypted Parquet objects named by the hash of their ciphertext. No agreement between peers is required to write one.

The transport and storage layers are meant to be replaceable.

iroh is the preferred live transport because it dials by public-key identity and traverses NATs, but it is a substrate choice, not the innovation. Any of the bottom two rows can be swapped without the rows above noticing, and that is the whole point of naming an object by the hash of its ciphertext.

Figure 2 · the layering
  1. DuckDB / Polars / Arrow / Sparkcompute
  2. Parquetdata format
  3. SOPobject semantics
  4. SSP / MLStrust / membership
  5. iroh / HTTP / S3 / IPFS / local FStransport / storage
  6. disk / NAS / cloud / edge / USBphysical custody

specified by sovminterchangeable

Figure 2: The transport and storage layer should be replaceable. iroh is one option, not the main innovation.

These are not competitors. Most of them compose with sovm.

IPFS

Content addressing, discovery, transport and distributed persistence are solved there, and solved well. sovm adds the layer above: authorization, identity, membership, key evolution, dataset and object semantics, and an analytical representation.

SOP objects can be replicated over IPFS.

Databricks / Delta Lake

ACID analytical tables, a central catalog, governance and large-scale compute are what a lakehouse is for. sovm addresses the stage before or outside centralization: disconnected producers, blind replicas, no global write coordination, trust that does not depend on the storage layer.

Edge nodes → sovm → ingestion → Delta Lake or Databricks is a coherent architecture.

Blockchain and Web3

Blockchains solve consensus between mutually distrustful parties. sovm assumes the dataset belongs to one authority even where the storage infrastructure does not. There is no reason to pay for consensus between devices that already answer to the same owner.

No token, no chain and no global ledger appear anywhere in the specification.

Syncthing

Syncthing already does encrypted replication through untrusted devices, and does it well. It is organized around files, folders and device synchronization. sovm is organized around immutable analytical objects, producer identity, dataset membership, explicit authorization and independently generated streams.

Either can move bytes; they answer different questions about what those bytes are.

S3 + KMS + Parquet + DuckDB + Tailscale + glue code

For a centrally managed organization with reliable connectivity, this is often simpler and better, and we would rather say so than pretend otherwise. sovm earns its complexity when the topology includes disconnected producers, heterogeneous networks, storage you want to use but not trust, several interchangeable replication paths, or datasets that should outlive the current vendor and application.

The same object stores and the same query engines sit underneath sovm.

7.1 Near-term wedge

Security telemetry and agent telemetry.

Thousands of laptops, servers and gateways produce highly sensitive event streams: process execution, authentication, network events, endpoint observations, forensic logs. The traffic is append-heavy, the producers go offline, several backup copies are desirable, and a compromise of the storage layer must not expose the telemetry.

A security dataset can exist in ten places without creating ten readable copies.

AI agents have the same shape. They produce conversations, tool calls, observations, decisions, actions and outputs continuously, and that record is the audit trail for everything the agent did on someone's behalf.

Models are replaceable. A decade of personal or organizational context is not.

7.2 Architectural showcase

Industrial, remote and fleet telemetry.

Mines, factories, ships, farms, oil platforms and scientific stations share one topology: intermittent connectivity, local autonomy, append-only sensor data, opportunistic replication, cheap local caches, and central analytics that happens later rather than live.

A vehicle fleet draws the same picture in miniature — vehicle, then phone, then depot Wi-Fi, then regional cache, then cloud. Every intermediate system in that chain holds encrypted objects without read authority, which is what makes the chain safe to build out of whatever hardware is already there.

7.3 Long-term vision

Sovereign human and organizational memory.

A personal data lake whose operator is the individual: devices, messages, activity, files, wearables, financial records and the actions their agents took for them. The same structure describes an organization that wants its own record rather than a vendor's copy of it.

A person's data should be able to outlive the current cloud provider, the current model, the current application, the current laptop and the current agent implementation.

The honest boundary

One authority holds the keys, and that authority can be a person or a household. Lose the key material and the backups and the lake is gone — see section 10.

8.1

Retail and branch analytics

Sites generate their own data and stay useful when the link to head office is down. DuckDB owns the computation, Parquet owns the representation, sovm owns custody and replication.

8.2

Research data collection

Irreplaceable observations gathered in poor-connectivity environments, replicated opportunistically into several encrypted copies without breaking confidentiality obligations.

8.3

Archival to untrusted storage

Including rented capacity, with cryptographic rather than contractual assurance that the host cannot read what it holds.

8.4

Small-team operational data

Stays inside one trust boundary and survives any single machine failing, without a central service to run.

A vector index is a materialization, in exactly the way DuckDB is an execution engine.

Neither of them is the identity of the data. An embedding index over sovm objects is derived state: it carries provenance back to the objects it was built from, and it can be thrown away and rebuilt.

The vector database is never the source of truth. It is a rebuildable, provenance-bearing index over sovm data.

Private AI falls out of that rather than motivating it. When valuable historical data stays inside a sovereign trust boundary, compute can be brought to the authorized copy instead of exporting another plaintext dataset to another provider.

Rebuildable

The index is derived from objects named by the hash of their ciphertext, so rebuilding it from the same input set is a defined operation rather than a migration.

Provenance-bearing

Every object carries an immutable producer signature, so a row in a derived index can be traced to the node that produced it.

No export step

SOP objects are encrypted Parquet. Once an authorized node decrypts, Arrow, Polars, DuckDB and standard loaders read them directly, column pruning intact.

Selective by construction

The encrypted catalog carries time ranges and value bounds, so a working set is selected on those and only the matching objects are decrypted.

Being explicit about this matters more than the feature list.

sovm owns custody, identity and replication. It does not own computation, and it makes no claim about what happens after an authorized node decrypts. Anything in the list on the right is somebody else's layer, and the specification says so rather than leaving a reader to discover it.

Note — No models and no inference

sovm provides no inference, no models, no embeddings, no vector search and no vector database. It carries the data those things are built on.

Note — No statistical privacy

There is no differential privacy and no secure aggregation. A node that trains must decrypt, and sovm makes no claim about what a model reveals afterwards.

Note — No consensus between strangers

The dataset belongs to one authority. sovm is not a protocol for sharing data between mutually distrusting parties and provides no collaborative multi-writer semantics.

Note — Revocation does not reach back

A machine that was authorized cannot be made to forget what it already read. Revocation stops future access only.

Note — Metadata still leaks

Storage providers observe object sizes and transfer timing. Traffic analysis is reduced, not eliminated.

Note — No recovery key

If the trust authority loses its key material and its backups, the data is gone. That is the design, not a gap in it.

Project status · Experimental

Specification published. Implementation not yet released.

SOVM/1 is at Revision 1, published 2026-09-03. A revision versions this document, not the protocol: nothing normative moved, and none of the ten decision gates that would promote SOP/1 to a normative production protocol has passed. There is no released implementation. The changelog has what a revision is and does; the status page has where every part of the specification stands today, gate by gate.

See open technical gates →