Resource version store
A retained version of a resource — a revision of the working copy, or a published snapshot — is a row in resourceVersions and a content-addressed object in the resource assets container. The object is a zstd keyframe, compressed on its own, or a delta compressed against exactly one keyframe with that keyframe's plaintext as the dictionary. Successive versions of one document differ only by the edit between them, so a version usually costs a few hundred bytes, and the storage meter moves by what the owner actually added rather than by another copy of everything they already had (storage quotas).
Every trigger, channel and retention rule stays what resource snapshots describes. This page is only what a version is on disk, and what that changes for the code that writes, reads, charges and evicts one.
The codec
The codec is the keyframe-store workspace package, a general library that never sees a schema: it takes bytes and returns bytes, addresses an object by the SHA-256 of its plaintext, and decides per write whether a delta or a keyframe is cheaper. The keyframe-store README is the reference for the object format, the promotion ratio, the segment budget and the compression parameters. Two properties of it are what the resource layer is built on:
- Reconstruction is at most two reads. A delta's own header names its keyframe, so there is no chain to walk and no version depends on the one written before it.
- An object is write-once and content-addressed. A version whose content the store already holds is a row pointing at the existing object — nothing written, nothing charged. That is the case that makes the meter honest about a save that changed nothing, and about a restore back to a version the store still has.
The store's only contact with Azure is createSnapshotObjectStore: one resource's objects under {id}/objects/{hash}, written create-only so two writers of the same content cannot race, and deleted through the blob deletion event rather than directly, because that path already deletes and releases the ledger entry as one retried unit.
Objects are scoped by resource rather than by owner. Successive versions of one document are where deduplication fires, an object then has exactly one payer, and every path that already takes {id}/ wholesale — purge, and the ledger release behind it — takes the objects with it, so nothing outside the app process has to know the store exists.
The same codec works on the wire as well: a large document's save is compressed against the bytes the server stores rather than uploaded whole, with the window and level derived once in @esposter/shared for both encoders (delta content saves).
Taking a version
flowchart TD
trigger["A trigger takes a version — revision interval, restore, import, publish"] --> anchor["Read the channel anchor — the newest keyframe row, and the bytes anchored above it"]
anchor --> store["Write to the store — dedupe, delta, or promote to a keyframe"]
store --> row[("Upsert the version row — hash, base, both sizes, reason, summary")]
row --> commit{"Inside a publish transaction?"}
commit -->|yes| claim["The publication claim commits with the row"]
commit -->|no| charge
claim --> charge["Charge the stored bytes to the owner — nothing for a deduplicated write"]
charge --> evict{"A revision past 30 days, or past the hundred?"}
evict -->|no| done["Save proceeds"]
evict -->|yes| shed["Delete those rows"]
shed --> collect["Subtract every hash a surviving row still names, as its own or as its base"]
collect --> publish["Publish the remainder for deletion — the handler deletes and releases each"]
publish --> done
writeSnapshotVersion is the one way a version is taken, whichever channel it lands in. It reads the channel's anchor, writes the content to the store against it, and upserts the row that makes the version visible. The object is durable before the row exists, so a failure between the two leaves no version rather than a row naming nothing — the correct outcome for a version that was never stored. An orphaned object is adopted by the next write of the same content, never swept.
A write and a collection of the same resource never overlap. Both run inside a transaction holding the resource's advisory lock (lockSnapshotObjects), because the gap between an object landing and its row existing is one a concurrent eviction could otherwise fall into: it reads the surviving rows, finds nothing naming the hash the write is about to record, and frees the object — or the anchor it decodes against — under a row that then names something unreadable. Under the lock a collection either sees the row or runs before the write reads its anchor, in which case the write finds the anchor gone and honestly starts a new lineage. The lock is keyed on the resource id, so unrelated resources never wait on each other, and it is transaction-scoped, so a failed write releases it with its rollback.
A write that lands second learns it lost. Two writers of the same content both find nothing stored, and each encodes against its own anchor. The blob write is create-only, and its refusal is passed back as the object store's answer rather than swallowed: the loser reads the header the winner stored and records that base, deduplicated and uncharged, because a row carrying the loser's own base would name a keyframe the stored delta never decodes against, and a charge for its bytes would be for bytes it never stored.
The anchor is derived, never stored. A channel's anchor is its newest row whose base is empty, and the bytes anchored to it are the stored sizes of the rows above it — readSnapshotAnchor answers both from the primary key. A column holding either would be a second source of truth for what one indexed read answers, and one that could disagree with the rows after a failed write.
The charge is separate from the write, because a publish takes its version inside the transaction that claims publishVersion and the charge takes the ledger row's lock and then the user's — a transaction held open across those waits on locks it is itself holding. chargeSnapshotVersion runs after the transaction and charges exactly the stored bytes, against the object's own key, so the object's BlobCreated reconciles to the same figure and its eventual deletion releases it.
The one rewrite is a publish repairing its own snapshot at the version it already claimed (publishing). The upsert moves the row to the new object, and the object it named before is collected once the row has committed, if no other row still names it — published inside the transaction, a commit that failed would leave the row naming an object already gone.
Reading a version
readSnapshotVersionContent looks the row up by resource, channel and version, reconstructs the plaintext through the store, and parses it with the type's content schema. A version whose row is gone — evicted, or swept by an unpublish between the listing and the click — reads as no content, which the public read turns into the 404 page rather than an internal error. A row still standing over objects the sweep already deleted is the same absent version wearing a different mask, and answers the same way: the store raises ObjectNotStoredError for exactly that case, and nothing else, so the read can convert absence into a 404 without swallowing a truncated or mismatched object, which stays the internal error it is. The history listing is a query over the channel's rows, ordered by version; reason, summary and the taken-at clock are columns, so the listing never opens an object.
Collection
flowchart TD
released["Rows an eviction or an unpublish deleted"] --> named["Every hash they named — their own, and their base"]
survivors["Every surviving row of the resource"] --> retained["Every hash those still name, as own or as base"]
named --> subtract{"Named by a survivor?"}
retained --> subtract
subtract -->|yes| keep["Keep — a delta still decodes against it"]
subtract -->|no| drop["Publish for deletion — bytes released when the handler lands"]
An object survives while any row of the resource names it, as its own hash or as its base. collectSnapshotObjects answers that with one query over the resource's rows, which is what the base hash is denormalised onto the row for: without it, deciding whether a keyframe is still needed would mean reading every surviving delta's header. Because a delta is anchored to a keyframe that is itself a version, a keyframe is collectable only once its own row and every delta anchored to it are gone — eviction therefore sheds the oldest segment whole.
Eviction hands the deletion path the exact set of objects nothing references, rather than a version number computed a fixed distance behind the newest. An unpublish deletes the published channel's rows the same way, and keeps its prefix sweep for the asset clones under {id}/published/, which are not objects.
Key files
| File | Role |
|---|---|
packages/keyframe-store/src/createKeyframeStore.ts | the codec — dedupe, delta or keyframe, two-read reconstruction, collect |
packages/db-schema/src/schema/resource/resourceVersionsInResource.ts | the version row — hash, base, both sizes, reason, summary |
apps/web/server/services/resource/snapshot/createSnapshotObjectStore.ts | the store's one contact with Azure — objects under {id}/objects/ |
apps/web/server/services/resource/snapshot/writeSnapshotVersion.ts | the one way a version is taken, in either channel |
apps/web/server/services/resource/snapshot/readSnapshotAnchor.ts | the channel's anchor, derived from its rows |
apps/web/server/services/resource/snapshot/chargeSnapshotVersion.ts | the charge for the stored bytes, after the transaction |
apps/web/server/services/resource/snapshot/readSnapshotVersionContent.ts | reconstruct and parse one version |
apps/web/server/services/resource/snapshot/collectSnapshotObjects.ts | what an eviction or unpublish may delete |
apps/web/server/services/resource/snapshot/lockSnapshotObjects.ts | the per-resource lock a write and a collection both hold |
apps/web/server/services/resource/snapshot/takeResourceRevision.ts | where a revision is taken and the expired and the oldest evicted |
apps/web/server/trpc/procedure/resource/createResourceProcedures.ts | where publishing, the public read and unpublishing are wired |
Notes
- Existing snapshots written as full blobs under
{id}/revisions/and{id}/published/were discarded rather than converted: a resource published before the store shipped has to be published again, and the blobs stay charged until purge takes the directory. Migration was scoped out because the history was recoverable by a republish and a conversion would have paid a full read of every snapshot for versions the ring buffer evicts within a session. - Collection publishes the deletion of what it freed, and the publish is best-effort once the rows are gone: a publish that fails leaves the objects stored and charged until purge takes the directory. That is a harsher term than the write-path orphan above, which was never charged in the first place — this one was live and billed, and nothing hands the charge back until purge clears the whole directory. The order is deliberate — publishing first and then failing to delete the rows would free objects a surviving row still names, which is a version that cannot be read rather than bytes that cost. What is not built for it is an outbox of released hashes: that is a second write path and a sweep to drain it, for a failure the SDK's own retries already make rare, and the same trade the ledger's charges settled (/docs/resource/storage-quotas).
- A write always encodes twice — standalone and against the anchor — because the promotion ratio compares the two. The two overlap on the threadpool, so the save path pays the slower of them rather than their sum: on a multi-megabyte document that is still hundreds of milliseconds, where a full copy was an upload of the same size; the committed bench beside
createKeyframeStoreis the record of what each shape costs, and the compression level is an option so a sweep of it never touches the implementation. - The second phase the same substrate unlocks — addressing assets by content so a publish references them instead of cloning them — stays a proposal, because its correctness rests on a scan of content rather than on a column.