Navigation

Resource snapshots

Every resource type has restorable point-in-time versions of its working copy, whether or not it can be published. A snapshot is addressed by channel: published is what a publish writes, revisions is where the working copy's own recovery points live, and everything between them — the version row, the listing, the reconstitution, the restore and the ledger — is one mechanism rather than one per channel.

A rollback reaches either channel, on any type, from a panel over the thing being restored — never a nav blade of its own, and never restricted to a deliberate publish on a publishable type. Restoring is an operation on one resource, so it belongs beside that resource rather than in a surface of its own.

The mechanism: channels

{id}/content.json                        working copy
{id}/objects/{hash}                      every retained version's content, either channel — see the version store
{id}/published/{publishId}/files/…       the immutable channel's asset clones
{id}/files/…                             binary assets, FileAssets types only

A version is a resourceVersions row keyed by resource, channel and version number, pointing at a content-addressed object; how a version is encoded, charged and collected is the resource version store. The revision channel is reference kind and owner-only; the published channel is immutable kind and publicly served.

SnapshotChannelDefinitionMap is the one place a channel says what it is: its kind, its retention, and the title its rows wear.

A channel is an address space, not a workflow. The two differ on six axes — kind, counter, retention, visibility, whether an unpublish sweep takes them, and whether taking one is an outward act with a public url, a view count, a notification and an activity entry. A createSnapshot(id, channel) driving all of that from a map would be a branch with indirection between it and its reader, which is the over-generalization the capability admission rule forbids one level up. So the split is drawn where the operations are genuinely the same one:

Shared by both channelsOwned by each caller
the version row, and the object it namestaking a snapshot
the counter in Postgres, and the history listing beside itpublish's transform, version claim, succession check and repair
reconstitution — read, re-apply live state, hand backthe revision's age and count eviction
restore — reconstitute, then saveResourceContentpublish's activity entry, notification and view counting
the ledger charge on write and release on evict

Reconstitution is the row that matters, because sharing it fixed a defect rather than saving lines — see the boundary below.

The counter lives in Postgres and is never derived from the listing: the listing answers which snapshots exist, the row answers what the next version is and which one is live, and the two are allowed to disagree, because a number is claimed before its write and a write that fails leaves the number with no row. Revisions take two columns on resources rather than a table — revisionVersion numbers them and revisionTakenAt is the clock the automatic trigger reads. A revision's reason and its one-line summary are columns on its version row, which the listing returns.

Two snapshot kinds

The clone is the expensive half of a snapshot, and making it a property of the channel is what makes a second channel affordable:

KindWhat is writtenCostSurvivesUsed by
Immutablecontent + a clone of every referenced asset, urls rewritten to the clonesone version, plus one round trip per assetthe working copy deleting or replacing an assetpublished
Referencecontent only, urls untouched, resolving to live {id}/files/…one version, usually a deltanothing — an asset the owner deletes is gone from itrevisions

Revisions take the reference kind: {id}/files/ is only emptied by purge, which destroys the revisions in the same sweep, so the window in which one can rot is exactly "the owner deleted an asset and then rolled back past the deletion". A rolled-back revision with one broken image beats no rollback, and a per-asset clone on every revision would mean no revisions at all.

It follows that the revision channel holds no assets, only JSON — so parseResourceAssetPath needs no new segment, and the rule that {id}/published/… assets serve anonymously while a publication row exists cannot accidentally extend to revisions.

flowchart LR
  WORK[("{id}/content.json<br/>working copy")]
  FILES[("{id}/files/…<br/>binary assets")]

  WORK --> TAKEREV["revision take<br/>serialize, bump counter and clock,<br/>evict past 30 days or past 100"]
  WORK --> TAKEPUB["publish take<br/>transform, claim in txn, succession repair"]

  TAKEREV --> REV[("revision n")]
  TAKEPUB --> CLONE["cloneContentAssets<br/>published/{publishId}/files/…"]
  CLONE --> PUB[("published version n")]

  REV -.->|"urls resolve to the live assets"| FILES
  PUB -.->|"urls rewritten to its own clones"| CLONE
  REV -->|restore| WORK
  PUB -->|restore| WORK

The snapshot boundary

Survey draws a line through its content blob that no other part of the system knows about: model is snapshot state, settings is live state, re-read on every public read so closing a survey takes effect without re-publishing every participant link already sent.

That line has to be declared in both directions, or restore gets it wrong: a restore that copies the snapshot's content wholesale into the working copy carries settings with it, silently reopening a closed survey or flipping the response mode between Anonymous and Identified — a setting the write boundary makes authorization decisions on (survey response modes).

So the boundary is a two-way declaration the mechanism owns: ResourceLiveContentMap says which parts of a type's content are live rather than frozen, and reapplyLiveResourceContent applies it on every path that reconstitutes a snapshot — the public read, the version preview, and the restore. Writing those paths apart is what lets them disagree; one shared reconstitution makes it impossible.

sequenceDiagram
  actor Owner
  participant R as resource router
  participant SNAP as snapshot channel
  participant WORK as working copy

  Owner->>R: restoreSnapshotVersion(id, channel, n)
  R->>SNAP: read version n of the channel
  R->>R: take a BeforeRestore revision of the working copy
  R->>R: re-apply live state over the snapshot
  Note over R: the boundary — survey collection settings, and anything else a type declares live
  alt immutable channel
    R->>WORK: clone the snapshot's assets back into {id}/files
  else reference channel
    Note over R: the snapshot already points at the working copy's own assets
  end
  R->>WORK: saveResourceContent — contentVersion++, after-save hooks, activity trail

The revision taken before the write is what makes the mechanism append-only: a rollback is not a rewind but an append whose content happens to equal an earlier state, so undoing one is simply the next append. It is taken once the snapshot is known to exist — a restore that was never going to land does not spend a ring-buffer slot on its way to failing — and it is allowed to throw, because a restore whose undo silently did not happen is the defect this exists to close.

Two things sit outside the invariant: retention ends a revision at its age or its count, so append-only holds over recent history rather than all of it, and the unpublish sweep deletes {id}/published/ outright. The revision channel survives an unpublish untouched, so nothing recoverable goes with it.

When a snapshot is taken

Every trigger is the system's, never the owner's. A recovery point that exists because somebody remembered to ask for one is not a recovery path — it is the same failure mode no manual recovery rejects one layer down, and it fails exactly when the owner was too absorbed in the work to think about losing it. So there is no Save version command, and nothing in the product asks an owner to take a snapshot.

TriggerChannelRationale
Publish commandpublishedthe outward act, unchanged
Before a restorerevisionsthe undo, taken by the restore itself
Before a Sheet importrevisionsthe other write that replaces a draft wholesale; the import is refused if it fails
First save an interval after the last onerevisionsat most one per interval, so a working session leaves a handful of recovery points
Every other autosave—never

The interval is measured from the last revision, not from the last save. saveResourceContent fires on every coalesced keystroke batch, and Sheet and Dashboard put real data in the content blob — so a per-save revision copies the whole artifact each time, charges the owner's quota while they type, and grows a listing that has no limit. Throttling is the whole of the cost control. But throttling on the save clock inverts the feature: updatedAt moves with every autosave, so a resource being actively edited never looks idle, and the session that most deserves recovery points is the one that leaves none. revisionTakenAt is a clock only a revision moves, so a working session leaves one point per interval however continuously it is edited.

The interval is claimed in the row, not checked before it. A save reads its Resource before it writes, so two concurrent saves both hold a revisionTakenAt from before either took a revision and both pass a caller-side comparison — leaving two revisions inside one interval and two copies of the artifact charged to the owner. The automatic take's UPDATE therefore carries the interval in its own WHERE, and a claim that returns no row takes no revision: losing that race is not a failure, because the interval already has its point. The caller-side check stays as a filter, so the saves that obviously have nothing to take skip the content download. A deliberate take — before a restore, before an import — claims unconditionally, since there the revision is what makes the act undoable.

A resource's first content write takes none — there is no prior state to keep, and the blob a revision would copy is the one that write is creating. The take is also the one trigger that swallows its own failure: a save must never fail because a revision could not be taken. A blueprint deploy takes none either — it creates resources rather than overwriting one, so there is no draft to hand back.

Retention

The published channel prunes nothing: publishes are deliberate and rare, and a retired public artifact is something an owner may need to point at.

Revisions are kept for 30 days, with at most 100 standing at once — the channel's maxAgeMs and maxRetained. Time is the standard the reference products retain by: Figma's free tier, Dropbox Basic and Google Drive's uploaded files keep 30 days, Notion keeps 7 to 90 by plan, and SharePoint, whose own guidance calls less than 30 days a risk of "inadvertent data loss", pairs an age with a count. The ceiling is Google Drive's hundred, so a resource edited all day — a revision every 15 minutes is up to a few thousand in 30 days — keeps a bounded history inside the window.

Expiry has no job behind it. A revision past its age is gone to every read the moment it expires: getSnapshotRetainedSince is the cutoff both the history listing and the version read compare createdAt against, so a resource nobody edits loses its old revisions on time like any other. The next revision taken deletes the rows past the age or the count, and hands the deletion event exactly the objects no surviving version still needs (resource version store), so the evicted bytes' ledger entries are released with them; a bare delete would make eviction a slow quota leak nothing reconciles (storage quotas).

flowchart TD
  expire["A revision passes 30 days"] --> hidden["Gone to every read — history lists it no more, a restore reads nothing"]
  hidden --> edited{"Is the resource edited again?"}
  edited -->|yes| take["The next revision take deletes its row and collects its bytes"]
  edited -->|no| held["Its bytes stay stored and counted until the next take or the resource's purge"]

The accepted cost is the last box: an expired revision of a resource nobody edits is unreachable but still stored, and counted against its owner, until the resource is edited again or purged. Deleting it on time would take a sweep, and a sweep is what this design exists not to run.

Versions the owner sees

contentVersion is never shown. It is an optimistic-concurrency token that increments once per autosave, so surfacing it tells an owner their document is at v437 because they typed 437 times.

The two counters that reach the UI are different axes: publishVersion is what the public sees, revisionVersion is what you can return to. Overview's Status row shows each only where it carries information:

TypeStateStatus row
Not publishable—that a restore point exists, once one does
Publishablenever publishedDraft chip, plus that a restore point exists once one does
Publishablepublished, draft unchanged sincePublished chip, v{publishVersion}, up to date
Publishablepublished, draft moved sincePublished chip, v{publishVersion}, and that changes are unpublished

"A restore point exists" holds while the newest revision is inside its 30 days: the row reads revisionTakenAt against the same cutoff every version read uses, since revisionVersion only ever counts up and would claim a restore point on a resource whose revisions have all expired.

The last row is a comparison rather than a guess: resourcePublications.publishedContentVersion records the contentVersion the publish was taken from, and updatedAt cannot answer it because a rename or a tag edit moves that too.

revisionVersion itself is never rendered — an owner picks a version by time, reason and summary, never by ordinal — but it remains what the mechanism counts with.

The rollback surface

Version history is a panel over whichever blade is open, not a nav blade: rollback is wanted where the damage happened, which is where every product that does this well puts it.

  • Opens from the page's overflow menu (Resource/Blade/Header), because Sheet and TodoList are blade-only types with no Editor blade — the page header is the one surface every type has. It is the only command the feature has: the history is a place you go, never a thing you maintain.
  • Deep-linkable by route — ?versions opens the panel, ?version={channel}|{n} names the version being previewed, so the back button, a refresh and a shared link all land in the same place.
  • One list, two address spaces. Both channels merge into one time-ordered timeline, because the owner has one question. Current is always the first row, so the list is never empty on a resource that has just been created; a Published only chip filters on publishable types.
  • The timeline is filed under the resource it was read for. A read names its resource and writes that resource's slice of a keyed list, never "the one on screen": reads of two resources do not supersede each other, so one issued before a switch can land after it, and filed ambiently it would list, preview and restore the resource left behind.
  • A row is choosable: its channel and version spelled out rather than a bare ordinal, a relative time with the absolute one on hover, its reason from SnapshotReasonTitleMap, and a one-line summary from SnapshotSummaryMap — 12 items, 3 columns · 40 rows. The summary is computed where the snapshot is taken and carried on its version row, so the listing stays one query for the whole history.
  • Preview in place renders a published version through the type's own public renderer where the blade was, under a banner carrying Restore this version and Back to current. A revision has no rendered form of its own — that would be a read-only renderer per type, publishable or not — so its row restores rather than previews.
  • Restore notifies with an Undo that restores the BeforeRestore revision it had just taken, naming the resource it was offered for rather than whichever is open when it is clicked. Single-use: a second fire would restore a draft the first already replaced.

A restore never re-points the publication — it produces a draft to review and re-publish, mirroring the recycle bin rule — and it lands through saveResourceContent like any other content write, so after-save hooks re-derive what the restored content declares and the activity trail records a Restored entry.

Because it replaces content underneath an already-open blade, the content stores re-read themselves through ResourceContentHookMap.Reload rather than the blade being keyed on a counter something bumps, and the editors that hold the live document themselves are then handed it through a second stage (third-party document adapters).

Why not Azure Blob versioning

A resource version is not a blob version. It is content.json plus the assets it references, and Azure versions each blob on its own timeline with no cross-blob consistency point — so "restore this resource as of Tuesday" resolves to Tuesday's JSON pointing at asset blobs since replaced or deleted. Cloning the assets alongside the content is a thing the application does and the storage account cannot.

Three lesser reasons, each independently sufficient: versioning is a blob-service property, so enabling it for resource-assets enables it for every container on the account; non-current versions are billed but invisible to the ledger, which charges the current blob and reconciles off BlobCreated (storage quotas); and a version is addressed by an opaque id nothing in the app stores, lists or hands to a restore.

Procedures

ProcedureAuthInputPurpose
resource.readSnapshotHistorygetOwnerProcedure{ id }both channels merged, newest-first by time
resource.restoreSnapshotVersiongetOwnerProcedure{ channel, id, version }reconstitute a snapshot into the working copy
resource.saveResourceRevisiongetOwnerProcedure{ id }the pre-import safety net, and the only take a client asks for
{type}.readPublishedVersionContentgetOwnerProcedure{ id, version }owner-only read of one published snapshot, for preview

The channel rides with the version on every command, because a version alone names one snapshot per channel. The reason a revision carries is never a client's to name: Automatic is decided by the save path from a clock and BeforeRestore by the restore itself, so a caller choosing between them would be choosing what its own write is called.

Key files

FileRole
apps/web/shared/services/resource/SnapshotChannelDefinitionMap.tswhat a channel is — kind, retention, title
apps/web/shared/services/resource/SnapshotSummaryMap.tsthe per-type one line a history row carries
apps/web/server/services/resource/snapshot/takeResourceRevision.tsthe revision take, its eviction and its ledger charge
apps/web/shared/services/resource/getSnapshotRetainedSince.tsthe cutoff a version is gone past, which every read compares
apps/web/server/services/resource/snapshot/readSnapshotHistory.tsa channel's version rows as history rows
apps/web/server/services/resource/ResourceLiveContentMap.tsthe boundary — what a type declares live
apps/web/server/services/resource/reapplyLiveResourceContent.tsthe reconstitution every snapshot read goes through
apps/web/server/trpc/routers/resource.tshistory, restore and save-version procedures
apps/web/server/trpc/procedure/resource/createResourceProcedures.tsthe publish take, its version claim and its succession repair
apps/web/app/components/Resource/VersionHistory/the panel, its rows, the preview banner and the restore dialog
apps/web/app/store/resource/versionHistory.tsthe timeline, the restore and its Undo
packages/db-schema/src/schema/resource/resourcesInResource.tsrevisionVersion, revisionTakenAt
packages/db-schema/src/schema/resource/resourceVersionsInResource.tsone row per retained version, in either channel
packages/db-schema/src/schema/resource/resourcePublicationsInResource.tspublishedContentVersion

Notes

  • After an unpublish the publish numbering restarts at 1, because unpublish deletes the publication row and the published channel's version rows with it, and sweeps the asset clones best-effort. The live version is read from the publication row rather than inferred from which versions exist: the listing answers what can be returned to, the row answers what is being served.
  • Purge and soft delete need no step of their own: purge takes {id}/ wholesale, which is already every channel.
  • Named checkpoints was rejected for the Sheet editor because undo/redo already traverses prior states, and the resource-level version of the same idea is rejected in owner-named versions — a row is chosen by its time, its reason and what it holds, none of which the owner has to supply.
  • Whether a resource's edits are durable is save state, not this page: version history is where an owner goes to undo, and the title row is where they see that there was nothing to undo in the first place.

Sources

Details

Command palette

Keyboard shortcuts