Navigation

Datasets

The standard for serving tabular data across products. Whenever one product needs to read another product's data (survey responses in a dashboard, file rows in an import), it goes through this contract — never through product-specific reads.

Serving is the DatasetProvider capability (resources): a resource type opts in via ResourceDefinitionMap, and DatasetProviderType keys the provider map so a declared-but-unregistered provider is a compile error. Consuming is not a capability — it is just calling dataset.readDataset from a component (dashboard binding form, email merge fields).

Contract

Shared models in apps/web/shared/models/dataset/ (one type + schema per file, interface-first). ColumnType and ColumnValue are reused from the Sheet resource models — they are already the canonical cell-type vocabulary.

interface DatasetColumn {
  name: string;
  type: DatasetColumnType; // ColumnType minus Computed — a computed column is served as the type it yields
}

interface Dataset {
  columns: DatasetColumn[];
  rows: Record<string, ColumnValue>[];
  totalRows?: number; // when present, the uncapped count — truncation is detectable only for providers that supply it
}

enum DatasetProviderType {
  ProgramStatus = "ProgramStatus",
  Sheet = "Sheet",
  SurveyResponses = "SurveyResponses",
}

interface DatasetReference extends ItemEntityType<DatasetProviderType> {
  id: string; // the resource id
}

A reference is just a resource id — a Sheet resource is the dataset, so a reference carries no sub-item selector. DatasetProviderType (server-resolvable references) is a different axis from DataSourceType (Csv/Json/Xlsx — file formats parsed client-side). Do not merge them: one describes where data lives, the other how a file is encoded.

Serving

One procedure resolves every reference:

ProcedureAuthInputPurpose
dataset.readDatasetResource ownerDatasetReferenceResolve a reference to a Dataset

Public viewers never call this: published resources bake resolved datasets in at publish time.

flowchart LR
  DASH["Dashboard binding form"] -->|DatasetReference| RD["dataset.readDataset"]
  EMAIL["Email merge fields"] -->|DatasetReference| RD
  IMPORT["Sheet import (one-time row copy)"] -->|DatasetReference| RD
  RD --> OWN["readDataset<br/>requireOwnedResource"]
  OWN --> MAP["DatasetProviderMap[type]"]
  MAP --> SR["readSurveyResponsesDataset"] --> AT[("SurveyResponseEntity<br/>Azure Table")]
  MAP --> FR["readSheetDataset"] --> BLOB[("Sheet content blob")]
  MAP --> PR["readProgramStatusDataset"] --> PT[("ProgramParticipantEntity<br/>Azure Table")]

Server structure (server/services/dataset/): DatasetProviderMap.ts maps DatasetProviderType → the provider and the resource type it reads, one provider per folder. readDataset is the only way into a provider: it resolves the reference to a resource the caller owns, of that type, before the provider runs — so a provider is handed an owned resource and has no auth check of its own to forget. Each provider owns its column/row derivation:

  • readSurveyResponsesDataset — columns from the survey model's questions (name + question-type → ColumnType mapping); rows flattened from SurveyResponseEntity JSON in Azure Table, non-primitive answers JSON-stringified.
  • readSheetDataset — reads the Sheet resource's content blob and converts content.data via dataSourceToDataset.
  • readProgramStatusDataset — participant/addedAt/responded rows from ProgramParticipantEntity in Azure Table, keyed by the non-secret publicId (never the token — dashboards bake dataset snapshots into public publishes).

Rules

  • Row cap — AZURE_MAX_PAGE_SIZE (1000) on every provider; datasets are for visualization and import, not bulk export. Truncation is totalRows > rows.length, never the presence of totalRows — a provider that can count cheaply always reports the uncapped total, which equals the row count on an uncapped read. getDatasetTruncation is the one place that comparison lives, so a warning can never disagree with the rows on screen; a provider that cannot count omits the field and its consumers simply never warn, because truncation is then unknowable rather than absent. Add pagination only when a real consumer hits the cap (deferred).
  • Consumers choose copy or reference. Import (Sheet resource) copies rows once. Binding (dashboard visuals, email editor merge fields) stores the DatasetReference and re-resolves on load. All call the same procedure.
  • Fetch on load + manual refresh. No live subscriptions through this layer (deferred).
  • No external providers (HTTP APIs, SQL) until secret storage and injection-safety work is scoped (deferred) — the enum grows one value per new provider, nothing else changes.

Details

Command palette

Keyboard shortcuts