Datasets
The standard for serving tabular data across products. Whenever one product needs to read another product's data (survey responses in a dashboard, file rows in an import), it goes through this contract — never through product-specific reads.
Serving is the DatasetProvider capability (resources): a resource type opts in via ResourceDefinitionMap, and DatasetProviderType keys the provider map so a declared-but-unregistered provider is a compile error. Consuming is not a capability — it is just calling dataset.readDataset from a component (dashboard binding form, email merge fields).
Contract
Shared models in apps/web/shared/models/dataset/ (one type + schema per file, interface-first). ColumnType and ColumnValue are reused from the Sheet resource models — they are already the canonical cell-type vocabulary.
interface DatasetColumn {
name: string;
type: DatasetColumnType; // ColumnType minus Computed — a computed column is served as the type it yields
}
interface Dataset {
columns: DatasetColumn[];
rows: Record<string, ColumnValue>[];
totalRows?: number; // when present, the uncapped count — truncation is detectable only for providers that supply it
}
enum DatasetProviderType {
ProgramStatus = "ProgramStatus",
Sheet = "Sheet",
SurveyResponses = "SurveyResponses",
}
interface DatasetReference extends ItemEntityType<DatasetProviderType> {
id: string; // the resource id
}
A reference is just a resource id — a Sheet resource is the dataset, so a reference carries no sub-item selector. DatasetProviderType (server-resolvable references) is a different axis from DataSourceType (Csv/Json/Xlsx — file formats parsed client-side). Do not merge them: one describes where data lives, the other how a file is encoded.
Serving
One procedure resolves every reference:
| Procedure | Auth | Input | Purpose |
|---|---|---|---|
dataset.readDataset | Resource owner | DatasetReference | Resolve a reference to a Dataset |
Public viewers never call this: published resources bake resolved datasets in at publish time.
flowchart LR
DASH["Dashboard binding form"] -->|DatasetReference| RD["dataset.readDataset"]
EMAIL["Email merge fields"] -->|DatasetReference| RD
IMPORT["Sheet import (one-time row copy)"] -->|DatasetReference| RD
RD --> OWN["readDataset<br/>requireOwnedResource"]
OWN --> MAP["DatasetProviderMap[type]"]
MAP --> SR["readSurveyResponsesDataset"] --> AT[("SurveyResponseEntity<br/>Azure Table")]
MAP --> FR["readSheetDataset"] --> BLOB[("Sheet content blob")]
MAP --> PR["readProgramStatusDataset"] --> PT[("ProgramParticipantEntity<br/>Azure Table")]
Server structure (server/services/dataset/): DatasetProviderMap.ts maps DatasetProviderType → the provider and the resource type it reads, one provider per folder. readDataset is the only way into a provider: it resolves the reference to a resource the caller owns, of that type, before the provider runs — so a provider is handed an owned resource and has no auth check of its own to forget. Each provider owns its column/row derivation:
readSurveyResponsesDataset— columns from the survey model's questions (name + question-type →ColumnTypemapping); rows flattened fromSurveyResponseEntityJSON in Azure Table, non-primitive answers JSON-stringified.readSheetDataset— reads the Sheet resource's content blob and convertscontent.dataviadataSourceToDataset.readProgramStatusDataset— participant/addedAt/responded rows fromProgramParticipantEntityin Azure Table, keyed by the non-secretpublicId(never the token — dashboards bake dataset snapshots into public publishes).
Rules
- Row cap —
AZURE_MAX_PAGE_SIZE(1000) on every provider; datasets are for visualization and import, not bulk export. Truncation istotalRows > rows.length, never the presence oftotalRows— a provider that can count cheaply always reports the uncapped total, which equals the row count on an uncapped read.getDatasetTruncationis the one place that comparison lives, so a warning can never disagree with the rows on screen; a provider that cannot count omits the field and its consumers simply never warn, because truncation is then unknowable rather than absent. Add pagination only when a real consumer hits the cap (deferred). - Consumers choose copy or reference. Import (Sheet resource) copies rows once. Binding (dashboard visuals, email editor merge fields) stores the
DatasetReferenceand re-resolves on load. All call the same procedure. - Fetch on load + manual refresh. No live subscriptions through this layer (deferred).
- No external providers (HTTP APIs, SQL) until secret storage and injection-safety work is scoped (deferred) — the enum grows one value per new provider, nothing else changes.