An incremental model is useful only if it notices the inputs that can change its answer. In marketing reporting, changes may arrive through historical backfills, account reconnects, new attribution settings, or corrections to existing source rows.
A filter on the last few calendar days does not necessarily capture those changes.
I have worked on changed-date refresh and transformation architecture in BigQuery, including a Dataform implementation developed in parallel with an existing production system. This article describes design patterns from that work, not a claim that the whole migration has completed production cutover.
Collection and recomputation solve different problems
API recovery retrieves missing inputs. Recomputation rebuilds outputs whose assumptions or source data changed.
A source table can be complete while a mart remains wrong because it was materialized before the correction, uses an old attribution setting, or depends on an intermediate model that has not refreshed.
Dataform provides dependency-aware transformations and incremental table configuration. Those features still require you to define which existing output rows should change. See the Dataform table documentation.
Identify changed history explicitly
Consider this synthetic example:
| Input change | Affected reporting dates |
|---|---|
| A reconnect loads corrected history | Two dates from earlier months |
| An attribution setting changes | The agreed recomputation range |
| A recent daily sync completes | The dates whose source rows changed |
A collection timestamp or append-only input-change record can identify historical dates changed since the last successful refresh. A settings event needs its own scope because a configuration change may not update raw row timestamps.
The dirty scope may include account, date, model family, or settings version. Avoid treating the entire warehouse as one success/failure checkpoint.
Keep exact dates and partition bounds separate
A sparse set of changed dates determines correctness. Bounds around those dates can help limit the partitions a query reads.
Conceptually:
-- Schematic filter; names are synthetic.
WHERE report_date BETWEEN dirty_date_min AND dirty_date_max
AND report_date IN UNNEST(dirty_dates)The membership condition prevents rebuilding unrelated dates inside a wide range. Whether a particular bounds expression enables pruning needs verification against the generated SQL and actual execution. Do not assume a dynamic filter performs like literal dates.
Google documents qualifying filters and pruning limitations in Query partitioned tables.
Choose replacement behavior for the model
An upsert is not sufficient for every recomputation. If recalculation makes an output row disappear, a merge that only inserts and updates may leave the old row behind.
For a model that owns a complete date slice, one option is to replace the affected slice after preparing and validating the new result. The following is a simplified SQL illustration, not a complete Dataform implementation:
-- Build the replacement from a defined input snapshot first.
CREATE TEMP TABLE replacement AS
SELECT * FROM prepared_output
WHERE report_date IN UNNEST(dirty_dates);
BEGIN TRANSACTION;
DELETE FROM reporting_output
WHERE report_date IN UNNEST(dirty_dates);
INSERT INTO reporting_output
SELECT * FROM replacement;
COMMIT TRANSACTION;This requires agreement on the model's ownership boundary, compatible schemas, and what concurrent jobs may change. BigQuery supports transactional DML; see Multi-statement transactions.
Advance checkpoints after successful output
A checkpoint describes work completed, not work attempted.
If a base model succeeds but a dependent rollup fails, the rollup must retain enough state to retry its own scope. Overlapping jobs also need coordination so an older snapshot cannot overwrite newer completed work.
During parallel migration, use refresh state isolated from production. Advancing production checkpoints from a comparison build could cause the live workflow to skip dates it has not processed.
Validate the change, not just the execution
For a bounded recomputation, compare expected scope, row grain, source totals, and affected output. Exercise interrupted runs, reruns, empty output, historical changes, and overlapping work.
The relevant proof is in the marketing-platform case and the separate Dataform reporting warehouse.
A focused first delivery can investigate one model that misses historical corrections and establish a validated refresh path.