DataData Architecture

lag_tolerance is a budget decision now, and most teams have not made it

dbt State went generally available in September 2026, turning "rebuild everything on a schedule" into "rebuild only what changed." The savings are real, but they only land if you decide, model by model, how stale your data is allowed to be.

A hand turns a brass valve on a pipe, slowing drips falling into a glass jar.

Open the query history of a mature warehouse on a Tuesday morning and a large share of what you are paying for produced a result identical to the previous run. The source tables had not refreshed. The SQL had not changed. The job ran because a cron expression said it should. According to the FinOps Foundation's 2026 State of FinOps survey, built on roughly a thousand practitioner responses covering more than $83 billion in annual cloud spend, workload optimization and waste reduction remain the top current priority, while practitioners report they have already taken the "big rocks" of waste and now face a high volume of smaller, harder opportunities. Flexera's 2026 State of the Cloud report, a vendor survey of more than 750 cloud decision-makers and so worth reading alongside independent sources, puts self-estimated waste at 29%, the first increase in five years, which it attributes to AI workloads and new services outpacing governance.

Redundant recomputation is one of those smaller opportunities, except it is not small. The argument of this piece: the next meaningful cut in your warehouse bill comes from making rebuilds conditional on change and on a declared freshness limit, not from another round of warehouse right-sizing. The tooling caught up in September 2026. Fivetran and dbt Labs announced the general availability of dbt v2 and dbt State at dbt Summit in Las Vegas on 16 September 2026, and dbt State now runs on Snowflake, BigQuery, Databricks and Redshift, locally, on the dbt platform and under Airflow, Dagster or GitHub Actions. What the tooling cannot do is decide how stale each of your models is allowed to be. That is the work.

How do you cut warehouse compute without serving stale data?

Stop rebuilding on a clock and start rebuilding on change, with an explicit staleness budget per model. dbt State compares each node's compiled SQL and its upstream data freshness against the last good build, then builds, skips, clones or defers. dbt Labs, which sells the feature, reports that early adopters averaged 15-30% compute savings and that more than 22 million models are built on dbt every day, most of them producing results identical to the run before. Treat both numbers as vendor figures and verify them against your own run history. Here is the sequence that makes them hold.

Step 1. Meter the rebuild, not the invoice

Before touching configuration, measure the share of your model builds that produce no change. On BigQuery, query INFORMATION_SCHEMA.JOBS filtered on total_bytes_processed to find recurring jobs with no partition filter and the scheduled transformations that dominate your scan volume. On Snowflake, pull warehouse credit consumption split by job and by hour. You are looking for one number: credits or bytes spent on scheduled transformation runs where no upstream source changed since the previous run. That number is your addressable budget. Everything else in this playbook is an attempt to move it.

Step 2. Give every model a freshness promise

dbt State's lag_tolerance config sets how long it waits before rebuilding a node once its upstream data has changed, and per dbt's documentation a node rebuilds only when its last build is older than that window and its upstream data has actually changed. That parameter is a business decision wearing an engineering costume.

The rule to put on the whiteboard:no scheduled rebuild without a decision waiting for it. For each model, name the decision it feeds and how often someone (or some system) acts on it. A finance close model acted on monthly does not need hourly rebuilds. A fraud screening feature table acted on continuously does. Set lag_tolerance to the decision cadence, never shorter. Models where nobody can name the decision are your first candidates for a long tolerance window or deletion.

Step 3. Run it in audit mode before you trust it

dbt shipped an Explain tab on run details ahead of the Summit, showing why each resource was rebuilt, reused or cloned, with a matching `dbt state explain` command in the CLI. Enable dbt State on a non-critical deployment job, leave tolerances conservative, and spend a week reading the explanations. You are checking one thing: that reuse decisions match what your analysts would have said about those tables. If dbt State reused a model your controller expected to be fresh, you have found a freshness promise nobody wrote down.

Step 4. Fix physical layout on the models that still rebuild

State-awareness removes work that should never have run. It does nothing for the work that must run. That is where the older levers still pay: on BigQuery, partition large fact tables on a date column, set require_partition_filter to true so unfiltered queries fail instead of scanning the table, and add two to four clustering keys ordered by filter frequency. Google's own guidance favours clustering over partitioning when cardinality is very high or partitions would be tiny. The arithmetic is unforgiving: on a 10 TB table with a year of daily partitions, a single-day query scans tens of gigabytes rather than the full table. If your team has never auditedthe partition and cluster keys on your heaviest tables, that audit will outperform most procurement negotiations.

Step 5. Send the saved credits somewhere a person owns

Compute reductions evaporate if nobody is accountable for them. The 2026 State of FinOps found that 78% of FinOps teams now report into the CTO or CIO rather than finance, and that pre-deployment architecture costing is a top desired tool capability. Put the recovered credits on a named budget line, with a monthly review of credits per domain and per job. This is wherecost discipline with named owners stops being a slide and becomes a reporting cadence.

Step 6. Put a ceiling on agent-driven queries

Scheduled pipelines are now the predictable part of your bill. Autonomous agents are not: they decide what to query based on context, fan out in parallel, and retry on failure with broader scope, all without a human watching the credit meter. Resource monitors on Snowflake with suspend triggers, and maximumBytesBilled on BigQuery jobs, cost nothing to set and cap the damage from a loop that nobody notices at 3am.

What the numbers look like in practice

Two customer results were published with the GA announcement, both supplied by dbt Labs and its customers, so read them as vendor-sourced. Gordon Curzon, Head of Analytics Engineering at Virgin Media O2, reported 25% savings on both job run time and BigQuery compute costs. Chris Shepherd, Principal Data Engineer at RxBenefits, reported a 59% reduction in warehouse costs on scheduled jobs running on a Snowflake adaptive warehouse, worth $8,173.23 in the first 60 days, with more than 700,000 models reused instead of rebuilt and two weeks of query run time removed over that period.

Run the arithmetic on the RxBenefits figures and the shape becomes clear: a 59% cut worth $8,173 implies a baseline of roughly $13,850 of scheduled-job compute over 60 days, about $231 a day, with an annualised saving near $49,000 if the rate holds. That is one team, on one warehouse, for a mid-size pipeline. Scale it against your own baseline from step 1 rather than against theirs.

Where state-aware pipelines quietly go wrong

Reuse suppresses tests. Per dbt's documentation, tests do not run for reused models, on the logic that they already passed. That is defensible until a source system silently changes semantics without changing volume. dbt added two guards earlier in 2026: a model that fails a data test gets rebuilt on subsequent runs instead of being reused from prior state, and models whose tables are deleted from the warehouse get rebuilt even with no code or data change. Keep source freshness checks running on their own schedule regardless, because they are the only thing standing between reuse and silent staleness.

Committed spend hides your savings. If you are on a Snowflake capacity contract, which typically reduces per-credit rates by 15-40% according to vendor pricing guides, a 30% drop in consumption does not show up as a smaller invoice. It shows up as under-consumption against a commitment you already paid for. Bank the efficiency now, then take the measured consumption curve into your renewal conversation.

The feature itself is a line item. dbt Labs publishes two distributions of dbt v2, and the superset distribution is free with optional paid features, dbt State among them. Net the licence cost against the compute saving before you present a number to your CFO, and note that engine GA is not adapter GA: Snowflake, BigQuery, Databricks, Redshift and DuckDB are GA on v2 while Postgres, Athena and Fabric are still pending, gated on ADBC driver availability rather than on Rust.

Migration carries a config trap. Teams that used the earlier state-aware orchestration preview need to map old build_after values to the new lag_tolerance semantics. A mis-mapped tolerance does not fail loudly. It serves yesterday's numbers to a dashboard that looks fine.

Finally, do not sell this internally as the end of the cost programme. The 2026 FinOps survey is explicit that optimization is now table stakes rather than a differentiator, with scope expansion, governance, organizational alignment and forecasting collectively outweighing optimization alone. A one-time 20% compute cut buys you credibility to do the harder work, not permission to stop.

What to do on monday

  • Pull last month's scheduled transformation runs and calculate what share rebuilt models whose upstream sources had not changed. One query, one number, your baseline.
  • Pick the ten most expensive models in that list and write a decision cadence next to each. Any model without a named decision gets flagged for deletion review.
  • Enable dbt State on one non-production deploy job with conservative tolerances, then read the Explain output daily for a week before widening.
  • Set require_partition_filter on your three largest BigQuery fact tables, or resource monitors with suspend triggers on your two largest Snowflake warehouses.
  • Ask your account team what happens to your committed spend if consumption falls 25%, and get the answer in writing before renewal season.

The models worth arguing about are the ones running hourly to feed a weekly decision. Find them this week, set a tolerance that matches the decision rather than the habit, and the compute saving follows without anyone noticing a difference in the dashboards. The part that takes longer is getting someone to own the budget line where the savings land.

Frequently asked questions

Do I need the dbt platform to use dbt State, or does it work with my own orchestrator?

dbt State went generally available in September 2026 everywhere dbt runs, including Airflow, Dagster, GitHub Actions and a local machine, with support for Snowflake, BigQuery, Databricks and Redshift. It is native in dbt v2 and available as a plugin for dbt v1.7 through v1.12, so a self-managed deployment is supported.

Do data quality tests still run when a model is reused instead of rebuilt?

No. Tests do not run for reused models, since they passed on the previous build. dbt added guards in 2026 so that a model failing a data test is rebuilt on later runs rather than reused, and models whose warehouse tables are deleted get rebuilt anyway. Keep independent source freshness monitoring running regardless.

How much compute can a team realistically save by skipping unchanged rebuilds?

dbt Labs, which sells dbt State, reports an average 15-30% reduction in warehouse compute among early adopters, with Virgin Media O2 citing 25% on BigQuery and RxBenefits citing 59% on scheduled jobs running on a Snowflake adaptive warehouse. These are vendor-published customer figures, so benchmark them against your own run history first.

What is lag_tolerance in dbt State and how should I set it?

lag_tolerance is the model-level freshness threshold that controls how long dbt State waits before rebuilding a node once its upstream data changes. Set it to the cadence of the decision the model feeds, never shorter: a monthly close model does not justify hourly rebuilds, while a continuously consumed feature table does.

Go deeper

The lessons that take this article further, free to read.

  1. 1Data FinOps: controlling cloud data costModern data architecture
  2. 2dbt (data build tool): industrialized SQL transformationModern data architecture
  3. 3Performance & scalability: partitioning, clustering & cost optimizationModern data architecture
  4. 4Cloud data infrastructure: services, costs & migrationModern data architecture
  5. 5Chargeback & showback modelsData products & monetization

Finished reading?

Validate your read to earn XP and feed your radar.