If agents are the new primary consumer of your data, is your infrastructure built for the wrong audience?
At dbt Summit 2026, Fivetran and dbt Labs announced a cluster of new products designed to make enterprise data consumable by AI agents rather than human analysts. CDOs need to separate the genuine architectural shift from the vendor positioning.
Claude VectorData & Analytics LeadSeptember 23, 2026Listen to the podcast
4 min
Chapters
Key takeaways
- Pull last month's query logs and calculate what share came from service accounts rather than named humans to find your real agent load.
- Treat dbt Labs' 30% machine-generated query figure as directional, since the vendor sells the tooling that raises it.
- If agent traffic is small, spend the quarter writing business definitions down instead of changing architecture.
- Build the semantic layer so terms like active customer are defined in writing, because an agent has no memory to fill the gaps.
- Set a spending cap on agent query volume before a pilot touches production data to avoid a five-figure compute surprise.
Read the full transcript
Host:This is Leaders Insights. On the table. If agents are the new primary consumer of your data, is your infrastructure built for the wrong audience?
Expert:The mistake I keep seeing this month is CDOs rebuilding their entire stack because a keynote told them agents are the new users.
Host:Panic-driven architecture. And the keynote was loud.
Expert:At DBT Summit last week, DBT Labs and 5tran rolled out a cluster of products all pointed at the same idea. That the thing querying your data warehouse is no longer a human analyst but an AI agent asking questions on its own.
Host:Translate that. What actually changes when the consumer is a bot instead of a person?
Expert:A human analyst runs maybe 40 queries a day and gets bored. An agent will fire off thousands, chain them together, and never sanity check a nonsense result, so the failure mode flips. A person notices when revenue comes back as negative. An agent just passes it downstream to the next agent.
Host:So this is real, not just a slide.
Expert:The shift is real. The urgency is being sold to you. Those are different things.
Host:Give me a number that tells me which is which.
Expert:DBT Labs put out a figure that something like 30% of the queries hitting their customers' warehouses are now machine-generated rather than human. Worth flagging, they sell the tooling that makes that number go up. So treat it as directional, not gospel. Cross-check against your own query logs before you believe your stack is under siege.
Host:And if I check my logs and it's 3% not 30?
Expert:Then you don't touch your architecture this quarter. You watch the trend line. The independent read matters here. O'Reilly's radar surveys through 2026 show agent adoption climbing fast in engineering but still thin in the data layer specifically. Plenty of companies are experimenting. Far fewer are in production.
Host:Here's the uncomfortable one. Vendors always say the audience changed so you need to buy the new thing. Why is this any different from the last five waves?
Expert:It mostly isn't, and that's the honest answer. But there's one piece that isn't marketing. The metadata problem. Metadata is just the labels on your data. What this column means. Where it came from. Whether it's trustworthy. A human analyst fills those gaps from memory. They know the revenue table excludes refunds because they've been burned before. An agent has no memory and no scar tissue. So if your definitions live in someone's head instead of written down, the agent guesses and it guesses confidently.
Host:So the real work isn't buying an agent-ready warehouse. It's documentation.
Expert:It's the semantic layer. The written down business definitions that say active customer means logged in within 30 days, not ever. MIT Sloan had a piece this year making exactly this point. The organizations getting value from agents aren't the ones with the fanciest models. They're the ones who did the boring work of defining their terms first. That's not a product you purchase. It's a project.
Host:Second number. What's the cost angle because thousands of automated queries can't be free.
Expert:That's the one nobody put on a slide. If agents run 10 to 100 times the query volume of humans, your cloud compute bill scales with it. These are walking into pilots with no cap on how much an agent can spend and getting a five-figure surprise. So before you let an agent loose on production data, you set a query budget the same way you'd set a spending limit on a corporate card.
Host:You've now told people not to rebuild, not to trust the vendor number, and to write documentation. That's a lot of not doing.
Expert:Because the expensive mistake is motion. Reacting to a keynote by re-platforming cost you a year and a budget. The cheap, useful move is small.
Host:Then land it. One thing my listener does Monday morning.
Expert:Pull last month's query logs and calculate what share came from a service account instead of a named human. That's your real agent load, not dbt's 30% yours. If it's small, you've got time to write your definitions properly. If it's already large and undocumented, you've found your fire. And it isn't the one the vendors were pointing at. Measure your own audience before you rebuild for someone else's.
Host:That'll do it. What we read for this one?
Expert:The new stack, dbt labs, vendor, data tooling, O'Reilly radar, MIT Sloan Management Review.
Host:End of episode. The CDO calculators are running at MBA-training.com.
This week's announcements from dbt Summit 2026 share a single premise: the primary consumer of enterprise data is changing, and the infrastructure built to serve human analysts is increasingly inadequate for AI agentsAI agentsAgentic AI refers to AI systems that pursue goals autonomously by planning, taking actions through tools, and adapting based on results, with minimal step-by-step human direction.View full definition → that query, reason, and act at machine speed. Whether that premise becomes a durable architectural reality or remains a marketing frame depends on decisions CDOs are making right now about metadata, transformation logic, and what "ready" actually means.
What shipped in dbt Core v2 and dbt State?
According to dbt Labs (a vendor with a direct commercial interest in the adoption of these tools), dbt Core v2 and dbt State both reached general availability at dbt Summit 2026. dbt State is the more practically significant of the two for most teams today. It addresses a straightforward and expensive problem: in standard dbt runs, every model rebuilds from scratch even when the underlying data has not changed. dbt State tracks what has changed between runs and rebuilds only what needs to be rebuilt.
dbt Labs claims that Fanatics, the sports merchandise company, used dbt State to cut warehouse compute materially by eliminating unnecessary rebuilds. The company presents this as a compute cost and development speed story. It is worth treating that case study as a directional signal rather than a benchmark: vendor-published case studies select for positive outcomes.
For CDOs, the practical question is not whether dbt State works in principle (incremental computation is well-established) but whether their current dbt estate is structured in a way that makes state tracking reliable. Teams with poorly documented dependencies or inconsistent model tagging will not capture the full benefit without remediation work first.
The dbt v2 release also formalises the direction of travel forindustrialised SQLSQLSales Qualified Lead: a prospect the sales team has validated as ready for direct outreach and a proposal, having passed clear qualification criteria.View full definition → transformation at enterprise scale: the tool is explicitly positioningpositioningThe mental space you want your brand to occupy in your target customer's mind relative to alternatives.View full definition → itself as the definition layer for what data means, separate from where it is stored or computed.
What does Fivetran's Context Layer mean by agent-ready data?
The headline announcement, from Fivetran and dbt Labs jointly (both vendors, both with commercial stakes in this framing), is the Fivetran Context Layer. The claim is that it makes enterprise data "agent-ready" by providing AI agents with structured context about data assets: what a table represents, how it was built, what business rules govern it, and how it relates to other assets.
This matters because agents querying a raw warehouse without that context will produce unreliable outputs. A model that sees a column named "revenue" without knowing whether it is recognised, billed, or collected revenue will hallucinate a definition. The Context Layer is presented as the mechanism that closes that gap.
Take the commercial framing seriously here. Fivetran and dbt Labs have a shared incentive to position their joint stack as the agent-readiness standard. The architectural concept they are describing, giving agents access to rich metadata and semantic context rather than just raw tables, is sound and increasingly discussed across the industry. But the specific implementation is one vendor's answer to the problem, not a settled industry standard.
The more durable insight is aboutactive metadata and what a data catalogdata catalogA centralized inventory of an organization's data assets, enriched with metadata, that helps people find, understand, and trust the data they need.View full definition → must do when agents, not analysts, are the audience. Agents need metadata that is machine-readable, always current, and attached to the data at the point of query, not buried in a separate catalog that humans consult occasionally. CDOs who have not yet addressed that gap should not wait for a single vendor to solve it for them.
dbt Charts and the open lakehouselakehouseA hybrid architecture combining the flexibility of a data lake with the analytical capabilities of a data warehouse, on a single storage layer.View full definition → vision are worth watching, not adopting
dbt Labs also announced dbt Charts, a lightweight visualisation capability, alongside what it describes as an "open lakehouse vision." Both are early-stage claims from a single vendor. dbt Charts addresses a real friction point: teams that define metrics in dbt today still need a separate tool to surface them, and that handoff creates inconsistency. Whether a visualisation layer built inside a transformation tool displaces existing BIBITechnologies and processes that turn raw data into actionable insights via reporting, dashboards and analysis, so teams can decide based on facts rather than intuition.View full definition → investments is a question that will take quarters to answer.
The open lakehouse framing is more strategic than concrete. dbt Labs is signalling that it intends to compete at the storage and architecture layer, not just transformation. CDOs with existing Databricks or Snowflake commitments should note this as a positioning move and watch how it develops, particularly the retirement of the dbt Snowflake Native App announced for November 2026, which is a concrete change requiring attention for any team running that configuration.
Modeling for agents is a tokentokenA token is the basic unit of text that language models process, often a word fragment, whole word, or punctuation mark rather than a single character.View full definition → economy problem
The development most people will underrate this week is a short piece from dbt Labs describing how one team cut token costs by twenty times by modeling Gong call transcripts in the warehouse before feeding them to an AI, rather than pushing raw transcripts to the APIAPIApplication Programming Interface: a standardised interface that lets applications communicate and exchange data without knowing each other's internal workings.View full definition →. The reduction came from restructuring the data so the agent received only what it needed, in a format it could process efficiently.
This is a different design discipline than traditional data modeling. Analysts tolerate wide tables, redundant columns, and verbose schemas because they are good at filtering. Agents are not. They consume tokenstokensA token is the basic unit of text that language models process, often a word fragment, whole word, or punctuation mark rather than a single character.View full definition → and their cost scales with the verbosity of what they receive. The implication for CDOs is that data products built for agent consumption need explicit modeling decisions around token economy, not just semantic clarity. Teams that treat agent-readiness as a metadata labeling exercise will hit this problem when they move from prototype to production workloads.
The O'Reilly Radar framing of "intelligent data orchestration" as the successor to dashboards points in the same direction: the question shifts from "can a human read this?" to "can a machine reason from this reliably and cheaply?"
The practical priority for most CDOs this quarter is not adopting any specific tool from this week's announcements. It is auditing whether their current transformation and metadata layers would give an agent enough context to produce a trustworthy answer, and being honest about the gap between where they are and what that requires.
Go deeper
The lessons that take this article further, free to read.
- 1dbt (data build tool): industrialized SQL transformationModern data architecture
- 2Active metadata and the data catalogModern data architecture
- 3The metrics & semantic layerAnalytics, BI & decision intelligence
- 4Data lineage & metadata management: knowing where your data was bornData governance & compliance
- 5Data products: definition, design & lifecycle managementModern data architecture
Sources
- AI coding agents need a secrets-safe context boundary
- Fivetran + dbt Labs Announces New Capabilities to Make Enterprise Data Agent-Ready at dbt Summit 2026
- Everything we announced at dbt Summit and why it matters
- We built dbt State to stop rebuilding what hadn't changed
- Celebrating the 2026 dbt partner of the year winners
- Enterprise Analytics Beyond Dashboards: Intelligent Data Orchestration with LLMs
- Spot New Tech Skills Emerging From the Workforce
- Building on AI’s Unfinished Foundation
- Databricks processes your data. dbt defines what it means
- dbt Core v1.12 is GA
- Model for the token, not the table
- How dbt State cuts warehouse compute and speeds up every run
- dbt Summit 2026: the keynotes and product sessions
- Retiring the dbt Snowflake Native App
Finished reading?
Validate your read to earn XP and feed your radar.