+150 XP

Mapping the SaaS data landscape: sources, systems, and owners

A single free-trial signup at a mid-size SaaS company can trigger data writes to six different systems in under ten seconds: a product analytics event, a CRM lead record, a billing account shell, a marketing attribution touch, a support ticket eligibility flag, and a data warehouse sync job. None of these systems agree on what to call the customer. This is the SaaS data landscape, and mapping it is the first job of anyone doing analytics, finance, or ops in the sector.

Why source mapping matters before you touch a dashboard

Most SaaS "data problems" are not analytics problems. They are ownership problems. A churn number looks wrong not because the SQL is bad, but because billing defines "customer" as an active subscription while product defines it as a logged-in workspace.

A source map is a simple artifact: for each dataset, who owns it, where it lives, what it's used for, and what breaks it. Before building metrics, build this map. It prevents the single most common failure mode in SaaS analytics: two teams presenting different numbers for "the same" metric in the same board meeting.

The five core data sources in SaaS

1. Product telemetry (usage data)

This is event-level data: logins, clicks, feature adoption, API calls, session length. Captured via instrumentation tools like Amplitude, Mixpanel, or Segment (a customer data platform, CDP, that routes events to multiple destinations).

  • Owner: typically Product or Data Engineering.
  • Typical breakage: instrumentation drift. A developer renames an event ("signup_complete" becomes "user_created") and three months of trend lines silently break. No error is thrown, the data just quietly means something different.
  • Grain: usually event-level, timestamped, tied to a user ID or account ID.

2. Billing and subscription data

Lives in systems like Stripe, Chargebee, or Zuora. This is the system of record for MRR (monthly recurring revenue), plan tier, seat count, discounts, and payment status.

  • Owner: Finance, sometimes RevOps.
  • Typical breakage: mismatched customer identity across billing and CRM (a company renamed in Salesforce but not in Stripe), and manual discounting that doesn't flow back into reporting systems.
  • Grain: account/subscription level, usually monthly or event-triggered (upgrade, downgrade, cancellation).

3. CRM data (customer relationship management)

Salesforce or HubSpot records: leads, opportunities, deal stages, contract terms, renewal dates. This is the sales-side narrative of the customer.

  • Owner: Sales Ops / RevOps.
  • Typical breakage: stale or manually-entered fields. A rep forgets to update deal stage; forecasts built on CRM data become unreliable. Also duplicate account records when multiple reps create separate entries for the same company.

4. Support and success data

Zendesk, Intercom, or Gainsight tickets, NPS (Net Promoter Score) surveys, health scores. This captures customer sentiment and friction.

  • Owner: Customer Success / Support.
  • Typical breakage: health scores built on stale usage snapshots, or ticket volume conflated with dissatisfaction (a power user files many tickets because they use the product heavily, not because they're unhappy).

5. Marketing and acquisition data

Ad platforms (Google Ads, LinkedIn), web analytics (GA4), and attribution tools. Captures CAC (customer acquisition cost) inputs: spend, channel, campaign, conversion.

  • Owner: Marketing / Growth.
  • Typical breakage: attribution model disagreements (last-touch vs. multi-touch), and cookie/consent restrictions under GDPR (General Data Protection Regulation, the EU's data privacy law) or CCPA (California Consumer Privacy Act) that shrink trackable traffic, especially in Europe post-2018 and increasingly in the US since 2020.

Building the source map: a practical template

A usable source map has five columns. Here's a compressed example:

DatasetSystemOwnerGrainCommon failure
Product eventsAmplitude/SegmentProduct EngEvent-levelRenamed/untracked events
SubscriptionsStripe/ChargebeeFinanceAccount-levelID mismatch with CRM
Deals/accountsSalesforceRevOpsAccount-levelManual entry, duplicates
Tickets/healthZendesk/GainsightCSTicket-levelStale scores, sentiment conflation
Campaigns/spendGA4/Ad platformsMarketingSession/campaignAttribution model conflict, consent gaps

The critical column most teams skip is "common failure." Naming the failure mode in advance is what lets you build monitoring for it, rather than discovering it during a board review.

The identity resolution problem

The reason these five systems don't naturally agree: each uses a different key.

  • Product telemetry keys on a device or user ID.
  • Billing keys on an account/subscription ID.
  • CRM keys on a company/contact record.

Reconciling these is called identity resolution, and it's usually done through a customer ID mapping table maintained in the data warehouse (Snowflake, BigQuery, Databricks are the common platforms in 2026).

A simplified mapping query looks like this:

sql
-- Simplified identity resolution: join product usage to billing account
SELECT
    p.user_id,
    p.account_id AS product_account_id,
    b.subscription_id,
    b.crm_account_id,
    c.salesforce_account_name
FROM product_events p
LEFT JOIN billing_accounts b
    ON p.account_id = b.external_account_id
LEFT JOIN crm_accounts c
    ON b.crm_account_id = c.account_id
WHERE p.event_date >= CURRENT_DATE - INTERVAL '30 days';

If product_account_id and crm_account_id don't reliably map 1:1, every usage-based churn or expansion metric downstream is suspect. This join failing silently is one of the most common root causes of "the dashboard numbers don't match" incidents in SaaS companies.

For a deeper technical reference on building this kind of pipeline discipline, the dbt Labs glossary is a solid, free primer on analytics engineering concepts referenced throughout this module.

Knowledge check

1. A churn metric looks different between the billing team and the product team. According to the lesson, what is the most likely root cause?

2. Why should a team build a source map before building metrics or dashboards?

3. A developer renames the event 'signup_complete' to 'user_created' in the product analytics tool, and no error is thrown. What does this scenario illustrate?

MULTIPLE CHOICE

4. Select ALL correct answers about a 'source map' as described in the lesson.

Select all the correct answers.

MULTIPLE CHOICE

5. Select ALL correct answers about product telemetry (usage data) in SaaS companies.

Select all the correct answers.

Who owns what: a governance reality check

Ownership disputes are the norm, not the exception, in SaaS data. A few patterns worth knowing:

  • Product-led growth (PLG) companies (Notion, Figma-style motion) tend to have Product own the most authoritative customer definition, since usage drives the business. Billing is often downstream and thinner.
  • Sales-led enterprise SaaS (traditional Salesforce-style motion) tends to have CRM as the "source of truth" for account identity, with product telemetry treated as supplementary.
  • Data governance councils, common at scale (Series C+ or public companies), formalize this by assigning a data steward per domain: a named person accountable for definitions, not just data quality tooling.

Governance frameworks worth knowing by name: the DAMA-DMBOK (Data Management Body of Knowledge) is the standard reference framework for data governance roles and responsibilities, useful if you want a formal vocabulary for these ownership conversations.

🎬 [VIDEO: "Data Governance Explained" - youtube.com - search for recent DAMA or Data Council talks explaining stewardship and ownership models in modern SaaS data stacks]

Where this breaks in practice: three real patterns

  1. The renamed event. Engineering ships a refactor, renames trial_started to trial_activated. Growth's activation dashboard flatlines. Nobody notices for two weeks because no alert was tied to that event's volume.
  1. The billing/CRM identity drift. A customer is acquired by another company, renamed in Salesforce, but the Stripe account still shows the old legal name. Revenue reporting by "customer" undercounts the true logo count.
  1. The attribution war. Marketing reports CAC using last-touch attribution; Finance reports CAC using fully-loaded spend divided by new logos. Both are "correct" by their own definition. Without a documented source map, this becomes a recurring, unproductive argument rather than a five-minute reconciliation.

Key Takeaways

  • SaaS data comes from five core systems: product telemetry, billing, CRM, support/success, and marketing. Each has a distinct owner, grain, and typical failure mode.
  • Build a source map (dataset, system, owner, grain, common failure) before building metrics. It's the fastest way to prevent conflicting numbers in leadership meetings.
  • Identity resolution, reconciling different ID schemes across systems, is the single most common root cause of "the numbers don't match" incidents.
  • Ownership follows business model: PLG companies anchor on product data, sales-led companies anchor on CRM data. Know which model you're in before deciding whose number is authoritative.
  • Governance frameworks like DAMA-DMBOK give you formal vocabulary (data steward, system of record) to resolve ownership disputes rather than re-litigate them each quarter.