Foundations & core concepts of digital analytics
Open any analytics report and three numbers sit side by side: users, sessions, events. They describe the same traffic, they never match the numbers in the finance system, and almost every argument about "the analytics being wrong" turns out to be an argument about what those three words mean. The object at the centre of digital analytics is built, not found. Someone decided what counts as an event, when a session expires, and which identifier stands in for a person. This lesson defines that object, then shows where the data physically comes from: the browser, or your own servers.
What digital analytics actually is
Digital analytics is the collection and interpretation of records of interaction with digital products and channels: websites, apps, connected devices, emails, ad units. What separates it from the rest of your data estate is the unit of record. A CRMCRMCustomer Relationship Management: software and strategy to manage and analyse customer interactions throughout their lifecycle.View full definition → stores rows about customers. A ledger stores rows about transactions. Digital analytics stores a time-ordered stream of small behavioural facts, each stamped with a moment, an identity and a context, and then aggregates them upward on demand.
The vocabulary, stated once and precisely:
- Event: one recorded interaction. A name plus a set of parameters (Adobe's Web SDK calls them XDM fields; other tools say properties).
add_to_cartwithitem_id,price,currencyis an event. - Session, called a visit in Adobe Analytics: a group of events from the same identity, closed by an inactivity timeout, conventionally 30 minutes.
- User, called a visitor: an identifier, not a human being. Usually a cookie or device ID, sometimes a login ID.
- Dimension: an attribute you slice by. Country, device type, traffic source, page name.
- Metric: a number you aggregate. Events, sessions, revenue, average order value.
- Hit: the older primitive, a single call to the collection server. Adobe Analytics counted page views and custom link calls this way for two decades; the event-based model absorbed it.
Sub-concept 1: the event, and what hangs off it
An event has three parts: what happened (the name), the details (parameters), and who and when (identity and timestamp). Everything else in analytics is arithmetic performed on that.
Take a streaming service such as Spotify. A play is an event. Its parameters carry the track, whether the play came from search, a saved playlist or an algorithmic recommendation, the device type, and how far into the track the listener got before skipping. None of those are separate reports. They are fields on one record, and the reports are built by counting and grouping them afterwards.
This is why event design decides what questions you can ever answer. If the play event does not carry the source of the play, no dashboard, no matter how expensive, will tell you whether recommendations beat search. You cannot recover a parameter you never sent. The construction method for that schemaschemaA schema is the formal blueprint that defines how data is structured, named, typed, and related within a database, file, or message.View full definition → belongs to a later lesson; what matters here is understanding that the event is the atom and that its parameters are the limit of your future curiosity.
Sub-concept 2: sessions and users are conventions, not facts
A session does not exist in nature. It is a grouping rule. GA4 opens one with a session_start event and closes it after 30 minutes without activity; Adobe Analytics applies a comparable visit timeout, configurable per report suite. Universal Analytics used to start a fresh session whenever the campaign source changed mid-visit, which inflated session counts for anyone clicking through several ads. GA4 dropped that rule. Same behaviour, different number, because someone changed a definition.
"User" is looser still. What you are counting is an identifier: a first-party cookie in a browser, an app instance ID on a phone, or a login ID if the person authenticates. One person on a laptop, a phone and a smart speaker is three identifiers until a login stitches them together. A shared household account is the reverse: several people compressed into one user. Safari caps cookies written by JavaScript at seven days, so a returning visitor in month two can arrive as a brand new user.
Treat a user count as a count of identifiers under a stated stitching rule. Write the rule down. When two tools disagree by 20%, this is usually why.
Sub-concept 3: dimensions versus metrics, and scope
Dimensions describe, metrics measure. "Organic search" is a dimension value; "1,482 sessions" is a metric. Every report is one or more metrics broken down by one or more dimensions, which is why the pairing has to be legal.
The trap is scope. A dimension can be attached to an event (the page a click happened on), to a session (the channel that started the visit), or to a user (first acquisition source, country of registration). Metrics have scope too. Ask for bounce rate, a session-level metric, broken down by an event-level dimension, and the tool will return something. It will be nonsense.
Adobe made this distinction explicit long before most tools did, with props (traffic variables that describe the hit and do not persist) and eVars (conversion variables that persist for a set expiry and take credit for later success events). Same value in both, different attributionattributionA framework for assigning credit to the touchpoints that contributed to a conversion, so you can measure which channels and interactions actually drive results.View full definition → behaviour, different report. High-cardinality dimensions are the other hazard: put a raw URL with query strings into a dimension and you generate hundreds of thousands of unique values, most tools bucket the tail into "(other)", and your report quietly stops adding up.
Google Analytics 4 Tutorial for Beginners
Sub-concept 4: client-side and server-side collection
Every event reaches your analytics from one of two places.
Client-side collection runs in the user's browser or inside your mobile app: JavaScript or an SDK observes an interaction and sends a request to a collection endpoint. It sees things no server ever will: scroll depth, hovers, form abandonment, screen size, the referrer, whether the video actually played. It is also fragile. Ad blockers drop it, browser privacy limits shorten its identifiers, weak mobile connections lose beacons, and anyone can forge a request to the endpoint.
Server-side collection sends the event from infrastructure you control, usually your application backend, after the fact has been confirmed. Orders, refunds, subscription renewals, payment failures and fraud reversals belong here, because your server knows the truth and the browser only knows what it was told. Adobe (which sells an analytics suite, so read its guidance accordingly) supports both paths: a browser SDK for interaction data, and a server-side data insertion APIAPIApplication Programming Interface: a standardised interface that lets applications communicate and exchange data without knowing each other's internal workings.View full definition → for events that originate in your systems.
A third arrangement, server-side tagging, keeps the browser as the origin but routes the first hop to a container you run, which then fans the data out to vendors. It restores some resilience and gives you a checkpoint where consent and PII rules can be enforced before anything leaves.
Most serious stacks run both sides and then have to reconcile them. Decide in advance which side is authoritative for each number. Revenue from the server. Interaction behaviour from the client. Never sum the two and call it a total.
Real-world cases
Spotify Wrapped, run each December since 2016, is the clearest public demonstration of what an event stream is. Nothing in it comes from a survey or a dashboard. It is a year of individual play events, each carrying a track, an artist, a timestamp and a context, replayed per listener and counted. The product is only possible because the atom was designed to hold enough detail.
Adobe's own trajectory shows the object being redefined underneath an industry. Adobe Analytics was built around the hit: a page view or a custom link call, described by props and eVars. Its later Web SDK sends XDM-formatted experience events instead, closer to the schema-first model the rest of the field converged on. Same vendor, same customers, a different unit of record, and every historical comparison had to be rebuilt around it.
CMO action items
- Publish a one-page definitions sheet: what closes a session, what an identifier is, how logged-in and anonymous activity are stitched, what "active user" means in your company. Circulate it to finance and to your agencies. Most measurement disputes end there.
- For each of your top ten reported numbers, write down whether it is collected client-side or server-side. If revenue is coming from the browser, move it.
- Check the scope of the dimensions in your standard reports before you act on a breakdown. A session metric split by an event dimension will render happily and mislead you.
Common mistakes that kill results
- Reading "users" as "people". It is a count of identifiers under a stitching rule, and that rule shifts every time a browser vendor changes cookie policy.
- Comparing session counts across two tools with different timeout and campaign rules, then asking your team to explain the gap. There is no gap to explain, only two definitions.
- Assuming server-side data is neutral truth. It is blind to everything the interface did and only carries the identity you remembered to pass in.
- Letting raw URLs, search strings or IDs into dimensions unbucketed, then wondering why a large "(other)" row swallows the tail of every report.
Resources
- 🔗Google Analytics 4 Help Center
Official documentation for GA4 including event setup, conversion tracking, and attribution model configuration — the most authoritative reference for implementation questions.
- 🔗Measure What Matters by John Doerr — OKR Framework
Free resources from John Doerr's OKR methodology that provide the strategic framework for connecting business objectives to the metrics your analytics infrastructure should be measuring.
Related articles
Recent articles from the blog that build on this lesson.