Attribution after iOS 14: what actually works.
The privacy-first environment is no longer a disruption to recover from, it is the permanent condition. The harder question is which of your four conflicting numbers to trust.
People still search for this by the name of the event, which is understandable and slightly misleading. iOS 14 shipped in 2020. Treating it as the disruption implies there is an after, a settling, a version of measurement that eventually comes back. There is not. What that release started is now simply the operating environment, and the useful question has moved on from what broke to something harder: when four systems give you four different numbers for the same week, which one do you spend against.
That is the question this piece answers, because it is the one that actually costs money. Most brands we meet are not confused about whether tracking degraded. They know. They are confused about which of the numbers in front of them is a measurement, which is a claim, and which is a model, and they are making budget decisions across all three as though the words were interchangeable.
The rest of this is the stack we build, the four kinds of number it produces, and the order to build it in. The order matters more than any individual component, and it is the thing most brands get backwards.
What actually changed, stated plainly
For roughly a decade, paid media measurement was close to free. A pixel saw nearly every conversion, last-click attribution was crude but precise, and automated bidding worked well because the signal reaching it was clean. That era ended through an accumulation rather than a single event: app-level tracking permission, browser tracking prevention, consent requirements in Europe and Brazil, third-party cookie deprecation, and platforms increasingly filling the resulting gaps with modeling rather than observation.
The observable effect is that the pixel now sees a partial and non-random slice of reality. Across the consumer accounts we have rebuilt, browser-only capture typically lands somewhere between half and four-fifths of true conversions, varying with browser mix, geography, and how the consent banner is designed. That range is our own observation across the accounts we work on rather than a published industry figure, and any agency quoting a precise number for your business before looking at it is quoting somebody else’s account.
The less obvious effect, and the more consequential one, is that platforms stopped merely reporting and started estimating. A meaningful share of the conversions in your Google Ads or Meta interface today were not observed. They were modeled from patterns across consenting users and applied to non-consenting ones. Those are legitimate estimates and they are frequently better than the alternative of counting nothing. They are not observations, and a dashboard that displays both in the same column with the same styling is asking you to forget the difference.
Four numbers, four different meanings
This is the section that changes how people run their reporting, so it is worth being pedantic. When you ask what drove revenue last month, four separate kinds of answer exist, and they are not competing estimates of the same quantity. They answer different questions.
- Platform-reported conversions. A claim, made by an interested party, using its own attribution window and its own rules for what counts as a touch. Useful for judging one campaign against another inside the same platform. Not additive across platforms, and not a measurement of contribution.
- Observed correlation in your own data. What your first-party events actually recorded: this visitor arrived from this source, then purchased. Genuinely observed, and silent on causation. It tells you what happened alongside the sale, not what caused it.
- Modeled contribution. The output of an attribution model applied to your observed events, distributing credit across touches by a rule you chose. Internally consistent, defensible, and dependent on the rule. Change from last-click to position-based and the answer changes without any underlying fact changing.
- Experimentally measured incrementality. What happens to revenue when you turn something off for a comparable group and leave it on for another. The only one of the four that answers what advertising caused. Expensive, slow, and the only one that can overturn the other three.
Almost every serious budget mistake we are asked to diagnose comes from treating one of these four as another. The most common is reading modeled contribution as measured causation, then defending a channel to finance on the strength of a rule somebody picked in a settings menu.
Precision about which number you are holding is not pedantry, it is what makes a channel decision defensible. Two of these four can be improved by engineering. The third is a choice you should make deliberately and document. The fourth is the only one that settles an argument, which is why it is worth the cost on the small number of questions that genuinely matter.
The four-layer stack
Each layer produces one of those numbers, and each depends on the one beneath it. Built in order, they compound. Built out of order, the upper layers inherit the errors of the missing lower ones and produce confident nonsense.
Layer one: server-side event delivery
Your server sends the conversion to the platform directly, rather than relying on a browser that may be blocked, degraded, or closed. The Conversions API for Meta, Enhanced Conversions and the Measurement Protocol for Google, and equivalents elsewhere. This is table stakes and most brands still have it either missing or subtly broken. Done properly it recovers a substantial share of the conversions the browser failed to report, in our experience commonly in the region of fifteen to thirty percent of the missing volume, though the figure depends almost entirely on how much of your traffic was being lost and why.
The trap here is that a broken implementation reports better numbers than a working one, because double counting looks like recovery. The setup detail matters more than the decision to do it, and the single check that separates the two is whether platform-reported conversions match the order count in your own database.
Layer two: your own first-party event data
A tracking endpoint on your own domain, writing to a store you control. Not a conversion pixel: the full journey, including everything that happens before a purchase, page views, product views, add to cart, return visits, and the session that first introduced the customer weeks earlier. This is the layer that turns attribution from something you receive into something you can compute, and it is also the layer brands skip most often, because it is the one that requires engineering rather than configuration.
The strategic argument for owning it is that it is the only measurement asset that survives platform policy changes, because it is yours. The practical argument is that nothing above it is possible without it: you cannot model attribution over events you do not have, and you cannot design a clean holdout if you cannot segment your own audience. We built ours into the platform for exactly that reason. If you are not working with us, the established third-party tools in this category are a reasonable substitute, with the honest caveat that you are renting the layer rather than owning it.
Layer three: attribution modeling over your own events
With a real event stream you can apply models rather than accept one: first-touch, last-click, linear, time-decay, position-based, and data-driven. The point is not to find the true model, because there is not one. The point is that comparing models tells you how sensitive your conclusion is to the rule. A channel that looks strong under every model is genuinely strong. A channel that only looks strong under last-click is a channel sitting close to the transaction, which is a different finding.
This is also the layer that resolves the arithmetic problem every multi-channel brand runs into. Consider a month with two hundred purchases where a typical buyer sees a Meta ad, later searches a branded term, and converts. Meta counts it inside its view-through window. Google counts it as a search conversion. Both are behaving correctly by their own rules, and the two dashboards together will describe considerably more than two hundred conversions. There is no version of that arithmetic where the platform numbers reconcile, because they were never intended to. Modeling over your own events is how one purchase becomes one purchase again.
Layer four: incrementality testing
The top layer is the only one that produces causal knowledge. You withhold advertising from a comparable group, run it for another, and read the difference. Everything below this layer describes association; only this layer answers whether the revenue would have arrived anyway.
Most brands skip it, and the stated reason is usually the discomfort of switching spend off. The real reason is more often that a badly designed test produces a null result, the null gets read as proof the channel does not work, and nobody wants to be responsible for that. Which is a good argument for designing the test properly rather than for avoiding testing, since the alternative is deciding the same question on the basis of a credit rule.
How to treat modeled conversions
Modeled conversions deserve their own position rather than a blanket verdict, because both available blanket verdicts are wrong. Dismissing them means ignoring conversions that genuinely happened. Accepting them at face value means treating an estimate produced by a party with an interest in the outcome as an observation.
The workable stance is to let them inform bidding and to keep them out of your own ground truth. Modeling exists in part to give automated bidding a fuller signal, and on that job it does help. But your internal view of what the business earned should reconcile to orders in your commerce platform or closed deals in your CRM, and the gap between that figure and the sum of platform claims should be tracked as a standing quality signal rather than closed by argument. When the gap moves sharply, something changed in the measurement rather than the market, and that is worth knowing within a day.
The order to build it in, and the order most brands use
The dependency runs strictly upward: delivery, then data, then modeling, then experiments. Each layer is meaningless without the one below. An attribution model over an event stream missing a third of its purchases will produce a clean, confident, wrong answer, and it will produce it in a nicely designed dashboard.
The order brands actually attempt is close to the reverse. They buy an attribution tool first, because it is purchasable and produces a dashboard within two weeks. Then they discover the tool is reading a partial event stream. Then they fix server-side delivery, at which point every historical comparison in the tool is broken. Then, eighteen months in, somebody proposes a holdout test. The sequence works out roughly twice as expensive as doing it in order, and the intermediate reporting is trusted for exactly as long as it takes for someone to check it against the bank.
If you can only fund one layer this year, fund the second one. Server-side delivery improves what the platforms know about you. Your own event data is the only layer that improves what you know about you, and it is the prerequisite for everything above it.
What good looks like
A brand with all four layers running does not have perfect attribution, and would be suspicious of anyone claiming to. What it has is a defensible answer and a known error bar. Revenue reconciles to a single source of truth. The gap between that truth and platform claims is monitored rather than argued about. Model comparison shows which conclusions are robust and which are artifacts of a credit rule. A small number of high-stakes questions get settled experimentally rather than rhetorically.
The commercial consequence is narrower than the marketing for these systems suggests, and more valuable. You stop scaling the campaign that was over-claiming credit. You find the channel that was quietly working and under-credited, which is nearly always an upper-funnel one. And the conversation with whoever controls the budget changes character, because you can show a number they can verify against the accounts rather than a number from a platform that sells the thing being measured.
What is still genuinely unsolved
Honesty about the limits is part of the argument. Cross-device journeys remain partially unresolvable without a login, because the deterministic link simply is not there and probabilistic matching is a judgment call rather than a fact. Long consideration cycles defeat most attribution windows, so a considered purchase researched over three months will under-credit whatever started it. Genuine view-through effect on brand search is real, hard to measure, and the single most common place spend is either quietly wasted or quietly under-funded.
None of that argues for giving up and going back to the platform dashboard. It argues for holding your numbers with the right amount of confidence, which is more than none and less than the interface implies, and for spending your experiment budget on the questions where being wrong is expensive.
The measurement environment is not going to loosen, and the brands that treat that as a permanent condition rather than a temporary inconvenience are already making better decisions than their competitors on the same data. Building the stack is eight to twelve weeks of unglamorous engineering for most brands. Most agencies will not propose it, because it delays the work that looks like work. That is precisely why it remains an advantage rather than a hygiene factor.
Is attribution after iOS 14 still a real problem in 2026?
The disruption is over in the sense that nothing is going back, so it is better understood as the permanent operating environment than as a problem awaiting a fix. What has changed since 2020 is the nature of the difficulty. The early issue was missing data. The current issue is that platforms fill those gaps with modeled estimates presented alongside observed conversions in the same column, so the practical challenge is knowing which of the numbers in front of you is a measurement, which is a claim, and which is an estimate.
Why do my Meta and Google conversion numbers add up to more than my actual orders?
Because each platform counts any conversion it touched inside its own attribution window, and a single customer frequently touches several platforms before buying. Someone who sees a Meta ad, later searches your brand name, and then purchases will be counted by both. Neither platform is malfunctioning and neither is lying; they are answering the question of what they were involved in rather than the question of what caused the sale. The only way to make one purchase count once is to model attribution over your own first-party event data.
Should I trust modeled conversions?
Use them for bidding, exclude them from your ground truth. Modeling exists partly to give automated bidding a fuller signal and it does improve that job. But your internal view of what the business earned should reconcile to orders in your commerce platform or closed deals in your CRM, and the gap between that figure and the sum of platform claims should be tracked as a standing quality signal. A sudden move in that gap indicates a measurement change rather than a market change, which is something you want to know within a day.
What should I build first if I can only afford one measurement project?
Your own first-party event data, on your own domain, in a store you control. Server-side conversion delivery is more commonly recommended and it improves what the platforms know about you, which is valuable but bounded. First-party event data is the only layer that improves what you know about you, and it is the prerequisite for attribution modeling and for designing a clean incrementality test. Buying an attribution tool before the event data exists is the most common and most expensive sequencing error in this field.
Do I need incrementality testing if I have multi-touch attribution?
They answer different questions and one cannot substitute for the other. Multi-touch attribution distributes credit for conversions that happened, according to a rule you selected, which makes it internally consistent and silent on causation. Incrementality testing measures what would not have happened without the advertising. You need attribution modeling for routine allocation because you cannot run an experiment every week, and you need experiments for the small number of decisions where being wrong is expensive, typically upper-funnel channels and brand search.
How long does it take to build a full attribution stack?
For most brands, eight to twelve weeks of engineering to reach the point where the first three layers are trustworthy, then an ongoing cadence for experiments. The realistic constraint is rarely technical difficulty; it is that the work is invisible to everyone outside the project while it is underway, which makes it politically fragile. Building it in dependency order matters more than building it fast, because a layer added on top of a broken one below inherits the error and hides it behind a better interface.
Written by Sam Nouri, founder, adsrunner. If this resonated and you want to apply it to your own account, you can book a strategy call or run a free audit.
How we research, source figures, and handle corrections: editorial policy.