learn-marketing-attribution-with-phoebe / Leader session 3 of 6
Learn Marketing Attribution with Phoebe · Leader session 3 of 6

Data-driven attribution, decoded

Markov. Shapley. GA4's black box. These are the models that finally stop guessing and read credit off your actual journeys - and the ones your vendors wave around to sound sophisticated. This session decodes the two big ideas behind them - the removal effect and fair-share marginal contribution - in plain language, no math, and hands you the three questions that puncture any "data-driven" claim. Because smarter is not the same as ground truth.

🟡 Core Leaders · CMO / CDAIO No tech required 45 min
0-3 · Recap 3-22 · Markov & Shapley 22-40 · DDA & the 3 questions 40-45 · Q&A
Part 0

Why leave the heuristics behind

Last session ended on a ceiling: every heuristic is a fixed rule with a knob a human turned. If the webinar in your funnel is really worth 45%, no heuristic will ever find that out - it has no way to read your data. Data-driven attribution breaks that ceiling. Instead of asserting the weights, it learns them from the whole population of journeys. Two ideas do almost all the work - the removal effect (Markov) and fair-share marginal contribution (Shapley). You do not need the math. You need to understand what question each one asks, and where each one still falls short. Because "learned from data" and "true" are not the same thing.

Live - presented in session Self-study - read after class ★ Work-along - everyone does it The brand: Lumen Skincare
★ What you walk out with today A plain-language grip on the two ideas behind every "AI attribution" pitch - the removal effect and marginal contribution - plus three questions that will tell you, in any meeting, whether a data-driven number is trustworthy or a black box being sold as ground truth.
Part 1 · the removal-effect idea

Markov: what happens if we remove this channel? 7 min live

The Markov idea is disarmingly simple. Picture every customer journey as a walk between states - Start, then each channel, ending at either Convert or "Null" (left without buying). Across all your journeys, you can measure how often each move happens. Then you ask the one question that gives Markov its power: if we deleted this channel entirely, how much would conversion drop? That drop is the channel's credit.

Start Display Paid social Email Paid search ✅ Convert Null ✂ Remove paid social from every path... ...conversion drops 34%. That drop = its credit. Credit is the conversion you would lose without the channel. Removal effects do NOT sum to 1 - they must be normalized before you read them as shares.
🔍 Click to zoom - Markov: credit is the conversion you would lose if the channel vanished
LiveThe removal effect, in one plain sentence3 min

Markov gives the most credit to the channels whose absence would collapse conversion. Delete a channel from every journey, re-measure the conversion rate, and the size of the drop is that channel's importance. A channel nobody misses gets little; a channel that quietly holds the funnel together gets a lot - even if it never grabs the last click.

  • It reads the whole population, not one path. Unlike a heuristic, Markov does not split this $92 journey - it learns from all of Lumen's journeys at once what each channel is worth.
  • It rewards the connectors. On Lumen's data, paid social and email often score high on removal effect - they are the touches that keep people moving toward the purchase, even though last-touch handed them almost nothing.
  • Critical caveat: removal effects do not sum to 1. Each is measured independently, so they must be normalized before you read them as budget shares. A vendor showing raw removal effects that add up to 130% is showing you an un-normalized model.
Real world

A retailer ran the removal effect and found their email channel - dead last under last-touch - was the single biggest drop when removed. Pulling email out of the journey cost more conversions than pulling out paid search. They had been about to cut the email team. Removal effect saved the channel, because it measured what was lost in its absence, not what showed up at the finish line.

Part 2 · the fair-share idea

Shapley: everyone's fair average contribution 7 min live

Shapley comes from game theory - Lloyd Shapley, 1953 - and it answers a fairness question: if channels are teammates cooperating to win a conversion, what is each one's fair share of the prize? The answer is each channel's average marginal contribution: across every possible order the team could have formed, how much did adding this channel move the result?

Order A Display +$14 P.social +$41 P.search +$37 Order B P.search +$60 Display +$9 P.social +$23 Average each channel's added value over ALL orders → its fair Shapley credit Display Paid social Paid search Shares synergy: if two channels together beat their solo sum, the surplus is split between them - nobody hogs the joint win. The "fair" model: efficiency, symmetry, null-player, additivity. Shares always sum to the full conversion value. Exact Shapley cost is O(2^n) - fine for ~10-15 channels; beyond that, teams approximate it.
🔍 Click to zoom - Shapley: each channel's fair share is its average added value across every ordering
LiveMarginal contribution and shared synergy3 min

The mechanic in one line: line the channels up in every possible order, and each time a channel joins, note how much it added to the result so far. Average those additions over all orderings - that average is the channel's Shapley credit.

  • It is provably "fair". Shapley is the only credit split that satisfies four common-sense axioms at once - efficiency (shares sum to the whole), symmetry (equal channels get equal credit), null-player (a useless channel gets zero), and additivity. That is why it is the reference standard for fairness.
  • It shares synergy. If display and social together drive more than the sum of what each does alone, that surplus gets split between them - neither one hogs the joint win. Heuristics cannot do this; they have no concept of two channels lifting each other.
  • The cost is the catch. Exact Shapley is O(2^n) - it considers every subset of channels. Fine for the ~10-15 channels most brands run; beyond that, teams switch to Monte-Carlo approximations. For Lumen's nine channels, exact Shapley is comfortable.
How to hold Markov vs Shapley in your head Markov asks "what breaks if this channel disappears?" Shapley asks "what does this channel add, on average, to every team it could join?" Both learn from your data, both reward the channels heuristics ignore - and both are still correlational, not proof of cause. Keep that last part close.
Part 3 · data-driven in the wild

GA4 DDA, and the black-box warning 6 min live

You may never run Markov or Shapley yourself - but you almost certainly already use a data-driven model, because Google made it the GA4 default. GA4's Data-Driven Attribution (DDA) is Shapley-based with a time-decay element, and Google removed the old rules-based models from GA4 reporting entirely. It is powerful. It is also a black box - and that combination is exactly where leaders get fooled.

Inputs it can see Google-observable Consented touches Web + app only Last 50 interactions ⬛ Black box Shapley + time-decay 90-day lookback Weights not exposed Per-channel credit A number per channel No model shown No per-touch weights Take it or leave it ⚠ What DDA cannot see - and will never tell you it missed Non-Google channels · unconsented users · offline touches · CTV/print/word-of-mouth · the full cross-channel journey. It is not ground truth - it is one vendor's partial, unexplained view.
🔍 Click to zoom - GA4 DDA: a powerful model that only sees part of the journey and never shows its work
LiveWhy "data-driven" is not the same as "true"3 min

DDA is a genuine upgrade over last-touch. But two properties should keep you cautious every time you read one.

  • It is a black box. Google does not expose the model or the per-channel weights. You get a number and no way to audit how it was produced. You cannot ask it "why did display get 8%?" and get an answer.
  • It only sees what Google can see. DDA works on Google-observable, consented, web-and-app touches - it defaults to the last 50 interactions with a 90-day lookback. Your influencer, your CTV, your offline, your unconsented visitors, and every non-Google channel are simply invisible to it. It is not modeling the journey; it is modeling the slice of the journey Google happened to witness.
  • It is still correlational. Like Markov and Shapley, DDA finds channels that correlate with conversion. It cannot prove any channel caused it. Only incrementality experiments do that - Sessions 4 and 5.
Real world

A CMO presented GA4 DDA numbers to the board as "the truth, finally - it's AI". A sharp director asked one question: "Does it include our CTV spend?" It did not - CTV is not a Google-observable web touch, so DDA never saw it. Half the media budget was invisible to the model being sold as ground truth. The number was not wrong; it was answering a smaller question than the room assumed.

Self-studyThe family tree: heuristic → learned → causal2 min read
ModelCore ideaStill just correlation?
MarkovRemoval effect - what conversion you'd lose without itYes - correlational
ShapleyAverage marginal contribution, fairly sharedYes - correlational
GA4 DDAShapley + time-decay, on Google-observable touchesYes - correlational + partial view
Incrementality (Session 4-5)Randomized holdout / geo experimentNo - this one is causal

Data-driven is a real step up from heuristics - it learns instead of asserting. But it is still standing on correlation. The jump to causation needs an experiment, not a smarter split. That is the whole next arc of the course.

Work-along 1 of 2

Read a data-driven report - and ask the three black-box questions ★ 10 min · on your reports

No math. We put a data-driven attribution report on screen - Lumen's GA4 DDA output - and practice the three questions that puncture any "it's AI, trust it" claim. These three questions are the entire leader skill for this session.

Question 1 - What window? Ask what lookback and interaction limit the model uses. Lumen's DDA defaults to a 90-day lookback and the last 50 interactions. A journey that started 100 days ago, or the 51st touch back, is invisible. If your real cycle is longer than the window, the model is cropping your journeys.

Question 2 - What touches can it see? Ask which channels the model actually observes. DDA sees Google-observable, consented, web-and-app touches only. For Lumen that means influencer, CTV, and offline are missing entirely. Name every channel the model is blind to.

Question 3 - Can you see the weights? Ask to see the per-channel model, not just the output. With GA4 DDA the answer is no - it is a black box. That is not a dealbreaker, but it means you cannot audit why a channel got its credit, so you treat the number as an informed opinion, not a proof.

Write the verdict: "This is a data-driven model, which beats last-touch - but it runs a 90-day window, is blind to CTV and influencer, and won't show its weights. I'll use it for direction, not as the final word on budget."

The move that earns respect Memorize the three questions - what window, what touches, can I see the weights. Asking them turns "the AI says paid search wins" into a conversation about what the model can and cannot know. A leader who asks them cannot be sold a black box as ground truth.
Work-along 2 of 2

When is data-driven worth it - and when is it overkill? ★ 10 min · pen and paper

Data-driven is not automatically the right choice. It needs volume, a real multi-touch journey, and a decision big enough to justify it. Decide, out loud, when Lumen should reach for it and when a heuristic is honestly good enough.

Take a low-volume, short-path case - a new Lumen product with 40 conversions a month, mostly one-touch. Is Markov/Shapley worth it? (No - there is not enough journey data to learn from, and barely a path to split. A simple model is more honest here.)

Take a high-volume, multi-touch case - Lumen's core serum line, thousands of 5-touch journeys a month across nine channels. Worth it? (Yes - this is exactly where learned credit beats a fixed rule, because the interactions between channels are real and measurable.)

Take a strategic budget case - reallocating the $4M media budget. Is data-driven enough? (No - it is better than a heuristic, but it is still correlational. A budget of that size needs an incrementality test on top. That is the honest limit of everything in this session.)

Write your rule of thumb: data-driven attribution earns its complexity when you have volume, genuine multi-touch journeys, and a decision worth the effort - and even then, it informs the budget rather than deciding it.

Real world

A team spent three months standing up a Shapley model for a product line that got 30 conversions a month. The model's outputs swung wildly week to week because there was not enough data to stabilize them - and they made real budget cuts on the noise. Sophistication without volume is not rigor; it is expensive guessing with a fancier label.

Before Session 4

This week ◐ 25 min total

Check yourself

Three questions before you go 🎯 ◐ 90 seconds

1 · In Markov attribution, the "removal effect" measures...

Removal effect = the conversion you would lose without the channel. Channels whose absence collapses conversion get the most credit - and the raw effects must be normalized, since they do not sum to 1.

2 · What makes Shapley value attribution distinctive among the models?

Shapley averages each channel's marginal contribution across all orderings and splits joint synergy between the channels that created it. Heuristics have no concept of channels lifting each other.

3 · A director says "GA4 DDA is AI, so it's the true attribution". Your correction?

DDA does not expose its model or weights, and it only sees Google-observable web/app touches - so CTV, offline, influencer, and non-Google channels are invisible. It is a better opinion than last-touch, still correlational, and far from ground truth.

Source material

What this session covers

This session distills the data-driven-attribution chapter of the leading courses into leader-first concepts - the removal effect, fair-share marginal contribution, and the GA4 black box - without any math derivation. Certificates and full video courses stay with their official sources.

Marketing Attribution and Mix Modeling (LinkedIn Learning) - ch.3Markov removal effect + Shapley fairness, conceptually - Parts 1-2
Google Analytics Data-Driven Attribution (Skillshop / GA4 docs)DDA as configured, the black-box caveat - Part 3
The math of Markov & Shapley from scratchbuilt in Python in Builder Sessions 4-5
MMM vs MTA vs incrementalitythe same question three ways in Leader Session 4
Causation via holdout / geo experimentsthe causal jump in Leader Session 5

Leader Session 3 cheat sheet · pin this

Data-driven =Credit learned from the whole population of journeys, not asserted by a fixed rule. It reads the weights off your data.
Markov · removal effectCredit = the conversion you'd lose if the channel vanished. Rewards connectors. Raw effects don't sum to 1 - normalize them.
Shapley · fair shareEach channel's average marginal contribution across all orderings. Shares synergy fairly. Exact cost O(2^n) - fine for ~10-15 channels.
GA4 DDAThe GA4 default. Shapley + time-decay, 90-day lookback, last 50 interactions. A black box - model and weights not exposed.
The three questionsWhat window? What touches can it see? Can you see the weights? Ask all three of any data-driven report.
Still correlationalMarkov, Shapley, and DDA all find correlation, not cause. Only incrementality experiments prove causation - Sessions 4-5.