learn-marketing-attribution-with-phoebe / Builder session 10 of 10
Learn Marketing Attribution with Phoebe · Builder session 10 of 10

The unified 2026 measurement stack

You have built every piece: SQL journeys, heuristics, Markov, Shapley, ML attribution, GA4 DDA, MMM with adstock and saturation, and geo-lift incrementality. This capstone assembles them into one thing - a calibrated loop, not a single tool. MMM anchors the quarterly budget, incrementality tests give it causal ground truth, and modeled attribution steers daily tactics. You will write Lumen's full measurement plan, build the reconciliation dashboard where the three methods agree, and run the eleven-pitfall audit on your own stack. This is the session that turns a pile of models into a system you can defend.

🔴 Hardest Builders Capstone 45 min
0-4 · The loop 4-22 · Stack + pitfalls 22-40 · Plan, reconcile, audit 40-45 · Ship it
Part 0

From a pile of models to a system

Nine sessions gave you nine ways to answer "what worked?" - and they disagree, on purpose. The mistake is to pick a winner. The 2026 answer is to orchestrate them: each method is good at a different question and a different cadence, and you wire them into a loop where the causal method calibrates the correlational one and the fast method steers the day-to-day. No single tool is the truth. The system is.

Live - built in session Self-study - read after class ★ Build-along - run the code yourself Capstone: the whole Lumen stack
★ What you walk out with today A one-page Lumen measurement plan (which method answers which question, at which cadence), a reconciliation dashboard that flags where MMM, MTA and lift agree - and where they do not - and a working audit against the eleven pitfalls that break real attribution stacks.
The one idea of the whole course MMM for the strategic budget, incrementality for causal truth, DDA/MTA for daily tactics - a calibrated loop, not a single tool. Everything else this session is detail on how to run that loop without fooling yourself.
Part 1 · the architecture

The calibrated loop: MMM + incrementality + DDA 6 min live

Three methods, three roles, wired in a cycle. MMM sets the strategic budget. Incrementality experiments give MMM its causal calibration. Data-driven attribution steers the daily tactical moves inside that budget. Each feeds the next - and the feedback arrows are the whole point.

MMM role · STRATEGIC quarterly budget across channels Incrementality role · CAUSAL geo-lift / holdout experiments the ground truth DDA / MTA role · TACTICAL Markov / Shapley / GA4 DDA daily channel steering lift → calibrates MMM priors budget → sets tactical ceiling tactics surface channels worth testing next Not a hierarchy - a cycle. The causal leg keeps the correlational legs honest; the fast leg keeps the slow leg relevant.
🔍 Click to zoom - the calibrated loop: MMM (strategic) ← incrementality (causal) → DDA (tactical), each feeding the next
LiveEach method's job, cadence, and question3 min
MethodAnswersCadence
MMM (strategic)How should we split the whole budget across channels?Quarterly / annual
Incrementality (causal)Does this channel actually cause sales, and by how much?A few big tests a year
DDA / MTA (tactical)Which campaigns and creatives to shift this week?Daily / weekly

The failure mode is using one for another's job - a daily MTA number to set the annual budget, or a slow MMM to pause a campaign today. Match the method to the decision's stakes and speed.

Part 2 · orchestration

The 2026 stack, wired on Lumen 6 min live

The loop becomes a pipeline: raw events land in the warehouse, three method families read from the same governed tables, and their outputs converge on one set of decisions. And it all sits inside the 2026 privacy reality - which is exactly why MMM and lift resurged.

Raw events touchpoints conversions spend_weekly geo revenue Warehouse governed tables, one source of truth MTA / Markov / Shapley tactical · daily MMM (adstock + Hill) strategic · quarterly Lift tests (geo / RCT) causal · a few a year Decisions budget split channel shifts what to test One warehouse, three method families, one decision layer. The orange arrow is lift calibrating MMM - the loop from Part 1.
🔍 Click to zoom - raw events → warehouse → three method families → decisions, with lift feeding back into MMM
LiveThe 2026 privacy reality this sits inside3 min

Why this shape and not a single MTA dashboard? Because of what broke:

  • Cookies partly survived Chrome's long deprecation saga - but MTA broke anyway. Safari and Firefox block third-party cookies, Apple's ATT gutted mobile signal, and SKAN / AAK aggregation strips user-level detail. MTA now misses an estimated 30 to 60% of touches.
  • MMM and lift resurged precisely because they run on aggregate and experimental data - which no browser update or OS privacy change can revoke. Privacy-proof by construction.
  • So MTA is demoted, not deleted: it is a tactical signal on the consented slice you can still see, kept honest by MMM and lift. That is the whole reason the stack is a loop and not a single tool.
Real world

The lazy 2026 take is "cookies came back, so MTA is fine again." It is wrong: cookies were never the only thing that broke MTA. Safari, Firefox and ATT never used third-party cookies to begin with, and their walls are still up. Any stack betting on MTA alone is measuring a shrinking, biased half of reality.

Part 3 · what breaks stacks

The eleven-pitfall checklist + ethics 6 min live

Every pitfall below has sunk a real attribution program. Read them as a pre-flight checklist - if your stack does any of these, fix it before you trust a number.

PitfallThe fix
1 · Calling heuristics "data-driven"Heuristics are fixed rules. Reserve "data-driven" for Markov / Shapley / DDA that learn credit from the data.
2 · Last-touch as default truthIt is a tactical read that over-credits demand-capture. Never set a budget on it alone.
3 · Treating correlation as causationOnly experiments are causal. Label MTA / Markov / Shapley / MMM as correlational everywhere they are reported.
4 · Not normalizing Markov removal effectsRemoval effects do not sum to 1 - always normalize before you split revenue, or credit leaks.
5 · Exact Shapley at O(2^n) costFor many channels use Monte-Carlo Shapley; exact coalitions explode past ~15 channels.
6 · Attribution windows silently changing resultsFix and document the lookback window. A 7-day vs 30-day window can flip which channel "wins".
7 · Missing-touchpoint bias (2026)MTA misses 30-60% of touches. Do not report it as complete - triangulate with MMM and lift.
8 · Double-counting across toolsGA4, the ad platforms and MMM all claim the same conversion. Reconcile to one governed number.
9 · Ignoring view-through vs click-throughAn impression is not a click. Keep them separate or you inflate awareness channels.
10 · MMM without incrementality calibrationAn uncalibrated MMM is a confident correlation. Anchor it with lift-test priors (Session 9).
11 · Overfitting deep models on sparse dataLSTM / attention MTA needs volume. On thin data a simpler model generalizes better - do not chase the fancy one.
LiveThe ethics layer you own as the builder3 min

Three responsibilities that are not optional in 2026:

  • Privacy and consent. Model on data you have the right to use. Aggregate and experimental methods (MMM, lift) are not just privacy-resilient - they are the privacy-respectful default.
  • Honesty about incomplete data. When MTA sees half the touches, reporting its number with false confidence is a form of lying. Carry the uncertainty forward - intervals, not just points.
  • Transparency to stakeholders. Tell the CMO which model produced a number, what it over-credits, and how sure you are. The builder who names the model's limits earns more trust than the one who hides them.
The builder's oath Name the method, carry the uncertainty, and never let a correlational number wear a causal costume. Do that and your stack is not just accurate - it is defensible.
Build-along 1 of 3

Assemble Lumen's measurement plan ★ 7 min · map it

Turn the loop into an actual plan: for each business question Lumen's CMO asks, which method answers it, and on what cadence? Encode it so it is a living config, not a slide.

Python · the plan as config# Lumen's measurement plan: question -> method -> cadence -> what it can/can't say plan = [ {"question": "How do we split the $4M annual budget?", "method": "MMM (adstock + Hill saturation)", "cadence": "quarterly", "caveat": "correlational - must be lift-calibrated"}, {"question": "Does paid_social actually cause sales?", "method": "geo-lift incrementality test", "cadence": "2-3 big tests / year", "caveat": "causal, but scoped + slow"}, {"question": "Which campaigns to shift this week?", "method": "GA4 DDA + Markov (data-driven MTA)", "cadence": "daily / weekly", "caveat": "misses 30-60% of touches"}, {"question": "Where does the next dollar earn most?", "method": "MMM response curve (marginal return)", "cadence": "quarterly", "caveat": "read the knee, not the average"}, ] import pandas as pd plan_df = pd.DataFrame(plan) print(plan_df.to_string(index=False))

Question first, tool second. The plan is organized by the CMO's questions, not by your models. That is what makes it a plan and not a tech inventory.

Cadence is a first-class field. Daily tactics and quarterly strategy run on different clocks. Writing the cadence down stops anyone using a weekly number for an annual decision.

Every row carries its caveat. The caveat column is the honesty layer - it travels with the number so no one forgets what the method cannot say.

Real world

The best measurement teams pin a version of this table to the wall. When a stakeholder asks "what's our ROAS?" the answer is "which decision are you making?" - and the table routes them to the right method. That reframe alone prevents most attribution arguments.

Build-along 2 of 3

The reconciliation dashboard ★ 8 min · run it

Put the three methods side by side for each channel. Where they agree, you have confidence. Where they diverge, you have your next investigation - or your next lift test.

Python · reconcile the three viewsimport pandas as pd import numpy as np # each channel's credit share under three methods (all sum to ~1 across channels) recon = pd.DataFrame({ "mta_dda": {"paid_search": .34, "paid_social": .18, "display": .06, "ctv": .04}, "mmm": {"paid_search": .22, "paid_social": .29, "display": .11, "ctv": .13}, "lift": {"paid_search": .19, "paid_social": .31, "display": np.nan, "ctv": .15}, }).round(2) # agreement score = 1 - spread across the methods we DO have for each channel def agreement(row): vals = row.dropna() return round(1 - (vals.max() - vals.min()), 2) recon["agree"] = recon.apply(agreement, axis=1) recon["flag"] = np.where(recon["agree"] < 0.85, "investigate", "confident") print(recon) # paid_social: MMM + lift both say ~0.30, MTA under-credits it -> trust the causal read # paid_search: MTA over-credits (last-click bias); MMM + lift agree it is smaller

Three columns, one channel per row. The dashboard's job is not to average the methods - it is to surface the disagreement, which is where the learning is.

Where they agree, ship it. When MMM and lift both land near 0.30 for paid social, that is a calibrated, near-causal number - act on it with confidence.

Where they diverge, route it. Paid search high on MTA but low on MMM and lift is the classic last-click bias - the disagreement itself is the finding, and it tells you what to test next.

Agreement is the real KPI A single fancy attribution number invites false confidence. Three methods and an agreement score invite the right question - "why do these disagree?" - which is how a measurement stack actually gets smarter over time.
Build-along 3 of 3

The pitfalls audit on your own stack ★ 7 min · run it

Turn the eleven pitfalls into a runnable checklist. Score your own stack, and let the failures become your roadmap.

Python · self-auditpitfalls = { "heuristics_called_data_driven": False, "last_touch_is_default_truth": True, # <- Lumen's old dashboard "correlation_reported_as_causal": True, # <- MMM shown without caveat "markov_not_normalized": False, "exact_shapley_too_slow": False, "attribution_window_undocumented": True, # <- nobody knows the lookback "mta_missing_touches_ignored": True, # <- reported as complete "double_counting_across_tools": True, # <- GA4 + ads + MMM all claim it "view_vs_click_not_split": False, "mmm_not_lift_calibrated": True, # <- no experiments feeding priors "deep_model_overfit_sparse": False, } open_risks = [name for name, tripped in pitfalls.items() if tripped] score = 1 - len(open_risks) / len(pitfalls) print(f"stack health: {score:.0%} ({len(open_risks)} open risks)") for r in open_risks: print(" - fix:", r)

Be honest in the booleans. The audit only works if you mark the pitfalls you are actually committing. Optimistic scoring defeats the purpose.

The open risks are your backlog. Each True is a concrete engineering or process fix - documenting a window, adding a lift test, reconciling double-counts.

Re-run it quarterly. Stack health is not a one-time score. As you close risks and add methods, the audit tracks whether your system is getting more defensible or just more complex.

Real world

Most real 2026 stacks score somewhere around 50-60% on first audit - and the two risks that hurt most are almost always "MMM not lift-calibrated" and "correlation reported as causal." Fix those two and you have leapfrogged the majority of marketing measurement programs.

After the course

Ship it ◐ 60 min total

Check yourself

Three questions before you go 🎯 ◐ 90 seconds

1 · The unified 2026 stack is best described as...

No single tool is the truth. MMM sets the budget, lift tests calibrate it causally, and modeled attribution steers tactics - each feeding the next in a loop. The disagreement between them is a feature, not a bug.

2 · In this whole stack, which methods are causal?

MTA, Markov, Shapley and MMM all describe correlation. Only incrementality experiments withhold a channel and measure the delta, which is why lift is the causal anchor the rest of the stack gets calibrated against.

3 · Name a top pitfall that breaks real attribution stacks.

Markov removal effects do not sum to 1 - skip the normalization and credit leaks. And GA4, the ad platforms and MMM will each claim the same sale unless you reconcile to one governed number. Both are on the eleven-pitfall checklist.

Source material

What this session covers

This capstone synthesizes the entire builder track - every method from B1 to B9 - into one orchestrated, auditable stack. It is the integration layer no single course teaches, drawn from the whole source universe of this program.

Synthesis of Builder Sessions 1-9SQL → heuristics → Markov/Shapley → ML → GA4 → MMM → lift - the whole loop
The 2026 cookieless reality (industry)why MMM + lift resurged - Part 2
Attribution pitfalls (research synthesis)the eleven-pitfall audit - Part 3, Build-along 3
Production orchestration (Airflow / dbt)named as the deployment path, not built here
Vendor certificatesGA4, Meta Blueprint badges stay with their official sources

Builder Session 10 cheat sheet · pin this

The unified stack =MMM anchor (strategic budget) + incrementality (causal calibration) + DDA/MTA (daily tactics). A calibrated loop, not one tool.
Match method to decisionMMM quarterly, lift a few times a year, MTA daily. Never use a daily number to set an annual budget.
2026 privacy realityCookies partly survived Chrome but MTA still misses 30-60% (Safari/Firefox/ATT/SKAN). MMM + lift = privacy-proof.
Only experiments are causalMTA, Markov, Shapley, MMM are correlational. Lift is the ground truth that calibrates the rest.
Reconcile, do not averagePut the three methods side by side. Agreement = confidence; disagreement = your next investigation or lift test.
Top pitfallsNot normalizing Markov · double-counting across tools · MMM without lift calibration · correlation dressed as causal. Audit quarterly.