The same deployment, two true stories
Klarna's support agent is quoted in board decks as proof that agents deliver, and in skeptics' posts as proof that the hype collapsed. Both citations are accurate. The deployment did produce the headline numbers - and the company did partially reverse course. What decides which story gets told is the choice of metrics. That choice is a leadership act, and it is the last skill this track teaches: how to demand the scorecard that tells you the truth before the truth gets expensive.
The believers' exhibits 13 min live
First, the honest version of the good news. These numbers are real, published, and repeatable enough to take seriously. If agents never paid back, this track would have been one session long.
LiveExhibit A: Klarna by the numbers5 min▶
The most-cited agent deployment in the world, and deservedly so at launch. Klarna's AI assistant, built on the LangChain stack, published these figures:
- 2.5 million conversations handled - roughly two-thirds of the company's support chats.
- Resolution time from 11 minutes to 2 minutes - the metric customers actually feel.
- Output equivalent to ~700 full-time agents - the line every headline ran with.
- A projected $40M profit improvement - the line every CFO ran with.
Hold these numbers with respect: they are what a successful large-scale agent deployment looks like. And hold your applause: Part 2 shows what the same deployment looked like a year later, and why both views belong in the same board pack.
LiveExhibit B: the quieter wins4 min▶
The less-headlined cases are in some ways more instructive - narrower, internal, and easier to measure honestly:
- Uber: agentic tooling for large-scale code migration - around 21,000 developer hours saved. Internal, countable, no customer emotions in the loop.
- AppFolio: a property-management copilot that roughly doubled response accuracy - a quality win, not a headcount story.
- LinkedIn: an internal SQL Bot that lets analysts self-serve data questions - value measured in unblocked analysts rather than press releases.
The pattern worth stealing: the most defensible ROI stories are internal-facing, have a clear human baseline, and measure a boring number well. If your first agent's business case looks like Uber's rather than Klarna's, that is a feature.
Self-studyThe benchmark rates - what "normal" payback looks like3 min read▶
Beyond the marquee names, the survey benchmarks give you calibration for your own plans:
- 41% of deployments report positive payback within 12 months - agents can pay for themselves, and most still do not within a year. Both halves of that sentence matter.
- Median time-to-value is 5.1 months - budget for two quarters of investment before the curve bends.
- Use case changes everything: SDR-style sales agents average 3.4 months to value; finance and operations agents average 8.9 months. If a vendor promises finance-agent value in six weeks, you now know to smile and check the scorecard.
The skeptics' exhibits 10 min live
Now the other half of the courtroom. Same industry, same years, same companies in some cases - and a very different set of numbers. A leader who can recite only Part 1 is a mark; a leader who can recite both is a buyer.
LiveKlarna, act two - the metric they were not watching4 min▶
A year after the headlines, Klarna adjusted course. Customer satisfaction had dropped - the agent handled routine cases fine but disappointed on complex and emotional ones. The company rehired human agents for exactly those cases and settled into a deliberate hybrid model.
- Read it precisely: the $40M-class efficiency numbers largely survived. What died was the "replace everyone" narrative - the deployment was oversold, not worthless.
- The lesson is metric selection: every celebrated number was a cost or speed metric. The dissenting signal - CSAT - was a quality metric, and it was not on the headline scorecard. The problem announced itself in the one place nobody had put a spotlight.
- The leadership takeaway: any scorecard composed only of cost and speed metrics will read as triumphant right up until the customers leave. Quality metrics are not decoration; they are the early-warning system.
LiveThe base rates: Gartner and MIT4 min▶
Zoom out from one company to the population and the skeptics hold strong cards:
- Gartner: forecasts over 40% of agentic AI projects canceled by end-2027 - the named causes are cost, unclear business value, and inadequate risk controls. Notice: all three are leadership failures from sessions a3-a5, not model failures.
- MIT: the GenAI Divide research found 95% of pilots show no measurable P&L impact - carried, as always in this track, with its methodology caveat: definitions of impact vary and the sample skews, so treat it as a warning shot rather than a precise census.
- The scale gap: roughly 78% of organizations report piloting agents while only about 14% have scaled them - the graveyard between pilot and production is where unclear metrics send projects to die.
Put Part 1 and Part 2 together and the resolution is not "agents work" or "agents fail". It is: agents pay back where somebody defined payback honestly in advance. Which is why this session ends with a scorecard and not a verdict.
Self-studyHow to read a vendor case study3 min read▶
Most ROI evidence you will ever see arrives via vendor marketing. Three filters before you believe a number:
- Company-published: the metrics were chosen by the party with the strongest incentive to look good. Not fake - curated. Ask what is NOT reported: no CSAT in a support-agent case study is a choice.
- Survivor bias: case studies are written about the deployments that worked. The 40% that get canceled do not get landing pages. Every case study is drawn from the winners' bracket.
- Post-hoc metric selection: when success is measured after the fact, the teller picks whichever number moved. The defense is the same discipline you will draft today: metrics defined before launch, including the ones that could dissent.
The honest scorecard 8 min live
Five metrics, demanded before any agent gets scale-up funding. Each comes with its gamed twin - the flattering version you will be offered instead, and must refuse by name.
LiveThe five metrics, and the gamed twin of each8 min▶
Demand these five before funding scale-up. Anything less is a story, not a scorecard.
| Metric | Honest version | Gamed version to refuse |
|---|---|---|
| 1 · Task resolution | Resolution rate vs the human baseline on the same case mix | A high rate earned by routing only easy tickets to the agent - "95% resolution" of the simplest 60% |
| 2 · Quality / CSAT delta | Satisfaction change across ALL cases, including complex and emotional ones | Surveys sampled from happy paths, or quality simply not measured - the exact Klarna gap |
| 3 · Cost per resolved task | All-in: tokens, infrastructure, monitoring, AND the humans handling what the agent could not | License-fee-only math with token bills and ops labor hidden in other budget lines |
| 4 · Escalation rate + workload shift | How often humans take over, and where that workload actually went | Silent deflection - customers who gave up counted as "resolved", staff absorbing overflow uncounted |
| 5 · Incidents + near-misses | Count of failures AND close calls, reviewed on a schedule | Only formally logged incidents - a quiet logbook presented as a safe agent |
And the scorecard's spine: every metric is measured against a human baseline captured before launch, and the whole card is defined before the pilot - so nobody gets to choose the flattering number afterwards. That single discipline is what separates your board pack from a vendor landing page.
Draft your scorecard ★ 15 min · everyone works
The track's final worksheet. Take one live or proposed agent - ideally the same use case from your a4 approval matrix and a5 scoresheet - and give it the scorecard it will be judged by.
Name the agent and its job in one sentence. Same use case as a4/a5 if you have been carrying one - you are about to complete its governance file.
Define the five metrics FOR YOUR CASE: what counts as a resolved task, whose satisfaction you measure, what goes into all-in cost, what counts as an escalation, what counts as a near-miss. Specific nouns, not categories.
Write the human baseline for each: what does the current human process score today? If nobody knows, your first funded work item is measuring it - BEFORE the agent launches.
Name the gamed twin you refuse for each metric, in writing. "We do not count deflected customers as resolutions" is a sentence that saves audits.
Write the kill criteria: "we stop if ..." - the CSAT floor, the cost ceiling, the incident count that ends the experiment. An agent without kill criteria is a project that can only ever be declared a success.
After the track ◐ 30 min total
- Finish the scorecard if you did not complete it live, run the stress-test prompt, and staple it to your a4 matrix and a5 scoresheet - one agent, one complete governance file.
- Find one agent ROI claim in the wild this week - a vendor page, a keynote, a LinkedIn post - and run the three case-study filters on it: who published it, who is missing from the sample, when were the metrics chosen?
- Read the benchmark self-study card and place your own use case: is it a 3.4-month SDR-shaped case or an 8.9-month finance-shaped one? Reset your own payback expectations accordingly.
- Ask the owner of any live agent in your organization for its current CSAT-class quality metric. If the answer is "we don't track one", you have found your Klarna gap - close it before it closes on you.
- What is the human baseline for this task, measured before the agent launched - and if we never measured it, how fast can we?
- Does our cost per resolved task include tokens, infrastructure, monitoring, AND the humans handling escalations - or just the license line?
- What is the escalation rate, and where exactly did that workload move - who absorbed it?
- How many incidents and near-misses did the agent log last month, and who reviews them on what schedule?
- What are the kill criteria - which specific number makes us stop, and who has the authority to call it?
Evidence covered
The leader track teaches from a verified evidence pack - published case studies, vendor material flagged as such, and independent research, sourced in the course map. This page covers:
Three questions before you go 🎯 ◐ 90 seconds
1 · Klarna is cited as both the best ROI story and the best cautionary tale. What actually decides which story gets told?
The $40M-class efficiency numbers largely survived; the "replace everyone" narrative did not. The dissenting signal lived in the one metric that was not on the headline scorecard - which is why quality metrics are the early-warning system.
2 · A team reports 95% task resolution and asks for scale-up funding. Your first scorecard question is...
The gamed twin of resolution is a high rate earned on easy tickets only. Baseline, case mix, quality delta, and escalation are exactly the columns that expose it - or confirm it honestly.
3 · Why do kill criteria ("we stop if ...") belong on the scorecard before launch?
Post-hoc metric selection is how the 40% cancellation statistic gets populated slowly and expensively. Pre-committed stop conditions are what let a neutral party call "stop" from the numbers - Gartner's named causes are all absences of this discipline.