Minimize output, maximize outcome.

Part 1 ended on a gap. Scrum gives you a cadence for building predictably, but nothing in the ceremonies asks whether the thing in the sprint should exist. That gap is where Marty Cagan's 50 to 80 percent lives.
This post is the other half. Same running class exercise, the Marvel Unlimited app, same breakout group. The last two sessions were where our team stopped listing complaints and had to commit to something.
The loop that fills the gap Scrum leaves: from a business problem to something you would actually release, killing work before you build it rather than after.

Who exactly
Jeff opened the discovery half with Clayton Christensen's milkshake study, which is the most useful ten minutes of product theory I know.
McDonald's wanted to sell more milkshakes. They did the obvious research: ask milkshake buyers what would make the milkshake better. Thicker, cheaper, more flavours. They made those changes. Sales did not move.
Then someone went and watched. A large share of milkshakes sold before 8am, to people alone, who took them to go. Those people were not buying a dessert. They had a long boring commute and one free hand, and they wanted something that lasted the whole drive and stopped them being hungry at 10am. They had hired the milkshake for that job.
Which means the milkshake's real competitors were a banana (gone in two minutes), a bagel (needs two hands), and a donut (crumbs, and you are hungry again by nine). None of that shows up if you only ask milkshake buyers about milkshakes.
The lesson: look for differences that make a difference. Not demographics. The differences in goal and context that would actually change what you build. A 34-year-old marketing manager and a 34-year-old marketing manager tell you nothing. A 20-minute commute and a 90-minute commute tell you everything.
Then personas, and Jeff's version is deliberately cheap. Proto-personas, built from your current assumptions in about ten minutes, given a name and a face. Their job is not accuracy, it is to stop the team arguing from personal preference. With a relevance test: only write down a characteristic if it would change what you build.
WHAT I'D TAKE BACK
When we say "the user" in a planning discussion, we usually each mean a different person, and we never find out because we never name them.
Map what they already do
Before designing anything, map the current journey as it exists today, one step per note. One rule governs the whole thing: verb phrases, not nouns. "Check calendar", not "calendar". Nouns are features; verbs are what people actually do, and a map built from nouns quietly turns into a feature list.
Then annotate with four things: where it hurts, what reward the person is chasing, how long and how often, and the workarounds they use outside your software. That last one is the gold. Nobody builds a workaround for a problem they do not have.
What "better" would even mean
You cannot improve everything, so you have to pick one outcome. Business metrics are lagging: revenue, NPS, churn. By the time they move, the cause is months old. Product metrics are leading, and Jeff gave two rules for picking them: measure how customers get value (for Zoom, meetings hosted and attended per user per week, not logins), and measure adoption of whatever you just changed.
He walked Dave McClure's pirate metrics, AARRR, with one addition of his own: rollback, for people who tried it and actively went back. Revenue sits at the end because it is a consequence, not something users do.
Then the move that made the whole workshop click: laddering down.

The business problem in the exercise was long-time customers leaving. And here is the thing: you cannot solve retention directly. There is no feature called retention. You have to find the specific customer problem causing the churn and solve that instead.
A product objective should describe a change in behaviour without prescribing the feature. The moment it names a solution, you have skipped discovery.
IN THE EXERCISE
Someone realises they are the churned user
Day four, we had to turn our pile of complaints into one objective. The debate ran between fixing search and fixing reading continuity. The search case was concrete and technical: open the app up to external indexing via universal links, so that searching Spider-Man on your phone surfaces Marvel Unlimited and deep-links you straight into it. Whoever proposed it already had the metric ready, too. Search a keyword, count how many comics come back before the optimisation and after.
Then someone stopped mid-sentence and said this:"I signed out, because it was a challenge for me. So there was no continuity for me. As a user. We need to make it easy for them to continue and complete those stories."
That is a person discovering, out loud, that they are the churn. Not a persona, not a survey response. Them. The room went quiet for a second and the objective basically wrote itself: make Marvel Unlimited the easiest platform to find your next favourite comic, so readers return and explore more.
Notice what it does not say. It does not say "build a recommendation engine."
Generate options, including stupid ones
Only now do you allow solutions. Jeff runs silent brainstorming first, and the reason is social rather than creative: if you open with discussion, the loudest and most senior person anchors the room and the quieter people never get their idea out. Everyone writes alone first, then you talk.
Start with the obvious ideas, just to clear them out of your head.
Build on someone else's sticky note.
Steal from an unrelated product.
Write down impossible ideas.
Number four is the one people skip and it is the point. He showed an old car alarm advert where the anti-theft device is a monkey that lives in your boot and beats up the thief. Obviously unbuildable. But dial it back and the seed is there: something that responds on your behalf, immediately, without you being present.
IN THE EXERCISE
AR glasses and a phone call from Stan Lee
Six minutes, as many ideas as possible. Ours included a Marvel timeline map showing the reading order across the multiverse, recommendations from Marvel-world influencers, curated Netflix-style rows by character, an onboarding survey asking your favourite character and your current mood, a "read my mind" idea that turned into AR glasses tracking your eye movement, and a simulated phone call from Stan Lee.
The Stan Lee call is the one worth keeping, because of what happened next. Someone pointed out it might be more intimidating than exciting. Which is a real product insight arrived at through a silly idea: a personalised intervention from a beloved figure can read as pressure rather than delight. You do not get that from a sensible ideas list.
Choosing: value against risk, not value against effort
Most prioritisation grids plot value against effort. Jeff plots value against risk and confidence, and argued it hard: cost is a secondary concern next to the possibility that you have misunderstood the problem entirely. An expensive thing that works is fine. A cheap thing built on a wrong assumption is pure waste, and it is cheap precisely because you did not check.

High value and high confidence: build it. High value and high risk: do not drop it, test it. That quadrant is your discovery backlog. The whole point of the grid is that it routes work to two different places rather than producing one ranked list.
IN THE EXERCISE
Watching a good idea land in the wrong quadrant
The nicest thirty seconds of the workshop was watching our team place the timeline map. Someone loved it on sight and said so. Then, before it had settled anywhere, the same person talked themselves leftward: it would clearly be valuable, but we had no idea whether readers would actually use it. It ended up top-left, in the test-it column, and everybody was fine with that.
Enthusiasm and confidence got separated in real time. That is the entire job of the grid, and an idea everyone loved still did not go straight to the backlog.
Time budgeting
A related idea I want to steal outright. Do not prioritise different kinds of work against each other, because strategic work always wins on paper and technical work never does. Instead budget time: something like 60 percent strategic, 20 percent tactical, 20 percent technical, with a fourth slice for regulatory work where it applies. Jeff was blunt that failing to budget technical time does not delay technical work, it eventually stops product work altogether.
Framing the bet
You have picked an idea. Before it becomes tickets it gets framed, and Jeff's tool is the opportunity canvas, which is mostly a set of questions that expose what you are assuming.
Left side: the problem | Right side: the bet |
|---|---|
Which users and customers have this problem | What solution we propose |
What problems they have now | How users behave differently afterwards |
How they solve it today, the workarounds | How we will measure that change |
What business problem this serves | How people will discover and adopt it |
He suggested filling it in on the stakeholders' behalf rather than asking them to author it, because people correct a draft far more willingly than they write one from scratch. That is a small piece of organisational cunning and it is completely true.
Then, before building, play good, better, best. Sketch three levels of solution and expect to eliminate about two thirds of the scope. Teams do not cut scope well when there is only one design on the table, because cutting feels like damage. With three, cutting is just choosing.
Draw the toast first
Before Jeff showed us a single story map, he stopped and played us a talk about toast.
https://www.youtube.com/watch?v=yAo8KLKZMW4
Tom Wujec, "Got a Wicked Problem? First, Tell Me How You Make Toast".
His premise is almost annoyingly simple. Ask a room of people to draw how they make toast, no words allowed, just a drawing. What comes back is never the same twice. Some people draw the loaf, some draw the inside of the toaster, some draw themselves standing in a kitchen. And every drawing, however scruffy, turns out to be the same kind of object underneath: a system of nodes and links.
Then he changes one thing. Do it again, but one step per sticky note. The models immediately get better, because the drawing stops being a picture and becomes something you can pick up and move. Change one more thing, do it in groups rather than alone, and they get better again.
That is the whole talk, and it runs nine minutes.
What Jeff pulled out of it. He was not showing it to us for the toast. He was showing us the method that had been quietly running underneath every exercise we had done all week, and he drew four rules out of it.
One step per note. A model you can rearrange beats a picture you cannot. This is exactly why story maps are sticky notes and not a document.
Three to five people. Bigger groups produce worse models, not better ones.
Generate in silence, converge out loud. Open discussion anchors on whoever speaks first, which is the same reason silent brainstorming exists.
Timebox it, and trust intuition over analysis. He called this one satisficing. Stop when it is good enough for now.
All of which compresses into the mantra he repeated on day one and again on the final morning: think it, write it, say it, place it. Get the thought down before you say it, then put it where everyone can see it. For a distributed team that also means one big persistent board acting as an information radiator, rather than a fresh blank board every meeting that nobody ever returns to.

Then he made us do it. The task was to map our own morning routine, from waking up to sitting down at work, one sticky note per step. No discussion first. Everyone built their own.
The lesson was in what came back. Some people finished in six notes. Some had twenty-five. The twenty-five had kids, or pets, or a long commute; the six had none of those.
Which is the differences that make a difference idea from earlier in this post, turning up in a room full of people who all assumed they were describing the same ordinary morning. Nobody was. And that is precisely the thing you are hunting for when you map a real user's journey ten minutes later, which is exactly what we did next.
The story map
This is Jeff's own territory and the most useful visual he showed. He had photos of a real one on a wall: a long strip of sticky notes with blue tape running across it in horizontal lines.

The backbone runs along the top: the big activities, in the order a person does them. Underneath each one hang the specific tasks, roughly in order of importance. Then the tape: each horizontal band is a release slice. Everything above the first line is the smallest thing that still tells a complete story end to end.
The two questions on his slides were exactly right: are these viable releases, and which customers will we focus on and how will we measure success. He also had a strong opinion here: developers need pictures, not text user stories, to estimate and plan. Sketch the UI on index cards before anything reaches Jira.
IN THE EXERCISE
The survey that could trap you
Our chosen feature was a short survey asking your favourite character and your current mood, feeding recommendations. We immediately hit two design questions that were really risk questions.
First: does the survey belong at the start of the journey or the end? Useful at the start for a brand new reader with no history, but a returning reader might want it later, if at all.
Second, and this is the one I liked, someone spotted a risk in our own idea: if we keep recommending against a stated preference, we might stop people ever discovering a character they would have loved. We ended up talking about Spotify's smart shuffle as the shape we wanted, familiar enough to trust, loose enough to surprise you.
We had built a story map for maybe eight minutes and it had already generated a real product risk. That is the map doing its job.
Testing before building
Now the part that separates this from ordinary agile. Before you build the thing, you run the validated learning loop, which came from Steve Blank and got popularised by Eric Ries as build, measure, learn.

Jeff was pointed about the word "build" in build, measure, learn. It does not mean production code. It means the cheapest possible test: a fake door (put the button in before the feature exists and count who presses it, as TripAdvisor and Atlassian both did), a paper prototype (a Hewlett-Packard team tested complicated data-mapping usability in a single afternoon), or a spike.
And a framing I had not seen before and immediately liked: stop saying low fidelity and high fidelity. Fidelity has three separate axes, visual, data and functional, and they move independently. A paper prototype can have almost zero visual fidelity and very high functional fidelity, which is exactly right for a usability test. The question is not high or low, it is the right fidelity to answer the question you are asking.
The three different things people mean by MVP
This section explained about six years of arguments I have watched.

The first one is not from any book. It is what most organisations actually practise, and Jeff had it on his whiteboard as the honest version: "the most I can get in the time, but I'm going to have to give up something." The second and third are close to opposites and both are in circulation, which is why a stakeholder hears a shippable product and a team means a fake door test.
Release to learn before you release to earn

Jeff's example of a good MVP was the original iPhone: a bad camera, no app store, no copy and paste, and it solved web browsing and texting so much better than anything else that none of that mattered to the people it was for.
Spotify's version, from Henrik Kniberg's engineering culture posts, is think it, build it, ship it, tweak it. Ship less than perfect solutions to a small subset of your audience, iterate until awesome, and do not be afraid to kill what does not work. Jeff's warning: teams get stuck in tweak it forever. At some point you decide to scale it or drop it.
Two tracks, one team
The last big idea, and the one Jeff drew on the final morning with "Two Tracks, One Team" written across the top.

Scope doesn't creep, understanding grows.
Jeff Patton
Which is the same built-in instability from the 1986 paper in Part 1. The mess is the method. The practical consequences are what I found most useful, because they are changes you can make on a Monday:
Standup stops being a round of individual status reports and becomes a walk through the work items themselves. When you talk about the item rather than the person, people swarm on whatever is stuck instead of listening politely. Split it into now (delivery) and next (discovery) so people can opt into the second half.
Backlog refinement stops being one long dreaded meeting per sprint and becomes short sessions twice a week, with optional attendance.
Sprint review covers not just what was built but what was learned. Otherwise discovery work is invisible and quietly gets deprioritised.
Engineers and QA join the discovery work: interviews, observation, crazy eights, design studios. Which loops back to Sherif Mansour's line from Part 1 about never running a customer interview without a developer in the room.
THE QUESTION I ACTUALLY ASKED
What if it is a small team?
Right as he was drawing this, someone in the chat said project standups are very different to business-as-usual standups. And I typed the thing I had been thinking all week:
"What if it is a small team?"
Because most of this is drawn for a team of eight or nine with a dedicated PM, a dedicated designer, and enough engineers to lose two to discovery for a week without delivery stopping. That is not most of us.
The answer I have landed on after thinking about it: the tracks are two kinds of work, not two groups of people. A small team does not run two tracks by splitting in half. It runs them by budgeting time, exactly like the strategic and technical split earlier. Some fraction of the week is learning work, it is on the board, and it is reviewed in the sprint review.
And the smaller the team, the more the cheap tests matter, because you cannot afford to build the wrong thing twice. A fake door test and three interviews cost almost nothing and are worth more to a small team than to a big one. The discipline is the scale, not the headcount.
Where AI fits
Both Jeff and Sherif Mansour covered this, and their view was more restrained than I expected from an AI conversation in 2026.
On using AI to work. Atlassian generated more than sixteen thousand prototypes in ten months. When the cost of a prototype drops to nearly zero, the bottleneck moves. What becomes scarce is judgement: deciding what to test, saying no, and keeping a shared understanding across product, design and engineering. Sherif's specific warning was that product managers should not disappear into writing code with AI at the expense of steering.
On putting AI in the product. Chat is a useful universal interface early, because watching what people type into an open box tells you which workflows they actually want. Then you build those as real workflows. He called it paving the cow path. But chat also confuses new users who do not know what to ask for, so it is a discovery instrument more than a destination.
On when not to use it. When deterministic code is cheaper and better. Basic spam detection does not need a model. And match model cost to stakes.
On opt-outs. The sharpest bit. When a customer asks for granular AI toggles, do not just build the switches. That way lies a product where every feature has a dependency matrix. Use the five whys to find the actual concern, usually data handling or lack of training, and solve that at a higher level.
What I am actually taking from all of this
Write the intended outcome down before building, and the metric that would show it. If we cannot, we do not understand the request yet.
Say which MVP we mean. Every time.
Put value against risk, not value against effort, and send the top-left quadrant to a test rather than to the backlog.
Take the cheapest possible test seriously. Fake doors and paper are not a lesser form of engineering, they are how you avoid spending a month on an assumption.
Get an engineer into every customer conversation.
Budget time for learning instead of hoping it fits in the gaps, because on a small team it never fits in the gaps.
Part 1: Scrum gave us a cadence, not a compass
Part 1 argued that Scrum gave us a cadence and not a compass. This is the compass. It is not a new framework to adopt. It is mostly a set of questions asked earlier than we are used to asking them, and a willingness to throw work away while it is still cheap to throw away.

