Picture of Sara Gallagher
Sara Gallagher

Why Do We Measure Decisions by Their Outcomes

The result is the scoreboard everyone can read: a good call is the one that worked. But what if grading decisions that way is what teaches your best people to make worse ones?


In my first chief exec role, I learned how rarely a good decision is sitting there waiting to be found. Usually everything on the table has something wrong with it, and the job is picking the option that sucks the least.

Almost every call I make runs on incomplete information. Facts that aren’t available or knowable, data that’s gone stale, events outside my control and my view. (A lot of the AI decisions on executive desks right now are exactly this: a big commitment made on information that’ll be out of date in weeks.)

Then it plays out, and the right answer looks obvious. Once you know how something turned out, it’s hard to remember it was ever unclear. That’s hindsight bias, and it makes grading a decision by its result feel like common sense. It worked, so it was a smart call. It blew up, so it was a dumb one.

The longer I lead, the less I trust that instinct. I’ve come around to thinking it’s the grading itself that trains good people to make worse ones.

 

Why does a good outcome feel like proof of a good decision?

One reason is dull and obvious. You can see an outcome and measure it.

But outcomes also carry emotional stakes, which makes them hard to dismiss. When a call works and real value shows up (or real disaster gets dodged), we want to reward whoever made it. We won a hard game. Of course we want to celebrate. When it goes the other way, the loss lands harder. (There’s a name for that: loss aversion. The pain of losing runs about twice as intense as the pleasure of gaining the same thing.)

Even in healthy cultures, we tend to call a bad outcome a bad decision. In unhealthy ones, we pin it on the person who made it.

It’s such an intuitive way to keep score that I don’t think most of us stop to ask what else we could measure.

Baron and Hershey put a name on the habit in 1988. Outcome bias works like this: you grade the call by how it ended, never mind what you knew going in. Which sounds obvious right up until you catch yourself doing it.

Give people two identical decisions, change nothing but the ending, and they rate the one with the happy ending as the better decision. Same information, same logic, same call…just different luck.

A 2023 replication found it running stronger than the original. It held up even among people who said outcomes shouldn’t count. They agree in principle, then do it anyway.

Annie Duke, the poker champion, has the cleanest name for it. In Thinking in Bets, she calls it “resulting”: treating the quality of a decision and the quality of its result as though they were the same thing.

They are not the same thing. A careful decision can end badly. A reckless one can get rewarded with a win. Keep score only on results and you can’t tell them apart. And if you can’t, nobody gets better at deciding.

When the outcome is the grade, people protect themselves

When people believe they’ll be judged on the outcome instead of the judgment, they get careful in expensive ways. Watch it move through a team.

They wait. A decision that could be made today slides to next week, because next week might bring more information, and more information feels like less risk. (I wonder how true that is. The literature on information overload says more information helps up to a point, then starts working against you. I’ve written before about how long teams take to decide, and why.)

They soften the escalation. Escalating means arriving with a recommendation and the data behind it. A team afraid of being wrong arrives with options instead.

They spread the accountability around. Nobody wants to be the single name on a bad outcome, so a call that needs one senior decision-maker ends up with three, four, or five signatures.

None of this is a character flaw. It’s the rational response to a plain fact: when you’re graded on a result you can’t fully control, protecting yourself and your team from blame is the smart play. Underneath it sits one fear: being the person holding the bag when a call goes sideways. When raising a risk feels personally dangerous, you hedge. You wait. You keep your head down.

I’ll take a defensible wrong call over a lucky right one

Outcomes still tell me something, particularly about decision quality over a long stretch. But decision by decision, it matters more to me that a call was defensible than that it turned out right.

Defensible means:

  • Based on principles (criteria) defined up front
  • Made at the right time
  • No bigger than it needs to be (decide what you need to move forward and learn…don’t decide what isn’t necessary)

Defensibility is what produces judgment. A lucky outcome teaches you nothing about deciding better. A defensible decision with a bad outcome teaches you plenty: adjust the criteria, fix the timing, right-size the next.

A lucky outcome teaches you nothing about deciding better. A defensible decision with a bad outcome teaches you plenty.

 

Outcomes stop being the grade. You start asking how you got there. The question moves from “did it work?” to “can you show me your reasoning?”

And no, this doesn’t let anyone off the hook. Accountability moves to the quality of the thinking, a harder place to hide than a lucky result.

You can’t grade judgment you never defined

That first item (principles defined up front) is the one I’ve had to get most deliberate about. If judgment is what you’re grading, somebody has to say out loud what good judgment looks like.

So when I can’t see the answer, I write the criteria down first. Otherwise I can’t recall why I chose what I chose, or work out how to choose better.

Take AI. I’ve stopped trying to make organizational decisions about AI on data + vibes, which is where I think a lot of leaders are stuck. A few of us here are defining the criteria we’ll use instead, even as the information under them changes weekly.

  • Reversibility: How easily can we undo this decision?
  • Cost risk: What’s the likelihood cost will hold or climb?
  • Security: How mature is this product or process at protecting our data?
  • Maintainability: Can this survive without one power user?
  • Lock-in: How expensive would it be to leave this vendor or product?
  • Burden on our people: What’s the learning curve, and how much does it disrupt what people do today?

They give us something consistent to reason from. What’s on the list will change. Having them is what speeds us up.

Grade only outcomes and you teach the people under you to hide risk and pad estimates. Ask for the reasoning and you get faster, more proactive teams.

Waiting for the picture to clear is a decision too

Timing and size are where defensibility falls apart most often. Timing first. The cost of deciding now is usually lower than the cost of waiting, and almost nobody prices the waiting.

Some decisions announce themselves. A hire, a vendor, a project you greenlight or kill: there’s a decision on the table, so you treat it like one.

The one that slips past is how long you wait before you commit. Waiting feels responsible, like doing your homework first. So it stays a non-decision, something you never quite notice making, while the days stack up.

I’ve made this mistake more than once. The clearest case came early in the pandemic. One of the first standing decisions to land on me as President of The Persimmon Group was whether to bring the team back to the office. It felt enormous. My team was divided, feelings ran hot, and I was brand new in the chair with everyone watching how I’d handle it. I badly wanted to get this right.

So the call never got made. We stayed virtual, month after month, while I held out for a clearer picture. It never came. My team got months of low-grade unease. Some got lonelier by the month, ready to be back with people and sick of the wait. Others were happy at home and still on edge, waiting to hear if we’d order them in.

We did get there. We built triggers pegged to local case counts, with red, yellow, and green statuses we could shift between as conditions changed. Green was come in, basic restrictions only. Yellow was come in under tighter ones. Red was stay home.

Good framework. Also one we’d thought of in week 1 and then sat on while we “weighed all our options.”

To me, what made the call feel that big was the audience. I was new, and people were watching. I froze. It cost the team months of peace of mind, for information I never needed.

Why do we keep hunting when “good enough, now” would do?

Now decision size.

There are two honest ways to make a decision, and we pick the wrong one all the time. You can hunt for the best option there is, and keep looking until you’re satisfied there isn’t a better one (that’s maximizing). Or set a bar, take the first option that clears it, and move (satisficing).

Either one can be the right call. You’d think the expensive mistake is the rushed one. In my experience it’s the surprising one: we maximize when we should satisfice, and that’s the more common error.

A major change in marketing strategy earns that hunt. Choosing the team’s default video tool does not.

I was too late on the back-to-office call. But that same time period, we got one decision exactly right. Like a lot of organizations, we took the whole company virtual almost overnight. That meant moving from Skype to Teams in weeks when we’d planned on months.

Instant chaos. When do we build a team versus a channel? When is something a chat versus an email? Why are we on this thing in the first place?

So a few of us sat down, put one page of guidelines together in a day or so, and rolled it out at a company all-hands days later. We said plainly that it was a first pass, something we’d revise as we learned the tool. And we did, many times over. But moving fast on a “good enough” version cut the chaos short when we couldn’t afford to be slow.

When I disagree with the call, I want to see the thinking

Look. I tell my own team this. When I’m out and you’ve got a call to make, I won’t grade you on whether it worked. If I come back and disagree with what you chose, what I want to see is the thinking.

We keep guiding principles for exactly this. Ours are written down and say what counts as a sound basis for a decision. An example: Put the customer first. So if someone adds scope to an agreement for a discount and I can see the reasoning, it goes in the “good decision” column…even if I’d have decided differently.

If I see a call that got rushed when it deserved real thought, that’s a coaching moment. The opposite counts too: the call that sat for weeks when it should have been made in the room. An organization that only coaches one of them never gets better at deciding.

I have people who coach me too.

I no longer believe leaders have some special lock on good decision-making. I still make calls that turn out wrong, more often than I’d like. But I have people, from the front lines to my peer leaders, who show me where my reasoning could have been tighter. That’s how I grow. And it’s how the organization gets better at making good calls, most of the time, on the information we actually have.

Judgment. Not luck.

If You Only Do One Thing

Pull up the last handful of decisions you reacted to on your team, praise and criticism both. For each one, ask yourself one question: was I reacting to the result, or to the thinking that got us there?

If it was mostly the results, you may be sending a message you didn’t intend: the goal is to be right rather than to decide well. Outcomes are often outside anyone’s control. Judgment is what compounds. It’s what makes the next call, and the hundred after it, faster and better.

Until next time,
Sara

Sara Gallagher

Sara Gallagher helps PMO Leaders, CIOs, and CTOs execute strategy smarter, faster, and kinder—by making valuable work easier to do. She leads The Persimmon Group, a consultancy focused on unsticking teams. 

More BIg Dumb Questions

Do PMOs still need leaders with project management experience?

Issue #53 · On the Difference Between a Named Gap and an Invisible One

Why Do We Measure Decisions by Their Outcomes

Issue #52 · On What Keeping Score by Results Is Teaching Our Teams

Have Executives Moved Past Project Management?

Issue #51 · On the Difference Between Managing Projects and Delivering Benefits

BDQ in your inbox

Never miss an issue.

The obvious questions no one wants to ask first — in your inbox every week.