How to Grade OKRs and What a Good Score Actually Is

OKR Scoring: How to Grade OKRs and What a Good Score Actually Is

Most teams meet OKR scoring for the first time in the last week of the quarter, which is exactly the wrong moment to be deciding what a number means. Someone asks what we should put against a key result that landed at 62 percent of target, three people give three different answers, and the score that gets recorded is whatever survives the meeting. The grade is then filed and never used again.

That is the failure mode worth naming before any of the mechanics. Scoring is not administration. It is the point in the cycle where an organization finds out whether it sets targets it can read.

What is OKR scoring?

OKR scoring is the assessment, at the end of a cycle, of how far each key result moved against the target it was set, expressed as a single figure. Google’s own guidance puts the mechanism plainly: OKRs are graded on a scale of 0.0 to 1.0, with 1.0 meaning the objective was fully achieved.

The score is not the output. What you do with it is. A grade that produces no change in how the next set is written has cost the organization a quarter of measurement and returned nothing, and that is the most common state we find scoring in when we arrive.

What scale should you use to score OKRs?

Three scales are in general use, and the choice between them matters less than using one consistently across every team. IBM sets out the same three: a percentage of completion, a colour-coded status, and a numerical value, usually between 0 and 1.

ScaleHow it worksReads well whenBreaks down when
0 to 1 decimalEach key result gets a decimal against target. Atlassian anchors it as 0.3 for missing by a lot, 0.7 for real progress short of target, 1.0 for hitting the stretch targetKey results are genuinely numeric and the organization is comfortable with partial creditTeams read 0.7 as a C grade rather than a success
Percentage completeThe key result is scored at its completion level, so a half-finished result scores 50%Leadership already reads everything in percentages and you want no translation costIt flatters milestone-shaped key results that finished on paper and changed nothing
Red, amber, greenA status colour rather than a figure, where green means on track or completed and red means off trackThe first cycle or two, when precision would be falseEvery quarter after that. Colours hide the distance between 0.4 and 0.6

Pick one and hold it for at least four cycles. The value of a score comes from comparing it to the last one, and an organization that changes scale every other quarter has thrown away its own baseline.

How do you score a key result, step by step?

Take the starting value, the target value and the value you finished on, and express the distance travelled as a fraction of the distance you committed to. A key result that moved support response time from 40 hours to 22 against a target of 16 travelled 18 of the 24 hours it promised, which is 0.75.

Two details decide whether that number is worth anything. The starting value has to have been recorded when the key result was written, not reconstructed at quarter end, because a baseline recalled from memory tends to move in whichever direction makes the score look reasonable. And the owner scores their own key result first. What Matters notes that self-assessment adds subjective judgment to the objective score, which is a feature here rather than a problem: the person closest to the work is the one who knows whether the number is honest.

Objective-level scores follow from the key results beneath them. Some organizations average them, others weight them. Either is defensible. Averaging an objective whose key results are not equally load-bearing is not, and that is worth settling before the quarter starts rather than during the scoring meeting.

What is a good OKR score?

Somewhere around 0.6 to 0.7 for ambitious key results. Google states the target range directly: the sweet spot for OKRs is somewhere in the 60-70% range, and Atlassian is more blunt about the implication, noting that scoring 0.7 on a key result is considered a success.

This is the single hardest thing to land with a leadership team, because 0.7 looks like a failing grade to anyone who went to school. It isn’t, and the arithmetic explains why. A team that consistently scores 1.0 has been setting targets it already knew it could hit, which means the target was a forecast rather than a commitment to stretch. Google’s own framing is that scoring higher may mean the aspirational goals are not being set high enough, while scoring lower may mean the organization is not achieving enough of what it could be.

One qualification that gets lost in the retelling. Not every key result is aspirational. What Matters separates the two: for committed OKRs, grading is pass/fail and should hit a score of 1.0, while for aspirational ones there is a spectrum between greatness and failure. A team that applies the 0.7 expectation to a compliance deadline has misread the instrument.

How often should OKRs be scored?

At the end of each cycle, which for most organizations means quarterly. Google grades organizational OKRs annually and quarterly, and What Matters places formal OKR sessions on a regular rhythm, quarterly, during each OKR cycle.

Weekly check-ins are a different conversation and should stay one. A check-in asks what we change in the next seven days. A score asks what we achieved and what that says about how we set targets. Collapse the two and every weekly meeting turns into a performance review, and teams respond to performance reviews the way people always have, by setting targets they are confident of hitting. You end up with precise scoring of unambitious goals, which is the most expensive possible outcome of a well-run process.

Should OKR scores be used in performance reviews?

No. Google’s guidance is explicit that OKRs are not synonymous with employee evaluations and that they are not a comprehensive means to evaluate an individual.

The reason is mechanical, not philosophical. The moment a score affects a rating or a bonus, the person writing next quarter’s key results is negotiating their own grade, and they will negotiate well. Ambition becomes a personal risk that nobody is paid to take. Organizations that want both usually try to firewall the two conversations by calendar or by owner. In our experience that holds for about two cycles, until the first person is visibly rewarded for a set of key results that were never in doubt, and then everyone recalibrates at once.

What does a full set of 1.0 scores tell you?

That the targets were forecasts. Atlassian’s version of the same warning is direct: if you are regularly scoring 1.0 on your results, the targets need to be more ambitious.

There is a second reading worth checking before you act on the first. Sometimes a perfect set means the key results were activity disguised as outcome, the kind that score 1.0 because the team did the thing, not because anything moved. Launch the portal, run the twelve workshops, publish the policy. All completed, all scored 1.0, and the behaviour the objective existed to change is exactly where it was in January. Look at what each key result measured before you conclude the team was playing safe.

Where does OKR scoring go wrong?

Scores get assigned to key results nobody has looked at since week three, using baselines nobody recorded, in a meeting where disagreement is socially expensive. The number that comes out is a negotiated artefact rather than a measurement.

The pattern our coaches see most often is an organization that scores diligently and never reads the scores back. Grading without reflection is bookkeeping.

What do you do with the scores once the quarter closes?

Read them as a set before writing the next one, and look for the pattern across teams rather than the result of any single key result. What Matters frames this as the point of the exercise: reflection helps us understand what happened and figure out how to improve things for the next cycle.

Three patterns are worth acting on. A team clustered at 1.0 is not stretching. A team clustered below 0.3 either took on too much or lost the resourcing argument somewhere in week four. And a team whose scores are scattered widely across its own objectives usually has a prioritisation problem rather than an execution one, because the spread says effort went where attention happened to be. None of that shows up while you are scoring one key result at a time.

Frequently asked questions

What is a good OKR score?

Around 0.6 to 0.7 for ambitious key results. Committed key results that must be delivered are graded pass or fail and should reach 1.0.

How do you calculate an OKR score?

Express the distance the metric travelled as a fraction of the distance committed to, using the baseline recorded when the key result was written. Moving from 40 to 22 against a target of 16 is 0.75.

Should OKRs be scored every week?

No. Score at the end of the cycle, usually quarterly, and keep weekly check-ins as a conversation about what changes next week.

Who scores an OKR?

The person who owns the key result scores it first. The team lead runs the conversation about what the set of scores means together.

Do OKR scores affect performance reviews?

They should not. Tying scores to ratings or pay makes ambition a personal risk, and teams respond by setting targets they already know they will hit.

What if a key result can’t be scored numerically?

Score it pass or fail and rewrite it next cycle. A key result that resists scoring is usually describing an activity rather than a measurable change.

How many key results should an objective have?

Two to five. Atlassian’s guidance is the same, and beyond five the set stops being memorable enough to score meaningfully.

If your scores get recorded and never read back, that is a reflection problem rather than a scoring one, and it is usually fixable inside one cycle. Our coaches rebuild the scoring and reset sequence as part of Implémentation des OKR, and it is the most common reason an organization calls us about OKR consulting rather than another training cohort.

About the author

Dirk Schmellenkamp is Founder and CEO of the OKR Institute, which has certified more than 70,000 practitioners and run OKR implementation programmes for over 1,000 organizations across 50+ countries since 2017. The OKR Institute is accredited by the International Coaching Federation, the Association for Coaching and HRD Corp.

PDG de l'Institut OKR