Engage Logo

Calibration

Calibration is the process of comparing proposed performance ratings across managers before they are finalised, so that a given rating means the same thing in every team. Done properly it corrects for differing manager standards; done badly it becomes quota allocation with a meeting attached.

The problem it solves

Two managers assess the same standard of work differently. One rates generously because they dislike difficult conversations, or because they believe advocating for their team is part of their job. Another rates conservatively because they hold a high bar or want to leave room to reward improvement.

Left alone, the result is that a rating tells you as much about the manager as about the employee. Pay, promotion and, eventually, termination decisions then rest on a measure that is not comparable across the organisation.

Calibration is the correction. Managers present their proposed ratings and the evidence behind them to peers, the group challenges the reasoning, and ratings are adjusted where the evidence does not support them or where the same evidence produced different ratings in different teams.

The important word is evidence. Calibration works when the discussion is about what people did. It fails when it is about how many of each rating there should be.

Running a session

  • Assemble managers at the same level, with a facilitator who is not one of them and who has no team in the room.
  • Have each manager present specific evidence for the ratings at the extremes first, since the top and bottom carry the most consequence and attract the least challenge.
  • Ask the comparison question directly: would this person get the same rating in another team here? Where the answer is no, one of the two ratings is wrong.
  • Challenge the reasoning rather than the number. A manager who cannot describe what the person did to earn a rating has not assessed them.
  • Watch for the standard biases as they appear: recency, the halo from one visible success, likeability, and the tendency to rate quiet contributors lower than voluble ones.
  • Record what changed, for whom, and on what basis. This is the part most often skipped and the part most needed later.

Sessions run to time when the facilitator is willing to stop a manager who is negotiating rather than evidencing. Sessions overrun when nobody does.

What it must not become

Calibration and forced distribution are different things that frequently occupy the same meeting.

CalibrationDistribution
AsksDoes this rating match the evidence and the standard elsewhere?How many can be rated at each level?
ConstraintConsistencyBudget or quota
Can ratings move up?Yes, and they should sometimesRarely, in practice
RecordEvidence and reasoningThe allocation

The diagnostic is simple. If ratings only ever move downward in your calibration sessions, it is not a calibration session. That is not necessarily wrong as a management practice, but it should be described accurately, because managers who are told they are calibrating and are in fact allocating learn to inflate their initial ratings to leave room, which corrupts the input.

If a budget constraint exists, apply it after the ratings are settled, to the money rather than to the assessment.

After the session

  • Managers communicate the final rating and own it, including where it changed. A manager who tells their report that they proposed higher and were overruled has transferred the difficulty and undermined the process.
  • The reason for a changed rating is given in the same terms as any other rating: what the evidence showed.
  • The aggregate distribution is reviewed by group as well as by team. Persistent differences in rating outcomes between demographic groups doing similar work are worth examining, and calibration data is where they show.
  • The record is retained, because a rating that later supports a pay decision, a promotion refusal or a termination will be examined, and the calibration note is part of what makes it defensible.

Employees should be told that calibration happens, in general terms. A process that adjusts ratings after the manager conversation, and that nobody has explained, reads as arbitrary when someone eventually discovers it.

Frequently asked questions

What is calibration in performance management?

Comparing proposed ratings across managers before they are finalised, so a given rating means the same thing in every team. It corrects for the fact that managers apply different standards to the same quality of work.

How is calibration different from forced distribution?

Calibration asks whether a rating matches the evidence and the standard applied elsewhere. Distribution asks how many people may be rated at each level. If ratings only ever move down in your sessions, you are running the second and calling it the first.

Who should be in a calibration session?

Managers at the same level, with a facilitator who has no team in the room. The facilitator's main job is to stop discussion that has turned into negotiation over numbers rather than examination of evidence.

Should employees be told that calibration happens?

Yes, in general terms. A process that adjusts ratings after the manager conversation and that nobody has explained reads as arbitrary when an employee eventually finds out about it.

What should be recorded from a calibration session?

What changed, for whom, and on what basis. That record is what makes a rating defensible later, when it supports a pay decision, a promotion refusal or a termination that someone questions.

How Engage supports calibration

Engage brings proposed ratings and the evidence recorded against them into the same view, so a calibration session can compare what people actually did rather than trade numbers. Changes are recorded with their reasons and retained against the employee record, and the resulting distribution can be read by team, manager and group, which is where inconsistency shows up before it becomes a grievance.

See performance management in Engage
WhatsApp