Guide · Feedback

Why people overrate themselves: the self–peer rating gap

Ask a person how they are doing and ask the people around them, and the two answers disagree in a predictable direction. This guide covers what the research measures, why the gap exists, why it is the single most useful number a feedback tool can produce — and how to use it without turning it into a weapon.

About a 8-minute read · Sources at the foot · Last reviewed

What the research actually measures

The cleanest evidence comes from comparing what people say about themselves with what a second party says about the same behavior. Heidemeier and Moser (2009) pooled studies of employees’ self-ratings against their supervisors’ ratings of the same performance. Two findings matter here.

  • The gap is real and it points up. Self-ratings sat above supervisor ratings by d = 0.32 (89 studies, 35,417 people, Western samples). Read as a probability, about 3 in 5 employees rate their own performance above their supervisor’s rating of it.
  • The two views barely correlate. Across 115 samples and 37,752 people, self- and supervisor ratings correlated at r = .22 (ρ = .34 after correcting for measurement error). Knowing how someone rates themselves tells you surprisingly little about how they are seen.

The pattern is not confined to performance reviews. McEwan and colleagues (2017) reviewed 72 teamwork-training interventions (8,439 participants) and found that the same training looked far more effective when outside observers scored teamwork behavior (d = 0.80) than when the team scored itself (d = 0.38) — a 71% versus 61% chance that a trained team out-scores an untrained one, depending only on who is holding the clipboard.

The plain version: you are not the best available witness to your own behavior. Neither is anyone else. The people who work with you are a second witness, and the disagreement between the two is information.
Three views of the same personThree horizontal bars on a one-to-five scale: a self-rating of 4.4, a prediction of 3.9, and a teammates' rating of 3.6. A bracket marks the gap between the self-rating and the teammates' rating. Illustrative values.12345You4.4Your prediction3.9Your teammates3.6the gapIllustrative values. In a real note the bars are yours alone.
Illustrative values. Heidemeier & Moser (2009): self-ratings sit about a third of a standard deviation above supervisor ratings. The second, smaller gap — between what a person predicted and what their teammates said — is the one worth talking about.

Three reasons the gap exists

  1. Different evidence. You see your intentions, your effort, and the obstacles. Others see what arrived. When a task slipped because of three things outside your control, you rate the effort; they rate the slip. Neither is lying.
  2. Different base rates. People judge themselves against what they could have done and others against what people usually do. The first comparison is more forgiving, because the imagined self is always available and always slightly better.
  3. Motivated reading. A generous self-view is comfortable and, in most workplaces, mildly rewarded — self-ratings are often the opening bid in a review. The pressure is upward. It is not a character flaw; it is what the incentives ask for.

None of these is fixed by telling people to be more honest. They are fixed by adding a second observer and comparing.

Why the gap is the useful number

A score on its own is a verdict. A gap is a question. “Your teammates rate you 3.6 on speaking up early about risks” invites defensiveness. “You rated yourself 4.4, predicted your teammates would say 3.9, and they said 3.6” invites a different response: what are they seeing that I am not, and why did I already half-know it?

That second gap — between what you predicted others would say and what they did say — is the more useful of the two. A large self–peer gap with an accurate prediction means a person who knows how they are seen and disagrees; that is a conversation about standards. A small self–peer gap with a wildly wrong prediction means a person who does not know how they land; that is a conversation about signal. Most feedback tools cannot tell these apart, because they never ask for the prediction.

Four patterns of self, prediction and peer ratings, and what each usually means
PatternSelfPredictionPeersWhat it usually meansUseful next step
Aware, disagreesHighLowLowKnows how they are seen, holds a different standardTalk about the standard, not the score
UnawareHighHighLowDoes not know how they landAsk for one specific example
UnderratesLowLowHighHarsher on themselves than the room isName the strength out loud
AlignedCalibrated — rare, and worth noticingNothing to fix

What not to do with it

  • Do not use it to rank. The gap is a property of one person’s self-knowledge, not a comparison between people. Two people with the same peer score and different self-scores are not “better” or “worse”; they are differently informed.
  • Do not put it in a review. The moment a self–peer gap can affect pay or promotion, the self-rating becomes strategic and the number is gone. Facets keeps the team survey structurally separate from the manager tools for exactly this reason — no table on one side references the other, apart from the record of a summary a person chose to send.
  • Do not act on one rater. Below 3 raters a peer score is close to a name tag. Below 5, question-by-question detail is noise with a decimal point.
  • Do not diagnose. The research says that the gap exists, not why this person’s does. The reasons above are the population’s, not the individual’s.

How Facets shows it

In a Facets team run each person rates every teammate on observable behavior from the last three months, rates themselves on the same items, and predicts how their teammates rated them. The private note each person receives puts the three views side by side, names what teammates see going well, names one thing to work on, and shows the gap. Nobody else can open it — not the team lead, not the workspace owner, not a manager. A named facilitator can see it before release; that is the only exception. The team-level pack shows the team’s climate as group numbers and never names or ranks an individual.

That structure is a deliberate answer to the research above: add a second witness, ask for the prediction, show the comparison, keep it private, and put the conversation — not the score — at the center. How the team survey works, step by step · The items themselves · The conversation it is built for.

Frequently asked questions

Do people usually overrate themselves at work?
On average, yes. Across 89 studies and 35,417 employees, self-ratings sat about a third of a standard deviation above supervisors' ratings — roughly a 3-in-5 chance that any given person rates themselves higher than their supervisor does.
Is a self–peer gap always a bad sign?
No. It measures self-knowledge, not performance. A person who predicts accurately how others see them and still disagrees is holding a different standard; that is a conversation about standards, not a problem to correct.
How accurate are peer ratings compared with self-ratings?
Peers are a second witness, not an oracle. Several peers aggregated cancel out one rater's bias in a way a single self-rating cannot, which is why below 3 raters a peer score should not be shown at all.
Should self-assessments be part of performance reviews?
That is a policy choice, but the evidence says self- and supervisor ratings barely correlate (r ≈ .22), and a self-rating that affects pay becomes an opening bid. Keep developmental self-ratings out of evaluative records.
How does Facets measure the gap?
Each person rates teammates, rates themselves on the same behaviors, and predicts how teammates rated them. The private note shows all three side by side. Nobody else can open it.

Sources and notes

Heidemeier, H., & Moser, K. (2009). Self–other agreement in job performance ratings: A meta-analytic test of a process model. Journal of Applied Psychology, 94(2), 353–370. · McEwan, D., Ruissen, G. R., Eys, M. A., Zumbo, B. D., & Beauchamp, M. R. (2017). The effectiveness of teamwork training on teamwork behaviors and team performance: A systematic review and meta-analysis of controlled interventions. PLoS ONE, 12(1), e0169604.

Probability-of-superiority glosses are computed as Φ(d/√2). The two Heidemeier & Moser figures are different analyses in the same paper and are reported separately on purpose. Full discussion on the research page.

See your team's three views.

Everyone rates everyone, rates themselves, and predicts how they were rated. Each person gets the comparison in a note nobody else can open.

All guides