Evidence

What the research says, and where it stops.

Everything below is evidence about the premises Facets is built on — peer versus self ratings, what feedback does, what makes a debrief work. None of it is evidence about Facets. That distinction is the point of this page.

Start with what we cannot claim

Facets has not been shown to improve team performance.

No validation study of this instrument exists. Demonstrating that it moves outcomes needs a randomized design with a comparison arm — partly because regression to the mean guarantees the lowest-scoring teams will appear to improve with no intervention at all. That study is planned, not done, and until it is finished the claim will not appear anywhere on this site.

The items are original drafts, not a validated scale.

They were written to the construct definitions and anchoring logic of published scales rather than copied from licensed instruments. They have not been through a psychometric pilot. The reliability figures used in the reports come from published estimates for peer ratings in general, and are labeled as provisional until real data replaces them.

Individual change is below the noise floor at this team size.

With two or three peers, the smallest change that could be reliably distinguished from measurement error is larger than any realistic six-month improvement. Team-level and cohort-level change are defensible; a personal score moving between waves is mostly not.

Why peers rather than self-report

d = 0.32

Leniency of self-ratings of job performance against supervisor ratings; Δ = .49 corrected for unreliability. The self–supervisor correlation is a separate analysis in the same paper — about .22 across 115 samples and 37,752 people.

Heidemeier & Moser (2009), 89 samples, 35,417 people, Western samples

.37

Interrater reliability between two peers rating the same person. Low enough that a single peer is weak evidence, which is exactly why the reports show a reliability band that widens when few people rated you.

Conway & Huffcutt (1997)

.80 vs .38

The same teamwork training, measured by third-party observers versus by self-report. Measurement source is not a detail.

McEwan et al. (2017), 72 interventions, 8,439 participants

~55–60%

Share of variance in multisource ratings attributable to the individual rater rather than the person rated. This is why scores are also reported perceiver-centered.

Scullen, Mount & Goff (2000)

The conclusion is not that self-knowledge is useless. It is that self and other are measuring partly different things, and the discrepancy between them carries information neither one has alone. That is why Facets asks you to rate yourself, and separately to predict how your teammates rated you.

A self–peer correlation of .2 to .3 is the normal finding, not a broken instrument. If a product tells you its self and peer scores agree closely, that is the thing worth questioning.

What feedback actually does

38%

Share of measured effects that were negative — feedback made performance worse. Average effect across all of them was d = 0.41, which is the number usually quoted on its own.

Kluger & DeNisi (1996), 131 papers, 607 effects, 12,652 participants

d ≈ .05

Peer-rated behavior change following multisource feedback with no structured follow-up. Effectively nothing.

Smither, London & Reilly (2005), 24 longitudinal studies

.25 vs .08

Effect when feedback is used for development only, against when it is mixed with administrative use. Tying feedback to a performance decision roughly thirds it.

Smither, London & Reilly (2005)

d ≈ .67

Effect of a debrief; .54 with outliers removed. The authors describe this as a 20–25% improvement, which is a percentile shift rather than a gain in output — the average debriefing team lands near the 75th percentile of teams that do not. Facilitated and structured outperforms unstructured, though the paper offers that as an explanation rather than as the basis of the headline figure. The average debrief ran 18 minutes.

Tannenbaum & Cerasoli (2013), 46 samples, N = 2,136

Read together, these say something uncomfortable about the category. A measurement instrument that delivers scores and stops is not a neutral product — it is an intervention with a real chance of making things worse, particularly when the feedback points at the person rather than the work.

Which is why the debrief protocol is not a paid tier and never will be, and why individual results are structurally unavailable to anyone who makes decisions about your job.

What the team-level questions measure

The climate scales are written to constructs with meta-analytic support: psychological safety, direction and role clarity, coordination, task cohesion, relationship conflict as distinct from task conflict, information flow and transactive memory, team efficacy, reflexivity, and viability. Trust is asked about each teammate rather than about the team, and pooled. Team processes as a whole relate to team performance at around ρ = .31 across 40 samples covering some 3,100 teams.

Worth knowing about that literature: the individual process dimensions are hard to separate empirically — they load onto a higher-order factor — so the reports treat them as a profile to discuss rather than as independent scores to optimize one at a time.

Things this product deliberately does not use

Team IQ or the “c-factor”. Independent replication attempts have partially or wholly failed, and one of the original papers has a published correction.

Project Aristotle. Unpublished, with no methods released. The underlying psychological-safety research is real and cited directly instead.

Faultline effect sizes from the 2011 meta-analysis. Retracted in 2016.

Belbin team roles, and type-based frameworks generally. The factor structures do not reproduce.

On the rest of the category

The published technical reports for the widely-sold instruments generally show decent internal consistency and decent short-interval test-retest for their continuous scores. The gap is elsewhere: reliability of a score gets reported in place of stability of the category that was actually sold to you — the four-letter type, the top five themes, the color. Categorical outputs are markedly less stable than the scores beneath them.

Two patterns are worth checking for yourself in any vendor’s technical documentation. First, whether peer or observer ratings appear anywhere, or whether every number traces back to self-report. Second, whether the outcome evidence is user satisfaction — people saying the result described them well — rather than behavior or performance. High self-endorsement is what a well-written horoscope produces too; it does not distinguish a valid instrument from a flattering one.

The honest version of the criticism is not that these tools do nothing. Structured team-building interventions do produce moderate effects on how teams feel and how they work together. It is that the personality debrief has never been isolated as the active ingredient — the structure and the conversation may be doing the work.

Sources

Every figure on this page is drawn from a published meta-analysis or primary study, named inline. The full research brief — every figure traced to its source with its verification status, and a list of what has not yet been checked against a primary source — is sent by email.

Get the research brief

Tell us who you are and we will send the research brief with every claim's verification status. The link is yours alone and works for a week.

Your address is used for this and, if you ticked the box, the notes. Nothing else.

This page is the argument. The instrument itself — every question, its anchors, keying, scoring and thresholds, and a plain list of what has not yet been measured about it — is published separately, so that the two cannot drift into each other.