Published methods

The instrument, item by item.

Every question Facets asks, with its anchors, keying, scoring and thresholds, read from the same code the surveys run on. Also what has been measured about it, which at this revision is nothing, stated in the same voice as everything else.

Revision 2026-09-15·Team survey 0.5.0 · Peer feedback (development review) peer-0.5.0 · Leadership 360 0.1.0 · Engagement cycle 0.1.0

Each instrument below shows its scales and first dimension to everyone. The rest is sent by email, so we know who is reading and can tell you when a version changes.

How to read this page

Every statement here is one of three kinds, and the labels beside each section say which. The distinction matters because the usual failure in this category is a claim about why an instrument is built a certain way being read as a claim about how well it works.

Specified

What the instrument is: items, keying, anchors, scoring steps, thresholds, versions. A description of the artefact, which claims nothing about it.

Reasoned

Why it is built this way: the construct definitions and literature each scale was written to, and the composition model behind aggregation. Each such statement carries its evidence on the research page.

Measured

How well it works: reliability, factor structure, invariance across waves, criterion validity, norms. No such evidence exists yet for any of these instruments.

The rule every version follows

Wording is the invariance guarantee. Items are only ever added under a new version, never edited under the old one. A construct that turns out to be badly covered gets a new item under the next version; it does not get a better version of an existing item.

Wording is never duplicated in the database: responses are stored against an item id and an instrument version, and the text lives only in code. That protects comparability across waves by construction. It has not yet been tested as an empirical property, and the section on what has not been measured says so.

Instrument · version 0.5.0

Team survey

A round-robin instrument for intact teams of three to twelve. Each member rates every teammate on observable behaviour, rates the team as a whole on 11 core climate scales (plus 2 conditional ones, administered only to teams that work apart or that opt in), names teammates on four network prompts, and rates themselves on the same behaviours they were rated on.

Response scales

Specified

frequency5Five-point frequency, with a not-observed option (1–5)

1
Rarely or never
2
Occasionally
3
About half the time
4
Usually
5
Almost always
NO
Not enough opportunity to observe

Each dimension also carries written behavioural anchors for the 1, 3 and 5 points, shown to the rater before rating.

agreement7Seven-point agreement (1–7)

1
Strongly disagree
4
Neither agree nor disagree
7
Strongly agree

progress7Seven-point progress (1–7)

1
Not at all
2
Barely
3
A little
4
About halfway
5
Mostly
6
Almost fully
7
Fully

Used only by the six-week pulse, for how far a recorded commitment has happened.

Peer ratings

Specified
Referent
Each teammate, rated by every other member
Scale
frequency5
Stem
In the last 3 months, how often did {name} …

For a team younger than the rating window the stem reads "Since this team formed, how often did {name} …". Core dimensions are always administered; the 3 extended dimensions rotate between raters when the team has fewer than 5 members so that every target still receives each extended item from at least two raters. A wave administers only the dimensions that existed at its pinned instrument version.

D1Dependable contribution

CATME-B 'Contributing to the team's work'; Marks et al. (2001) action processes.

1
Misses commitments or delivers work others must redo; waits to be assigned.
3
Delivers most commitments; occasionally needs reminding; takes a reasonable share.
5
Consistently delivers on time at high quality; volunteers for hard or unglamorous work.
D1adeliver what they committed to, on time and at the quality the team needed?
D1btake on a fair share of the work without being asked?

87 more items, the thresholds, the scoring steps and the change log for this instrument open with the emailed link — request it below.

Instrument · version peer-0.5.0

Peer feedback (development review)

The team survey's peer dimensions administered to raters the subject chose rather than to an intact team. The items, anchors, scale and self-rating pass are identical to the team survey and carry its version number; only the free-text prompts differ, because the team survey's were addressed to a teammate.

Response scales

Specified

frequency5Five-point frequency, with a not-observed option (1–5)

1
Rarely or never
2
Occasionally
3
About half the time
4
Usually
5
Almost always
NO
Not enough opportunity to observe

Each dimension also carries written behavioural anchors for the 1, 3 and 5 points, shown to the rater before rating.

Peer ratings

Specified
Referent
The subject, rated by each invited colleague
Scale
frequency5
Stem
In the last 3 months, how often did {name} …

Dimensions D1–D8 exactly as in the team survey, including the written anchors.

D1Dependable contribution

D1adeliver what they committed to, on time and at the quality the team needed?
D1btake on a fair share of the work without being asked?

19 more items, the thresholds, the scoring steps and the change log for this instrument open with the emailed link — request it below.

Instrument · version 0.1.0

Leadership 360

Upward feedback. A leader invites the people who work with them; each rater rates the leader on seven observable leadership behaviours, two overall indices and four free-text prompts. The leader completes a self-rating and a meta-perception on the same behaviours.

Response scales

Specified

frequency5Five-point frequency, with a not-observed option (1–5)

1
Rarely or never
2
Occasionally
3
About half the time
4
Usually
5
Almost always
NO
Not enough opportunity to observe

Each dimension also carries written behavioural anchors for the 1, 3 and 5 points, shown to the rater before rating.

agreement7Seven-point agreement (1–7)

1
Strongly disagree
4
Neither agree nor disagree
7
Strongly agree

Leadership behaviours

Specified
Referent
The leader, rated by each invited rater
Scale
frequency5
Stem
In the last 3 months, how often did {name} …

Self-rating stem: “In the last 3 months, how often did you …”. Meta-perception stem: “How do you think your team rated you: how often did you …”. Raters record their relationship to the leader (direct report, peer, other); it is used for nothing that could identify them.

L1Direction and clarity

Hackman's compelling direction; contingent-reward clarity (Judge & Piccolo 2004); role clarity as a psychological-safety antecedent (Frazier et al. 2017).

1
Priorities shift without explanation; people learn what mattered after the fact.
3
Goals are usually clear; trade-offs and the 'why' are sometimes missing.
5
Everyone can say what matters most this quarter and why; changes are explained when they happen.
L1amake clear what the team's priorities were and why they mattered?
L1bset expectations for your work that were specific enough to act on?

19 more items, the thresholds, the scoring steps and the change log for this instrument open with the emailed link — request it below.

Instrument · version 0.1.0

Engagement cycle

An organisation-wide cycle runs the team survey's core climate scales inside each reporting unit (team referent, unchanged, so unit scores are comparable with a team run's; the conditional scales are never administered in a cycle) plus the six items and two open questions below at the individual and organisation referent. The token “{org}” is replaced with the organisation's display name at render; the stored text never changes.

Response scales

Specified

agreement7Seven-point agreement (1–7)

1
Strongly disagree
4
Neither agree nor disagree
7
Strongly agree

recommend11Eleven-point likelihood (0–10)

0
Not at all likely
10
Extremely likely

Reported as a distribution and a mean. No promoter or detractor index is computed or labelled.

Engagement, intent to stay and recommendation

Specified
Referent
The respondent, about themselves and their organisation

Individual-referent. These items aggregate as plain means with error bars. No agreement statistic is computed for them, because “the organisation agrees it is engaged” is not a sentence this instrument can support.

engagementWork engagement

EN1I put sustained effort into my work, even on difficult days.

Vigour, per Schaufeli's definition: willingness to invest effort and persistence in the face of difficulties. Original wording — UWES's vigour items lead with felt energy, this leads with invested effort under difficulty.

EN2The work I do feels worth doing.

Dedication, specifically the sense of SIGNIFICANCE. The one facet in the published definition that UWES-9 carries no dedicated item for, which is why it is both farthest from their wording and better coverage than another enthusiasm item would be.

EN3I get deeply involved in the work I am doing.

Absorption: being fully concentrated and deeply engrossed. Tracks the definition rather than the items. KNOWN WEAKEST OF THE THREE — absorption is the likeliest to collapse into dedication under CFA and the least actionable for a manager; the remedy if it does is EN4 under v0.2, never a reword.

5 more items, the thresholds, the scoring steps and the change log for this instrument open with the emailed link — request it below.

Request the full specification

One link, sent to you, unlocks every item on this page plus the PDF and JSON export for thirty days in this browser.

Tell us who you are and we will send the full item bank as a PDF, plus the machine-readable JSON and the rest of the page. The link is yours alone and works for a week.

Your address is used for this and, if you ticked the box, the notes. Nothing else.

How scores are aggregated, and why

Reasoned

Climate scales are written and aggregated under a referent-shift consensus model (Chan, 1998): every item asks about the team, so the team mean is meaningful only when members agree. Agreement is tested two ways on the member × item matrix — rwg(j) under a uniform null (James, Demaree & Wolf, 1984; LeBreton & Senter, 2008) and the average deviation index AD_M(J) against its A/6 criterion (Burke, Finkelstein & Dusig, 1999; Burke & Dunlap, 2002) — and a scale that fails either is reported as a spread of views rather than a shared one. The between-team ICC(1) of that literature (Bliese, 2000) needs many teams and is not computable for one; it will be reported for the pooled sample once there is one.

Self-state blocks and the pooled trust item are individual- and dyad-referent respectively and aggregate as plain means; no agreement statistic is reported for them.

Peer ratings of an individual follow the Social Relations Model (Kenny, Kashy & Cook, 2006): each rating carries a perceiver effect, a target effect and a relationship effect. Full estimation needs groups of four or more and two or more items per construct; at the sizes this instrument serves, the perceiver effect is removed by centring and the target effect is the reported score.

Reliability of an aggregate of raters is projected with the Spearman–Brown formula from a single-rater reliability of .37 (Conway & Huffcutt, 1997). Two raters give about .54, three .64, four .70, six .78. This is arithmetic on a published figure, not a measurement of this instrument.

The engagement items are individual-referent and aggregate as plain means; no agreement statistic is reported for them.

The evidence behind each of these choices, with its verification status, is on the research page and in the research brief it offers.

What has been measured

Measured

Nothing yet. Specifically:

  • No reliability estimates from real data: no alpha, no omega, no standard error of measurement.
  • No factor structure. 13 correlated climate constructs in 40 items is a discriminant-validity bet that a confirmatory factor analysis may not support.
  • Invariance across waves is protected by construction (wording is never duplicated in the database) and has never been tested. Configural, metric and scalar invariance all remain open.
  • No criterion validity. Nothing links any score to any outcome.
  • No fairness or adverse-impact analysis.
  • No norms, by decision: benchmarks against self-selected samples with untested invariance are not published. Teams and organisations are compared with their own history and their own internal spread.
  • The items added in team survey 0.4.0 were chosen by reasoning from construct facets, not selected from pilot data.

This section is left empty rather than filled with assertions, and it will be filled under new revisions as data arrive. The counts a reader can check against the tables above: 16 peer items in seven dimensions, 40 climate items in 13 scales, 13 internal-state items, 8 calibration vignettes, 14 leadership items, 6 engagement items.

Known limitations and open questions

Reasoned
  • Nothing in the measured section can be filled in until enough teams have run the survey. Confirming that the climate scales measure distinct things needs roughly 500–700 respondents across 60–100 teams, by the usual rules of thumb for a 40-item, 13-factor model. Until that sample exists, everything on this page is a description of the instrument or a reason for a choice in it, never a result.
  • Two climate items ask about you where every other climate item asks about the team: EF2 (“I would bet on this team…”) and VI2 (“If I could, I would choose to stay on this team”). Both are ordinary in the fields they come from — confidence measures commonly use a first-person item, and a member's own intention to stay is part of what team viability means — but the inconsistency is real. It is listed here rather than quietly corrected, because item wording is never edited once published: editing it would silently break comparison with every wave already collected.
  • Peer ratings use a five-point frequency scale with a written anchor at every point, following the CATME convention; climate uses a seven-point agreement scale. Fewer points make it harder to detect how much raters agree; more points make each anchor harder to word precisely. Which trade is right here is a judgement we made, not something data have settled.
  • The leadership items were written for a leader and the team reporting to them. Used with a senior team whose remit is the whole organisation, item D5 and the direction-and-clarity scale can read oddly — they ask about “the team's goals” where “the business” is what the group actually steers. A variant for that case may be needed; none exists yet.
  • The target reading level is grade 8, so that answering does not depend on fluent English. Some of the written anchors are above it. No systematic check of the whole bank has been run, so we cannot tell you how many.
  • Five climate scales changed composition between team survey 0.3.0 and 0.4.0, so results from either side of that line cannot be compared on those scales, and Facets does not yet re-score older waves onto the newer version for you. If you hold waves from both, treat the comparison as unavailable rather than as a change in the team.
  • Of the three engagement items, EN3 (absorption) is the one we expect to perform worst. Absorption is the facet most likely to turn out to be indistinguishable from dedication, and the least useful to a manager who wants to act on it. If that is what the data show, a replacement item would be added under a later version of the engagement instrument — currently 0.1.0 — rather than EN3 being reworded.
  • In an engagement cycle, a unit can be large enough to report in one wave and too small in the next. The suppression rules handle each wave correctly on its own, but they do not solve the comparison between them: what a reader sees as a change may be a change in who was reportable. We have no method for this, and would rather say so than publish a difference that looks solid.
  • The scoring engine is versioned separately from the instruments and is now at 0.2.0; on 2026-09-14 it corrected the rule that decides whether a team agreed about a climate scale. A report produced before that date can say views differ on a scale the team in fact agreed about. If you hold one, have its agreement labels re-scored before relying on them, and do not read them alongside the labels in a newer report.
  • The agreement statistic rwg(j) is computed against a single assumption about what disagreement would look like by chance: answers spread evenly. LeBreton & Senter (2008) recommend also reporting it against a mildly skewed assumption, which is the harder test when answers lean positive — as climate answers usually do. That second figure is not reported yet, so the agreement labels this instrument produces are the more permissive of the two.

Citation and reuse

Facets (2026). Facets instrument specification, revision 2026-09-15: Team survey 0.5.0, Leadership 360 0.1.0, Engagement cycle 0.1.0. https://facets.team/instrument

Published for inspection, citation and critique. Reuse terms for the item bank and the scoring method are being settled with counsel and will be stated here; a Creative Commons licence for the items is intended. Until then, please cite rather than copy, and send critique through the contact form at facets.team/contact.

Critique is what publication is for. If an item is ambiguous, an anchor is above a grade-8 reading level, or a threshold looks wrong, tell us through the contact form. Changes arrive as new versions, listed in each instrument’s change log; nothing is reworded in place.