5. Scoring spec

The scoring engine is deterministic, versioned, and test-covered. Two analysts exporting the same wave and cut get identical values to four decimal places. Displayed values are rounded at the edge, never in the pipeline.

5.1 Principles

  • The pack defines the math; the engine executes it. Weights, required sets, bands, and primary metric are pack-version properties.
  • One primary metric per pack. mean_0_100, pct_favorable, or frequency_pct (for "how often does this happen" practice models). Secondary metrics are computed but labeled.
  • Respondent-level first. Practice scores are computed per respondent, then aggregated. This keeps missing items from biasing units with different skip patterns.
  • Every Scorecard is replayable. It stores engine_version, pack_version_id, weighting_scheme, and input_hash.

5.2 Item normalization

For an item answered on a scale 1..S (S = 5 or 7):

r'  = (S + 1 − r)            if reverse_coded, else r
x   = 100 × (r' − 1) / (S − 1)          # 0–100
fav = 1 if r' ≥ S − 1 else 0            # top-2 box on 5-pt; top-2 on 7-pt (configurable to top-3)
  • N/A and skipped answers are missing, not zero.
  • Frequency scales (Never … Always) produce x the same way; frequency_pct uses "Often" + "Always" as the favorable set.
  • Forced-rank and max-diff items do not enter the index. They produce preference shares (max-diff via hierarchical Bayes or count-based scores in MVP) reported separately.
  • NPS-style 0–10 items are outcome items; never in the index.

5.3 Respondent-level scores

Practice score for respondent i, practice p:

answered = { items j in p where x_ij is present }
if |answered| < min_items_answered(p):  P_ip = missing
else: P_ip = Σ_j∈answered (w_j · x_ij) / Σ_j∈answered w_j

min_items_answered defaults to max(2, ceil(0.6 × items_in_p)).

Dimension score for respondent i, dimension d: weighted mean of that respondent's non-missing practice scores in d, requiring ≥ 60% of the practice weight present.

Respondent-level index is not computed. The index is an aggregate property (§5.4). This avoids a single person's missing dimension removing them from everything.

5.4 Aggregate scores for a cut

For a cut C with respondent set R_C:

Practice(p, C)   = Σ_i∈R_C  v_i · P_ip  / Σ_i∈R_C v_i        (over i with P_ip present)
Dimension(d, C)  = Σ_p∈d  ω_p · Practice(p, C) / Σ_p∈d ω_p
Index(C)         = Σ_d  Ω_d · Dimension(d, C) / Σ_d Ω_d      (only if all required dimensions computed)
  • v_i = respondent weight. Default 1 (census). For samples, or when the EM switches on non-response adjustment, v_i is a post-stratification weight so the weighted respondent mix matches the population on up to two variables (default: site × level). Weights are trimmed to [0.3, 3.0]. The unweighted index is always computed and shown in the method appendix. If they differ by more than 2 points, the Scoreboard shows a banner.
  • ω_p, Ω_d = practice and dimension weights from the pack. Equal weights by default.
  • pct_favorable aggregates as the (weighted) share of favorable responses per item, then averaged to practice and dimension the same way.

Uncertainty. Each aggregate carries a 95% CI. Means: ± 1.96 × sd / √n_eff with n_eff = (Σv)² / Σv² and a finite-population correction for census cuts (√((N − n)/(N − 1))). Proportions: Wilson interval. Index and dimension CIs: bootstrap (1,000 resamples of respondents, seeded by input_hash, so replay is identical).

5.5 Bands

A pack version defines one BandScheme.

Absolute scheme (default for new firms without a book):

Band Range (0–100) Default label
B1 0.0 – 49.9 Constraint
B2 50.0 – 64.9 Emerging
B3 65.0 – 79.9 Established
B4 80.0 – 100 Distinctive

Norm-relative scheme: bands by quartile of a named NormSet version (Bottom quartile / Second / Third / Top). The NormSet version is stamped on the pack version; changing norms is a pack minor version and is shown on every chart footnote.

Borderline rule. If the 95% CI of a score straddles a cut point, the ScoreCell is band_borderline = true. The UI shows the band with a hollow marker and the interpretation text is prefixed "Borderline:". Insight Objects cannot assert a band change on a borderline score.

Interpretations are the consultant-authored text for each (target, band). They are part of the pack, approved by a named partner, and rendered verbatim. The engine never rewrites them.

5.6 Comparisons and significance

Comparison Eligible when Test "Meaningful" requires
vs previous wave Same comparability_group; item versions marked comparable; hierarchy mapping ≥ 90% for unit-level cuts Welch's t (means); two-proportion z (% fav) p < 0.05 and
vs internal unit or rest of org Both cells pass k Same tests; unit vs rest-of-org, not vs total (avoids self-inclusion) Same thresholds
vs firm book NormSet eligible (§5.9) Percentile position; no significance test Shown as percentile band only
vs external reference Licensed or published set, unexpired Percentile position Always labeled "indicative · third party"

Multiple comparisons. Heatmaps and "movers" lists apply Benjamini–Hochberg FDR control at q = 0.10 across the cells displayed. Cells that fail after correction render in neutral colour even if their raw p < 0.05.

Minimum n for tests. No significance flag below n = 30 per side; the cell shows the value and "n too small to test."

Composition check (Wave Compare): if the share of any level or site in the respondent mix moves by more than 10 points between waves, a composition-adjusted delta (reweight wave 2 to wave 1 mix) is computed and shown alongside the raw delta.

5.7 Driver analysis

Question answered: which practices move the index in this client?

  • Method: relative weights analysis (Johnson, 2000) of dimension or index scores on the pack's practice scores, at respondent level. Outcome can alternatively be a pack-declared outcome item (e.g., intent to stay, confidence in strategy).
  • Uncertainty: bootstrap 1,000× for 95% CIs on each practice's share of explained variance.
  • Minimum: n ≥ 200 complete-case respondents and ≥ 10 respondents per predictor. Below that, the Drivers view shows correlations only, labeled "exploratory."
  • Language: results are labeled "association," never "impact" or "causes." Insight Objects inherit this constraint in their claim linter.
  • Priority matrix: practice score (x) × relative importance (y). High importance / low score = "fix first."

5.8 Text and video metrics (secondary signal)

  • Prevalence = share of respondents with ≥ 1 coded segment for the theme (not share of segments).
  • Intensity = mean of a 3-level severity code (mention / problem / blocker) assigned by the model and sample-checked by an analyst.
  • Segment skew = prevalence in segment ÷ prevalence overall; shown only where both cells clear k_text.
  • Confidence = high when a human-validated sample of ≥ 50 segments shows ≥ 85% agreement with model coding; medium at ≥ 70%; low otherwise or unvalidated.
  • Video emotion/conviction tags are never aggregated into any score and never shown to clients. They may only be used to help an analyst find clips.

5.9 Benchmarks

Layer Source Eligibility before any cut is shown Label
Client longitudinal This client's prior waves Comparable items; mapping ≥ 90% "vs Wave N"
Firm book Opted-in clients' k-safe Scorecards ≥ 5 contributing clients and ≥ 1,000 respondents in the cut; no single client > 30% of respondents "Firm benchmark · n clients · n respondents · built "
External reference Licensed or published norms Licence valid; items mapped "Indicative · third party · "

Book norms are built on item lineages and practice codes, so a firm's forks remain benchmarkable if they share lineage and are marked comparable.

5.10 Anonymity and suppression

The anonymity engine runs at three points: before launch (simulation), at compute (Scorecard), and at render (every client-facing request). A render-time failure blocks the view even if the Scorecard was previously safe.

Rules

  1. Primary suppression. Any cell with n < k_scores is suppressed. Default k = 5; works-council mode default k = 8, configurable to 10.
  2. Complementary suppression. If a parent cell is shown and exactly one child is suppressed, the value could be recovered by subtraction. The engine suppresses the next-smallest sibling as well, or rolls both into "Other " if the combined n ≥ k.
  3. Differencing across filters. For any two cuts A and B available on the same surface where B ⊂ A, if 0 < |A| − |B| < k the narrower cut is blocked. Enforced on ad-hoc filters and on saved cuts, and across pairs of waves when the population is unchanged.
  4. Filter depth. Client surfaces allow at most max_filter_depth_client stacked demographic filters (default 2; works council 1). Firm analysts may go deeper; results remain k-checked.
  5. Rare categories. Demographic values with fewer than k respondents are bucketed to "Other" or "Prefer not to say" for reporting.
  6. Response-rate display. Completion for units with population < k is shown only at the parent level. Nobody sees "4 of 4 have responded."
  7. Live data. During fieldwork, only firm users see scores, and only for cuts ≥ 2k. Client users see completion only until close.
  8. Text. Comments are shown by segment only if the segment has ≥ k_text respondents with text. All text is PII-redacted (names, emails, phone numbers, employee IDs, and unit-specific role titles that identify one person, e.g., "the Plant C night-shift quality lead"). Quotes carry a permission tier; only client_shareable quotes reach the Client Portal, and client managers never receive raw comment exports.
  9. Video. Faces and voices reach a client surface only with client_shareable_video consent. Otherwise the clip renders as transcript text. Consent is collected per video answer, separately from the Likert section, and can be withdrawn until wave close.
  10. 360. Rater groups (peers, direct reports) are shown only if the group has ≥ rater_group_k responses; otherwise merged into "Others." Manager ratings are shown as-is only if the ratee has been told this in advance (pack setting).
  11. Rounding. Client-facing scores display as integers; n is shown as exact only when ≥ 10, else "<10" (still ≥ k).

Suppressed-cell copy

  • Heatmap cell: — with tooltip "Fewer than 5 responses. Hidden to protect anonymity."
  • Manager view: "Your team had fewer than 5 responses. Results are shown for (n=48)."
  • Export: suppressed cells are blank with a footnote, never zero.

5.11 Worked example

Practice DIR-P2 Priority clarity, 3 items, 5-point agreement scale, item 2 reverse-coded, equal weights, min_items_answered = 2.

Respondent Item 1 Item 2 (rev) Item 3 x values P_i
A 4 2 → r' = 4 3 75, 75, 50 66.7
B 2 4 → r' = 2 N/A 25, 25 25.0
C 5 N/A N/A 100 missing (1 < 2)

Cut practice score (unweighted) = mean(66.7, 25.0) = 45.8 → Band B1 "Constraint". With n = 2 the cell is suppressed on every surface; the example is illustrative only.

5.12 Test obligations

  • Golden datasets per shipped pack with expected Scorecards to 4 d.p.
  • Property tests: reverse-coding symmetry; permutation invariance of respondents; suppression never yields a recoverable cell (fuzzed hierarchy and filter pairs).
  • Replay test: recompute any published Scorecard from input_hash and versions; mismatch blocks release.