5. Scoring spec
The scoring engine is deterministic, versioned, and test-covered. Two analysts exporting the same wave and cut get identical values to four decimal places. Displayed values are rounded at the edge, never in the pipeline.
5.1 Principles
- The pack defines the math; the engine executes it. Weights, required sets, bands, and primary metric are pack-version properties.
- One primary metric per pack.
mean_0_100,pct_favorable, orfrequency_pct(for "how often does this happen" practice models). Secondary metrics are computed but labeled. - Respondent-level first. Practice scores are computed per respondent, then aggregated. This keeps missing items from biasing units with different skip patterns.
- Every Scorecard is replayable. It stores
engine_version,pack_version_id,weighting_scheme, andinput_hash.
5.2 Item normalization
For an item answered on a scale 1..S (S = 5 or 7):
r' = (S + 1 − r) if reverse_coded, else r
x = 100 × (r' − 1) / (S − 1) # 0–100
fav = 1 if r' ≥ S − 1 else 0 # top-2 box on 5-pt; top-2 on 7-pt (configurable to top-3)
N/Aand skipped answers are missing, not zero.- Frequency scales (Never … Always) produce
xthe same way;frequency_pctuses "Often" + "Always" as the favorable set. - Forced-rank and max-diff items do not enter the index. They produce preference shares (max-diff via hierarchical Bayes or count-based scores in MVP) reported separately.
- NPS-style 0–10 items are outcome items; never in the index.
5.3 Respondent-level scores
Practice score for respondent i, practice p:
answered = { items j in p where x_ij is present }
if |answered| < min_items_answered(p): P_ip = missing
else: P_ip = Σ_j∈answered (w_j · x_ij) / Σ_j∈answered w_j
min_items_answered defaults to max(2, ceil(0.6 × items_in_p)).
Dimension score for respondent i, dimension d: weighted mean of that respondent's non-missing practice scores in d, requiring ≥ 60% of the practice weight present.
Respondent-level index is not computed. The index is an aggregate property (§5.4). This avoids a single person's missing dimension removing them from everything.
5.4 Aggregate scores for a cut
For a cut C with respondent set R_C:
Practice(p, C) = Σ_i∈R_C v_i · P_ip / Σ_i∈R_C v_i (over i with P_ip present)
Dimension(d, C) = Σ_p∈d ω_p · Practice(p, C) / Σ_p∈d ω_p
Index(C) = Σ_d Ω_d · Dimension(d, C) / Σ_d Ω_d (only if all required dimensions computed)
v_i= respondent weight. Default 1 (census). For samples, or when the EM switches on non-response adjustment,v_iis a post-stratification weight so the weighted respondent mix matches the population on up to two variables (default: site × level). Weights are trimmed to [0.3, 3.0]. The unweighted index is always computed and shown in the method appendix. If they differ by more than 2 points, the Scoreboard shows a banner.ω_p,Ω_d= practice and dimension weights from the pack. Equal weights by default.pct_favorableaggregates as the (weighted) share of favorable responses per item, then averaged to practice and dimension the same way.
Uncertainty. Each aggregate carries a 95% CI. Means: ± 1.96 × sd / √n_eff with n_eff = (Σv)² / Σv² and a finite-population correction for census cuts (√((N − n)/(N − 1))). Proportions: Wilson interval. Index and dimension CIs: bootstrap (1,000 resamples of respondents, seeded by input_hash, so replay is identical).
5.5 Bands
A pack version defines one BandScheme.
Absolute scheme (default for new firms without a book):
| Band | Range (0–100) | Default label |
|---|---|---|
| B1 | 0.0 – 49.9 | Constraint |
| B2 | 50.0 – 64.9 | Emerging |
| B3 | 65.0 – 79.9 | Established |
| B4 | 80.0 – 100 | Distinctive |
Norm-relative scheme: bands by quartile of a named NormSet version (Bottom quartile / Second / Third / Top). The NormSet version is stamped on the pack version; changing norms is a pack minor version and is shown on every chart footnote.
Borderline rule. If the 95% CI of a score straddles a cut point, the ScoreCell is band_borderline = true. The UI shows the band with a hollow marker and the interpretation text is prefixed "Borderline:". Insight Objects cannot assert a band change on a borderline score.
Interpretations are the consultant-authored text for each (target, band). They are part of the pack, approved by a named partner, and rendered verbatim. The engine never rewrites them.
5.6 Comparisons and significance
| Comparison | Eligible when | Test | "Meaningful" requires |
|---|---|---|---|
| vs previous wave | Same comparability_group; item versions marked comparable; hierarchy mapping ≥ 90% for unit-level cuts |
Welch's t (means); two-proportion z (% fav) | p < 0.05 and |
| vs internal unit or rest of org | Both cells pass k | Same tests; unit vs rest-of-org, not vs total (avoids self-inclusion) | Same thresholds |
| vs firm book | NormSet eligible (§5.9) | Percentile position; no significance test | Shown as percentile band only |
| vs external reference | Licensed or published set, unexpired | Percentile position | Always labeled "indicative · third party" |
Multiple comparisons. Heatmaps and "movers" lists apply Benjamini–Hochberg FDR control at q = 0.10 across the cells displayed. Cells that fail after correction render in neutral colour even if their raw p < 0.05.
Minimum n for tests. No significance flag below n = 30 per side; the cell shows the value and "n too small to test."
Composition check (Wave Compare): if the share of any level or site in the respondent mix moves by more than 10 points between waves, a composition-adjusted delta (reweight wave 2 to wave 1 mix) is computed and shown alongside the raw delta.
5.7 Driver analysis
Question answered: which practices move the index in this client?
- Method: relative weights analysis (Johnson, 2000) of dimension or index scores on the pack's practice scores, at respondent level. Outcome can alternatively be a pack-declared outcome item (e.g., intent to stay, confidence in strategy).
- Uncertainty: bootstrap 1,000× for 95% CIs on each practice's share of explained variance.
- Minimum: n ≥ 200 complete-case respondents and ≥ 10 respondents per predictor. Below that, the Drivers view shows correlations only, labeled "exploratory."
- Language: results are labeled "association," never "impact" or "causes." Insight Objects inherit this constraint in their claim linter.
- Priority matrix: practice score (x) × relative importance (y). High importance / low score = "fix first."
5.8 Text and video metrics (secondary signal)
- Prevalence = share of respondents with ≥ 1 coded segment for the theme (not share of segments).
- Intensity = mean of a 3-level severity code (mention / problem / blocker) assigned by the model and sample-checked by an analyst.
- Segment skew = prevalence in segment ÷ prevalence overall; shown only where both cells clear
k_text. - Confidence = high when a human-validated sample of ≥ 50 segments shows ≥ 85% agreement with model coding; medium at ≥ 70%; low otherwise or unvalidated.
- Video emotion/conviction tags are never aggregated into any score and never shown to clients. They may only be used to help an analyst find clips.
5.9 Benchmarks
| Layer | Source | Eligibility before any cut is shown | Label |
|---|---|---|---|
| Client longitudinal | This client's prior waves | Comparable items; mapping ≥ 90% | "vs Wave N" |
| Firm book | Opted-in clients' k-safe Scorecards | ≥ 5 contributing clients and ≥ 1,000 respondents in the cut; no single client > 30% of respondents | "Firm benchmark · n clients · n respondents · built |
| External reference | Licensed or published norms | Licence valid; items mapped | "Indicative · third party · |
Book norms are built on item lineages and practice codes, so a firm's forks remain benchmarkable if they share lineage and are marked comparable.
5.10 Anonymity and suppression
The anonymity engine runs at three points: before launch (simulation), at compute (Scorecard), and at render (every client-facing request). A render-time failure blocks the view even if the Scorecard was previously safe.
Rules
- Primary suppression. Any cell with
n < k_scoresis suppressed. Default k = 5; works-council mode default k = 8, configurable to 10. - Complementary suppression. If a parent cell is shown and exactly one child is suppressed, the value could be recovered by subtraction. The engine suppresses the next-smallest sibling as well, or rolls both into "Other
" if the combined n ≥ k. - Differencing across filters. For any two cuts A and B available on the same surface where B ⊂ A, if
0 < |A| − |B| < kthe narrower cut is blocked. Enforced on ad-hoc filters and on saved cuts, and across pairs of waves when the population is unchanged. - Filter depth. Client surfaces allow at most
max_filter_depth_clientstacked demographic filters (default 2; works council 1). Firm analysts may go deeper; results remain k-checked. - Rare categories. Demographic values with fewer than k respondents are bucketed to "Other" or "Prefer not to say" for reporting.
- Response-rate display. Completion for units with population < k is shown only at the parent level. Nobody sees "4 of 4 have responded."
- Live data. During fieldwork, only firm users see scores, and only for cuts ≥ 2k. Client users see completion only until close.
- Text. Comments are shown by segment only if the segment has ≥
k_textrespondents with text. All text is PII-redacted (names, emails, phone numbers, employee IDs, and unit-specific role titles that identify one person, e.g., "the Plant C night-shift quality lead"). Quotes carry a permission tier; onlyclient_shareablequotes reach the Client Portal, and client managers never receive raw comment exports. - Video. Faces and voices reach a client surface only with
client_shareable_videoconsent. Otherwise the clip renders as transcript text. Consent is collected per video answer, separately from the Likert section, and can be withdrawn until wave close. - 360. Rater groups (peers, direct reports) are shown only if the group has ≥
rater_group_kresponses; otherwise merged into "Others." Manager ratings are shown as-is only if the ratee has been told this in advance (pack setting). - Rounding. Client-facing scores display as integers; n is shown as exact only when ≥ 10, else "<10" (still ≥ k).
Suppressed-cell copy
- Heatmap cell:
—with tooltip "Fewer than 5 responses. Hidden to protect anonymity." - Manager view: "Your team had fewer than 5 responses. Results are shown for
(n=48)." - Export: suppressed cells are blank with a footnote, never zero.
5.11 Worked example
Practice DIR-P2 Priority clarity, 3 items, 5-point agreement scale, item 2 reverse-coded, equal weights, min_items_answered = 2.
| Respondent | Item 1 | Item 2 (rev) | Item 3 | x values | P_i |
|---|---|---|---|---|---|
| A | 4 | 2 → r' = 4 | 3 | 75, 75, 50 | 66.7 |
| B | 2 | 4 → r' = 2 | N/A | 25, 25 | 25.0 |
| C | 5 | N/A | N/A | 100 | missing (1 < 2) |
Cut practice score (unweighted) = mean(66.7, 25.0) = 45.8 → Band B1 "Constraint". With n = 2 the cell is suppressed on every surface; the example is illustrative only.
5.12 Test obligations
- Golden datasets per shipped pack with expected Scorecards to 4 d.p.
- Property tests: reverse-coding symmetry; permutation invariance of respondents; suppression never yields a recoverable cell (fuzzed hierarchy and filter pairs).
- Replay test: recompute any published Scorecard from
input_hashand versions; mismatch blocks release.