Validity
Response consistency, speed, straight-lining and social-desirability indicators. Computed before anything else, and able to invalidate a session outright.
No proprietary mystery. The model, the scoring approach and the current state of the evidence, written so a head of people and an I/O psychologist can both check it.
The model
Response consistency, speed, straight-lining and social-desirability indicators. Computed before anything else, and able to invalidate a session outright.
The stable bright-side facets: drive, discipline, curiosity, composure, sociability, openness to adaptation.
Nine behaviours that help until pressure arrives. In development: the items are collected, but this layer is not scored or reported yet. When it is, it will be reported as risk zones with base rates, never as a diagnosis.
What the person is motivated by. In development: collected but not yet scored or reported. When live, it will be compared against what the work demands to produce fit, not ranking.
Learning agility, tolerance for ambiguity, and measured AI fluency, the layer that makes this an instrument for this decade.
Five group-level indices report today, cohesion, safety spread, role coverage, cognitive diversity and risk density, gated at five respondents. Energy concentration depends on forced-choice driver items that need a calibration we have not collected, so it renders as not yet available.
Layers feed forward: L0 can invalidate everything after it, LA–LD describe the person, and LE is composed from LA–LD plus consensus items. Nothing in LE is ever traced back to an individual's item responses in a manager view.
Scoring philosophy
The engine scores polytomous items with a graded response model once calibrated parameters exist; until then it scores classical test theory and says so on every report. Item parameters live in versioned calibration data, never in code, and every score records which calibration produced it.
The long-form instrument uses ranked blocks to blunt impression management. Because ranked data is ipsative, it is scored into a normative space before anything is reported, and the platform structurally refuses to compare people on raw ipsative output.
Every norm table carries a population definition, a collection window, an n, and a version. A rebuild requires n ≥ 300 and produces a new version rather than overwriting the old one, so a score from last year can still be reproduced exactly.
All arithmetic happens in z-space. Percentiles, stens and 0–100 values are derived at the display edge only, and percentiles are never averaged, a rule the pipeline enforces rather than documents.
Every reported score is drawn with a likely range derived from its reliability. Today those reliabilities are declared provisional placeholders, and they are deliberately conservative, chosen to make the bands wider rather than narrower, so the range errs on the side of claiming less. A nightly job replaces them with observed estimates as live data accrues, and the bands re-draw from the observed values automatically.
The bias-audit workbench computes group impact ratios with both the four-fifths and 2-SD tests, intersectionally, exportable as a dated artefact. Differential item functioning screens are planned once calibration samples exist to run them on.
Evidence status
Under-claiming is the brand. This table is maintained as evidence accumulates, and we would rather lose a deal than round a status up.
Internal consistency of the disposition facets (LA)
The α and ω estimators run nightly against live valid responses. The reliability figures currently drawing score bands are documented placeholders in the provisional norm table, capped deliberately low so the bands come out wider rather than narrower; observed values replace them as pilot data accrues, and this row moves only when they have.
Graded-response calibration and SEM derivation
The GRM scorer and the calibration import path are built, with parameters exportable for independent replication in R (mirt). No calibration has been estimated yet, so every production score today is classical test theory, exactly as stated on each report footer.
Criterion validity against performance outcomes
Longitudinal collection under way with design partners. Until effect sizes are published here, we make no predictive-validity claim of any kind.
Thurstonian scoring of the forced-choice long form
The scorer is built and deliberately refuses to run without calibration data: forced-choice responses are collected, never scored, until the calibration exists. We regard that refusal as the feature.
Adaptability and AI-fluency construct validity (LD)
A new construct in a fast-moving domain. Convergent evidence against observed tool-use behaviour is being collected; treat LD as developmental.
Cross-cultural measurement invariance beyond en-AU / en-GB / en-US
Norms exist for English-language populations only. We will not ship a locale before the invariance testing that justifies it.
Observed reliabilities are published live on the evidence ledger, generated from the same data the engine reads. The full statement of what this instrument cannot yet do is public: limitations · standards and conformance · construct glossary.