Measurement · Scoring and regulation
The research that everyone cites in favour of scoring argues for structure over intuition. It does not argue for collapsing structure into one number, and those are different claims.
A single capability score is the most requested feature in this category and the one I most consistently decline to build. The methodological literature on composite indicators says aggregation can disguise serious failings in individual dimensions. The measurement literature says an indicator used for decisions gets corrupted. And European regulators have classified employment scoring as high risk. Three independent bodies of work, pointing the same way.
Every employer I have shown a worker record to has asked the same question within about ninety seconds. Can you just give me a number.
The request is not lazy. It rests on one of the better-established findings in applied psychology, and anyone arguing against scoring has to deal with it first.
Dawes, Faust and Meehl reviewed roughly 136 studies comparing actuarial prediction with expert clinical judgement and found the mechanical method matched or beat the expert in almost all of them.1 Dawes had earlier shown that even equal-weighted linear models, using no clever weighting at all, matched or outperformed experts.2 More recently, work on decision noise has documented how much variation there is between judges evaluating identical material, and how much of it structured scoring removes.3
Put together, the case runs: human judgement about people is unreliable, mechanical aggregation is more reliable, therefore reduce the evidence to a score and let the score decide. Add the practical arguments, which are real. A number is comparable across candidates. It can be sorted. It fits in a column. At any volume above a handful of hires, a hiring manager cannot hold four dimensions of evidence for forty candidates in their head.
I accept all of that. It is why my objection is narrower than it first appears, and why it turns on a distinction the literature itself is careful about and its popularisers are not.
Read what the actuarial literature actually compared. On one side, an expert forming an overall impression. On the other, a defined procedure applied consistently. The finding is that the procedure wins.
That is an argument for structure. It is not an argument for collapse.
A structured assessment that reports four dimensions separately, each with its evidence and its provenance, is fully structured. It removes exactly the inter-judge noise the research is concerned with. It is mechanical in precisely the sense Meehl meant. It simply does not add the four numbers together at the end.
So the question the industry faces is not the question the research answers. The real question is whether an employer decides better from one composite or from four disaggregated values with their sources attached. In preparing this paper I looked for a study testing that comparison in an employment context and did not find one.4 The most-cited evidence for scoring is evidence for something adjacent to scoring, and the gap between them is where the argument lives.
Which means the case has to be made on other grounds. Fortunately the methodological literature on building composite measures is extensive, and unusually candid about what they cost.
The OECD and the European Commission's Joint Research Centre publish the Handbook on Constructing Composite Indicators, which is the methodological reference for this kind of work. It is not a document opposed to composites. It is a manual for building them, written by people who build them.
Its stated warnings are therefore worth more than an outside critique would be. Composite indicators, it notes, may disguise serious failings in some dimensions and increase the difficulty of identifying proper remedial action.5 On the technical side it observes that principal component methods may actually mask the taxonomic information in the underlying data.5
And on the joint where every composite is weakest, it is blunt. The handbook refers to the arbitrary nature of the weighting process, and calls weighting one of the most criticised characteristics of composite indicators.5
Apply that to capability. To produce a single readiness score you must decide how much cooking is worth relative to childcare, and garment care relative to both. There is no empirical basis for that ratio, because it depends entirely on the household doing the hiring. The weights are not a technical parameter. They are a substantive judgement about what the employer wants, made by the platform, on the employer's behalf, invisibly.
The handbook also states the condition under which a composite is appropriate: it should measure multi-dimensional concepts which cannot be captured by a single indicator.5 Capability in an occupation is multi-dimensional, and each dimension can be captured perfectly well on its own. It fails the test in the second clause.
| The warning | As published | In a capability score |
|---|---|---|
| Concealment | May "disguise serious failings in some dimensions" | An untested capability and an assessed one average into the same figure |
| Remedial difficulty | Increases "the difficulty of identifying proper remedial action" | The worker cannot tell which unit to improve to raise her score |
| Arbitrary weighting | "The arbitrary nature of the weighting process"; among the most criticised features | The platform silently decides what each household values |
| Wrong application | Appropriate for concepts "which cannot be captured by a single indicator" | Each capability can be reported directly, so the condition is not met |
The concealment problem is static. There is a dynamic one that is worse, and it was described precisely in the 1970s.
Goodhart's original formulation, from 1975, is that any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes.6 Campbell put the social-measurement version more fully in 1979: the more any quantitative social indicator is used for social decision-making, the more subject it will be to corruption pressures and the more apt it will be to distort and corrupt the social processes it is intended to monitor.7
Campbell's wording is the one that matters here, because it identifies two distinct harms. The indicator gets corrupted, and the underlying process gets corrupted. Both would happen to a capability score, and quickly.
Picture the mechanism concretely. A readiness score decides who appears in an employer's shortlist. Training providers notice which units move the score most. Assessment centres face commercial pressure from the agencies that send them candidates. Within about two cycles the score measures who has optimised for the score, and the capability it was built to represent has drifted away underneath it.
Disaggregated evidence is not immune to gaming. But gaming a specific claim requires faking a specific assessment, which leaves a specific trace. Gaming a composite requires only finding the cheapest dimension to inflate, and the composite hides which one you inflated.
The theory is old. The cases are recent and they come from three different sectors.
Amazon scrapped an experimental recruiting tool in 2018 after finding it downgraded applications containing the word women's.8 The relevant point is not that the tool was biased, which is well known. It is that the bias lived inside a score that presented as a neutral ranking.
In 2022 a Columbia mathematics professor published an investigation showing the university's second-place national ranking rested on data he characterised as inaccurate, dubious or highly misleading; the university subsequently stopped submitting data to the ranking.9 A composite that aggregated across many institutions had no mechanism for noticing that one input series was wrong.
Most directly relevant, a 2026 study of algorithmic hiring at scale examined roughly four million job applications processed through shared vendor scoring. Four per cent of applicants were rejected by all ten jobs they applied to, a correlation far beyond chance. And while aggregate fairness metrics looked acceptable, 10.6 per cent of individual postings showed significant adverse impact against Black applicants.10 The aggregate concealed the distribution, which is the handbook's first warning arriving in a live hiring system.
That last study is a recent preprint and should be treated accordingly, but its finding is the one I would most want tested further rather than the one I would most want to discard.
| Case | The composite | What it concealed | How it surfaced |
|---|---|---|---|
| Amazon, 2018 | Experimental applicant ranking | Downgrading of applications containing the word women's | Internal testing; the tool was discontinued |
| Columbia, 2022 | National university ranking | Input data characterised as inaccurate, dubious or highly misleading | One professor audited the submitted figures |
| Hiring at scale, 2026 | Shared vendor applicant scoring | 10.6% of postings with significant adverse impact, while aggregate fairness metrics looked acceptable | Study of roughly 4 million applications |
Anyone building a capability score for a global market should know that one major jurisdiction has taken a position.
The EU AI Act places AI systems used in employment into Annex III, its high-risk category. Point 4(a) covers recruitment and selection, including filtering applications and evaluating candidates. Point 4(b) covers decisions on promotion and termination, task allocation based on individual behaviour or personal traits, and monitoring or evaluating performance and behaviour.11 A capability score used to rank candidates sits inside that definition rather than near it.
The Act also requires employers deploying such systems to inform workers' representatives and the affected workers before use.11 That obligation is awkward for a composite, because informing someone about a score they cannot decompose is thin notice.
On timing, the position moved recently and anyone working from older material will have it wrong. The prohibitions applied from February 2025 and general-purpose model duties from August 2025. The Annex III obligations covering employment were originally due to apply from 2 August 2026, and a Digital Omnibus regulation that entered into force on 27 July 2026 pushed that to 2 December 2027.12 The classification did not change. Only the deadline moved.
For the markets this series is about, the honest statement is that no equivalent rule exists. I found no GCC or Saudi regulation addressing algorithmic or composite scoring in hiring.13 That is an absence of constraint rather than a licence, and building to the stricter standard now is cheaper than retrofitting to it later.
The practical objection to everything above is that employers need to sort, and you cannot sort on four dimensions. That objection is correct about the need and wrong about the solution.
An employer hiring for an elderly parent does not want a general capability score. They want candidates ranked against that brief. Those are different operations. The first invents a universal weighting and hides it; the second applies the employer's own stated weighting, openly, to a specific role, and can show which capability drove the ordering. The employer gets their sorted list either way. Only one of them survives the question why is this candidate above that one.
So the position, stated as a commitment rather than a preference: rank against a declared brief, never publish a standing composite, and show provenance beside every claim so that assessed, attested and untested remain visibly different things.
I should be straight about what this costs. It is slower to explain, harder to put in a marketing comparison, and it loses deals to competitors whose number looks cleaner. Those are real costs and I am choosing to pay them.
The question in the title has a short answer. Skills should have evidence, provenance and a date. A score is what you build when you have decided the reader will not look any further, and the entire premise of this series is that the reader should.
Disclosure
The case above stands on its own evidence; nothing in it depends on what follows. I am the founder of UpSkillMe, and the no-composite-score position described in section seven is a product commitment we have made. This paper argues for a design decision my company has already taken, which is the least neutral position a writer can occupy, and the reader should weigh section seven accordingly. The OECD and JRC handbook, the Campbell and Goodhart formulations and the EU AI Act are public and stand independently of us.
Next in this series: verification has a cost, and so does uncertainty.
Written in British English. The Goodhart wording used here is his own 1975 formulation, not the widely circulated paraphrase about measures becoming targets, which is a later restatement by a different author. EU regulatory dates are stated as at September 2026 and moved once already in July 2026; reverify against EUR-Lex before relying on them. The algorithmic monoculture study is a preprint and is labelled as such in the text.