The Capability Layer Frontline labour markets · No. 07 · 2026

Measurement · Scoring and regulation

Should skills have a score? No, and the confusion is worth naming

The research that everyone cites in favour of scoring argues for structure over intuition. It does not argue for collapsing structure into one number, and those are different claims.

A single capability score is the most requested feature in this category and the one I most consistently decline to build. The methodological literature on composite indicators says aggregation can disguise serious failings in individual dimensions. The measurement literature says an indicator used for decisions gets corrupted. And European regulators have classified employment scoring as high risk. Three independent bodies of work, pointing the same way.

If you have one minute

Structure beats intuition. One number is not structure.

S1The received view

The case for scoring people is better than its critics usually admit, and it starts with a genuine finding

Every employer I have shown a worker record to has asked the same question within about ninety seconds. Can you just give me a number.

The request is not lazy. It rests on one of the better-established findings in applied psychology, and anyone arguing against scoring has to deal with it first.

Dawes, Faust and Meehl reviewed roughly 136 studies comparing actuarial prediction with expert clinical judgement and found the mechanical method matched or beat the expert in almost all of them.1 Dawes had earlier shown that even equal-weighted linear models, using no clever weighting at all, matched or outperformed experts.2 More recently, work on decision noise has documented how much variation there is between judges evaluating identical material, and how much of it structured scoring removes.3

Put together, the case runs: human judgement about people is unreliable, mechanical aggregation is more reliable, therefore reduce the evidence to a score and let the score decide. Add the practical arguments, which are real. A number is comparable across candidates. It can be sorted. It fits in a column. At any volume above a handful of hires, a hiring manager cannot hold four dimensions of evidence for forty candidates in their head.

I accept all of that. It is why my objection is narrower than it first appears, and why it turns on a distinction the literature itself is careful about and its popularisers are not.

S2The crack

Every one of those studies compares structure with intuition, not one number with several

Read what the actuarial literature actually compared. On one side, an expert forming an overall impression. On the other, a defined procedure applied consistently. The finding is that the procedure wins.

That is an argument for structure. It is not an argument for collapse.

A structured assessment that reports four dimensions separately, each with its evidence and its provenance, is fully structured. It removes exactly the inter-judge noise the research is concerned with. It is mechanical in precisely the sense Meehl meant. It simply does not add the four numbers together at the end.

So the question the industry faces is not the question the research answers. The real question is whether an employer decides better from one composite or from four disaggregated values with their sources attached. In preparing this paper I looked for a study testing that comparison in an employment context and did not find one.4 The most-cited evidence for scoring is evidence for something adjacent to scoring, and the gap between them is where the argument lives.

Which means the case has to be made on other grounds. Fortunately the methodological literature on building composite measures is extensive, and unusually candid about what they cost.

Exhibit 01 / 04
The same worker, presented two ways, and only one of them survives a follow-up question
A capability record as a composite score, against the same record disaggregated with provenance.
PRESENTATION A · COMPOSITE 78 Readiness score Sortable. Comparable. Unanswerable. PRESENTATION B · DISAGGREGATED Meal preparation assessed, demonstration Garment care assessed, demonstration Childcare partner attested Elderly care not tested THE EMPLOYER'S ACTUAL QUESTION: "I NEED SOMEONE FOR AN ELDERLY PARENT. CAN SHE DO IT?" Presentation A cannot answer. 78 contains the answer and has destroyed it. Presentation B answers in one line. Not tested. Which is the honest answer, and actionable. The composite is more comparable and less useful. Both records contain identical information before aggregation.
Illustrative record, not a real worker. The capability names follow the unit structure of the Philippine domestic work qualification discussed in paper five.
S3What is actually true

The standard reference on building composite measures spends most of its length warning about them

The OECD and the European Commission's Joint Research Centre publish the Handbook on Constructing Composite Indicators, which is the methodological reference for this kind of work. It is not a document opposed to composites. It is a manual for building them, written by people who build them.

Its stated warnings are therefore worth more than an outside critique would be. Composite indicators, it notes, may disguise serious failings in some dimensions and increase the difficulty of identifying proper remedial action.5 On the technical side it observes that principal component methods may actually mask the taxonomic information in the underlying data.5

And on the joint where every composite is weakest, it is blunt. The handbook refers to the arbitrary nature of the weighting process, and calls weighting one of the most criticised characteristics of composite indicators.5

Apply that to capability. To produce a single readiness score you must decide how much cooking is worth relative to childcare, and garment care relative to both. There is no empirical basis for that ratio, because it depends entirely on the household doing the hiring. The weights are not a technical parameter. They are a substantive judgement about what the employer wants, made by the platform, on the employer's behalf, invisibly.

The handbook also states the condition under which a composite is appropriate: it should measure multi-dimensional concepts which cannot be captured by a single indicator.5 Capability in an occupation is multi-dimensional, and each dimension can be captured perfectly well on its own. It fails the test in the second clause.

Exhibit 02 / 04
The manual for building composites names the failure modes, and capability scoring hits all four
Warnings from the OECD and JRC handbook, against what each becomes when the composite is a worker's capability score.
The warningAs publishedIn a capability score
Concealment May "disguise serious failings in some dimensions" An untested capability and an assessed one average into the same figure
Remedial difficulty Increases "the difficulty of identifying proper remedial action" The worker cannot tell which unit to improve to raise her score
Arbitrary weighting "The arbitrary nature of the weighting process"; among the most criticised features The platform silently decides what each household values
Wrong application Appropriate for concepts "which cannot be captured by a single indicator" Each capability can be reported directly, so the condition is not met
Failure mode present
Source: Nardo, Saisana, Saltelli, Tarantola and others, OECD and European Commission Joint Research Centre, Handbook on Constructing Composite Indicators, 2008. Quotations as published. The handbook lists these among the disadvantages of composites and does not reject them outright; the application to capability scoring is this paper's argument.
S4The mechanism

An indicator that starts deciding things stops measuring the thing it was built to measure

The concealment problem is static. There is a dynamic one that is worse, and it was described precisely in the 1970s.

Goodhart's original formulation, from 1975, is that any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes.6 Campbell put the social-measurement version more fully in 1979: the more any quantitative social indicator is used for social decision-making, the more subject it will be to corruption pressures and the more apt it will be to distort and corrupt the social processes it is intended to monitor.7

Campbell's wording is the one that matters here, because it identifies two distinct harms. The indicator gets corrupted, and the underlying process gets corrupted. Both would happen to a capability score, and quickly.

Picture the mechanism concretely. A readiness score decides who appears in an employer's shortlist. Training providers notice which units move the score most. Assessment centres face commercial pressure from the agencies that send them candidates. Within about two cycles the score measures who has optimised for the score, and the capability it was built to represent has drifted away underneath it.

Disaggregated evidence is not immune to gaming. But gaming a specific claim requires faking a specific assessment, which leaves a specific trace. Gaming a composite requires only finding the cheapest dimension to inflate, and the composite hides which one you inflated.

S5The evidence

Three documented cases where the aggregate looked healthy and the underlying data did not

The theory is old. The cases are recent and they come from three different sectors.

Amazon scrapped an experimental recruiting tool in 2018 after finding it downgraded applications containing the word women's.8 The relevant point is not that the tool was biased, which is well known. It is that the bias lived inside a score that presented as a neutral ranking.

In 2022 a Columbia mathematics professor published an investigation showing the university's second-place national ranking rested on data he characterised as inaccurate, dubious or highly misleading; the university subsequently stopped submitting data to the ranking.9 A composite that aggregated across many institutions had no mechanism for noticing that one input series was wrong.

Most directly relevant, a 2026 study of algorithmic hiring at scale examined roughly four million job applications processed through shared vendor scoring. Four per cent of applicants were rejected by all ten jobs they applied to, a correlation far beyond chance. And while aggregate fairness metrics looked acceptable, 10.6 per cent of individual postings showed significant adverse impact against Black applicants.10 The aggregate concealed the distribution, which is the handbook's first warning arriving in a live hiring system.

That last study is a recent preprint and should be treated accordingly, but its finding is the one I would most want tested further rather than the one I would most want to discard.

Exhibit 03 / 04
In each case the aggregate read as healthy while the data underneath it did not
Three documented instances, across three sectors, of a composite concealing what it was built from.
CaseThe compositeWhat it concealedHow it surfaced
Amazon, 2018 Experimental applicant ranking Downgrading of applications containing the word women's Internal testing; the tool was discontinued
Columbia, 2022 National university ranking Input data characterised as inaccurate, dubious or highly misleading One professor audited the submitted figures
Hiring at scale, 2026 Shared vendor applicant scoring 10.6% of postings with significant adverse impact, while aggregate fairness metrics looked acceptable Study of roughly 4 million applications
Sources: Reuters reporting, October 2018; Thaddeus, Columbia University, 2022; Bommasani and others, FAccT 2026. The 2026 study is a recent preprint and is treated as indicative. Amazon has disputed that the tool was the sole factor in any decision.
S6The accelerant

European regulators have already classified employment scoring as high risk, and the clock is running

Anyone building a capability score for a global market should know that one major jurisdiction has taken a position.

The EU AI Act places AI systems used in employment into Annex III, its high-risk category. Point 4(a) covers recruitment and selection, including filtering applications and evaluating candidates. Point 4(b) covers decisions on promotion and termination, task allocation based on individual behaviour or personal traits, and monitoring or evaluating performance and behaviour.11 A capability score used to rank candidates sits inside that definition rather than near it.

The Act also requires employers deploying such systems to inform workers' representatives and the affected workers before use.11 That obligation is awkward for a composite, because informing someone about a score they cannot decompose is thin notice.

On timing, the position moved recently and anyone working from older material will have it wrong. The prohibitions applied from February 2025 and general-purpose model duties from August 2025. The Annex III obligations covering employment were originally due to apply from 2 August 2026, and a Digital Omnibus regulation that entered into force on 27 July 2026 pushed that to 2 December 2027.12 The classification did not change. Only the deadline moved.

For the markets this series is about, the honest statement is that no equivalent rule exists. I found no GCC or Saudi regulation addressing algorithmic or composite scoring in hiring.13 That is an absence of constraint rather than a licence, and building to the stricter standard now is cheaper than retrofitting to it later.

Exhibit 04 / 04
The regulatory direction is settled even though the deadline moved
EU AI Act application dates for the provisions relevant to employment scoring, as at September 2026.
Aug 2024 In force Feb 2025 Prohibitions apply Aug 2025 GPAI duties apply Aug 2026 Original employment deadline, superseded Dec 2027 Employment high-risk obligations apply Moved by Digital Omnibus regulation, in force 27 July 2026 EMPLOYMENT AI IS HIGH-RISK UNDER ANNEX III POINT 4. THAT CLASSIFICATION HAS NOT CHANGED. Recruitment and selection, filtering applications, evaluating candidates (4a) Promotion and termination, task allocation, monitoring performance and behaviour (4b)
Sources: Regulation (EU) 2024/1689, Annex III point 4 and Article 26(7); Digital Omnibus Regulation (EU) 2026/1744, in force 27 July 2026. Dates stated as at September 2026 and should be reverified against EUR-Lex before use in any compliance document.
S7The bottom line

Give the employer ranking without giving them a number, because the two were never the same request

The practical objection to everything above is that employers need to sort, and you cannot sort on four dimensions. That objection is correct about the need and wrong about the solution.

An employer hiring for an elderly parent does not want a general capability score. They want candidates ranked against that brief. Those are different operations. The first invents a universal weighting and hides it; the second applies the employer's own stated weighting, openly, to a specific role, and can show which capability drove the ordering. The employer gets their sorted list either way. Only one of them survives the question why is this candidate above that one.

So the position, stated as a commitment rather than a preference: rank against a declared brief, never publish a standing composite, and show provenance beside every claim so that assessed, attested and untested remain visibly different things.

I should be straight about what this costs. It is slower to explain, harder to put in a marketing comparison, and it loses deals to competitors whose number looks cleaner. Those are real costs and I am choosing to pay them.

The question in the title has a short answer. Skills should have evidence, provenance and a date. A score is what you build when you have decided the reader will not look any further, and the entire premise of this series is that the reader should.

Disclosure

UpSkillMe

The case above stands on its own evidence; nothing in it depends on what follows. I am the founder of UpSkillMe, and the no-composite-score position described in section seven is a product commitment we have made. This paper argues for a design decision my company has already taken, which is the least neutral position a writer can occupy, and the reader should weigh section seven accordingly. The OECD and JRC handbook, the Campbell and Goodhart formulations and the EU AI Act are public and stand independently of us.

Next in this series: verification has a cost, and so does uncertainty.

Sources

  1. Dawes, R. M., Faust, D. and Meehl, P. E. (1989). Clinical versus actuarial judgment. Science 243, 1668 to 1674.
  2. Dawes, R. M. (1979). The robust beauty of improper linear models in decision making. American Psychologist 34(7), 571 to 582.
  3. Kahneman, D., Sibony, O. and Sunstein, C. R. (2021). Noise: A Flaw in Human Judgment. Little, Brown Spark.
  4. No study was located testing employer decision quality with a single composite score against disaggregated structured evidence. Stated here as a gap rather than filled.
  5. Nardo, M., Saisana, M., Saltelli, A., Tarantola, S. and others (2008). Handbook on Constructing Composite Indicators: Methodology and User Guide. OECD and European Commission Joint Research Centre. Quotations from pages 13, 27 and 48 as published.
  6. Goodhart, C. (1975). Formulation as published in Goodhart's monetary policy work; quoted here in its original wording rather than the later popular paraphrase.
  7. Campbell, D. T. (1979). Assessing the impact of planned social change. Evaluation and Program Planning 2(1), 67 to 90.
  8. Amazon experimental recruiting tool, discontinued and reported October 2018. Amazon has disputed that the tool was the sole factor in any hiring decision.
  9. Thaddeus, M. (2022). An investigation of the facts behind Columbia's US News ranking. Columbia University Department of Mathematics, February to March 2022.
  10. Bommasani, R. and others (2026). Study of algorithmic monoculture in hiring, presented at FAccT 2026. Recent preprint; treated as indicative rather than settled.
  11. Regulation (EU) 2024/1689 (the AI Act), Annex III point 4(a) and 4(b), and Article 26(7). In force 1 August 2024.
  12. Digital Omnibus Regulation (EU) 2026/1744, in force 27 July 2026, deferring Annex III high-risk obligations to 2 December 2027.
  13. Author's search of Saudi and GCC regulatory sources, September 2026. No provision addressing algorithmic or composite scoring in hiring was identified. The search was not exhaustive.

Written in British English. The Goodhart wording used here is his own 1975 formulation, not the widely circulated paraphrase about measures becoming targets, which is a later restatement by a different author. EU regulatory dates are stated as at September 2026 and moved once already in July 2026; reverify against EUR-Lex before relying on them. The algorithmic monoculture study is a preprint and is labelled as such in the text.