Hiring signals · Selection validity
The number at the top of every job specification correlates with performance at .06. It survives because it is cheap to ask for, not because it works.
Experience is not a weak proxy for capability. It is close to no proxy at all. The best modern evidence puts the correlation between pre-hire experience and job performance at .06, and with staying in the job at .00. For most professionals this hardly matters, because the weak signal is one of several. For a frontline worker it is frequently the only signal that exists, and that asymmetry is where the damage lands.
Almost every job advertisement on earth contains a number of years. Five years' experience. Three to five in a similar role. Minimum two. It is the closest thing hiring has to a filter everyone agrees on, and the case for it is not stupid.
Experience is observable, cheap to state and cheap to check. It is legally safer than most alternatives, because it describes what someone did rather than what they are. It carries a plausible theory: people who have done a job for longer have seen more of its failure modes, made more of its mistakes at someone else's expense, and absorbed the tacit knowledge that no training programme transfers. And it is fair in an intuitive way, rewarding the person who put in the time.
That is the strongest version of the case, and I want it on the table before dismantling it, because the weak version is easy to knock down and proves nothing.
The problem is not that the theory is implausible. It is that it has been tested, repeatedly, for four decades, and it does not hold.
The largest modern meta-analysis on the question was published by Van Iddekinge and colleagues in Personnel Psychology in 2019. It pooled 44 studies covering 11,785 people and asked exactly the question a job advertisement is asking: does experience held before hire predict what happens after it?
The correlation with job performance came back at .06. With turnover, .00. With training performance, a slightly less dismal .11. The authors' own summary is that measures of prior experience are generally poor predictors of the outcomes employers care about.1
Two features of that finding matter more than the headline number. The first is that it measures pre-hire experience specifically, which is precisely what the job specification demands and precisely what a CV asserts. This is not a study of whether people improve on the job; it is a study of whether the years someone brings through the door tell you anything about what they will do once inside.
The second is the .00 against turnover. The most common defence of the experience filter is that it is not really about peak performance but about reliability: an experienced hire is a safer hire, less likely to leave, less likely to be a disaster. That defence predicts a negative correlation with turnover. The measured value is zero.
For twenty-five years, personnel selection ran on one table. Schmidt and Hunter's 1998 meta-analysis in Psychological Bulletin gave the field its reference values: work samples at .54, general mental ability and structured interviews at .51, and years of job experience already trailing at .18.2 A great deal of hiring content still quotes those numbers as current.
They are not current. In 2022, Sackett, Zhang, Berry and Lievens reanalysed the whole table in the Journal of Applied Psychology and argued that the 1998 corrections for range restriction had been applied too aggressively, inflating almost everything.3
Their revision pulls the field down across the board. General mental ability drops from .51 to .31 and loses its long-held top position. Structured interviews become the strongest single predictor at .42. Work samples fall from .54 to .33.
And years of job experience falls from .18 to .07.
The direction of that revision matters more than any individual coefficient. Everything got weaker, which means predicting job performance is harder than the field believed for two decades, and anyone selling certainty about hiring should be treated accordingly. But experience did not merely decline with the rest. It fell to a value that is, for practical purposes, indistinguishable from knowing nothing.
What does a validity of .06 actually buy a hiring manager?
The arithmetic is exact; the interpretation carries a caveat. A validity coefficient is not the same thing as decision usefulness, and a small correlation can still have value at scale when selection ratios are extreme. At the volumes a single household or workshop hires at, it does not.
If experience predicts almost nothing, its universal use needs an explanation that is not stupidity. There is one, and it is well developed.
Spence set out the logic in 1973: when a buyer cannot observe what they are actually purchasing, they attend to whatever correlates with it and is costly to fake.4 Arrow made the parallel argument about education as a screening filter, starting from the observation that the employer has no direct way of determining productivity before hiring.5 Phelps described the same behaviour at group level, where employers fall back on population averages precisely because individual information is expensive.6
The point common to all three is that the proxy is not chosen for its accuracy. It is chosen for its price, given that the accurate thing is unavailable.
Years of experience is the cheapest signal in the market. One integer, self-reported, verifiable by a phone call if anyone bothers, and socially acceptable to demand in a way that a cognitive test is not. Against a validity of .07 sits a cost of essentially zero, and for most of the history of hiring nothing better could be produced at scale.
The decisive evidence that employers know this comes from Altonji and Pierret, who found that as an employer accumulates direct observation of a worker, the wage return to that worker's schooling falls and the return to previously unobserved ability rises.7 The proxy is discarded the moment real information arrives. It was never the thing. It was a placeholder held in position by the cost of the thing.
| Signal | Validity | Cost to produce | Borne by | Why it holds its position |
|---|---|---|---|---|
| Years of experience | .07 | Essentially zero | Nobody | Free to demand, self-reported, legally and socially uncontroversial |
| Unstructured interview | .19 | An hour of untrained time | Employer | Feels informative; practitioners trust intuition over validated tools |
| Work sample | .33 | Task design plus supervised administration | Employer | Used where the job is standardised enough to justify building one |
| Structured interview | .42 | Protocol design plus trained interviewers | Employer | Strongest available, and the least used relative to its accuracy |
The proxy is discarded the moment real information arrives. It was never the thing.
Here is where the argument stops being an academic curiosity about coefficients and starts costing people money.
When a management consultant applies for a role, years of experience is one entry in a portfolio. There is a degree from a named institution, a named employer whose reputation transfers, a public network of colleagues who can be asked, a portfolio of visible work, and often a structured interview process at the end. The weak proxy is carried by four stronger ones. Its .07 is diluted to near-irrelevance.
When a housekeeper in Riyadh with ten years in one household applies for a role, the portfolio is: ten. A single integer, whose predictive validity is approximately zero, carrying the entire decision. Behind it sits one reference, from the employer who is losing her, which is a conflict of interest rather than a signal.
A bad proxy diluted among good ones is an inefficiency. A bad proxy carrying the full weight of a decision is something else, and the economics of that were demonstrated directly. Pallais found that employers systematically underinvest in inexperienced workers, and that when richer performance information about those workers was made available, their subsequent employment and earnings rose.9 The binding constraint was the information, not the capability.
A serious opponent does not dispute the coefficients. They dispute what the coefficients are correlated against.
Almost all validity research predicts supervisor performance ratings, and supervisor ratings are noisy, biased and only loosely connected to the value a worker actually creates. If the criterion is bad, a low correlation with it may say more about the criterion than the predictor. Experience might well predict the things ratings capture badly: judgement under pressure, knowing when to escalate, the absence of expensive mistakes that never happen and are therefore never recorded.
This is a real limit and I accept it. Validity research inherits the weaknesses of its criterion, and anyone who quotes these numbers as though they were measurements of worth rather than correlations with ratings is overreaching.
It does not rescue the filter, for two reasons. The turnover result does not depend on the criterion problem at all: whether someone left is a fact, not a rating, and the correlation there is .00. And Sturman's work found the relationship between time on a job and performance is not even a straight line but an inverted U, rising, then flattening, then declining.10 If experience were quietly accumulating value that ratings miss, more of it should not eventually be worse.
The honest statement of the limit is this: experience may carry information that current research instruments cannot see. What cannot be claimed is that a hiring manager can see it either. The manager has the same integer, and no better means of reading it.
Three things follow, in descending order of how uncomfortable they are.
The first is a straightforward substitution. Structured interviews at .42 and work samples at .33 outperform years of experience by a factor of six, and both are available to any employer willing to spend an hour designing them. The obstacle is not knowledge; the literature has been clear since 1998 and clearer since 2022. It is that a validated protocol costs something and an integer does not.
The second is a question worth asking before writing the next job specification. If your honest belief is that five years' experience tells you this person is unlikely to be a disaster, that is defensible and roughly what the evidence supports. If your belief is that it tells you who will be good at the job, the last twenty-five years of research disagrees, and has grown steadily more confident in disagreeing.
The third is the one that concerns the workers this series is about. Where a validated alternative does not exist for an occupation, there is no substitution to make, and the integer wins by default. Nobody has built assessment infrastructure for household work or informal trades, so employers use years because nothing else has been produced, and workers are priced on a number that explains 0.36 per cent of what they will actually do.
That is not a measurement problem dressed up as a moral one. It is a missing instrument, and instruments get built when someone decides the measurement is worth the cost.
Disclosure
The case above stands on its own evidence; nothing in it depends on what follows. I am the founder of UpSkillMe, which builds practical assessment and a worker-owned capability record for household work in Saudi Arabia and the trades in South Africa. The argument that work samples beat experience is the commercial premise of that company, and readers should weigh it knowing so. The coefficients are not mine and can be checked against the sources below.
Next in this series: why "unskilled" is a statement about schooling rather than capability.
Written in British English. Journal titles and paper titles are given as published, including American spellings. Validity coefficients are correlations with criterion measures, most often supervisor ratings, and inherit the limits of those measures; this is addressed directly in section six rather than footnoted. The 1998 figures appear only alongside their 2022 revision and should not be quoted alone.