Work samples vs interviews: what the evidence actually says
Assessment vendors claim work samples beat interviews. The current research does not support that. Here is what it does say, and what it means for cyber hiring.
Every assessment vendor tells you the same thing. Interviews are unreliable, work samples predict performance, the research is settled, here is a decimal number to prove it.
We sell assessments. So take this next part in that context.
The evidence does not say work samples beat interviews. The best current estimates put structured interviews ahead of work sample tests as a predictor of job performance. If you have been told otherwise by someone selling you a testing platform, you have been told something that is no longer supported.
The argument for assessing cyber candidates practically is still strong. It is just not the argument the industry keeps making, and the real one is more useful.
The numbers everyone quotes are out of date
For twenty-odd years, hiring content has run on Schmidt and Hunter’s 1998 summary of a century of selection research. It put work sample tests at a validity of .54 and general mental ability at .51, and it became the most-cited table in recruitment.
Two things have happened to it since.
First, Roth, Bobko and McFarland re-ran the work sample meta-analysis in 2005 with the studies published since 1984. Work sample validity came out at .33, not .54.
Second, and much more significantly, Sackett, Zhang, Berry and Lievens re-examined the whole field in 2022 in the Journal of Applied Psychology. Their argument is technical but the consequence is not: the standard corrections applied for range restriction had been systematically overcorrecting, which meant the validity of a wide range of selection methods had been substantially overestimated for decades.
When they redid the sums, the league table changed and the order changed with it.
| Selection method | Revised validity (Sackett et al., 2022) |
|---|---|
| Structured interview | .42 |
| Job knowledge test | .40 |
| Empirically keyed biodata | .38 |
| Work sample test | .33 |
| Cognitive ability | .31 |
Cognitive ability, which spent two decades being described as the single best predictor of job performance, is at the bottom of that list. Work samples sit below structured interviews. And the .33 figure for work samples matches what Roth and colleagues found independently seventeen years earlier, which is a decent sign it is about right.
Note that these are operational validities for single predictors, not a ranking of what to do first. That distinction matters and I will come back to it.
The word doing all the work is “structured”
Before anyone cancels their assessment plans, look at what .42 is actually measuring.
A structured interview means every candidate gets the same questions, in the same order, asked by trained interviewers, scored against anchored rating scales defined before anyone was interviewed, usually by multiple raters independently. It is a measurement instrument. It takes real work to build and real discipline to run.
That is not what happens in most cyber hiring.
What usually happens is a conversation. Different questions depending on how it flows, different interviewers with different bars, a scorecard filled in afterwards from memory if at all, and a decision that is substantially made in the first few minutes. Unstructured interviews are a well-documented weak predictor, and they are the format almost everyone actually uses.
So the honest statement is not “interviews beat work samples”. It is: a properly structured interview beats a work sample, and almost nobody runs one. If your process involves three people having three different conversations and then comparing impressions, the .42 figure has nothing to do with you.
What this means for cyber specifically
Three things follow, and they are more interesting than the vendor pitch.
Job knowledge tests do well, at .40. That is worth sitting with, because it partly rehabilitates the thing I have spent other articles criticising. Knowing things is not worthless. The catch is what “job knowledge test” means in the research: a properly constructed instrument measuring the knowledge the job actually requires. It does not mean forty multiple-choice questions about port numbers bought off the shelf. A cyber knowledge test built to that standard would be a reasonable predictor. Almost nothing sold as one is built to that standard.
The gap between describing and doing is unusually wide in security. These validity figures are averages across all occupations. In a job where the core skill is talking to stakeholders, an interview naturally samples the work well. In a job where the core skill is reconstructing an intrusion from logs, an interview samples someone’s ability to talk about reconstructing an intrusion from logs. Those are different competencies, and security has one of the larger gaps between them. General averages will understate what a work sample tells you here.
Verification is not the same as prediction. Validity coefficients measure how well a method predicts future performance across a population. That is a different question from “is this specific person’s CV true”. A work sample answers the second question directly and no interview does, which matters more now that CVs have stopped carrying much signal.
Single predictors are the wrong frame anyway
The table above ranks methods used alone. Nobody hires that way, and the research on combining predictors is where the actually useful finding lives.
Schmidt and Hunter’s original point, which survived their overestimates, was that the gains come from combining methods that measure different things. A structured interview and a work sample are not competing instruments. They are measuring different competencies and they add.
Which reframes the whole question. It is not “assessment or interview”. It is what order, and what is each one for.
That is where the practical argument sits, and it does not depend on work samples winning a league table:
- A work sample can be run at scale, on everyone, cheaply. A structured interview cannot. You are not running a calibrated, multi-rater, anchored-scale interview for two hundred applicants.
- So the work sample goes first, because it is the only rigorous instrument that survives contact with volume.
- Then the interview gets spent on the fifteen people who can demonstrably do the work, on the things it is genuinely best at: judgement, escalation behaviour, how they handle being wrong, whether they will still be here in two years.
The assessment is not there to beat the interview. It is there to make sure the interview is aimed at the right people and asking better questions.
What this does not excuse
A few things worth being straight about, since I have just spent a thousand words undermining the standard sales pitch.
A validity of .33 is not a crystal ball. It is a meaningful correlation, not a guarantee. Some candidates who do well will be poor hires, and some who do badly would have been fine. Any vendor implying certainty, including one selling you what we sell, is overselling.
A badly built work sample is worth less than the number suggests. These figures come from properly constructed instruments. A test of port-number trivia is not a work sample regardless of what it is called, and it will not perform like one. What separates a real assessment from a quiz matters more than the category label.
None of this measures fairness on its own. Validity is one property. Adverse impact is a separate one, and work sample tests have their own literature on it. Anyone telling you a method is fair because it is valid is conflating two different questions.
So what should you actually do
If you have the discipline to run a genuinely structured interview, do it. It is the strongest single instrument available and it is cheaper than most people think, because the cost is in the design work, which you do once.
If you are hiring cyber roles at any volume, put a practical assessment first, because it is the only rigorous method that scales to the top of the funnel, and because in security the distance between describing the work and doing it is unusually large.
Then do both, in that order, and stop pretending either one is a complete answer on its own.
One honest sentence
The best available evidence does not say that testing candidates beats interviewing them, and the case for assessing cyber candidates practically does not need it to.
Ready to do this on your next hire?
Or let us do it for you.
You can run this process yourself. Or send CyberHire the job spec and the applicant pool, and get back a shortlist ranked on demonstrated ability.