How to assess cybersecurity candidates on what they can do
Certifications tell you someone passed an exam. Here is how to build a cyber security assessment that shows what a candidate can actually do, and how to read the result.
A certification tells you someone passed an exam. A CV tells you what they would like you to believe. An interview tells you how well they talk about work.
None of those is the work.
The short version: a cyber security assessment is only worth running if the candidate cannot answer it without doing the work. Put a realistic artefact in front of them - logs, telemetry, a live environment - calibrated to the level you are actually hiring at, then score them on what they produced and how they got there. Everything below is detail on that one idea.
What the assessment is actually for
Get this wrong and everything downstream is wrong.
You are not trying to produce a perfect ranking of every applicant. You are not trying to replace the interview. You are trying to answer one question:
Who is worth an hour of a senior engineer’s time?
That is a much lower bar than “should we hire this person”, and holding the lower bar honestly is what keeps the process sane. An assessment that tries to make the final hiring decision will be too long, too punishing, and will lose you good candidates. An assessment that produces a defensible shortlist is doing its job.
The one test that tells you if your assessment is any good
Ask this of every question you plan to send:
Could a competent bluffer answer this without doing the work?
If yes, delete it. Three ways questions fail that test.
Recall questions. “What is Kerberoasting?” “Which port does DNS use?” “Name the phases of incident response.” These measure whether someone revised. Nobody investigates an incident from memory, and the answer is one search away.
Questions where the prompt contains the answer. “Review this log and identify the failed logon attempts” tells the candidate what they are looking for. The hard part of the job is not spotting the thing once someone has told you it is there. It is deciding what matters in a pile of events where nobody has labelled anything.
Questions with one obvious route. If there is exactly one sensible next step, you are testing whether they can follow instructions.
Here is what passing that test looks like in practice.
That is a discrimination task, and it is the single most useful shape for a technical assessment. The candidate is not asked to recognise a pattern. They are asked to separate a real signal from several plausible decoys, which is precisely what an analyst does on shift.
It also cannot be answered from training data. MITRE ATT&CK documents the technique and the RC4 indicator in public. Knowing the documentation does not tell you which account in this specific log set is the one, because that answer only exists in the artefact in front of them.
Build decoys deliberately. A challenge with one anomaly and forty obviously-benign events is a spotting exercise. A challenge with one anomaly and four things that look like anomalies but are not is an assessment.
Assess the role, not “cyber”
“Cyber security” is not a skill. It is a dozen jobs that share a conference circuit.
A SOC analyst spends their day triaging alerts and reading logs. A detection engineer writes and tunes the rules that generate those alerts. An incident responder takes over when something is confirmed and has to preserve evidence while the business asks when it will be fixed. An AppSec engineer reads code. These people are not interchangeable, and one generic “cyber test” tells you almost nothing about any of them.
| Role | What to actually put in front of them |
|---|---|
| SOC analyst | Alert triage against realistic telemetry, log correlation across sources, a phishing case with partial indicators |
| Detection engineer | Write or tune a rule against real data, then justify the thresholds and predict the false positives |
| Incident responder | A timeline to reconstruct from artefacts, plus a written summary for a non-technical stakeholder |
| Threat hunter | An open-ended hunt with a hypothesis, not a question with a known answer |
| Penetration tester | A live box with a real privilege-escalation path |
| Cloud security engineer | A misconfigured environment and an identity pivot to find |
| AppSec engineer | Code review with a genuine vulnerability and a plausible-looking non-issue |
The practical way to build this: take the job specification, strip out the buzzwords, and ask what this person will actually be doing in their first month. Then build that.
Someone who has genuinely hunted in a SIEM starts typing. Someone who listed KQL because the advert asked for it stalls at the schema tree. You learn that in about ninety seconds, and no interview question gets you there as reliably.
Calibrate to the level, not to your own ability
This is the most common way good assessments go wrong, and it is almost always well-intentioned.
You ask your best senior engineer to write the test. They write something they find interesting. Your best senior engineer finds Tier 3 problems interesting. You are now screening Tier 1 candidates against a Tier 3 bar, everybody fails, and the conclusion drawn is that the market is empty.
The market is not empty. The test is wrong.
Calibrate against what the person will own on day thirty, not what your strongest engineer can do on a good day. For a Tier 1 SOC role that is triage and escalation, not malware reverse engineering. For a senior detection engineer it is rule design and false-positive economics, not whether they can recite the ATT&CK matrix.
A useful sanity check: give the assessment to someone already doing the job at the level you are hiring for. If they find it uncomfortable, it is too hard. If they finish it without thinking, it is too easy. That takes twenty minutes and saves you a failed hiring round.
What to score, and what to leave alone
Score the outcome and the route. Whether they got the right answer matters. How they got there matters as much, because it generalises. A candidate who found the right account by writing a sensible query has shown something transferable. A candidate who guessed correctly has shown you nothing, and the two look identical if you only score the final answer.
Give partial credit for method. Someone who identified the right technique but misread one field is closer to hireable than someone who got the answer by accident. Binary right-or-wrong scoring throws that distinction away.
Do not score speed, except as a tiebreak. Fast and wrong is worse than slow and right, and time pressure disproportionately penalises the careful, which in security is not the trait you want to select against.
Avoid false precision. A candidate on 71 and one on 68 are the same candidate. Publishing a score to two decimal places implies a resolution the instrument does not have. Band the results and treat the bands as the output.
Reading the output
Two things to look for beyond the total.
Spiky profiles. A candidate who was excellent on log analysis and poor on cloud is not “average”. They might be exactly right for a SOC role and wrong for a security engineering one. A flat, uniformly middling candidate is genuinely average. A single total ranks those two identically, which is why the per-skill breakdown matters more than the headline number.
The near-misses. The candidate who took the right approach and made one wrong turn is often more interesting than the one who scored slightly higher by playing it safe. Read the actual answers for anyone near your cut line. That is fifteen minutes well spent.
Integrity: assume the LLM, design around it
Some candidates will use AI. Design on that assumption rather than hoping.
The most effective control is not detection, it is task design. An LLM will tell you what Kerberoasting is instantly. It cannot tell you which of six accounts in your specific log set is doing it, because that answer is not in its training data. Investigation tasks degrade gracefully under AI assistance. Recall tasks collapse completely. That is a design property, and it is worth more than any monitoring you bolt on afterwards.
Instrumentation is the second layer: tab switching, paste behaviour, whether the window left fullscreen. Those signals give you context, not verdicts. A strong score from a clean session and a strong score with heavy pasting are different results and you should know which you are holding.
Treat flags as a reason to probe in interview, never as an automatic rejection. The failure mode of anti-cheat is falsely accusing a good candidate, and that cost is real and unrecoverable.
What an assessment cannot tell you
Being straight about this makes the rest of it more credible, not less.
A technical assessment will not tell you whether someone communicates well under pressure, whether they escalate at the right moment, whether they will still be there in two years, or whether they will make your team better or worse to work in. It will not tell you how they behave at 3am with a regulator waiting, and it will not tell you whether they can disagree with a senior stakeholder and hold their position.
Those are interview questions and reference questions, and they matter enormously.
The point of assessing technical ability first is not that it replaces any of that. It is that it stops you spending your interview slots discovering that someone cannot read a log. You get to spend the whole hour on the things an assessment genuinely cannot measure, with a candidate you already know can do the work.
On the evidence: work sample tests remain among the better predictors of job performance available, but the honest figure is lower than the one usually quoted in hiring content. Roth, Bobko and McFarland’s 2005 meta-analysis in Personnel Psychology put the validity at .33, meaningfully below the higher numbers from the older literature that vendors tend to cite. Better than a CV. Not a crystal ball. Anyone selling you certainty is selling.
The mistakes that make assessments useless
- Testing trivia. Port numbers and definitions. Measures revision.
- One generic cyber test for every role. A detection engineer and a pentester have almost nothing in common day to day.
- Calibrating to your best engineer. Everyone fails, and you conclude the market is empty.
- Making it too long. Three hours across a large pool gets you a self-selected sample of the least busy.
- Scoring only the final answer. Throws away the difference between reasoning and guessing.
- Treating integrity flags as verdicts. The cost of a false accusation is not recoverable.
- Using it as a gate rather than as evidence. The output is information for the interview, not a verdict that replaces one.
If you are staring at a large applicant pool right now, the sequencing question is covered separately in how to screen 200 cyber applicants without reading 200 CVs.
One honest sentence
Assessment does not tell you who to hire. It tells you who can actually do the work, which is the one thing a CV, a certification and a good interview manner will never reliably tell you on their own.
Ready to do this on your next hire?
Or let us do it for you.
You can run this process yourself. Or send CyberHire the job spec and the applicant pool, and get back a shortlist ranked on demonstrated ability.