Criterion-Related Validity: A Practical Guide for Hiring
Ensure smarter hiring with criterion-related validity. Learn how scientifically validated tests predict job success. Try Testlify for data-driven hiring!Criterion-related validity is the evidence that a test’s scores actually predict a real outcome you care about, like job performance. In hiring, it answers one blunt question: do the people who score well on your assessment really do better on the job? If they do, the test has strong criterion-related validity. If they don’t, it is just a quiz with a score attached.
This matters because most hiring signals are weak. Resumes get embellished, interviews get rehearsed, and gut feel rewards whoever interviews well rather than whoever performs well. A validated assessment is one of the few inputs you can check against real results. Get the validity right and your shortlist starts predicting performance instead of guessing at it. Get it wrong and you are paying to sort candidates by noise.
TL;DR
- Criterion-related validity measures how closely test scores line up with a real outcome, like sales numbers or a performance rating.
- It comes in two forms: predictive validity (does the score forecast future performance?) and concurrent validity (does it match performance you can already see?).
- You measure it with a correlation coefficient that runs from 0 to 1.0. The stronger the link between scores and the outcome, the more the test is worth trusting.
- It is the difference between hiring on proof and hiring on hope. Low-validity tests quietly add cost; high-validity tests cut bad hires.
- Structured, validated methods now out-predict old favorites. A 2022 study ranked structured interviews above cognitive ability tests for predicting performance.
Summarise this post with:
What is criterion-related validity?
Criterion-related validity is how well a test predicts a separate, real-world outcome called the criterion. In hiring, the criterion is usually job performance: revenue closed, tickets resolved, code shipped, or a manager’s rating. The test is valid when high scorers go on to perform well and low scorers tend to struggle. In psychometrics, that is the whole point of the idea.
Say you are hiring a sales executive and you screen with a sales aptitude test. How do you know the test works? You wait, then you check. If the people who scored high are closing more deals six months later and the low scorers are missing target, your test tracks the thing you actually want. That link, score to outcome, is criterion-related validity in one sentence.
Here is the catch most teams miss: a test can look polished, read well, and still predict nothing. A clean interface is not evidence. Only the correlation with real performance is. So before you trust any assessment, ask what outcome it has been checked against, and how strong that link is.
What are the two types of criterion validity?
Criterion-related validity splits into two types, separated by timing. Predictive validity checks whether today’s scores forecast future performance. Concurrent validity checks whether scores match performance you can already see. Same question, asked at different points on the clock. Both tell you if the test is measuring something that matters.
Predictive validity: does the score forecast future success?
This is the version hiring cares about most. You test candidates now, hire some, and track how they do later. If the early scores line up with later results, the test has strong predictive validity. The downside is patience: you only get the answer after months of real performance data, so it takes time to confirm a test is pulling its weight.
Concurrent validity: does the score match current performance?
Instead of waiting, you give the test to people already doing the job and compare scores with their current results. If your top performers also score highest, the test has strong concurrent validity. It is faster to run, but it carries a quiet risk: your current employees are a filtered group who already survived hiring, so a test validated only on them may not behave the same on raw applicants.
What does criterion validity look like in practice?
Picture a 50-person sales team that adds a sales aptitude test to its screening. Six months later, the recruiters line up each new hire’s test score against revenue closed, deals won, and client retention. A pattern shows up fast: high scorers are beating quota, low scorers need constant coaching to stay afloat. That is predictive validity earning its keep, the test called the outcome before the outcome happened.

Now flip it. A company wants to know if its analytical reasoning test really sorts good data analysts from weak ones. It gives the test to current analysts and matches scores to performance reviews. If the strongest analysts score highest, that is concurrent validity. Either way, the company stopped trusting the test on faith and started trusting it on results. That shift is the entire job of validity.
How do you measure criterion-related validity?
You measure criterion-related validity by correlating test scores with a real outcome. Pick the outcome, collect the scores, line them up against performance, and calculate the correlation. The coefficient runs from 0 to 1.0; the higher it climbs, the more confidently the test predicts the result. Here is the working sequence.
- Define the criterion. Decide what good performance actually means for the role: revenue, output quality, ramp speed, retention, or a manager’s rating. A fuzzy criterion gives you a fuzzy answer.
- Administer the test. Give the assessment to candidates or current employees and record their scores cleanly.
- Gather the outcome data. Track real performance over time for predictive validity, or pull current performance for concurrent validity.
- Run the correlation. Calculate the correlation coefficient between scores and the outcome. A clear positive number means the test is doing its job.
- Refine and recheck. If the link is weak, adjust the test, the cutoff, or the criterion, then measure again. Validity is maintained, not declared once.
Pro Tip: Watch your sample. Validating a test only on tenured staff who already passed hiring can flatter the numbers, because the weakest candidates were screened out years ago. Where you can, check the test against a wider, more recent group so the correlation reflects real applicants, not just survivors.
Why does criterion-related validity matter in hiring?
Criterion-related validity matters because a low-validity test costs real money while looking productive. When scores do not predict performance, you are sorting candidates by noise and calling it a process. The bill arrives later, as turnover, missed targets, and re-hiring. Validity is the cheapest insurance against that bill.
Gallup puts the cost of replacing one employee at one-half to two times their annual salary, and calls that conservative. For a 100-person organization paying an average of $50,000, it pegs the yearly turnover bill at $660,000 to $2.6 million, part of a roughly $1 trillion problem for US businesses. A meaningful share of that traces back to hires who looked right on paper and a screen that could not tell the difference.

Validity also changes which methods you should trust. A 2022 Journal of Applied Psychology study by Sackett and colleagues re-ran decades of selection research and found that many methods had been overstated, with corrected validity estimates falling by about 0.10 to 0.20 points. Structured interviews came out as the top-ranked predictor of job performance, ahead of cognitive ability tests, which earlier work had treated as the gold standard. The lesson is not that one method wins forever; it is that you should check the evidence, not the reputation.
The market is moving the same way. In LinkedIn’s Future of Recruiting data, 93% of talent professionals said accurately assessing a candidate’s skills is critical to quality of hire, companies running the most skills-focused searches were 12% more likely to make a quality hire, and the share of paid job posts dropping the degree requirement rose to 26% in 2023 from 22% in 2020. As credentials fade as a filter, the burden shifts to skills tests that hold up, and a skills test is only as good as its criterion-related validity.
This is where the Testlify Competency-to-Evidence Matrix turns validity from theory into a workflow. The matrix starts with the role, maps it to the competencies that predict success, then ties each competency to measurable evidence: a skills assessment, a work sample, a structured interview, or a reference. Because every assessment is anchored to a defined criterion up front, you can check, over time, whether the scores actually track the outcome, instead of hoping they do. AI helps score and surface the evidence; the hiring team still makes the call. For the broader business case, see how teams maximize ROI with employee assessments.
How does it compare to other validity types?
Criterion-related validity asks whether a test predicts an outcome. Other validity types ask different questions, and a strong assessment usually needs more than one. Content validity asks whether the test covers the actual job. Construct validity asks whether it measures the intended trait. Criterion-related validity asks whether the scores predict results. Here is how they line up.
| Validity type | Core question | Hiring example |
|---|---|---|
| Criterion-related | Do scores predict a real outcome? | Sales test scores track revenue closed |
| Content | Does the test cover the real job? | A coding test uses tasks from the actual role |
| Construct | Does it measure the intended trait? | A reasoning test really measures reasoning |
| Face | Does it look relevant to candidates? | Applicants see the test as fair and job-related |
Two related ideas often come up alongside it. Construct validity confirms the test measures the trait it claims to, while discriminant validity confirms it is not accidentally measuring something else, like a coding test that quietly rewards verbal skill. Criterion-related validity is the one that ties everything back to performance, which is why hiring teams lean on it hardest.

Hire on proven validity, not guesswork
If a test cannot show it predicts performance, it does not belong in your hiring decision. Testlify’s assessments are built and benchmarked to track real outcomes, so your shortlist is scored on evidence before the first interview, not sorted by a number nobody has checked. Book a demo to see how validated assessments fit your roles, or start a free trial and run one on your next hire.
Key Takeaways
- Validity is proof, not polish. Criterion-related validity is the link between test scores and real performance. A test that looks good but predicts nothing is a cost, so judge every assessment by its correlation with outcomes, not its interface.
- Know which type you are claiming. Predictive validity forecasts future performance and takes time to confirm; concurrent validity checks current performance and is faster but easier to flatter. Match the method to the decision you are making.
- Measure it, then keep measuring. Define the criterion, run the correlation, and recheck as roles change. Validity is maintained over time, which means a test validated once is not validated forever.
- Bad validity is expensive. Gallup ties replacing one employee to one-half to two times their salary. Weak screens feed that bill, so validity is one of the cheapest ways to protect a hiring budget.
- Check evidence, not reputation. The 2022 Sackett research showed long-trusted methods were overstated and structured interviews predicted best. Pick selection methods on current evidence, because the rankings have shifted.
- Anchor tests to a criterion up front. Mapping each competency to measurable evidence, as in the Testlify Competency-to-Evidence Matrix, lets you confirm scores track outcomes instead of hoping they do, while the hiring team keeps the final decision.
Chatgpt
Gemini
Claude
Grok























