Criterion-Related Validity: A Practical Guide for Hiring (2026)
Ensure smarter hiring with criterion-related validity. Learn how scientifically validated tests predict job success. Try Testlify for data-driven hiring!

Criterion-related validity is the evidence that a test’s scores actually predict a real outcome you care about, like job performance. In hiring, it answers one blunt question: do the people who score well on your assessment really do better on the job? If they do, the test has strong criterion-related validity. If they don’t, it is just a quiz with a score attached.
This matters because most hiring signals are weak. Resumes get embellished, interviews get rehearsed, and gut feel rewards whoever interviews well rather than whoever performs well. A validated assessment is one of the few inputs you can check against real results. Get the validity right and your shortlist starts predicting performance instead of guessing at it. Get it wrong and you are paying to sort candidates by noise.
TL;DR
- Criterion-related validity measures how closely test scores line up with a real outcome, like sales numbers or a performance rating.
- It comes in two forms: predictive validity (does the score forecast future performance?) and concurrent validity (does it match performance you can already see?).
- You measure it with a correlation coefficient that runs from 0 to 1.0. The stronger the link between scores and the outcome, the more the test is worth trusting.
- It is the difference between hiring on proof and hiring on hope. Low-validity tests quietly add cost; high-validity tests cut bad hires.
- Structured, validated methods now out-predict old favorites. A 2022 study ranked structured interviews above cognitive ability tests for predicting performance.
What is criterion-related validity?
Criterion-related validity is how well a test predicts a separate, real-world outcome called the criterion. In hiring, the criterion is usually job performance: revenue closed, tickets resolved, code shipped, or a manager’s rating.
The test is valid when high scorers go on to perform well and low scorers tend to struggle. In psychometrics, that is the whole point of the idea. Say you are hiring a sales executive and you screen with a sales aptitude test. How do you know the test works? You wait, then you check.
If the people who scored high are closing more deals six months later and the low scorers are missing target, your test tracks the thing you actually want. That link, score to outcome, is criterion-related validity in one sentence.
Here is the catch most teams miss: a test can look polished, read well, and still predict nothing. A clean interface is not evidence. Only the correlation with real performance is. So before you trust any assessment, ask what outcome it has been checked against, and how strong that link is.

Types of criterion validity
Criterion-related validity comes in two forms, distinguished mainly by when you measure performance. Predictive validity asks whether a test score can forecast future job performance. Concurrent validity asks whether test scores align with performance that can already be measured. Both answer the same underlying question: does the test score relate to outcomes that matter?
Predictive validity: does the score forecast future success?
Predictive validity is especially useful in hiring because it tests whether assessment results can predict how candidates will perform after they are hired. You assess candidates before hiring, then compare their scores with performance measures collected later. If higher assessment scores consistently correspond with stronger job performance, the test demonstrates strong predictive validity.
The trade-off is time. You need to wait months, or sometimes longer, to collect enough performance data to know whether the assessment actually predicts success.
Concurrent validity: does the score match current performance?
Concurrent validity takes a faster approach. You give the assessment to people who are already performing the job and compare their scores with their current performance. If high-performing employees tend to score higher than lower-performing employees, the assessment shows strong concurrent validity.
The advantage is speed: you can study the relationship without waiting for new hires to accumulate performance data. The catch is that your current employees may not represent the candidate pool. They have already been hired, trained, and shaped by the organization, so a test that distinguishes today's employees may not predict how well tomorrow's applicants will perform.
Criterion related validity vs construct validity
Both types of validity ask whether an assessment is measuring what it should, but they look for evidence in different places. Criterion-related validity focuses on whether test scores relate to meaningful, real-world outcomes, such as job performance. Construct validity focuses on whether the assessment genuinely measures the underlying skill, trait, or concept it claims to measure. The table below breaks down the difference across definition, focus, example, and validation method.

How do you measure criterion-related validity?
Criterion-related validity is measured by testing whether assessment scores are meaningfully related to a real-world performance outcome. The process is less about proving that a test is “valid” once and more about building evidence that its scores consistently relate to something that matters on the job.
1. Define the performance criterion
Start by deciding what successful performance actually looks like for the role. Depending on the job, this could be sales revenue, productivity, quality scores, customer retention, time to proficiency, or a structured manager rating. The criterion needs to be measurable, reliable, and genuinely relevant to the job.
2. Administer the assessment
Give the assessment to the group you want to study, whether that means current employees or candidates. Record their scores consistently and make sure everyone is assessed under comparable conditions. This gives you the test data you will later compare with performance.
3. Collect performance data
Next, collect the real-world outcome you defined as the criterion. For predictive validity, you track performance after people are hired and have had enough time to demonstrate their ability. For concurrent validity, you compare assessment scores with performance data that already exists for people currently doing the job.
4. Compare scores with performance
Calculate the correlation coefficient between assessment scores and the performance criterion. The coefficient typically ranges from −1 to +1. A positive correlation means higher assessment scores tend to be associated with better performance, while a result close to zero suggests little linear relationship.
The strength of the correlation matters, but it should not be treated as a simple pass/fail number. A useful validation study also considers the size and quality of the sample, the reliability of the performance measure, and whether the relationship is statistically meaningful.
5. Validate the result with a representative sample
A strong correlation in one group does not automatically mean the assessment will predict performance in every hiring situation. Look at whether the sample represents the people you actually intend to assess and whether the relationship holds across relevant groups and time periods.
6. Recheck the assessment over time
Validity is evidence you maintain, not a label you assign once. As roles, candidate pools, assessment content, and performance expectations change, repeat the analysis to make sure the relationship still holds. If the evidence weakens, revisit the assessment, scoring approach, or performance criterion rather than assuming the test is still doing its job.
Pro tip: Watch for range restriction. If you validate a test only on employees who already passed your hiring process, the results may not tell you how well the assessment distinguishes the broader applicant pool. Testing with a sufficiently varied and relevant sample gives you a more realistic picture of the assessment's usefulness in hiring.
Why does criterion-related validity matter in hiring?
Criterion-related validity matters because a low-validity test costs real money while looking productive. When scores do not predict performance, you are sorting candidates by noise and calling it a process. The bill arrives later, as turnover, missed targets, and re-hiring. Validity is the cheapest insurance against that bill.
Gallup puts the cost of replacing one employee at one-half to two times their annual salary, and calls that conservative. For a 100-person organization paying an average of $50,000, it pegs the yearly turnover bill at $660,000 to $2.6 million, part of a roughly $1 trillion problem for US businesses. A meaningful share of that traces back to hires who looked right on paper and a screen that could not tell the difference.
Validity also changes which methods you should trust. A 2022 Journal of Applied Psychology study by Sackett and colleagues re-ran decades of selection research and found that many methods had been overstated, with corrected validity estimates falling by about 0.10 to 0.20 points.
Structured interviews came out as the top-ranked predictor of job performance, ahead of cognitive ability tests, which earlier work had treated as the gold standard. The lesson is not that one method wins forever; it is that you should check the evidence, not the reputation.
How does it compare to other validity types?
Criterion-related validity asks whether a test predicts an outcome. Other validity types ask different questions, and a strong assessment usually needs more than one. Content validity asks whether the test covers the actual job. Construct validity asks whether it measures the intended trait. Criterion-related validity asks whether the scores predict results. Here is how they line up.
Validity type | Core question | Hiring example |
|---|---|---|
Criterion-related | Do scores predict a real outcome? | Sales test scores track revenue closed |
Content | Does the test cover the real job? | A coding test uses tasks from the actual role |
Construct | Does it measure the intended trait? | A reasoning test really measures reasoning |
Face | Does it look relevant to candidates? | Applicants see the test as fair and job-related |
Two related ideas often come up alongside it. Construct validity confirms the test measures the trait it claims to, while discriminant validity confirms it is not accidentally measuring something else, like a coding test that quietly rewards verbal skill. Criterion-related validity is the one that ties everything back to performance, which is why hiring teams lean on it hardest.
Hire on proven validity, not guesswork
If a test cannot show it predicts performance, it does not belong in your hiring decision. Testlify’s assessments are built and benchmarked to track real outcomes, so your shortlist is scored on evidence before the first interview, not sorted by a number nobody has checked. Book a demo to see how validated assessments fit your roles, or start a free trial and run one on your next hire.
Key Takeaways
- Validity is proof, not polish. Criterion-related validity is the link between test scores and real performance. A test that looks good but predicts nothing is a cost, so judge every assessment by its correlation with outcomes, not its interface.
- Know which type you are claiming. Predictive validity forecasts future performance and takes time to confirm; concurrent validity checks current performance and is faster but easier to flatter. Match the method to the decision you are making.
- Measure it, then keep measuring. Define the criterion, run the correlation, and recheck as roles change. Validity is maintained over time, which means a test validated once is not validated forever.
- Bad validity is expensive. Gallup ties replacing one employee to one-half to two times their salary. Weak screens feed that bill, so validity is one of the cheapest ways to protect a hiring budget.
- Check evidence, not reputation. The 2022 Sackett research showed long-trusted methods were overstated and structured interviews predicted best. Pick selection methods on current evidence, because the rankings have shifted.
- Anchor tests to a criterion up front. Mapping each competency to measurable evidence, as in the Testlify Competency-to-Evidence Matrix, lets you confirm scores track outcomes instead of hoping they do, while the hiring team keeps the final decision.
Frequently asked questions (FAQs)
Related resources
View all
HR & recruitment
How can employers assess candidates’ clerical skills?

HR & recruitment
Types of interviews: A complete HR guide

HR & recruitment
Job characteristics model 101: A complete guide

HR & recruitment
9 HR models every HR professional must know in 2026

HR & recruitment
What Is a Projective Assessment? Types, Uses, and Reliability in Hiring (2026)

HR & recruitment
What are personality hires and signs you might be making one
Get started.
Hire on proof, not resumes.
Run your first skills-based assessment free — no credit card required.