Construct Validity
Construct Validity is the extent to which a measure or test accurately reflects the concept or construct it is intended to measure.
Construct Validity is the degree to which a test measures the psychological construct it claims to – the foundational evidence for EEOC-defensible assessment use. Also called: construct-related validity.

Why construct validity matters in hiring
In selection assessment, construct validity is not an academic nicety – it is the defensibility argument for the test. Under EEOC Uniform Guidelines on Employee Selection Procedures (29 CFR Part 1607), if a selection procedure causes adverse impact on a protected class, the employer must demonstrate that the procedure is ‘job-related and consistent with business necessity.’ Construct validity is one of three formally recognised methods of doing so. The other two are content validity (the test items represent the actual job behaviours) and criterion-related validity (test scores predict job performance).
If you use an assessment whose construct validity is unknown – a test bought off the shelf with no validation evidence, or an interview question someone in your team wrote because it ‘felt right’ – the assessment cannot be defended in an EEOC investigation. Adverse-impact exposure becomes a settlement, fast.
Construct validity vs content vs criterion-related validity
These three forms of validity are the three legs of EEOC-recognised job-relatedness. They are not interchangeable; they answer different questions and require different evidence.
Best practice: layered evidence. A strong assessment is supported by construct validity (the right thing is being measured), content validity (the items map to the job), and criterion-related validity (scores predict actual outcomes). Single-strategy validation is acceptable but weaker.
The two sides of construct validity: convergent and discriminant
Cronbach and Meehl’s 1955 framework decomposed construct validity into two complementary evidence streams. Both are needed.
Convergent validity
The test correlates with other measures of the same or related constructs. A new cognitive ability assessment should correlate strongly (typically r >= 0.50) with established cognitive measures (Wonderlic, Raven’s, GMA composites). If a ‘cognitive ability test’ shows weak correlation with established cognitive measures, the construct claim is suspect.
Discriminant validity
The test does not correlate strongly with measures of unrelated constructs. An emotional intelligence assessment should not correlate strongly with cognitive ability – otherwise it might just be re-measuring intelligence under a different label. Discriminant evidence is what prevents ‘construct contamination’ – measuring something other than what the test claims.
In selection practice, weak discriminant validity is the more common failure. Many ‘leadership potential’ or ‘soft skills’ assessments turn out, on inspection, to be substantially measuring cognitive ability or extraversion – useful constructs in their own right, but not what the marketing claims.
How construct validity is established
There is no single test for construct validity. It is built over time through accumulated evidence. The standard evidence types:
- Theoretical framework. Define the construct precisely, hypothesise how it should relate to other constructs, and predict observable behaviours.
- Factor analysis. Exploratory (EFA) and confirmatory (CFA) factor analysis tests whether items group into the expected dimensions. A personality test claiming to measure five traits should produce five factors.
- Convergent correlations. Administer the new test alongside validated measures of the same construct. Correlations r >= 0.50 are typical convergent benchmarks; r >= 0.30 is the lower bound for any meaningful relationship.
- Discriminant correlations. Administer alongside unrelated constructs. Correlations should be small (r <= 0.30) and non-significant.
- Known-groups validation. The test should distinguish groups that are theoretically expected to differ. A leadership potential test should produce higher scores for known successful leaders than for individual contributors.
- Predictive criterion correlations. Construct validity is strengthened when scores predict the outcomes the construct is theorised to predict – supervisor ratings, training outcomes, tenure. Track downstream impact with a quality of hire calculator.
- Cross-cultural and group-fairness evidence. Modern construct validation includes evidence that the test measures the same construct equivalently across demographic groups. EEOC and EU Pay Transparency Directive both treat this as part of the validity argument.
Construct validity in common selection assessments
Cognitive ability tests
Among the most extensively researched constructs in I-O psychology. General Mental Ability (GMA) shows strong construct validity through 90+ years of accumulated evidence – convergent across hundreds of cognitive measures, discriminant against personality, and predictive of job performance across virtually all role types. Buying a cognitive ability test without construct-validity evidence is the most common preventable selection-validity mistake.
Personality assessments
Strong construct validity exists for Big Five (OCEAN) personality dimensions through extensive research. Caution: many proprietary personality tests claim to measure unique constructs (‘cultural fit’, ‘innovation orientation’) that on inspection are repackaged Big Five facets. Demand the factor structure evidence.
Situational judgment tests (SJTs)
SJTs are construct-heterogeneous – a single SJT often measures cognitive ability, personality, and job knowledge simultaneously. This is not invalid per se but requires careful construct definition. The strongest SJTs are designed against an explicit construct framework, not just plausible scenarios.
Structured interviews
Structured interviews with behavioural questions and defined rubrics show stronger construct validity than unstructured interviews. Unstructured interviews often measure interviewer impression rather than candidate competence, which is why their predictive validity is roughly half that of structured interviews.
Work samples and simulations
Generally strongest on content validity (the work is the job). Construct validity argument is whether the simulation captures the cognitive and behavioural processes used in the actual role.
Where construct validity fails
- Construct underrepresentation. The test only measures part of the construct. A ‘communication skills’ test consisting only of grammar items misses listening, audience adaptation, and synthesis.
- Construct-irrelevant variance. The test measures something extra. A timed cognitive test that depends on reading English at native speed measures both cognition and English fluency, conflating two constructs.
- Theory drift. The construct is poorly defined. ‘Grit’, ‘EQ’, and ‘growth mindset’ have all been criticised for under-specification – tests measuring them often have unclear construct boundaries.
- Method bias. Two tests with the same format (both self-report) correlate because of shared method, not shared construct. Strong evidence comes from multi-method, multi-trait designs.
- Cultural bias. The construct itself may not transfer across cultures. ‘Assertiveness’ is differently valued in collectivist vs individualist cultures; a high score has different meaning.
How to evaluate a vendor’s construct validity claims
1. Ask for the technical manual. Every reputable assessment publisher provides one. If they don’t, the assessment isn’t ready for high-stakes use.
- Check the construct definition. Vague constructs (‘cultural fit’) without operational definitions are red flags. The manual should define exactly what is being measured.
- Look for convergent evidence. Correlations with established measures of the same construct, ideally r >= 0.50.
- Look for discriminant evidence. Correlations with unrelated constructs should be small. If ‘leadership potential’ correlates r = 0.75 with cognitive ability, you may be buying a cognitive test.
- Check factor structure. CFA or EFA results should show the items group into the claimed dimensions.
- Group fairness data. Differential item functioning (DIF) and measurement-invariance analyses across demographic groups. Missing or weak evidence here is an EEOC liability.
- Predictive evidence in your context. Construct validity is necessary but not sufficient – the skills assessment must also predict performance in roles like yours. Local validation or transportability evidence closes the loop.
Frequently asked questions
Construct validity is the degree to which a test actually measures the abstract psychological characteristic it claims to measure. A cognitive ability test claims to measure cognitive ability; construct validity is the accumulated evidence that it does, not something else. It is the most foundational form of psychometric validity.
Get started.
Hire on proof, not resumes.
Run your first skills-based assessment free — no credit card required.