Behaviourally Anchored Rating Scale (BARS): Behaviourally Anchored Rating Scale (BARS) : Behaviourally Anchored Rating Scale (BARS): Behaviourally Anchored Rati.
Summarise this post with:
Behaviourally Anchored Rating Scale (BARS) is a performance appraisal method that anchors each point on a rating scale to a specific, observable behavioural example describing what performance at that level looks like in the context of a defined job. Developed by Smith and Kendall in 1963. Also called: BARS, behavioral expectations scale (BES), behavioral rating scale.

The smith and kendall framework: where BARS came from
BARS was developed by psychologists Patricia Cain Smith and Lorne M. Kendall to address a persistent problem with traditional rating scales: managers were rating abstract traits using subjective and inconsistent reference points. Two managers could give the same employee very different ratings on the same trait, not because the employee’s behaviour differed but because the managers held different mental models of what each rating level meant.
Smith and Kendall’s 1963 paper proposed anchoring each rating-scale point with a specific behavioural example. They originally called the instrument a Behavioural Expectations Scale (BES); the term BARS became the dominant label as the method spread through industrial-organisational psychology.
BARS sits within what I-O psychology researchers call the “behavioural tradition” – a body of methods that ground performance evaluation in observable behaviour rather than inferred personality traits. Decades of subsequent research have shown BARS produces more consistent inter-rater agreement and stronger psychometric properties than unanchored Likert-style scales.
How a BARS scale is structured
Each BARS scale has the same structural components, applied to a single performance dimension at a time. A typical scale for a customer service representative on the dimension “handling escalated customer issues”:
- Rating 5 (Outstanding): Proactively identifies the underlying customer concern, resolves it on first contact, and proactively addresses related issues that the customer has not yet raised but is likely to encounter.
- Rating 4 (Above expected): Resolves the escalated issue on first contact within the SLA window, with high customer-satisfaction feedback. Coordinates with other functions where needed.
- Rating 3 (Meets expected): Resolves routine escalated issues within SLA window. Misses opportunities to proactively address related concerns.
- Rating 2 (Below expected): Resolves issues but frequently breaches SLA or requires multiple contacts. Escalates to supervisor more often than peers.
- Rating 1 (Unsatisfactory): Routinely fails to resolve escalated issues; customer-satisfaction feedback consistently below target. Requires supervisor intervention on most cases.
A complete BARS appraisal for a role typically covers 5-8 dimensions.
Building a BARS scale: the five-step methodology
BARS development is time-intensive – building a complete scale for one role typically takes 40-80 hours of subject-matter-expert time across multiple sessions. The standard methodology:
1. Identify performance dimensions. Subject matter experts (SMEs) – typically high-performing job holders and their managers – generate the list of behaviourally distinct dimensions of effective performance for the role. Aim for 5-8 dimensions per role.
- Generate critical incidents. Using John Flanagan’s 1954 critical incident technique, SMEs document specific examples of effective and ineffective performance they have observed in the role. Aim for 30-100 incidents per dimension.
- Retranslate the incidents. A second group of SMEs (not the original generators) sorts the incidents back into the original performance dimensions. Incidents that this second group cannot reliably classify into the same dimension are discarded as ambiguous.
- Rate the surviving incidents. SMEs rate each retained incident on a 1-7 (or similar) effectiveness scale. Incidents with high inter-rater agreement (low standard deviation across raters) become the anchor candidates.
- Construct the scale. For each dimension, select 5-7 incidents with low standard deviation and broad spread across the effectiveness range. These become the anchored scale points.
Benefits and limitations of BARS
- Benefit: reduces subjectivity. Anchored behavioural examples replace abstract trait labels, narrowing the interpretive variance between raters.
- Benefit: legally defensible. EEOC Uniform Guidelines (1978) require that performance evaluation methods be job-related and validated. BARS, built from job analysis, meets the standard better than most alternatives. Used widely in regulated industries (healthcare, financial services) where appraisal defensibility matters in litigation.
- Benefit: developmental value. The behavioural anchors themselves teach raters and employees what each level looks like, supporting clearer feedback and development conversations.
- Benefit: consistent across raters. Multiple validation studies have shown higher inter-rater agreement on BARS than on unanchored Likert-style scales.
- Limitation: time-intensive to build. 40-80 hours of SME time per role. For organizations with many distinct roles, the total development cost is substantial.
- Limitation: requires maintenance. As roles evolve, the behavioural anchors become outdated. Scales need periodic refresh, typically every 2-3 years.
- Limitation: rater training required. Without training, raters revert to halo effects or central-tendency bias, defeating the purpose of the anchored scale.
BARS vs other appraisal methods
| Method | Mechanism | Best for | Trade-off |
| BARS | Behaviour-anchored rating per dimension | Defined operational roles; regulated industries | Time-intensive to build |
| Unanchored Likert | 1-5 rating on trait labels | Low-stakes evaluation; quick deployment | High rater subjectivity |
| MBO / OKRs | Achievement against pre-set goals | Outcome-driven roles; agile contexts | Misses how the outcome was achieved |
| 360-degree feedback | Multi-source input | Development; leadership coaching | Less reliable for compensation decisions |
| Forced ranking (bell curve) | Distribution-constrained ranking | Pool-allocation in large workforces | Damages collaboration and morale |
See bell curve for forced-distribution alternatives and balanced scorecard for the strategic measurement framework.
Implementing BARS at scale
- Start with high-volume roles. Customer service, sales, manufacturing operators – roles with many incumbents amortize the development cost.
- Configure in the HRIS / performance management system. Workday, BambooHR, Lattice, 15Five, Culture Amp, Darwinbox, and similar platforms support BARS-style appraisal forms.
- Train raters before rollout. Frame-of-reference training calibrates raters against sample cases. Without it, BARS yields little advantage over unanchored scales.
- Pilot with one function. Run a complete cycle in one function before company-wide rollout. The pilot surfaces ambiguous anchors that should be revised.
- Refresh every 2-3 years. Re-run SME panels to update behavioural anchors as roles evolve.
Anchor selection decisions in demonstrated capability with Testlify’s validated assessments, which use behaviourally anchored criteria to evaluate candidates pre-hire. See also psychometric tests for complementary pre-hire assessment frameworks.
Chatgpt
Gemini
Claude
Grok









