CozyHR
Menu
Products
Docs
Resources
Compliance
Company
Support
Blog
Performance ManagementHR PoliciesCompensationHR Analytics

Performance Calibration: Running Fair Rating Reviews

How to run performance calibration meetings that make ratings comparable and defensible, without a crude bell curve: preparation, facilitation scripts, documentation and follow-...

CozyHR editorial team 09 September 2026 38 min read
CozyHR Blog
Performance Calibration: Running Fair Rating Reviews

Performance calibration is the step most Indian companies skip and then regret. Managers submit ratings, HR aggregates them, someone notices that one team has rated 80% of its people "Exceeds Expectations" while another has rated nobody above "Meets", and by then the increment letters are being drafted. A calibration meeting is the structured conversation between manager ratings and final ratings, where managers at the same level review their proposed ratings side by side, defend them with evidence, and adjust where the evidence does not hold. Done well it makes ratings comparable and defensible; done badly, the appraisal cycle becomes a lottery decided by which manager you report to.

What Performance Calibration Actually Is

Calibration is a moderation step. Managers arrive with draft ratings and leave with final ratings tested against a shared standard and against each other's people. The question is always the same: does the evidence support the rating proposed, compared with everyone else at this level?

The output is three things — final ratings, a written rationale for every rating at the extremes or that moved, and a shared understanding of what "good" means at each level.

The Problems It Solves

Rating inflation. Managers rate generously because generosity is cheaper than a difficult conversation. When 70% of a company sits in the top two boxes, the scale has stopped carrying information.

Harsh and lenient managers. Two engineers doing comparable work in adjacent teams get different ratings because one manager thinks top ratings require heroics and the other gives them for reliability. The employees compare notes, and the harsher manager loses their strong performer.

Recency and halo effects. The manager remembers the last eight weeks vividly and the previous ten months hazily. One salient trait then colours everything: the articulate person is credited with delivery they did not do, while the quiet person is marked down for "lacking impact" because impact was measured by visibility.

Ratings that cannot be compared. A "3" in Sales is calibrated against quota; a "3" in Design is calibrated against the design lead's private taste. One increment matrix then gets applied to both.

Increment budgets that break. If the compensation plan assumes a spread and submitted ratings cluster at the top, the budget does not stretch. HR then asks for more money, dilutes the top-band increment until it stops feeling like recognition, or sends ratings back for "moderation" — a forced curve applied late and without evidence.

Calibration is not a performance review, and the employee is never in the room. It is not a promotion committee, though the two often run back to back — keep them separate, because performance in the current role and readiness for the next are different questions. And it is not a compensation meeting: the moment a manager knows that moving someone from 3 to 4 is worth Rs 40,000, the conversation stops being about evidence.

Why the Forced Bell Curve Fell Out of Favour

For two decades, large Indian IT services firms and multinational captives ran forced distribution: a fixed percentage in each band, and managers had to make the numbers fit. Several very large employers publicly moved away from it, and the practice is far less common now.

The reasons are practical. A bell curve describes a large random population, and a twelve-person team that has been selectively hired, trained and already pruned is not one — forcing a curve manufactures a bottom performer who does not exist. It degrades collaboration, because if a fixed share must be rated low, helping a colleague succeed worsens your own position. It punishes good managers, since the one who built a strong team must downgrade someone while the one who tolerated weak performance has easy low ratings to hand out. And it produces reverse-engineered documentation, the worst kind of record to hold if a rating is later challenged.

"Guidelines, Not Quotas" in Practice

The workable middle is a distribution guideline: an expected shape that prompts a conversation when reality diverges, with no requirement to hit the numbers. Under a quota, a manager must move someone. Under a guideline, a manager must explain.

In the room it sounds like this. A team of 14 has six people proposed at the top band against a guideline suggesting two or three. The facilitator does not say "cut it to three", but "the guideline suggests two or three at this level and you have six — walk us through each one and tell us what they did that a solid performer did not do." Sometimes all six survive. Usually two or three do not, because the manager cannot articulate a difference between them and the people rated a band lower.

SituationGuideline helpsGuideline harms
Large cohort (40+ at one level and function)Yes — the shape is meaningful and drift is visible
Small team viewed in isolationYes — creates artificial bottom performers
Managers with very different rating habitsYes — surfaces the gap for discussion
A team that genuinely outperformed on a measurable outcomeYes — mechanical use suppresses real excellence
First cycle after visible rating inflationYes — gives the facilitator a neutral reference
Deciding an individual's ratingYes — an individual is never a distribution
Company-wide view after calibration closesYes — a health check on the process
New team with heavy backfillingYes — tenure mix makes the shape meaningless

Apply the guideline to populations, never to people, and look at the shape after the evidence conversation rather than before it.

Companies that drop forced ranking usually land on one of three models. Evidence-led calibration with a soft guideline is the default. Anchored group review uses reference employees per band ("this is what a 4 looks like at Level 3 in Engineering") with no distribution reference at all — higher quality, slower, dependent on a skilled facilitator. Relative ranking merges managers' ranked lists and assigns ratings at natural breaks rather than by percentile, which helps when a genuinely limited promotion budget must be allocated on relative merit.

Prerequisites for a Useful Calibration Meeting

Calibration cannot fix a broken performance system; it can only moderate a working one. Without the following, the session becomes an opinion contest.

Clear goals set at the start of the period. If goals were written in the last month of the year, the cycle is not measurable.

A rating scale with behavioural anchors. "Exceeds Expectations" means nothing until it is described in observable behaviour, per level.

Evidence captured through the year — goal data, project outcomes, written feedback, stakeholder input. If the only evidence is what the manager recalls in March, recency bias is guaranteed.

Manager training: ninety minutes per cycle on the scale, the evidence standard and how a session runs, including two or three practice cases.

A published appraisal timeline, and leadership willing to be moderated. If a senior leader can overturn a calibrated rating by email afterwards, the process will not survive a second cycle.

A Realistic Appraisal Timeline

Most Indian companies work to a financial-year rhythm with increments effective from April.

WeekActivityOwner
Jan, week 3Cycle announced; goal and evidence review reminderHR
Feb, week 1Manager refresher on scale, anchors and evidenceHR / L&D
Feb, weeks 2-3Self-reviews open and closeEmployees
Feb w4 - Mar w1Draft ratings and written rationale submittedManagers
Mar, week 1Pre-read packs circulated (48 hours minimum)HR
Mar, week 2Calibration sessions by cohortFacilitator + managers
Mar, week 3Cross-cohort review; sign-off; ratings lockedHR + leadership
Mar, weeks 3-4Compensation modelling against final ratingsHR + Finance
Apr, week 1Appraisal conversations and lettersManagers
Apr, weeks 2-3Appeal windowEmployees
Apr, week 4Appeals resolved; cycle retrospectiveHR

Half-yearly cycles compress this, and many companies run a lighter mid-year calibration covering only the extremes plus anyone whose rating moved by more than one band.

Designing the Rating Scale

A 3-point scale is simple and hard to game, but it compresses an outstanding contributor and a solidly-above-average one into the same box, making differentiated increments difficult. It suits companies under about 150 people. A 4-point scale removes the safe middle and forces every rater to place a person on the strong or the developing side of the line. A 5-point scale is the most common choice in India and distinguishes solid from strong from exceptional, which matters when increments span three or four bands — its risk is central tendency, where the middle box becomes the default for anyone the manager does not want to think hard about.

The argument for an even number is that it kills fence-sitting. The argument against is that fully meeting a demanding standard is the most common honest outcome in a well-run company, and a scale with no home for it forces managers to flatter or insult most of their people. The practical resolution is a 5-point scale with the middle box defined positively — "fully meets a demanding standard; most strong performers will be here". Some companies rename it ("Performing Well", "On Track") to break the association between the midpoint and mediocrity. Do not change the number of points between cycles without publishing a mapping, or drift analysis becomes impossible.

Behavioural Anchors

Anchors turn a label into a testable claim. Write them per level, describe observable behaviour, and avoid adjectives that cannot be evidenced.

RatingLabelDelivery and outcomesBehaviour and collaborationOwnership and judgement
5ExceptionalOutcomes materially beyond agreed scope; results visible outside the teamRaised the performance of others; sought out hard feedbackTook on ambiguous problems unasked and resolved them
4StrongMet all goals and exceeded on the hardest onesActively unblocked peers; gives and takes feedback wellAnticipated problems and escalated early with options
3Performing wellMet agreed goals at the standard expected for the levelDependable; constructive in reviewsOwns their scope end to end; asks for help appropriately
2DevelopingMet some goals; missed others within their controlNeeds prompting to collaborate; feedback partly acted onNeeds close direction on work that should be independent
1Below expectationsDid not meet core responsibilities of the roleCreated friction or rework for othersRepeated escalation on routine work

Two rules make anchors work. A rating of 5 or 1 always requires named, dated evidence in the session — "everyone knows" is not evidence. And evidence must describe what the person did, not how the manager feels about them.

Performance Rating Versus Potential Rating

Performance is backward-looking: what was delivered in this period, in this role. Potential is forward-looking: how likely the person is to succeed at materially larger scope. Keep them separate, or you get two predictable errors — the brilliant specialist with no interest in management is marked down, and the under-delivering but ambitious employee is marked up on promise.

The 9-box grid plots one against the other and can be a useful talent-review artefact, but treat it carefully in a smaller company. Cell counts get tiny, so the grid looks precise while carrying almost no information. Potential is the least evidenced and most bias-prone judgement in the process, yet the grid gives it equal visual weight. Labels leak, and once someone is described aloud as "low potential" it follows them. Calibrate performance properly, and run a separate, smaller talent review for the population where potential decisions genuinely matter.

Preparing the Pre-Read Pack

The biggest determinant of quality is whether managers walked in prepared. Circulate the pack at least 48 hours ahead, and state plainly that managers who have not read it will have their people deferred.

IncludeWhy
Name, level, role, function, managerIdentification and cohort placement
Date of joining; tenure in roleSets expectations; drives proration
Goal attainment summaryThe primary objective input
Self-review summary (3-5 lines)Often surfaces work the manager forgot
Proposed manager ratingThe starting point
Written rationale (100-150 words, mandatory)Evidence submitted before the room, so it cannot be improvised
Two or three specific evidence pointsNamed projects, outcomes, dates
Peer or stakeholder input where collectedCross-check on collaboration and impact
Last cycle's ratingDetects unexplained swings either way
Promotion candidacy flagSignals a separate discussion
Days on leave, with typeNeeded for fair proration
Internal transfer during the periodTriggers the previous manager's written input
Open PIP or disciplinary statusContext that must not surprise the room

Equally important is what stays out of the first pass.

Deliberately excludeReason
Current salary and CTCAnchors the rating to cost; the expensive person gets defended, the cheap one downgraded
Proposed increment or bonus figuresTurns an evidence conversation into a budget negotiation
Age, gender, marital status, caste, religion, region, photographIrrelevant to performance and actively invites bias — exclude entirely
Personal circumstances beyond documented leaveNot performance data
Unverified anonymous complaintsHandle through the correct process, not calibration

Salary data has a legitimate place — in compensation modelling after ratings lock, where you check for anomalies such as a top-rated performer sitting well below market. That is a valid conversation, just not the same one.

A Sample Pre-Read Row

Priya S. | Level 3 Software Engineer | Platform | Manager: R. Nair | 19 months in role Goals: 4 of 5 fully met, 1 partial (latency target missed by 12%). Last cycle: 3. Proposed: 4. Promotion candidate: No. Leave: 6 days. Manager rationale: Owned the billing reconciliation rewrite end to end after the original owner exited in June; delivered two weeks ahead of the revised plan with zero post-release incidents. Redesigned the on-call rota, cutting weekend pages from around 9 a month to 2. Missed the search latency target — root cause was an upstream dependency she flagged in July with a workaround that was deprioritised. Coaches two juniors informally; both named her unprompted in their self-reviews. Peer input (Product, SRE): "Explains trade-offs clearly and does not oversell." "The only person who reads the postmortems properly."

That takes ten minutes to write and saves the room fifteen minutes of interrogation. "Priya is a great team player and always delivers, strong 4" saves nothing, and should be sent back before the session rather than debated in it.

Who Is in the Room

The managers whose people are being calibrated — all of them, for the whole session, not dropping in for their own people. The comparison only works if every manager hears every case.

A facilitator, usually the HR business partner for that function, with no employees in the cohort. The skip-level leader, to provide context and arbitrate deadlocks — but not to speak first on every case, or everyone will simply agree. And a scribe, capturing decisions live. Nobody else: not employees, not managers from unrelated functions, not Finance during the rating pass.

The facilitator has one job — protect the standard. That means anchoring every discussion to evidence, managing airtime so the loudest manager does not calibrate the cohort, naming bias patterns without accusing anyone of prejudice, drawing out quiet managers, recording dissent, and closing each case explicitly.

Slicing Cohorts

By level within function is the default and most defensible: all Level 3 engineers compared together regardless of sub-team, because expectations are consistent. By function across levels works only for small functions, and needs the facilitator to keep restating the level expectation — a Level 2 doing excellent Level 2 work should out-rate a Level 4 doing adequate Level 2 work. Cross-function suits levels above senior manager, where per-function populations are thin. A short skip-level cross-check then reviews only the extremes across all cohorts, to confirm a 5 means the same thing in Sales as in Engineering.

Practical sizing: 6 to 10 managers covering 40 to 80 employees. Fewer than four managers gives too little comparison; more than twelve turns into rubber-stamping.

The Calibration Session Agenda

A session for roughly 50 employees across 8 managers needs about three hours. Ninety minutes produces a rubber stamp.

TimeSegmentWhat happens
0:00-0:10FramingPurpose, scale reminder, ground rules; state that ratings may move up as well as down
0:10-0:20Anchor settingAgree one reference person per band at this level
0:20-0:35Distribution snapshotProposed distribution by manager and cohort; no names, no decisions
0:35-1:20Top band reviewEvery proposed top-band rating, one by one, with evidence
1:20-1:30Break
1:30-2:00Bottom band reviewEvery bottom-band rating, plus a check that feedback was given during the year
2:00-2:35Flagged middle casesTwo-band moves, self-review disagreements, promotion candidates, new joiners, transfers, long leave
2:35-2:50Final distributionRevised shape reviewed; remaining outlier managers asked to explain
2:50-3:00CloseConfirm decisions, assign rationale write-ups, agree communication timing

Review the top band before the bottom — that is where inflation and budget pressure concentrate. And do not walk through every middle rating; a cohort of 50 may have 30 uncontroversial ratings that need no discussion at all.

Ground Rules to State Aloud

  1. Everything discussed here is confidential and stays in this room.
  2. Every top- or bottom-band rating needs specific, dated evidence.
  3. Describe what the person did, not what they are like.
  4. Ratings can move up as well as down. This is not a cost-cutting exercise.
  5. The distribution guideline prompts a question; it does not dictate an answer.
  6. Compensation is not discussed in this room.
  7. If you disagree, say so here — do not relitigate it by email afterwards.
  8. Challenge each other. Silence will be read as agreement.

Facilitating the Hard Moments

What separates real calibration from a formality is how the facilitator handles the predictable difficult moments. The scripts below are prompts, not lines to recite.

The Manager Who Cannot Produce Evidence

"You've proposed a 5 for Arjun. Give the room two specific things he delivered this year that a strong performer at this level would not have."

If the answer stays general — "he's always available", "the team loves him" — do not argue about the person. Redirect to the standard: "everything you've described sounds like a reliable, well-liked colleague, which is what we expect at a 3 or 4. For a 5 I need something that changed an outcome."

If nothing comes, park it: "come back after the break with two concrete examples, or we settle at 4 and move on." Parking gives the manager a face-saving route, keeps the session moving, and makes clear the evidence bar is real. Most parked cases resolve themselves.

Effort Rather Than Outcome

"She worked every weekend in Q3 and was on calls till midnight during the migration."

Dismissing effort makes managers defensive. Acknowledge, then separate: "that commitment is real and it belongs in her written feedback. But the rating measures outcomes at her level. What did the effort produce — and is there something in how the work was planned that meant it needed weekends?"

The second question matters. Sustained heroics are usually a symptom of poor planning upstream, and rewarding them teaches the organisation that firefighting pays better than prevention. The "difficult year" defence takes the same shape: recognise the circumstance in the conversation and the support offered, but do not encode it in the rating.

The Rating That Moved Because the Budget Is Tight

"I've moved Sameer from 4 to 3 because we can't afford four 4s in my team."

The response should be flat and immediate: "we do not discuss budget in this room, and a rating never changes for budget reasons. Give me the evidence-based rating. If the compensation model has a problem, I'll take it to Finance after this session."

Then follow through. If calibrated ratings genuinely exceed the budget, the honest options are to adjust the increment matrix, sharpen differentiation between bands, or go back to leadership. Bending ratings to fit the money destroys trust in the whole cycle.

Protecting the Quiet High Performer

Some of the strongest contributors have no advocate, because their manager is not a natural presenter or their work is invisible infrastructure. Look for them in the pre-read: high goal attainment with a middling proposed rating, strong peer input with no matching manager enthusiasm, or someone repeatedly named in other people's self-reviews.

"Deepa has full goal attainment, two peers named her unprompted as the person who unblocked them, and she's proposed at a 3. Help me understand what's missing."

A useful general prompt: "whose work would we only notice if it stopped, and are those people rated correctly?" The employees who do unglamorous, essential work and land in the middle year after year are precisely the ones whose resignation surprises everyone.

The Manager Who Rates Everyone the Same

Ten people at 4 and one at 3, or all twelve at 3. Do not challenge the ratings first — ask for differentiation.

"If you had to name the two who had the strongest year and the two who had the weakest, who are they and why?"

Almost every manager can answer, because they do know the difference; they are avoiding the consequence of writing it down. Then: "so Nikhil and Farida are clearly ahead, and Rohit had a harder year. Does the rating set reflect what you've just told us?" If the manager genuinely cannot differentiate, log it as a development issue — coaching fixes that before the next cycle, argument in this one does not.

New Joiners, Long Leave and Transfers

"She joined in December, three months isn't enough to rate."

Apply the published policy rather than deciding in the room: "under three months is 'not rated' with no increment and a first review next cycle. She has four months, so we do rate her — against four months of expectations, not twelve."

For long leave, including maternity leave, rate the period actually worked at the normal standard, with no penalty for absence.

"He was on leave for four months. We rate the eight he worked, at the normal standard for eight months. We do not scale the rating down because he was away, and we do not scale expectations up either."

Say this explicitly and repeatedly, because the drift is real — managers unconsciously mark down people who "weren't around for the big project", particularly after long maternity leave. Beyond the legal exposure, it is unfair and highly visible to the rest of the team. Leave is not performance.

For transfers, insist on written input from the previous manager, submitted with the draft rating: "Vikram moved to your team in September. Do you have R's input for April to August? If not we park him — we're not rating four months of work as if it were twelve."

The Manager Who Will Not Let Go

Let disagreement stand rather than manufacturing consensus.

"The room's view is 3, yours is 4, and we've heard both. I'm recording 3 with your dissent noted, and the final call sits with the skip-level leader. If you want to escalate, do it before ratings lock on Friday — not after the employee has been told."

Recorded dissent gives the manager a legitimate outlet and creates an honest record. Relatedly, when a name arrives with a reputation attached and managers who have never worked with the person volunteer views, cut it off: "let's hear only from people with direct evidence. Reputation isn't evidence, and it travels much further than it should."

Manager Rating Bias and Counter-Moves

Bias is not a character flaw; it is a predictable feature of judgement under time pressure. Naming the types gives the room a shared, non-accusatory language for challenging each other.

BiasHow it shows upCounter-move
RecencyRationale draws entirely on the last quarterRequire one first-half example for every top-band rating
HaloOne strength inflates every dimensionScore competencies separately before the overall rating
HornsOne visible failure defines the whole yearRequire a named strength for every bottom-band rating
LeniencyThe manager's whole team clusters at the topAsk for an internal ranking, then test whether ratings match it
SeverityTop ratings must be "earned the hard way"Read their band definitions against the anchors; check team attrition
Central tendencyEveryone lands in the middle boxAsk who is strongest and weakest, and why ratings do not show it
Similar-to-mePraise centres on style, college, language or hours workedRedirect to "what did they deliver?"; check whether top ratings share a profile
VisibilityClient-facing work over-rated, maintenance under-ratedAsk whose work would only be noticed if it stopped
Anchoring on last yearRating carried forward with little thoughtFlag any rating unchanged for three cycles for a fresh look
Attribution errorIndividual credited for a team outcome, or blamed for a systemic oneAsk what this person did that a competent peer would not have
Sympathy or hardshipRating inflated after a difficult personal yearRecognise it in the conversation and support offered, not the rating
Tenure biasLong service treated as a proxy for performanceCompare against level expectations, not years served
ProximityIn-office staff rated above equally productive remote colleaguesCompare output evidence; check whether remote staff cluster lower
Gendered languageWomen described as "supportive", men as "strategic"Read rationales aloud; replace adjectives with outcomes
Cost anchoringExpensive people defended, cheap people downgradedKeep salary out of the pre-read entirely

A technique that costs nothing: before the top-band discussion, read three or four rationales aloud without naming the employee and ask whether the language describes outcomes or impressions. It resets the standard more effectively than any slide on unconscious bias.

Proration and Special Cases

Ambiguity here is where perceived unfairness concentrates. Publish the rules with the cycle announcement and apply them without exception.

SituationRating treatmentIncrement treatment (illustrative)
Joined under 3 months before period endNot rated this cycleNo increment; first review next cycle
Joined 3-6 months before period endRated against period worked, at new-joiner expectationsProrated, or deferred if hired recently at market rate
Joined 6-12 months before period endRated normally against period workedProrated by months worked
Internal transfer during the periodOne rating, requiring written input from the previous managerFull increment against the calibrated rating
Promoted during the periodRated against the level held for most of the periodFull increment; check overlap with the promotion increase
Long leave (medical, maternity, sabbatical)Rated on the period worked, at normal standard, no penalty for absenceFull increment; do not prorate protected leave without legal advice
Statutory maternity leave for most of the periodNot rated where the worked period is insufficientPer policy and applicable law; apply consistently and document the rule
Serving notice at rating timeRate normally for the recordNo increment; excluded from promotion
On a formal improvement planRate on actual performance; PIP status is context, not the ratingPer policy, usually no increment

Two principles hold this together: rate the period worked at the standard for that period, and never let absence translate into a lower rating. For anything touching statutory leave entitlements or a termination decision, take specific advice from employment counsel rather than relying on a policy table.

Documentation and the Audit Trail

For every employee, record the final rating and the rating it started from; who raised any change, what evidence drove it and what the group concluded; any dissent with the manager named; and the decision date and session reference. For every session, record the cohort definition, attendees, facilitator, distribution before and after, and parked cases with their resolution.

The written rationale earns its keep in three ways. It survives the manager, so when a new manager must explain last cycle's rating, the file is the only source of truth. It supports difficult follow-through — if a low rating leads to an improvement plan or eventually an exit, a contemporaneous record made before the outcome was known is far more credible than one written afterwards. Employment disputes in India can surface long after the event, and exposure depends on the contract, applicable state legislation and the facts, so involve counsel early on any termination; the general principle is that contemporaneous documentation helps and reconstructed documentation does not. It also makes appeals answerable — "the group reviewed your rating alongside 47 peers at your level and concluded X on the basis of Y" is an answer; "that's what the process produced" is not.

Four sentences is enough: what they were responsible for; what they delivered, with dates and results; how they delivered it; and what the group concluded, including any change from the proposed rating.

"Proposed 5, calibrated to 4. The group recognised the vendor migration and the quality of the runbook documentation. Against the Level 4 anchor for Exceptional, however, the group did not find evidence of impact beyond the immediate team, which is what distinguished the two employees calibrated at 5 in this cohort. Manager R. Nair recorded disagreement on the basis that the migration prevented a business-critical outage; the skip-level leader confirmed 4, with a commitment to revisit at mid-year if the platform work extends across teams as planned."

Linking Calibrated Ratings to Increments

Ratings become consequential when they touch money, which is where the process is most easily corrupted. The rating pass must be complete and locked before the numbers appear.

  1. Managers submit draft ratings with written rationale.
  2. Calibration sessions run by cohort; ratings moderated on evidence alone.
  3. Cross-cohort review of top and bottom bands; leadership sign-off.
  4. Ratings locked — no changes without a documented exception.
  5. Compensation modelling begins against locked ratings.
  6. Increment matrix applied; anomalies reviewed for market position, parity and retention risk.
  7. Letters prepared and managers briefed.
  8. Outcomes communicated.

A Worked Increment Example

The figures below are illustrative and chosen to make the arithmetic easy to follow. Every company should set its own matrix based on budget, market benchmarks and business performance.

Take a company of 400 employees with an average salary of Rs 12,00,000 and a salary bill of Rs 48,00,00,000. Leadership approves an increment budget of 9%, or Rs 4,32,00,000.

RatingEmployeesShareIllustrative incrementTotal cost
5 — Exceptional246%18%Rs 51,84,000
4 — Strong9223%13%Rs 1,43,52,000
3 — Performing well23659%8%Rs 2,26,56,000
2 — Developing4010%3%Rs 14,40,000
1 — Below expectations82%0%Rs 0
Total400100%Blended 9.09%Rs 4,36,32,000

The model lands Rs 4,32,000 above budget, about 0.1% of the salary bill, closed by trimming the mid-band increment from 8% to roughly 7.85%.

Now suppose calibration had not run and the submitted ratings were inflated, with 15% at the top band and 40% at the next.

RatingEmployees (uncalibrated)IncrementTotal cost
56018%Rs 1,29,60,000
416013%Rs 2,49,60,000
31608%Rs 1,53,60,000
2163%Rs 5,76,000
140%Rs 0
Total400Blended 11.22%Rs 5,38,56,000

That is an overrun of Rs 1,06,56,000, roughly 25% over plan. Faced with it, most companies cut the top-band increment from 18% to around 13% — meaning the genuinely exceptional performer receives almost the same increase as a solid one — or send ratings back for "moderation".

At the individual level: two employees on Rs 12,00,000, both proposed at 5. One is confirmed at 5 on the strength of a platform initiative adopted across three teams, and her 18% takes her to Rs 14,16,000, up Rs 2,16,000. The other is calibrated to 4 because the group could not distinguish him from the strongest 4s, and his 13% takes him to Rs 13,56,000, up Rs 1,56,000. The Rs 60,000 gap is exactly why the top-band evidence bar has to be defended: let the top band get generous and the gap collapses.

Ratings typically feed three levers, and the mapping should be explicit. Base increment is driven by the rating, adjusted for market position and parity. Variable pay usually combines an individual multiplier from the rating with a company multiplier from financial results — state the formula openly. Promotion is a separate decision: a strong rating is necessary but never sufficient, since promotion also needs evidence of already operating at the next level and a role to fill. Where a strong performer sits well below market, handle it through a separate market correction budget after ratings lock, not by inflating the rating.

Communicating Outcomes

Calibration quality is invisible to employees. What they experience is the conversation with their manager, and a well-calibrated rating delivered badly does more damage than a mediocre process communicated honestly.

Managers should own the decision — "this is your rating and here is the reasoning", never "I fought for you but HR overruled me". The moment a manager disowns the outcome, trust moves from the company to that individual and the process loses its authority. They should lead with the same specifics the group discussed, and say plainly that calibration happened: "your rating was reviewed alongside everyone else at your level across the function, so it reflects a consistent standard rather than only my view."

What managers should not do: blame HR, the committee or the budget; name other employees or reveal anyone's rating; describe the guideline as a quota ("there were only so many 4s available" undoes months of design in one sentence); promise a future rating; or deliver the increment number in the same breath as the rating. Give the performance conversation its own space before moving to compensation.

The Downgrade Conversation

Prepare the manager with three things: the specific evidence, the specific gap against the anchor, and one concrete route to close it.

"Your rating this cycle is 3. Last cycle it was 4, so I want to explain the difference properly. Last year you led the payments integration end to end, which is the scope we expect at a 4. This year your delivery was solid — the reporting module shipped on time and the quality was good — but the scope was closer to the standard expectation for your level, and two goals were partially met. This isn't a judgement about your ability. It reflects what these twelve months looked like against a standard applied to everyone at your level. What would move it back to 4 is owning something with a dependency outside your own team. The migration work starting next quarter is exactly that shape. Do you want it?"

It names the previous rating rather than pretending nothing changed, separates the year from the person, and ends with a specific route rather than generic encouragement. Expect an emotional response and do not rush past it. If the employee disagrees, acknowledge, restate the evidence once, and point them to the appeal process.

The Appeal Process

A workable design: an appeal window of 7 to 10 working days after communication; a written appeal stating what the employee believes was not considered; review by the skip-level manager and the HR business partner, both able to see the calibration record; a written response within 10 working days; and a change only where new evidence emerges or a process error is found, not simply because the employee disagrees.

Track appeal volumes by manager and function. A consistently high rate for one manager usually indicates a communication problem rather than a rating problem; a spike across a function suggests the standard was applied inconsistently in that cohort. Most appeals do not change the rating, and that is not failure — the point is that the employee gets a considered, evidence-based answer from someone other than their own manager.

Measuring Whether Calibration Worked

MeasureWhat it tells you
Distribution drift across cyclesWhether inflation is creeping back
Share of ratings changed in sessionUnder 5% suggests rubber-stamping; over 30% suggests unprepared managers
Ratio of upward to downward changesOne-directional movement suggests a purpose other than accuracy
Manager average and spread over cyclesIdentifies persistently lenient or severe raters for coaching
Gap between the most and least generous manager in a cohortShould narrow over cycles if calibration is working
Voluntary attrition by rating band, 6-12 months onHigh attrition among top-rated staff signals a recognition or pay problem; low attrition among bottom-rated staff means the message is not landing
Correlation between this cycle's rating and the nextVery high suggests anchoring; near-zero suggests the scale is noise
Appeal rate by manager and functionCommunication quality and cohort-level inconsistency
Post-cycle pulse on clarity and fairnessThe outcome that ultimately matters

Do not over-index on any single number. Distribution alone says very little; distribution plus movement rate plus appeal rate plus attrition-by-band says a great deal. It is also worth comparing distribution by demographic group at an aggregate level, handled carefully so individuals cannot be identified.

Running Calibration Remotely or Across Locations

Keep sessions shorter and split them — two 90-minute blocks on consecutive days beat one three-hour video call. Always share the cohort view on screen; without it managers lose the thread and stop contributing.

Cameras on, and call on people by name, because the failure mode of remote calibration is silent agreement: "Rajesh, you worked with three of Meera's team on the integration — does the 4 for Kunal match what you saw?" Use a shared document rather than slides, so decisions get recorded live, and ask for questions in the main room rather than private chats. In multi-location companies, run the distribution by location as a routine check — where leadership sits in one city and delivery in another, the headquarters population is often rated higher, and proximity bias should be named openly if the gap persists.

How an HRMS Supports Calibration

Calibration can run on a spreadsheet, and plenty of companies do. It gets painful above roughly 150 employees and error-prone above 400. The specific load a system should take off you: goal and evidence capture through the year, so evidence exists in March because it was recorded in July; configurable rating scales with anchors visible at the point of rating; automated pre-read generation with excluded fields genuinely hidden; a live calibration view showing proposed and revised distribution by manager, level and function; rationale capture with a minimum standard enforced; a locked audit trail; proration applied automatically from joining, transfer and leave records; increment modelling against locked ratings; and cycle-over-cycle analytics on drift, manager consistency, attrition and appeals. The point is not the tooling — it is that HR time moves from assembling data to facilitating the conversation.

Common Mistakes

Calibrating without an evidence standard. The session becomes a debate about persuasion, and the most confident manager wins.

Running a forced curve and calling it calibration. People can tell within one cycle, and denying it is worse than doing it openly.

Discussing salary in the rating session. It contaminates every judgement that follows.

Reviewing every single employee. Concentrate on the extremes and the flagged exceptions.

Letting the most senior person speak first. Whatever they say becomes the anchor.

Skipping anchor setting. Ten minutes agreeing what a 4 looks like saves an hour of circular argument.

No facilitator, or one who owns people in the cohort. The session drifts to whoever is loudest.

Allowing post-session changes by email. One overturned decision and nobody takes the next session seriously.

Calibrating only downward. Managers will learn to inflate their drafts as a negotiating position.

Communicating the guideline as a quota. One sentence undoes months of careful design.

No appeal route. Disagreement does not disappear; it turns up in the exit interview instead.

Treating leave as underperformance. Unfair, highly visible, and a legal exposure worth avoiding.

Locking ratings and never looking back. Without measurement you cannot tell whether calibration improved anything.

A Cycle-by-Cycle Improvement Plan

Cycle 1 — establish the mechanics. Publish the scale with anchors, train managers, produce pre-read packs, run facilitated sessions, record decisions. Expect uneven evidence and messy discussions.

Cycle 2 — raise the evidence bar. Reject weak rationales before the session rather than in it. Introduce anchor setting, pilot peer input in one cohort, and share manager-level distributions privately with each manager.

Cycle 3 — tighten and measure. Introduce a carefully framed distribution guideline, run the cross-cohort review of extremes, add attrition-by-band and appeal-rate analysis, and coach outlier raters using their own data.

Cycle 4 onwards — extend the value. Feed outputs into development and succession planning, shorten sessions as manager quality improves, and publish an anonymised process summary to employees.

Implementation Checklist

  1. Confirm the rating scale, with behavioural anchors written per level and published.
  2. Publish the appraisal timeline with dates for every stage.
  3. Verify every employee has written goals or a written scope for the period.
  4. Define cohorts: 6-10 managers and 40-80 employees per session.
  5. Name a facilitator per cohort with no people being calibrated in that session.
  6. Build the pre-read template and confirm the excluded fields are genuinely hidden.
  7. Set the minimum standard for a written rationale and communicate it before drafting begins.
  8. Train managers on the scale, the anchors, the evidence bar and two practice cases.
  9. Publish proration rules for new joiners, transfers, promotions, long leave and notice periods.
  10. Decide whether to use a distribution guideline, and how it will be framed in the room.
  11. Circulate pre-reads 48 hours ahead and chase weak rationales before the session.
  12. Run sessions to the agenda, with a scribe capturing decisions live.
  13. Hold the cross-cohort review of extremes, then lock ratings after sign-off.
  14. Model increments against locked ratings and review anomalies for market position and parity.
  15. Brief managers with talking points and the rationale for each of their people.
  16. Open and properly staff the appeal window.
  17. Run post-cycle measurement and hold a facilitator retrospective within two weeks.

Frequently Asked Questions

What is a calibration meeting in performance management?

A calibration meeting is a structured session where managers at the same level review their proposed ratings together, before ratings are finalised. Each manager presents evidence, particularly for the top and bottom of the scale, and the group tests whether the ratings are consistent with the same standard across every team. The output is a set of final ratings with written rationale, so a rating means the same thing regardless of who the employee reports to.

Is performance calibration the same as a forced bell curve?

No. A forced bell curve requires a fixed percentage in each band, so a manager must move someone regardless of evidence. Calibration is evidence-led, and any distribution guideline is a reference that prompts discussion, not a target. The test is simple: under a quota a manager must move someone, under calibration a manager must explain. If your sessions end with someone downgraded purely to make the percentages work, you are running a forced curve whatever you call it.

How long should a calibration session take?

About three hours for 50 employees across 8 managers. That averages three to four minutes per employee, but the average is misleading in a useful way — most middle-band ratings need no discussion, so the time concentrates on the extremes and the flagged exceptions. Remote sessions work better split into two 90-minute blocks on consecutive days.

Who should attend a calibration meeting?

The managers whose people are being rated, a facilitator with no employees in the cohort, the skip-level leader to arbitrate deadlocks, and a scribe. Employees never attend, and Finance should not be present during the rating pass. Managers should stay for the whole session rather than dropping in for their own people, because the comparison only works if everyone hears every case.

Should the increment budget be discussed during calibration?

No. Ratings should be finalised and locked on evidence alone, before compensation modelling starts. Once managers know the rupee value of a band, the conversation stops being about performance. If the calibrated ratings exceed the budget, the honest responses are to adjust the increment matrix, sharpen the differentiation between bands, or return to leadership for more funding — not to quietly revise ratings.

How do we handle new joiners or employees on long leave?

Publish the rule before the cycle and apply it without exception. A common approach is that under three months of service means not rated and no increment, while anyone above that is rated against the period actually worked, at the standard expected for that period. For long leave, including maternity leave, rate the worked period at the normal standard with no penalty for absence. Where statutory entitlements or termination decisions are involved, take specific advice from employment counsel.

What if a manager disagrees with the calibrated rating?

Let the disagreement stand rather than manufacturing consensus. Record the final rating, note the dissent with the manager's name, and let the skip-level leader make the final call, with a clear escalation route before ratings lock. What must not happen is a rating changed by email after the session, or a manager telling their employee that they personally argued for more — that disowns the decision and undermines the process for everyone.

How do we know whether our calibration process is working?

Track a small set of measures across cycles rather than one number: distribution drift, the share of ratings changed in session, the spread between the most and least generous manager in a cohort, appeal rates by manager, voluntary attrition by rating band after the cycle, and a short pulse on whether ratings were explained clearly. If the manager-to-manager spread narrows and appeal rates fall over three cycles, calibration is doing its job.

Making Calibration Routine

Performance calibration is not a compliance exercise. It is the mechanism that turns individual opinions into a company-wide standard — one where a top rating is scarce enough to mean something, a middle rating is honest rather than evasive, and every employee can be given a specific reason for their outcome.

The ingredients are unglamorous: a scale with real behavioural anchors, evidence captured through the year rather than remembered in March, a prepared pre-read, a facilitator willing to ask "what did they actually do?", and a written record that outlives the manager. None of it is complicated; all of it takes discipline in the weeks when everyone is busy.

Start with one cycle. Take your largest cohort, run one proper session with a real agenda and a real facilitator, and compare the distribution before and after. The gap between the two is the size of the problem you have been carrying.

If your appraisal cycle currently lives in spreadsheets that only one person understands, CozyHR can take the mechanical load off it — goals and evidence captured through the year, configurable rating scales with anchors, pre-read packs generated automatically, a live calibration view during the session, a locked audit trail, proration handled from your leave and joining data, and increment modelling against final ratings. Take a look at how CozyHR could handle your next appraisal cycle, and spend the time you save on the conversation rather than the spreadsheet.