happily.ai research All research
Feedback

The Invisible Struggle: Self-Assessment Barely Matches Peer Feedback

Every August, thousands of employees answer an annual check-in about whether they love their job and whether they feel able to perform it well. Their answers move. The performance feedback they receive from colleagues does not. Agreement between the two sits at chance level.

κ = 0.02
agreement between self-rated and colleague-rated performance, chance level
64%
of self-reported strugglers rated Exceeding Expectations or better by colleagues
16,379
colleague feedback items linked to 2,703 annual self-assessments

An annual check-in question asks employees a deceptively simple thing: "How would you describe your feelings towards your job and your ability to perform it well?" The five answers cross two dimensions: how you feel about the work (love, neutral, dislike), and how you are coping with it (performing well, struggling).

Tracking the 885 people who answered in more than one annual wave showed a distinctive pattern. Attachment to the job is durable: 88% of people who love their job still love it a year later, and direct jumps from love to dislike almost never happen (4 in 948 transitions). What moves is the other half of the answer. Each year, roughly a quarter of the people who say "Love & Perform Well" slip to "Love & Struggle," and a slightly larger share of strugglers recover. The feeling holds; the sense of being able to perform flips.

That raises the question this study answers: when a person's self-assessment flips, did their performance actually change? These workplaces offer an unusually direct way to check. Colleagues send each other structured feedback carrying a four-level assessment, from "Needs Improvement" to "Truly Outstanding," and managers write goal reviews on a four-point scale. We linked every self-report to every assessment the person received within 180 days on either side, 16,379 colleague feedback items and 1,553 manager reviews in all, and measured the match.

There barely is one. The half of the answer that moves from year to year is precisely the half no one else can see.

κ = 0.02 between self-rated coping and colleague-rated performance. Knowing what a person said about themselves tells you almost nothing about how colleagues rated them.
Why this matters

If self-reported struggle showed up in performance data, you could wait for reviews to catch it. It does not. A "struggling" answer is real information about confidence and strain, but it lives in a channel performance systems never see. The reverse mismatch is just as common: a third of self-assured performers carry external ratings below their company's norm. Self-assessment and external assessment are two different instruments, and organizations that treat one as a proxy for the other will misread both.

Methodology
Sample
4,834 answers from 3,698 employees across 108 companies; annual August waves, 2024–2026.
Instrument
Five-option check-in crossing feeling (love / neutral / dislike) with coping (perform well / struggle).
External signals
16,379 colleague feedback items (four-level assessment, self-feedback excluded) on 2,703 answers; 1,553 manager goal reviews (1–4 scale) on 896 answers. Window: ±180 days around each answer.
Longitudinal panel
885 repeat responders; 1,136 consecutive answer pairs; 358 pairs with feedback in both years.
Tests
Per-answer aggregation (each answer weighs equally), company-demeaned z-scores, Cohen's kappa on binarized ratings, within-person year-over-year deltas, and a wave comparison restricted to companies present in all three waves.
Concern threshold
External "concern" = mean received rating below "Exceeding Expectations," the norm in a feedback culture where 73% of items are Exceeding or better.

The annual dance: confidence moves, attachment holds

First, how people answer at all. Nearly half of the 4,834 answers are "Love & Perform Well." Almost one in three is "Love & Struggle": attached to the work, doubting their ability to do it well. Outright dislike is rare, 2% of answers in total, which is why this is mostly a story about the two love states.

One in three says they love the job but struggle to perform it All answers to the annual check-in question, August waves 2024–2026. n=4,834 answers from 3,698 employees, 108 organizations. Love & Perform Well 48.1% Love & Struggle 32.3% Neutral Feeling 17.6% Dislike & Perform Well 1.5% Dislike & Struggle 0.5% Source: Happily People Science, August 2026. 108 organizations.
Figure 1 How the 4,834 answers distribute across the five states. The two dislike states together are 2% of answers. Interactive version: chart.

Year over year, 60% of consecutive answers repeat exactly, and stability rises with the state: "Love & Perform Well" holds at 67%, "Love & Struggle" at 54%, Neutral at 51%. No one who answered again stayed in a dislike state two waves running, though dislikers also drop out of the panel at higher rates (more on that under Limitations).

When an answer does change, the change has a shape. Half of all changes are the performance half flipping while love holds. Direct flips of the feeling half are nearly nonexistent.

When the answer changes, the performance half is what moves What changed when a person's answer changed between annual waves. 450 changes among 1,136 consecutive answer pairs, 2024–2026. Performance half flips Love & Perform ↔ Love & Struggle 52% Moves to or from Neutral 45% Jumps between Love and Dislike 3% Attachment is the stable half: 88% of answers given from a love state still love the job a year later. What flips is “able to perform it well.” Source: Happily People Science, August 2026. 108 organizations.
Figure 2 Of 450 within-person answer changes, 234 flip between "Love & Perform Well" and "Love & Struggle." Interactive version: chart.

The traffic between the two love states is also asymmetric: 144 slips into struggle against 90 recoveries, because the perform-well pool is twice as large. The rates themselves replicate almost exactly across both wave pairs (a 23% slip rate and a 29% recovery rate in each), so this is a stable annual rhythm, not a one-off shock. The question is whether the rhythm reflects anything others can see.

Colleagues rate strugglers and performers alike

Each colleague feedback item carries the giver's assessment of the receiver: Needs Improvement, Meets Expectations, Exceeding Expectations, or Truly Outstanding. If self-reported struggle tracked observed performance, people answering "Love & Struggle" should receive weaker assessments than people answering "Love & Perform Well." They barely do.

Colleagues can't tell self-reported strugglers from performers Share of colleague feedback rated “Truly Outstanding,” received ±180 days around the answer. n=2,703 answers, 16,379 feedback items. Whiskers: 95% CI. Love & Perform Well 36.0% Love & Struggle 33.8% Neutral Feeling 30.2% Dislike & Struggle 29.5% Dislike & Perform Well 24.8% 0% 25% 50% Source: Happily People Science, August 2026. 108 organizations.
Figure 3 Top-box share of colleague assessments by self-report state. The perform-vs-struggle gap within the love states is 2.2 points and the confidence intervals overlap. Interactive version: chart.
Colleague feedback received, by self-report state (per-answer means)
Self-report stateAnswers w/ feedbackMean rating (1–4)Truly OutstandingNeeds Improvement
Love & Perform Well1,3393.0736.0%1.7%
Love & Struggle8823.0333.8%1.6%
Neutral Feeling4283.0130.2%1.6%
Dislike & Perform Well412.8524.8%0.7%
Dislike & Struggle133.0329.5%0.0%
Perform vs struggle gap0.045 [−0.004, 0.094]2.2 pp0.1 pp

The gap between self-assured performers and self-doubting strugglers is 0.045 points on a four-point scale, with a confidence interval that lets us rule out any true gap larger than about 0.1 points. Colleagues flag "Needs Improvement" at the same 1.6% rate regardless of what the person says about their own coping. Z-scoring ratings within each company removes any nice-rating-culture confound; the gap that remains is 0.06 standard deviations, borderline at conventional significance and negligible in size.

The little gradient that does exist runs along the other axis. Ratings soften slightly from love (36.0%, 33.8%) to neutral (30.2%) to dislike (24.8% for "Dislike & Perform Well," though only 41 answers). Colleagues seem to pick up traces of the feeling half of the answer, the half that shows on your face, and none of the coping half.

Agreement is no better than chance

Collapsing both instruments to a binary puts a number on the disconnect. Self-report: perform well vs struggle. External: mean received rating at or above "Exceeding Expectations" vs below it. Among 1,855 answers with at least two feedback items in the window, the two instruments agree 54.9% of the time, against 53.9% expected if they were statistically independent. Cohen's kappa: 0.02.

Self-rating and colleague rating agree no better than chance Answers with ≥2 colleague feedback items within ±180 days; neutral answers excluded. n=1,855. “Below” = mean received rating under “Exceeding Expectations,” the norm in this feedback culture. Colleagues: Exceeding+ Colleagues: below Says: performing well 41.5% 769 answers 21.6% 400 answers Says: struggling 23.5% 436 answers 13.5% 250 answers κ = 0.02 observed agreement 54.9% vs 53.9% expected by chance Source: Happily People Science, August 2026. 108 organizations.
Figure 4 The mismatch runs both ways: 436 answers say struggling while colleagues say exceeding, and 400 say performing while colleagues rate below the norm. Interactive version: chart.

Read the off-diagonal cells. 64% of self-reported strugglers (436 of 686) carry an average colleague assessment of Exceeding Expectations or better; their struggle is invisible. And 34% of self-reported performers (400 of 1,169) sit below their feedback culture's norm; their confidence is not corroborated either. This is consistent with five decades of self-other rating research: meta-analyses put the self-peer correlation for job performance around r = 0.36 (Harris & Schaubroeck, 1988), and self-evaluations of ability generally around r = 0.29 (Zell & Krizan, 2014). Our kappa is lower still, likely because this self-report captures felt coping rather than a considered performance estimate.

The year the answer slips, ratings stand still

The strongest test uses each person as their own control. Take everyone who answered in consecutive waves, and compare the colleague ratings they received around each answer. If slipping from "perform well" to "struggle" reflected a real performance dip, the received ratings should fall with it.

The year self-assessment slips, colleague ratings don't move Mean colleague rating received (scale 1–4), ±180 days around each annual answer. 358 within-person answer pairs with feedback in both years. Year 1 Year 2 Slipped into struggle · 3.18 3.17 Stayed performing · 3.07 3.11 Stayed struggling · 3.03 3.10 Recovered · 2.99 3.08 All four groups sit within 0.19 rating points of each other, both years. Source: Happily People Science, August 2026. 108 organizations.
Figure 5 Mean received rating in year 1 and year 2 for the four transition groups. No line moves more than 0.10, and the slippers' line is the flattest. Interactive version: chart.
Within-person change in received colleague rating (358 consecutive-answer pairs)
TransitionPairsRating, yr 1Rating, yr 2Δ (SE)
Stayed performing1833.073.11+0.04 (0.03)
Slipped into struggle703.183.17−0.02 (0.05)
Recovered442.993.08+0.10 (0.06)
Stayed struggling613.033.10+0.07 (0.06)

Two details deserve attention. The people who slipped into struggle were the highest-rated group before the slip (3.18), and their ratings did not move afterward. And the people who stayed in struggle two years running saw their ratings drift upward while they continued to report struggling. Whatever changed for these people between waves, their colleagues' assessments did not register it.

The same holds in aggregate. Across the companies that ran all three waves, the share reporting struggle rose 7.2 points between 2024 and 2025, from 26.4% to 33.6%. The mean colleague rating received by the people answering did not decline over that period: it went from 3.07 to 3.12, a difference within noise and pointing the opposite way. A visible slice of the workforce started reporting they were struggling, and the assessment record for those years looks the same as before.

The one modest leak: manager goal reviews

Manager reviews are the exception, and a faint one. Managers rate goal attainment on a four-point scale, and here self-reported strugglers do score lower: 81.9% are rated as having met their goals, against 87.4% of self-assured performers. Managers mark strugglers below goal 1.4 times as often (18.1% vs 12.6%).

Only managers' goal reviews register the struggle, barely Share of manager reviews rating goal attainment “met” or better (3+ of 4), ±180 days around the answer. n=867 answers with ≥1 manager review. Dislike states omitted (n≤20). Neutral Feeling 87.9% Love & Perform Well 87.4% Love & Struggle 81.9% Strugglers marked below goal 1.4x as often (18.1% vs 12.6%). Borderline (p ≈ 0.05); no gap in culture ratings. Source: Happily People Science, August 2026. 108 organizations.
Figure 6 Share of manager reviews rating goal attainment "met" or better. The 5.6-point gap is the only external signal that clears noise, and it is borderline. Interactive version: chart.
Read with care

The 5.6-point gap sits at p ≈ 0.05 after several comparisons, and manager culture ratings show no gap at all (3.32 vs 3.27). Treat this as suggestive: goal shortfalls may be the one place self-perceived struggle brushes against something observable, but the evidence is thin.

The pattern fits a simple account. What a person calls "struggling" is mostly an internal state: strain, self-doubt, the feeling of running harder to stand still. Colleagues rate output and collaboration, which apparently hold up. Managers track goals, where sustained internal strain might eventually surface. The instruments disagree because they measure different things, a gap psychologists have documented as the illusion of transparency: people systematically overestimate how visible their internal states are to others (Gilovich, Savitsky & Medvec, 1998).

What this means

The practical conclusion is not that one instrument is right and the other wrong. The two carry non-overlapping information, and each fails in the direction the other covers.

Decisions this study supports
If you currently…This study suggests
Wait for performance reviews to reveal who is strugglingThey will not. 64% of self-reported strugglers are rated Exceeding or better. Ask people directly and treat the answer as the primary signal.
Discount "I'm struggling" answers from strong performersStruggle concentrates among the well-rated: the group that slipped into struggle had the highest colleague ratings of any group. High ratings do not rule out strain.
Treat a confident self-report as evidence performance is fine34% of self-assured performers carry ratings below their company norm. Confidence is not corroboration.
Respond to a "struggling" answer with performance managementThe data shows no performance deficit to manage. Respond with support and workload conversation, not scrutiny.
Run engagement and performance systems separatelyKeep both, and route them to the same conversation: goal reviews are the one external channel where struggle faintly registers.

Limitations

  • Colleague feedback is heavily positive-skewed (73% of items are Exceeding or better) and often requested by the receiver, which selects toward favorable raters. A blunter assessment channel might detect more.
  • The "concern" threshold (mean rating below Exceeding) is a norm-relative choice; results are similar with top-box definitions, but any binarization loses information.
  • People answering "Dislike" are rare (2% of answers) and drop out of the panel at higher rates (31–33% next-wave continuation vs 50% for the happiest state), so dislike-state estimates are indicative only, and some of the apparent recovery from dislike is exit.
  • Repeat responders skew toward engaged panel stayers; the transition analysis describes people who kept answering.
  • Manager reviews cover 19% of answers and skew to companies running structured reviews.
  • Ratings cluster within raters and companies; the company-demeaned check controls rating norms, not who chooses to rate whom.
  • The 2026 wave was still in the field when this analysis ran, so any level reported for that year is provisional. The aggregate comparison above uses only the two complete waves, 2024 and 2025.

References

  1. Harris, M. M., & Schaubroeck, J. (1988). A meta-analysis of self-supervisor, self-peer, and peer-supervisor ratings. Personnel Psychology, 41(1).
  2. Zell, E., & Krizan, Z. (2014). Do people have insight into their abilities? A metasynthesis. Perspectives on Psychological Science, 9(2).
  3. Gilovich, T., Savitsky, K., & Medvec, V. H. (1998). The illusion of transparency: Biased assessments of others' ability to read one's emotional states. Journal of Personality and Social Psychology, 75(2).
  4. Clance, P. R., & Imes, S. A. (1978). The imposter phenomenon in high achieving women: Dynamics and therapeutic intervention. Psychotherapy: Theory, Research & Practice, 15(3).
  5. Happily Research (2026). Job feelings and ability-to-perform study: annual self-reports (4,834 answers, 3,698 employees, 108 companies, 2024–2026) linked to 16,379 colleague assessments and 1,553 manager reviews.
See the signal ratings miss

Performance systems catch shortfalls. They do not catch strain. Happily.ai reads daily check-ins, feedback, and recognition together, so a quiet confidence dip gets support long before any review would notice.

Get in touch
Free pilot for qualifying teams