Why do engagement scores drop in a hard year — and what should you measure instead?

Whatever engagement survey you run this year, the numbers will come back far below reality. And it is not because your work is bad.
Conditions outside your control change your people’s mood. They answer based on how the last few days felt—not the reality of the whole year. So you get low scores. A whole year of your work goes unseen. And next year’s plan is built on wrong information.
The short answer: a rating scale asks people to judge a whole year, and the answer leans on how the last few days felt. In a hard year that pulls scores down across the board, whatever you did. A forced choice asks people to choose between two things they value, and a general dip in mood cannot pull that choice down in the same way.
What does a score below reality cost you?
Three things, and they arrive together.
- Numbers that sit below reality. The score reflects the weeks around the survey as much as the year being judged.
- A whole year of your work left unseen. Programmes that changed people's working lives disappear under a mood that has nothing to do with them.
- Next year’s plan built on wrong information. Budgets and priorities move toward the wrong problem, because the data said so.
None of this means the survey was run badly. It means the instrument — a rating scale — is measuring two things at once: how people judge their work, and how they feel on the day they answer.
What does a rating scale actually measure in a volatile year?
More of the present moment than most survey reports admit. Research on mood and judgement has shown for decades that when people are asked a broad question about a long period, how they feel at the moment of answering leaks into the answer.
- Job satisfaction rises and falls with mood within the same person. In an experience-sampling study, 27 employees rated their job satisfaction four times a day for four weeks. About 36% of the variation in their ratings happened within the same person, and mood explained 29% of that within-person variation (Ilies and Judge, 2002; a small, correlational study).
- Mood over recent weeks predicts the overall verdict. In a study of 24 managers who kept mood diaries for 16 working days, their average pleasant mood over the period predicted their overall job satisfaction, alongside what they believed about the job itself (Weiss, Nicholas and Daus, 1999).
- Feelings stand in for judgement. In a classic experiment, people phoned on sunny days rated their lives as more satisfying than people phoned on rainy days — and the rainy-day dip disappeared when the caller first asked about the weather, because respondents then set their mood aside (Schwarz and Clore, 1983; 84 participants).
The weather example needs a caveat. A study of more than a million Americans found that day-to-day weather had only very small effects on life-satisfaction ratings (Lucas and Lawless, 2013). The lesson is not that the weather decides your score — it is that a judgement about a whole year leans on how the respondent feels when the question arrives.
In a hard year, that feeling is heavy. Gallup's latest country data shows that 57% of employees in Egypt reported experiencing a lot of stress the previous day, against 48% across the Middle East and North Africa (three-year averages to 2025; Gallup's country figures carry margins of up to about seven points). That is the backdrop before this year began. A rating scale fielded into that backdrop picks up the backdrop.
That is why at GoldinKollar, we developed the GK Engagement Survey.
What does a forced choice reveal that a rating hides?
Instead of asking how satisfied they are, we ask them to choose.
Does your environment guarantee you stability, or give you room to experiment?
The answer has to be realistic, because you cannot do a thing and its opposite.
More importantly, choosing brings us closer to them. It reveals what actually moves them. Their true priorities. And the exact gap between what they want and what you deliver.
Survey research supports the format.
- Choosing can reduce rating biases. Multidimensional forced-choice questionnaires can significantly reduce many of the response biases associated with rating scales (Brown and Maydeu-Olivares, 2011).
- Comparing people takes the right scoring. One person's forced-choice scores can be compared with another's only when the questionnaire is designed and scored for it (Brown and Maydeu-Olivares, 2013).
- Forced-choice manager ratings held up better. When the same line managers rated their staff's competencies in both formats, a personality predictor correlated .38 with the forced-choice ratings and .25 with the rating-scale ratings (Bartram, 2007). That suggests the forced-choice ratings carried less leniency and halo. The author worked for a publisher of forced-choice tests.
- Ranking reduced language bias. In a study of about 3,700 undergraduate and MBA students in 16 countries, ranking rather than rating reduced both response bias and language bias (Harzing and colleagues, 2009). Language bias means answers shifting with the language of the questionnaire. The participants were students, not employees, but the finding is worth weighing for any workforce answering in Arabic and English.
| What you ask | What the answer reflects | Why it matters in a hard year |
|---|---|---|
| "How satisfied are you with your growth opportunities?" (1–5) | A judgement of the year, mixed with how the respondent feels this week | Mood pulls every item down together, so the whole profile sinks |
| "If you had to choose, which of these two matters more to you at work right now?" (for example, security and stability, or learning and variety) | A priority between two things people value | A choice between two goods is less exposed to a general dip in mood, because the mood touches both options |
| "Which of these two does your company actually give you more of today?" | Delivery, as the employee sees it | The gap between want and get points to the action, not just the problem |
What should next year's survey measure?
The things you will act on, in one instrument short enough to finish on a phone.
And this is one survey, not five. Engagement, culture, EVP, experience, and retention. One fieldwork, one invoice. You stop buying the same people’s attention three times a year.
It is built on scientific foundations and published research you can verify yourself. And the questions are written specifically for your culture, in your people’s everyday language. Not translated.
Every team gets three numbers, one clear action, an owner, and a date. Ten minutes on mobile. Arabic or English. Typed or spoken.
In practice the timetable is fixed: a sponsor report within 14 days of the survey closing, the headline to everyone within 30, and one action per group within 45 — and no result is ever reported for a group of fewer than seven people. The details are on the GK Engagement Survey page.
So your annual engagement numbers are your real numbers. GK Engagement Survey from GoldinKollar. Built with you—not sent to you.
What should you tell the leadership team about this year's number?
Tell them what the number is, before they decide what it means.
- Name the conditions. List what changed outside the organisation during fieldwork — prices, workload, a restructuring, a hard season. A score is read differently once the backdrop is on the table.
- Show the spread, not just the average. Which groups moved, which held, and where the gaps are largest.
- Compare with care. Compare with last year only if the questions, the timing, the population and the method were the same.
- Add behaviour. Put the score beside resignations, absence and internal moves for the same period. If people say they are less engaged but are not leaving, that is information.
- Promise one action per team — not a better score. Chasing next year's number teaches managers to manage the survey instead of the work.
If this year's number has already landed, the most useful next step is a second reading that mood cannot move as easily. That is what the forced choice is for.
Frequently asked questions
Why did our engagement score drop this year?
Often because the rating scale measured the moment as well as the year. If conditions outside the organisation were hard during fieldwork, scores tend to fall across the board. Check whether the drop is broad, which points to the backdrop, or concentrated in some teams, which points to something local.
Do engagement surveys still work in a hard year?
Yes, if you read the number against its backdrop and act on what people choose. A survey that ends in one action per team, with an owner and a date, is worth running in any year.
Are rating scales unreliable?
They measure what they ask: how people judge something at the moment they answer. The problem is reading that moment as a verdict on the whole year. Ratings work best alongside questions that ask people to choose.
What is a forced-choice question in an engagement survey?
A question that offers two things people value and asks which matters more, such as stability or room to experiment. Because both options are good, the answer reflects priorities more than mood.
Can we compare this year's engagement score with last year's?
Only if the questions, the timing, the population and the method stayed the same. Even then, report what changed outside the organisation before you read the difference.
How long does the GK Engagement Survey take?
About ten minutes on a phone, with no login, in Arabic or English, typed or spoken. No group under seven people is ever reported.
Sources
- Ilies, R. and Judge, T.A. (2002), "Understanding the dynamic relationships among personality, mood, and job satisfaction", Organizational Behavior and Human Decision Processes 89(2). sciencedirect.com
- Weiss, H.M., Nicholas, J.P. and Daus, C.S. (1999), "An examination of the joint effects of affective experiences and job beliefs on job satisfaction", Organizational Behavior and Human Decision Processes 78(1). sciencedirect.com
- Schwarz, N. and Clore, G.L. (1983), "Mood, misattribution, and judgments of well-being", Journal of Personality and Social Psychology 45(3). doi.org
- Lucas, R.E. and Lawless, N.M. (2013), "Does life seem better on a sunny day?", Journal of Personality and Social Psychology 104(5). doi.org
- Gallup, State of the Global Workplace 2026: Egypt country data (three-year average, 2023–2025). gallup.com
- Gallup, State of the Global Workplace 2026: regional data for the Middle East and North Africa (three-year average to 2025). gallup.com
- Brown, A. and Maydeu-Olivares, A. (2011), "Item response modeling of forced-choice questionnaires", Educational and Psychological Measurement 71(3). doi.org
- Brown, A. and Maydeu-Olivares, A. (2013), "How IRT can solve problems of ipsative data in forced-choice questionnaires", Psychological Methods 18(1). doi.org
- Bartram, D. (2007), "Increasing validity with forced-choice criterion measurement formats", International Journal of Selection and Assessment 15(3). doi.org
- Harzing, A.W. and colleagues (2009), "Rating versus ranking: what is the best way to reduce response and language bias in cross-national research?", International Business Review 18(4). doi.org