The Manager Gap: why two teams in the same company diverge
Nearly half the variation in team engagement sits between managers inside the same company, so the gap between two managers in one building is typically wider than the gap between two companies. Of the five responsiveness behaviors we measured, exactly one survives controlling for company. It is not speed, and the most plausible reason it works is trust.
Every manager development budget rests on two assumptions that almost nobody tests. The first is that the manager is the thing that matters, rather than the company having already set the outcome before anyone opens a laptop. The second is that the behaviors being trained are the behaviors that move it. Most manager training is bought on faith in both.
This study tests both on one panel: 424 managers across 24 companies, behavior measured from July to December 2025 and team engagement measured from January to June 2026. The windows do not overlap. That is the study's main defence against reading the arrow backwards.
On the first assumption, the verdict is split. Company explains 54.0% of the variation in team engagement. The individual manager, working inside that same company, explains 46.0%. Culture is ahead. But the manager's share is large and comparable, and it has a concrete meaning: the gap between two managers in one building is typically about as consequential as the gap between two companies.
On the second, the answer is narrower than we expected and more useful for being narrow. Of the five responsiveness behaviors we could measure, exactly one survives controlling for company: consistency, the share of weeks in which a manager replies at least once. Speed is a null. Length is a null. Measured skill volume is a null. Cadence beats craft.
Cadence also predicts trust, which is the most plausible account of why it works. A team reads regular presence as evidence that the manager is dependable.
Manager development and culture work compete for the same budget, and the case for each is usually made with an assertion rather than a number. This gives you the split: both matter at broadly comparable magnitudes, with culture slightly ahead.
It also tells you which manager behavior to buy. Fund one thing this year and make it a weekly response habit. Every other responsiveness behavior we tested either failed to survive or never registered.
boss field. Never from licence tier or job title.The conflict this study resolves
One thing has to be cleared up before any of that lands. Two earlier studies in our own corpus read, on their face, as arguments against investing in managers at all.
The response time study found that the correlation between manager responsiveness and team outcomes shrinks from r = 0.228 raw to r = 0.061 once company is controlled. The Power Skills study found no significant skill predictor of team engagement at all.
Neither result says managers barely matter, because neither measured that. Adjusting for company removes the level of team engagement, which culture genuinely sets. It leaves the spread between managers working inside the same culture untouched, and the spread is what a development program actually acts on. This study measures the spread directly. Both earlier nulls survive intact here, which is the point: the picture was always consistent, it was being read against the wrong quantity.
Finding 1: 46% of the variation lives between managers in the same company
Team engagement varies a great deal. The first question is simply where that variation sits. A random-intercept model answers it by giving every company its own baseline, then asking how much of what remains sits between the managers who share that baseline. Fitted with no predictors at all, it splits the variance in two.
| Component | Variance | Share |
|---|---|---|
| Between companies | 0.023649 | 54.0% |
| Between managers, same company | 0.020165 | 46.0% |
The gap between two managers in the same building is typically wider than the gap between two companies. The median company spans 0.386 between its highest and lowest scoring manager, against an overall standard deviation of 0.199. Figure 2 shows every company in the panel on that measure, one line each.
One large account is not driving this. The largest customer contributes 100 of the 424 managers, 23.6% of the panel; refit the decomposition without that company and the manager share moves from 46.0% to 43.5%. Two and a half percentage points, staying inside the same interpretive band.
It does not show that managers matter more than culture. On this measure they do not. Company is ahead, 54.0% to 46.0%, and removing the largest customer widens that lead rather than closing it. Any claim that the manager is the single biggest driver of engagement is unsupported by this data, including if we were the ones making it.
It is also not causal. A variance partition describes where variation sits, not what produced it.
Finding 2: cadence is the only behavior that survives
So manager identity matters. Which part of what a manager does? Each behavior was measured in the first window and tested against second-window team engagement, one behavior at a time. Every measure is first turned into a z-score inside its own company, which means each manager is graded against their own colleagues rather than against the platform. That is the step that makes "company held constant" literal rather than rhetorical.
| Behavior | n | Beta | p | R² |
|---|---|---|---|---|
| Consistency, share of weeks with a reply | 424 | +0.267 | 2.2 × 10⁻⁸ | 0.072 |
| Reply quality, AI-scored index | 290 | +0.156 | 0.006 | 0.026 |
| Reply rate | 424 | +0.091 | 0.062 | 0.008 |
| Reply depth, log characters | 285 | +0.024 | 0.69 | 0.001 |
| Reply latency, log hours | 283 | −0.021 | 0.73 | 0.000 |
Only consistency is both significant and stable. Fit all five behaviors together, on the 209 managers for whom all five are defined, and consistency strengthens to +0.496 (p = 5.5 × 10⁻⁸) while reply quality collapses to +0.040 (p = 0.58). That model explains 15.1% of within-company variance. Reply quality matters until cadence is in the room, and then it does not.
Share of weeks with at least one manager reply, over a rolling 26-week window. It is the one behavior here that survives every specification, it has a clean dose-response, and it is not sensitive to where the participation floor was set.
The nulls are nulls, not weak trends
These are not small effects waiting for a bigger sample. Reply latency comes in at beta −0.021, p = 0.73; reply depth at +0.024, p = 0.69. Each explains essentially none of the within-company variance, 0.000 and 0.001 respectively. The estimates are indistinguishable from zero and their signs carry no information.
Speed in particular replicates as a non-finding. The earlier response time study reached the same conclusion on a different sample by a different route. Two independent analyses now agree that how fast a manager replies carries no signal about their team's engagement six months later, and neither does how much they write.
The Power Skills null, stated precisely
Measured skill fares no better. Six Power Skills, entered as volume, explain 4.7% of within-company variance in team engagement, with no coefficient reaching p = 0.228. That replicates the earlier Power Skills null on a different design.
One caveat is essential here, and we are not burying it. Happily measures each Power Skill on two dimensions: a score, which is how much evidence of that skill has accumulated from written feedback, and a rating, which is the AI's evaluation of how well the skill is expressed. Only the score is available in BigQuery. So this null says that how much a manager expresses a skill in writing does not predict their team's engagement. It does not test whether how well they do it predicts anything. That question stays open, and it is the more interesting one.
One coefficient we are not reporting as a finding
Reply rate is +0.091 on its own and −0.355 in the multivariate model. Read naively, the second number says that replying more often lowers team engagement. It does not, and this study does not make that claim.
Reply rate and consistency correlate at r = 0.727. They are near-substitutes for the same underlying behavior: a manager who answers a high share of messages is usually a manager who answers in most weeks. Put both in one model and consistency absorbs the shared signal, leaving reply rate with a residual dominated by managers who answer a lot of messages in a few concentrated bursts. Statisticians call the resulting sign flip suppression: when two near-duplicate measures compete, the leftover part of one can take on a sign that has nothing to do with the behavior its name suggests. Consistency is the better-behaved of the two, it survives in both specifications, and it is the one we report.
Finding 3: the dose response, including where it breaks
A coefficient tells you a relationship exists. It does not tell you what the relationship looks like, and the shape is what a practitioner actually needs. Sorting managers into consistency bins and reading off mean team engagement gives the shape, including one stretch of it that misbehaves.
| Consistency | n | Team engagement |
|---|---|---|
| 0.000 | 180 | 0.278 |
| 0.077 | 49 | 0.228 |
| 0.154 | 40 | 0.281 |
| 0.269 | 35 | 0.319 |
| 0.385 | 36 | 0.317 |
| 0.615 | 41 | 0.360 |
| 0.923 | 43 | 0.488 |
From the third bin upward the line climbs steadily, and the top bin stands well clear of the rest. Managers who reply in roughly nine weeks out of ten have teams scoring 0.488, against 0.278 for managers who never reply. That is a gap of 0.210 on a 0 to 1 scale, slightly more than one full standard deviation.
Managers with the lowest non-zero cadence, 0.077, score 0.228, which is below the 0.278 of managers who never reply at all. We do not have an explanation we trust. It could be that sporadic, token responsiveness reads worse to a team than visible absence, or it could be noise in a 49-manager bin. We report it rather than smoothing it away, and no recommendation here rests on it.
Nor is the shape an artifact of where the participation floor was set. Across floor thresholds of 3, 5 and 10 routed messages, the reply-rate coefficient moves from 0.0855 to 0.0997, a swing of about 15% of its own magnitude and well inside the tolerance fixed before the analysis ran.
Finding 4: consistency buys trust, which is probably why it works
Findings 2 and 3 establish what separates managers. They do not explain why. The most plausible account is that consistency is not a dose of effort at all but a signal. A team watching a manager across six months is answering one question about that person: can I count on them to show up? Cadence is the evidence they have to go on.
That account is testable. If consistency works by building trust, then consistency should predict trust and not just engagement. It does.
Trust here is the composite defined by our trust networks study, reused unchanged so the two are directly comparable: 0.6 times the standardized count of colleagues who asked this manager to give them feedback, plus 0.4 times the standardized count of colleagues who exchanged recognition with the manager in both directions, the whole thing standardized within company. Being chosen as a feedback giver is a deliberate act of trust by the asker, which is why it carries the larger weight. Requiring recognition to run both ways filters out one-sided visibility. Neither component asks anyone whether they trust their manager. Both count what people actually did.
The trust sample is 308 managers across 19 companies. It is smaller than the 424-manager main panel because a manager needs peer-feedback or recognition activity in the second window to have a trust score at all.
| Model | n | Beta | p | R² |
|---|---|---|---|---|
| Trust composite | 308 | +0.285 | 2.9 × 10⁻⁷ | 0.083 |
| Peer feedback only, fallback | 308 | +0.201 | 3.4 × 10⁻⁴ | 0.041 |
Consistency predicts trust at beta +0.285 (p = 2.9 × 10⁻⁷). The first thing to check is whether that is carried by the sparse half of the composite. Reciprocal recognition is non-zero for only 46% of these managers, while peer-feedback solicitation is non-zero for 96%, and a composite built from one dense component and one thin one can easily turn out to be the thin one in disguise. This one does not: rerun on peer feedback alone and the relationship holds at +0.201 (p = 3.4 × 10⁻⁴). Weaker, and still there.
Trust also carries information about engagement that cadence does not. Add trust to the consistency model and consistency drops from +0.267 to +0.192, a reduction of 28.4%, while trust itself holds at +0.266 (p = 4.8 × 10⁻⁶). Both stay significant.
Consistency is measured in the first window and trust in the second, so that arrow is lagged. Trust and engagement, though, are both measured in the second window. The 28.4% attenuation therefore shows only that the two variables carry overlapping information about the same teams. It is not evidence of a causal chain running from cadence to trust to engagement, and we are not claiming one.
The threshold we went looking for and did not find
If trust is a judgment about dependability, it might be categorical rather than incremental: partial reliability buys nothing, and the payoff arrives only once a manager is nearly always there. That would hand a CHRO an actual number for a scorecard. It is the most useful thing this section could have produced, and we could not produce it.
Testing a single cut point invites the objection that the cut was chosen to fit, so we swept nine, from 20% to 90%, each entered as a step change alongside the straight-line term. Three clear an uncorrected p < 0.05. None survives Bonferroni correction, which sets the bar at 0.05 divided by nine, or 0.0056. The signs flip too, negative at 0.4 and positive at 0.8 and 0.9. That is the signature of fitting noise, not of a real discontinuity.
The test is underpowered in any case. Only 40 of 308 managers sit above 75% consistency, and 16 above 95%. There is very little data in the region where a threshold would live.
What we can say is that consistency predicts trust in a form we cannot distinguish from a straight line at this sample size. A threshold may exist, and a panel this shape would not see it. Answering the question needs a larger cohort of high-consistency managers, not a cleverer model.
The same caution runs backwards, to Figure 4. Reading that dose-response curve as a step is tempting and it does not hold up: under fixed, interpretable consistency bands rather than deciles, the pattern is noisy in both engagement and trust, and the top band holds few managers. The climb is real. The step is not established.
Finding 5: what employees ask for, and what managers actually write
The quantitative half says cadence matters and content does not, at least not in any way that reply length or a model-scored quality index can detect. That is a good reason to read the content directly rather than model it. What follows is drawn from two corpora on opposite sides of the same exchange: what managers write, and what employees say they want more of.
What managers actually write
We coded 1,999 manager replies, drawn from a deduplicated corpus of 26,711, against a taxonomy frozen before the main pass began.
| Reply act | Share | Median length |
|---|---|---|
| Acknowledge only | 49.8% | 16 chars |
| Commit to action | 23.0% | 154.5 chars |
| Thank or praise | 13.9% | |
| Coach or develop | 6.5% | 114 chars |
| Explain context | 4.5% | |
| Redirect offline | 1.0% | |
| Unintelligible | 1.0% | |
| Deflect or defend | 0.4% |
The biggest category is also the shortest one, by a wide margin. The median manager reply in the full corpus is 80 characters. 19.1% of all manager replies are pure emoji, with no text at all; within the acknowledge-only bucket, that rises to 38.4%.
Two independent coders agreed on 78.0% of cases, Cohen's kappa 0.682, which is substantial agreement on the Landis and Koch scale of 0.61 to 0.80. Their single largest disagreement lands directly on this category. They drew the line differently on contingent offers of help, phrasings like "let me know if you need anything": one read them as bare acknowledgment, the other as naming an actionable channel. On the same 150 rows, one coder put acknowledge-only at 52.7% and the other at 37.3%.
So the honest claim is that somewhere between 37% and 53% of manager replies carry no substantive content. Two things survive the disagreement and can be stated flatly: bare acknowledgment is the modal category under either coder, and the 19.1% pure-emoji share is an objective character count rather than a coding judgment.
Defensiveness is nearly absent, and that is probably selection
Deflecting or defending accounts for 0.4% of the corpus, 8 replies out of 1,999. The tempting read is that managers on this platform do not get defensive. That is one explanation and probably not the best one. A manager inclined to push back may simply not reply, leaving no text to code, and the participation floor removes low-activity managers before the corpus is drawn. A measured absence of defensiveness is exactly what defensiveness would look like if it expressed itself as silence. This corpus describes managers who engage, not all managers.
What employees ask for
The other side of the exchange is a census rather than a sample: 3,899 structured competency tags attached to feedback directed at 637 managers across 97 companies. Employees choose these from a fixed vocabulary, so there is no coder here and no reliability risk.
| Competency | Share of asks |
|---|---|
| Initiative-making | 9.2% |
| Detail | 7.6% |
| Leadership | 7.5% |
| Communication | 7.5% |
| Clarity | 7.1% |
| Attitude | 7.1% |
| Empathy | 5.9% |
The gap
Comparing the two lists is harder than it looks, because they are written in different vocabularies. Asks are competencies; replies are speech acts. Forcing them onto one axis would be a fabrication. So the gap below is drawn only across an explicit seven-pair crosswalk of competencies that a feedback reply can plausibly demonstrate, and that crosswalk covers 33.6% of all asks. Competencies with no defensible reply-act counterpart, Timeliness and Creativity and Detail among them, appear on the ask panel and are excluded from the gap rather than forced into a contrived pairing.
| Competency asked for | Ask | Nearest reply act supplied | Supply | Gap |
|---|---|---|---|---|
| Communication | 7.5% | Explain context | 4.5% | +3.0pp |
| Clarity | 7.1% | Explain context | 4.5% | +2.7pp |
| Leadership | 7.5% | Coach or develop | 6.5% | +1.1pp |
The two competencies employees most want more of, communication and clarity, map onto the reply act managers supply least: explaining the reasoning behind a decision, at 4.5% of replies. That is the clearest actionable signal in the qualitative half. It is suggestive rather than a measurement of unmet demand, since a third of asks is not all asks and the crosswalk is a judgment call. But it points at a specific teachable behavior instead of at "communicate better".
What this means
The behavior worth buying turns out to be the least glamorous one on the list. Cadence is cheap to measure, cheap to coach, and it is the only thing here that survives every specification. It is also the only one with a plausible mechanism attached to it. Frame the intervention accordingly: the ask is not "reply more", it is "be someone your team can predict".
What follows is ranked by evidence strength, with the strength stated rather than implied.
| Decision | Evidence | What the data says |
|---|---|---|
| Buy cadence, not polish | Strong | Consistency is the only behavior that survives every specification, at beta +0.267 (p = 2.2 × 10⁻⁸) alone and +0.496 in the full model. Make the intervention a recurring weekly commitment and measure the share of weeks with at least one reply. |
| Sell cadence as dependability, not effort | Strong | Cadence predicts the trust composite at beta +0.285 (p = 2.9 × 10⁻⁷), and the relationship holds on the dense component of that composite alone. The most plausible account of why cadence works is that a team reads regular presence as evidence the manager is dependable. |
| Stop training managers on response speed | Strong null | Latency is beta −0.021, p = 0.73, and two independent studies now agree. An SLA-style "respond within 24 hours" program optimizes a variable that does not move the outcome, and it may crowd out the one that does. |
| Target the explaining gap | Moderate | Communication (7.5%) and clarity (7.1%) are among the most-flagged coachable competencies, and explaining context is the rarest substantive reply act at 4.5%. Specific and teachable, with a visible deficit. |
| Raise the floor on empty replies | Moderate | Between 37% and 53% of replies are bare acknowledgment, and 19.1% of all replies are nothing but an emoji. Setting the standard at "acknowledge plus one concrete thing" is a low-cost change to a very common habit. |
| Do not treat manager development as a substitute for culture work | Strong, and it cuts against us | Company still explains more of the variance than the manager does, 54.0% to 46.0%. Manager development is a strong investment inside a viable culture. It is not a repair for a broken one. |
| Do not rank managers on measured skill volume | Open question | Six skills entered as volume explain 4.7% of within-company variance with a smallest p of 0.228. Only the volume dimension was available, so skill quality is untested rather than disproved. |
| Do not set a consistency target yet | Not established | The obvious next question is what number belongs in a scorecard: 8 weeks in 10, or 9, or every week. We swept nine cut points and none survived correction. Only 40 of 308 managers in the trust sample sit above 75% consistency and 16 above 95%, so the region a target would live in is nearly empty. Set the direction, not a cut line. |
The one-sentence version: company culture explains slightly more of the variation in team engagement than the individual manager does, the manager's share is large, and the only manager behavior that reliably tracks it is showing up week after week. Not speed, not eloquence, not measured skill. Cadence also predicts trust, which is the most plausible reason it works.
Limitations
These belong in the body rather than an appendix, because several of them change how far the findings travel.
- Nothing here is causal. The lagged design, behavior in the first window and outcomes in the second, mitigates reverse causality but does not eliminate it. A team that was engaged in 2025 tends to be engaged in 2026, and an engaged team may be more rewarding to reply to. Read every effect as association.
- The outcome is an activity-weighted index, not a feeling. It averages three components, DPSRRS, OFRS and RPRS. A team that answers more prompts scores higher. Reading it as wellbeing or satisfaction would be a category error, and it is not DEBI.
- Panel size and concentration. 424 managers in 24 companies, with the largest contributing 23.6%. Robustness was checked by exclusion, but 24 companies is a thin base for generalizing across industries or regions.
- Survivorship. Managers must appear in both windows to enter the panel, which selects for stability, and the participation floor removes low-activity managers by design. The question here is what separates good managers from mediocre ones, not present ones from absent ones, so effect sizes are smaller than in studies that compare participants to non-participants.
- Trust is a network measure, not a survey. It counts who asked this manager for feedback and who exchanged recognition with them. That is behavioral and hard to game, but it is not the same as asking employees whether they trust their manager. The trust sample is also smaller than the main panel, 308 managers in 19 companies, because it requires peer-feedback or recognition activity in the second window.
- The consistency-to-trust arrow is lagged; the trust-to-engagement arrow is not. Trust and engagement are both measured in the second window, so the 28.4% attenuation is a statement about shared information, not a mediation chain.
- The quality signal is weak. Reply quality is our own model's score on a 0 to 100 scale, available for 318 of 424 managers, and 52.5% of those non-null values are exactly zero (167 of 318) even after the floor. That pile-up at zero, which statisticians call zero-inflation, is why we treat the measure as unstable. Its univariate effect is real in this sample and it does not survive controlling for consistency.
- The multivariate model describes repliers only. Latency, depth and quality are undefined for a manager who never replied, and 138 of 424 panel managers (32.5%) sent zero replies in the behavior window. Those coefficients say nothing about the never-repliers.
- The ask corpus uses a wider window, 2024-01-02 to 2026-07-21, because the free-text field it originally read collapsed after 2023 and was replaced by a structured competency tag. It is descriptive and never enters the causal analysis.
- Coding rests on one primary coder with a second on a 150-row subset. Kappa is substantial, not near-perfect, and the largest disagreement lands directly on the headline category, which is why that number is a band.
- Two implementation bugs were found and fixed, both of which would have produced wrong published numbers. Routing originally keyed off a column that exists almost only on replied threads, which pinned reply rate at 1.0 for 90% of managers. The source table also carries duplicate rows biased toward replied threads, which inflated reply rate roughly twofold and made the reply corpus 48.7% duplicates. Deduplication alone moved the acknowledge-only share from 39.8% to 49.8%. The quantitative findings survived both corrections unchanged; the qualitative sample had to be redrawn and recoded.
Happily Research (2026). The Manager Gap: Why Two Teams in the Same Company Diverge. happily.ai/research/manager-responsiveness-skills/
References
- Landis, J. R. & Koch, G. G. (1977). The Measurement of Observer Agreement for Categorical Data. Biometrics, 33(1). Source of the 0.61 to 0.80 "substantial agreement" band used to interpret this study's Cohen's kappa of 0.682.
- Happily Research (2026). Response Time Is a Red Herring: Reply Quality Beats Reply Speed. Internal analysis. The company-adjusted correlation of r = 0.061 that this study revisits.
- Happily Research (2026). Power Skills: Why Optimism Drives Peer Recognition. Internal analysis. The earlier skill null replicated here on the volume dimension.
- Happily Research (2026). Trust Networks: Why Most of Your Most-Trusted People Aren't Managers. Internal analysis. Source of the peer-feedback and reciprocal-recognition trust composite reused unchanged in Finding 4.
- Happily Research (2026). Manager Responsiveness and Skills. Internal analysis, 424 managers across 24 companies, behavior July to December 2025 and outcomes January to June 2026.
Happily shows you how consistently every manager shows up for their team, week by week, so you can coach the one behavior that tracks engagement instead of the ones that sound like they should.
Get in touch