← Evidence hub (all four sources, triangulated)

The Consequence Network of Abusive Supervision and Toxic Work Environments: A Systematic Review

Lead+D Lab research synthesis · leadership-anchored · 3-source triangulation · 2026-07-28

Abstract

Objective: To synthesize the consequence network of bad management and toxic work environments, treating abusive, destructive, and despotic supervision as the spine and workplace bullying as a supplementary, environment-level signal. Methods: We integrated four abusive-supervision syntheses (Mackey et al., 2017, k=140; Zhang & Liao, 2015, k=119; Schyns & Schilling, 2013, k=57; Tepper, 2000), the Frazier et al. (2017) psychological-safety meta-analysis (k=136), three bullying meta-analyses (Nielsen & Einarsen, 2012; Verkuil et al., 2015; Nielsen et al., 2016), Hershcovis (2011) construct-reconciliation evidence, and field-wide metaBUS correlations. Screening was on definition and measure, not label. Results: Effects are strongest for relational and justice rupture (interpersonal justice ρ = −.66; leader-member exchange −.54), supervisor-directed deviance (ρ = .53), and emotional exhaustion (.36), and weakest for actual or objective turnover (near zero) and objective health. Psychological safety adds incremental variance beyond leadership. Bullying converges on mental health and sickness absence but its perpetrator is not necessarily the manager. Conclusion: The boss leaves the deepest scar on justice, deviance, and strain; behavioral withdrawal is under-supported and high power-distance generalizability remains thin.

What are we actually measuring? Construct clarity

The aggression-at-work literature suffers from severe jingle-jangle drift: abusive supervision, destructive and despotic leadership, workplace bullying, incivility, social undermining, and corporate psychopathy are label-proliferating constructs that all describe hostile interpersonal treatment at work. Hershcovis (2011) demonstrated that this proliferation is largely nominal. At the item level, the scales share literal content (insulting, silent treatment, rumor-spreading, derogatory remarks, and rudeness recur across measures), and across 25 pairwise construct-outcome comparisons only 7 were statistically distinct while 18 (72%) showed overlapping confidence intervals. The theoretically predicted superiority of the more severe constructs failed: abusive supervision was not significantly stronger than incivility on any outcome, and bullying exceeded incivility on only one (physical well-being). No construct showed a uniformly stronger pattern.

We therefore screened on definition and measure, not label. We treat the family as hostile leadership or workplace aggression, but preserve member distinctions only where a study operationally isolated the distinguishing feature (perpetrator identity, intensity, intent, frequency, power). The spine of this review is constructs whose harm source is the supervisor (abusive, destructive, despotic supervision); bullying, incivility, and social undermining enter only as a clearly labelled environment-level supplement and are never allowed to masquerade as leadership evidence. This follows Hershcovis's three screens (item overlap, whether the distinguishing feature was actually measured, outcome convergence) and her recommendation to code presumed distinctions as study-level moderators rather than separate constructs.

We also report two grains throughout. The abusive-supervision-specific grain uses the Tepper (2000) 15-item scale with the supervisor as perpetrator; the broader destructive-leadership umbrella pools petty tyranny, despotic, tyrannical, and aversive leadership instruments. They diverge systematically. For turnover intention, Schyns and Schilling (2013) found r = .222 under the Tepper scale versus r = .339 under other instruments (Qb = 14.306, p < .001), attributable to specificity matching: the Tepper items carry a personal connotation and correlate more with affectivity, well-being, leader-directed attitudes, and counterproductive behavior, whereas the broader instruments align more with resistance, justice, turnover, and performance.

ConstructDefinitionTypical measureSource of harmRole
Abusive supervisionSubordinates' perceptions of the extent to which supervisors engage in the sustained display of hostile verbal and nonverbal behaviors, excluding physical contact (Tepper, 2000, p. 178).Tepper (2000) 15-item scale (or short forms), 5-point frequency anchors; mean alpha .92 across adaptations (Mackey et al., 2017).Immediate supervisor (the boss).core
Destructive leadershipSystematic behavior by a leader that violates the legitimate interests of the organization by undermining its goals, tasks, resources, and effectiveness (Einarsen et al., 2002).Heterogeneous: petty tyranny (Ashforth, 1997), tyrannical/destructive leadership (Einarsen et al., 2002), aversive leadership (Bligh et al., 2007), coercive power (Elangovan & Xie, 2000).Leader.core
Despotic leadershipSelf-serving, autocratic, and controlling leader behavior that dominates subordinates for personal rather than organizational ends (De Hoogh & Den Hartog, 2008).Despotic leadership scale (De Hoogh & Den Hartog, 2008).Leader.core
Workplace bullyingA person repeatedly and over time exposed to negative acts (abuse, offensive remarks, ridicule, social exclusion) on the part of coworkers, supervisors, or subordinates (Einarsen, 2000).Negative Acts Questionnaire (behavioral-experience) or self-labelling with definition; behavioral method yields larger effects (Nielsen & Einarsen, 2012).Any organizational insider (perpetrator not necessarily the manager).context
Workplace incivilityLow-intensity deviant acts, rude and discourteous verbal and nonverbal behaviors with ambiguous intent to harm (Andersson & Pearson, 1999).Incivility scale (Cortina et al.).Any organizational member.context
Social underminingBehavior intended to hinder, over time, the ability to establish and maintain positive relationships, work-related success, and favorable reputation (Duffy et al., 2002, p. 332).Undermining scale (Duffy et al., 2002), separately identifying supervisor versus coworker source.Supervisor or coworker.context

Shaded rows = the leadership spine of this review. Unshaded = related constructs used only as labelled context.

Throughout the review we report two grains and they diverge in a systematic, interpretable way. The abusive-supervision-specific grain (Tepper 2000 15-item scale, supervisor perpetrator, sustained nonphysical hostility) is the conservative estimate of what the boss does; the broader destructive-leadership umbrella (petty tyranny, despotic, tyrannical, aversive, and coercive instruments) pools a wider behavioral band. For turnover intention, Schyns and Schilling (2013) found r = .222 under the Tepper scale versus r = .339 under other instruments (Qb = 14.306, p < .001), a specificity-matching effect: the personally toned Tepper items correlate more with affectivity, well-being, leader-directed attitudes, and counterproductive behavior, whereas the broader instruments correlate more with resistance, justice, turnover, and performance. The independent metaBUS abusive-supervision node sits at r = 0.253 for turnover intention, close to the Tepper-scale grain, which validates using the conservative AS-specific estimate when the question is specifically about the supervisor. The umbrella grain is useful for the toxic-environment picture but it inflates the turnover and performance links by importing despotic and tyrannical constructs, and so it should not be read as a pure estimate of supervisory abuse.

The consequence map (strongest → weakest)

Justice perceptions and relational rupture

Strongest domain; the proximal fairness and trust violation, with corrected effects in the −.50 to −.66 range.

Abuse is read first as a fairness and exchange violation, and the relational-rupture effects are the largest in the entire matrix. Mackey et al. (2017) report interpersonal justice ρ = −.66 (k = 5), interactional justice −.55 (k = 7), procedural justice −.36, and distributive justice −.25, alongside leader-member exchange ρ = −.54, ethical leadership −.50, and perceived organizational support −.40. Zhang and Liao (2015) converge on interactional justice r = −.51, procedural −.34, distributive −.31, and the original Tepper (2000) study found interactional justice r = −.53, procedural −.48, and distributive −.39, with interactional justice emerging as the comparatively strong predictor of both turnover and distress. The independent metaBUS node confirms the magnitude: interpersonal justice r = −0.554, procedural −0.273, justice overall −0.238, and trust in management −0.139 (k_articles = 23). This is where the supervisor does the most measurable damage: the target's sense of fairness, trust, and standing in the relationship collapses before behavior or health do.

Counterproductive work behavior and retaliation

Very strong; supervisor-directed retaliation is the single largest behavioral consequence and the densest metaBUS evidence (k_articles = 38).

Retaliation aimed back at the supervisor is the dominant behavioral response. Mackey et al. (2017) report supervisor-directed (interpersonal) deviance ρ = .53 (k = 14, N = 5,223), counterproductive work behavior ρ = .41 (k = 6), organization-directed deviance ρ = .41 (k = 11), and interpersonal deviance ρ = .35. Zhang and Liao (2015) replicate the pattern: supervisor-directed deviance r = .51, organization-directed r = .38, interpersonal-directed r = .37, and Schyns and Schilling (2013) pool counterproductive work behavior at r = .377 (k = 19). The field-wide metaBUS node is both the strongest and the densest source of triangulation here: counterproductive behavior r = 0.431 (k_articles = 38), counterproductive behavior toward individuals r = 0.458 (k_articles = 20), and toward the organization r = 0.388 (k_articles = 16). The displaced-aggression and social-learning mechanisms proposed by Schyns and Schilling (2013) fit a target who retaliates against the boss because direct resistance is unsafe, and who displaces hostility onto the organization.

Job attitudes

Strong and consistent across all sources; attitudes cluster around ρ = −.30 to −.40.

Attitudinal damage is robust and convergent. Mackey et al. (2017) report perceived organizational support ρ = −.40, job satisfaction −.34 (k = 20, N = 6,469), and affective commitment −.26. Zhang and Liao (2015) report job satisfaction r = −.35, affective commitment −.30, and organizational identification −.22, and Schyns and Schilling (2013) report job satisfaction r = −.336 (k = 21). The metaBUS node confirms the cluster: job satisfaction r = −0.301 (k_articles = 13), affective commitment −0.198, organizational commitment −0.204, and trust in management −0.139. Notably, the strongest single attitudinal effect in Schyns and Schilling (2013) is attitudes toward the leader at r = −.571, locating the attitudinal collapse at the supervisor relationship itself before it generalizes to the organization. These effects are consistent in sign and rank order across meta-analysis, primary study, and the independent metaBUS database.

Wellbeing and strain

Moderate-to-strong on strain (exhaustion ρ = .36); weak on objective and clinical health (metaBUS health r = −0.064).

The abusive-supervision spine shows clear strain effects but thinner clinical evidence. Mackey et al. (2017) report emotional exhaustion ρ = .36 (k = 22, N = 7,761), work-to-family conflict ρ = .35, job tension ρ = .24, and depression ρ = .24 (k = 5). The metaBUS node places emotional exhaustion at r = 0.374, stress at 0.336, exhaustion at 0.335, and anxiety at a weaker 0.173, while objective health is near-zero at r = −0.064 (k_articles = 8). Psychological safety supplies the mechanism that transmits these effects: Frazier et al. (2017) show positive leader relations build safety (ρ̂ = .44), and safety in turn predicts satisfaction (ρ̂ = .53), engagement (.45), and task performance (.43). The strain column is therefore real and convergent, but it is dominated by self-reported exhaustion and tension; the harder clinical and physiological outcomes remain under-powered, and the bullying supplement (below) is what fills the mental-health gap.

Performance and citizenship

Moderate but methodologically fragile; task performance ρ = −.19, with a supervisor-rating confound.

Performance damage is real but modest, and the estimates are methodologically fragile. Mackey et al. (2017) report task performance ρ = −.19 (k = 13, N = 2,872) and organizational citizenship behavior ρ = −.24 (k = 5). Zhang and Liao (2015) report engagement r = −.29, overall citizenship r = −.24, individual-directed citizenship −.21, organizational-directed citizenship −.17, and subordinate work performance −.16. The metaBUS node, despite dense sampling, returns weak performance effects: task performance r = −0.128 (k_articles = 32), in-role performance −0.126 (k_articles = 18), extra-role or citizenship −0.137 (k_articles = 14), and engagement −0.282 from only k_articles = 2. Two caveats bite here. First, supervisor-rated performance is partly a vehicle through which abuse is delivered, confounding the predictor with the criterion (Tepper et al., 2017). Second, Schyns and Schilling (2013) found the hard organizational-performance composite essentially null at r = .039 (k = 2, N = 333), so the performance column should be read as attitudinal and self-reported rather than bottom-line.

Withdrawal and turnover

Weakest-supported consequence; turnover intention is moderate (r ≈ .25-.31) but actual or objective turnover is near zero.

This is the weakest link in the consequence network. Turnover intention is moderate: Zhang and Liao (2015) report r = .30 (k = 13, N = 5,950) and Schyns and Schilling (2013) report r = .313 overall, but split sharply by instrument (Tepper scale r = .222 versus other destructive-leadership instruments r = .339, Qb = 14.306, p < .001), with the independent metaBUS node at r = 0.253 (k_articles = 6). Actual or objective turnover, however, is barely supported: the only primary evidence is Tepper (2000), who found self-reported voluntary turnover predicted at standardized beta = −.37 with organizational justice fully mediating the path, while the lone hard-criterion composite in Schyns and Schilling (2013) is r = .039 and non-significant. metaBUS engagement and voice are also thin (k_articles = 2 and 4). The honest reading is that abuse reliably produces the intention to leave but the evidence that it actually drives people out the door, on objective criteria, is close to absent.

Independent triangulation: metaBUS field-wide correlations

metaBUS is a curated database of the field's meta-analytic correlations, indexing "abusive supervision" as a single de-duplicated node. It is an independent check on the meta-analyses above, and its article count (k) shows where the evidence base is thick versus thin.

OutcomemetaBUS rk (articles)
Turnover intention+0.2536
Job satisfaction-0.30113
Affective commitment-0.1989
Organizational commitment-0.20410
Counterproductive behavior (CWB)+0.43138
CWB toward individuals+0.45820
CWB toward organization+0.38816
Interpersonal justice-0.5546
Procedural justice-0.2739
Justice (overall)-0.23820
Emotional exhaustion+0.3747
Stress+0.33614
Exhaustion+0.33511
Engagement-0.2822
Task performance-0.12832
In-role performance-0.12618
Extra-role/OCB-0.13714
Trust in management-0.13923
Voice behavior-0.1354
Health-0.0648
Anxiety+0.1734

Convergence across the abusive-supervision meta-analyses, the original primary study, and the independent metaBUS field-wide node is strong and reassuring, and the rank order of consequences holds in all three sources. metaBUS counterproductive behavior r = 0.431 (k_articles = 38, the densest cell) confirms Mackey (.41 to .53) and Zhang (.38 to .51); metaBUS interpersonal justice r = −0.554 confirms Mackey ρ = −.66; metaBUS job satisfaction r = −0.301 confirms Mackey −.34, Zhang −.35, and Schyns −.336; and metaBUS emotional exhaustion r = 0.374 confirms Mackey ρ = .36. Evidence density, read from metaBUS k_articles, is thickest for counterproductive behavior (38), task performance (32), trust in management (23), job satisfaction (13), and exhaustion or stress (11 to 14). It is thinnest for turnover intention (k_articles = 6), anxiety (4), voice (4), engagement (2), and objective health (8, with r = −0.064 near zero). Actual or objective turnover is the thinnest cell of all: a single primary study in the synthesis and a non-significant composite (r = .039), so behavioral exit is the least-supported claim in the network while deviance, performance, satisfaction, justice, and exhaustion are thickly and consistently established.

Mechanisms

Three mechanisms carry the consequence network, with differential empirical support. Organizational justice is the dominant tested mediator: Tepper (2000) showed that justice perceptions fully mediated the abusive-supervision to voluntary-turnover path (the incremental abusive-supervision effect was non-significant once interactional, procedural, and distributive justice were controlled), and both Mackey et al. (2017) and Zhang and Liao (2015) document that abuse sharply erodes all three justice forms (interactional ρ up to −.55; r −.51). Psychological safety is the theoretically implied but under-tested complement: Frazier et al. (2017) show positive leader relations build psychological safety (ρ̂ = .44, k = 30, N = 10,180; transformational leadership .42; trust in leader .39; leader-member exchange .38; inclusive leadership .36), and psychological safety in turn predicts satisfaction (ρ̂ = .53), engagement (.45), task performance (.43), citizenship (.32), and voice (.31); critically, safety added incremental variance in task performance beyond leadership (ΔR² = .18 to .23 versus a reverse ΔR² of .01 to .08), implying abuse destroys the very safety state that transmits leadership effects, although no abusive-supervision meta-analysis has formally modeled safety as a mediator. Emotional exhaustion and social-exchange or conservation-of-resources routes round out the picture (emotional exhaustion ρ = .36), but these remain theorized rather than meta-analytically quantified. Honest caveats apply: Frazier et al. (2017) ran no formal indirect-effect or longitudinal test, only 13% of correlations were cross-source (same-source ρ = .45 versus different-source .29), objective task performance was only ρ = .07 versus subjective .40, and file-drawer effects appeared for performance, citizenship, and learning behavior.

Toxic environment beyond the boss (supplementary)

Two bullying mental-health meta-analyses and one sickness-absence meta-analysis fill the wellbeing and behavioral-withdrawal columns that the abusive-supervision spine under-measures, with the explicit caveat that their perpetrator is not necessarily the manager. Verkuil et al. (2015, k = 48, N = 115,783) report bullying to mental health r = .36 cross-sectionally, dropping to r = .21 longitudinally (mean lag 28 months), with post-traumatic stress symptoms r = .46 and burnout r = .51 the largest subclusters. Nielsen and Einarsen (2012, k = 66, N = 77,721) converge (mental health r = .34, depression r = .34, post-traumatic stress r = .37) and add prospective mental-health OR = 2.33 alongside a reciprocal vicious circle in which baseline mental health predicts later bullying (r = .19). Nielsen et al. (2016, k = 10) link bullying to subsequent sickness absence OR = 1.58 (95% CI 1.39-1.79), rising to a risk ratio of 2.26 for long-term absence exceeding eight weeks in a single primary study, with the association robust to demographic and work-exposure controls and a symmetric funnel plot. These effects roughly halve moving from cross-sectional to longitudinal designs, signaling common-method inflation. Read this evidence as describing a toxic environment, not specifically the boss.

Headline findings

StatisticWhat it saysSource
ρ = −.66, k = 5, N = 1,111; 80% CV −.66/−.66, SDρ = .00Strongest association in the abusive-supervision matrix: interpersonal justice collapse.Mackey et al. (2017)
ρ = .53, k = 14, N = 5,223; ρXP U.S. = .60Largest behavioral consequence: supervisor-directed retaliation.Mackey et al. (2017)
Tepper scale r = .222 vs other instruments r = .339, Qb = 14.306, p < .001Two-grains divergence on turnover intention by measurement instrument.Schyns & Schilling (2013)
standardized beta = −.37, p < .01; Δχ²(2) = 11.71Only abusive-supervision to actual turnover result, fully mediated by justice.Tepper (2000)
ρ̂ = .44, 95% CI [.39, .50], k = 30, N = 10,180Leadership builds the psychological-safety state abuse destroys.Frazier et al. (2017)
ρ̂ = .62, 95% CI [.51, .73], k = 15, N = 4,648Psychological safety to learning behavior, the largest safety outcome.Frazier et al. (2017)
r = .36, k = 48, N = 115,783 (CS); r = .21, k = 22 longitudinalBullying to mental health, cross-sectional and prospective.Verkuil et al. (2015)
OR 1.58, 95% CI 1.39-1.79, k = 10Bullying to subsequent sickness absence, robust to controls.Nielsen et al. (2016)
18 of 25 (72%) pairwise construct-outcome comparisons had overlapping CIsEmpirical redundancy of the aggression constructs.Hershcovis (2011)
r = 0.431, k_articles = 38Densest independent field-wide triangulation point: counterproductive behavior.metaBUS

Evidence gaps & future directions

  1. Psychological safety is the theoretically implied mediator of leadership effects yet no abusive-supervision meta-analysis has formally tested it; Frazier et al. (2017) provide only incremental-validity evidence, with 87% same-source data and no causal or longitudinal path test.
  2. Actual or objective turnover is near-absent: turnover intention is moderate but the lone hard-criterion composite is r = .039 (ns), so the claim that abuse drives employees out the door remains attitudinal rather than behavioral.
  3. Almost all abusive-supervision evidence is cross-sectional self-report (Mackey et al., 2017; Zhang & Liao, 2015), precluding causal inference and leaving reverse and reciprocal paths largely unmodeled.
  4. High power-distance generalizability is weak: samples are predominantly United States and Scandinavian, Zhang and Liao (2015) find the abusive-supervision to turnover link null in high power-distance Asia (r = .04), and Mackey et al. (2017) flag lower abusive-supervision means and cultural heterogeneity in the international pool.
  5. Mediation is piecemeal: primary studies test one or two mechanisms at a time (Tepper et al., 2017), so the relative explanatory power of justice, social exchange, exhaustion, and safety remains unknown.
  6. Objective and clinical health outcomes are under-powered: metaBUS health r = −0.064 and Mackey depression ρ = .24 rest on few samples, and sleep is null and non-robust in Nielsen and Einarsen (2012) (r ≈ .10, fail-safe N = 0).
  7. Construct drift persists: most distinct aggression labels never measured their claimed distinguishing features (intensity, intent, frequency, power), so the field still aggregates nominal distinctions (Hershcovis, 2011).
  8. Voice and engagement, key healthy-organization signals, are thinly studied under abusive supervision (metaBUS k_articles = 4 and 2).
  9. Supervisor-rated performance confounds abuse with the rating act itself (Tepper et al., 2017), biasing the performance column in an unknown direction.

Key references

Generated through the Lead+D Research-OS pipeline (OpenAlex + Unpaywall + annas retrieval; GLM-worker digests of full-text meta-analyses; metaBUS field-wide correlations from the eemm build). Leadership held as the analytic spine; workplace-bullying evidence included only as labelled context. Effect sizes reported verbatim from source; none estimated.