This study examines whether the effect of hybrid work on innovative work behavior (IWB) among public employees is heterogeneous across individuals and organizational contexts. The specific question is whether subgroups for whom hybrid work is especially conducive to IWB can be identified. Methodologically, the study also asks how such heterogeneity claims can be validated rather than assumed.
Using observational survey data from 6,075 Korean public employees in 2024, the study applies uplift modeling with the X-Learner (CatBoost base learner) to estimate conditional treatment effects. The heterogeneity results are then validated with permutation (placebo) inference, grouped-ATE calibration, resampling across 100 train–validation splits, four alternative base learners and covariate-adjusted average-effect estimation.
The X-Learner ranked employees no better than chance (uplift AUC = 0.019; permutation p = 0.245), and the bivariate hybrid–IWB association did not survive covariate adjustment (adjusted difference = 0.022; p = 0.13). These results offer a lens for interpreting the divergent findings in prior hybrid-work–IWB research, suggesting that much of the apparent inconsistency may reflect confounding and unstable subgroup estimation rather than substantive, individually targetable heterogeneity.
The cross-sectional, self-reported design limits causal inference and may introduce common method bias, but the findings underscore the importance of examining heterogeneous effects in hybrid work research.
Because the analysis did not recover reliably targetable subgroups, hybrid work is best provided broadly under transparent, procedurally fair conditions. The policy focus shifts from “who is predicted to benefit” to “how hybrid work is designed and supported”.
This study redirects attention from individual targeting to structural job design. The null on individual-level heterogeneity, together with prior evidence that design-level antecedents such as job autonomy are direct predictors of IWB, motivates broad, procedurally fair provision rather than selective allocation. Methodologically, the study adapts permutation-based inference and grouped-ATE calibration from causal-ML econometrics to uplift modeling, offering a transferable template for reliable heterogeneity analysis.
