This study examines how the FATAA ethical principles of fairness, accountability, transparency, accuracy and autonomy relate to the usage of generative artificial intelligence (GenAI) and, through usage, to organizational performance, with ethical leadership as a contextual moderator.
Drawing on institutional theory and behavioral reasoning theory, the study tests a model linking FATAA principles to GenAI usage and, in turn, to organizational performance. A hybrid analytical design combines partial least squares structural equation modeling (PLS-SEM) with three machine-learning methods and is applied to survey data from 301 AI-literate professionals.
Fairness and accuracy are most strongly associated with GenAI usage, whereas accountability, autonomy and transparency show weaker or non-significant effects. GenAI usage, in turn, is strongly associated with organizational performance, with bootstrapped indirect effects confirming that only fairness and accuracy carry through usage to performance. Ethical leadership is tested as a moderator at both the adoption stage and the performance stage and, notably, neither conditioning effect is supported.
The study advances the literature in three ways. First, it shows that ethical principles are not equally consequential for GenAI adoption, which calls into question the universalist framing of AI ethics. Second, the hybrid PLS-SEM and machine-learning design surfaces both explanatory effects and feature-level importance, an integration that single-method studies cannot deliver. Third, ethical leadership is examined as a moderator at both the adoption and performance stages, with neither conditioning effect supported, which documents a dual-stage boundary condition for leadership in AI-enabled environments and positions governance-by-design as a candidate explanation for future testing.
1. Introduction
Generative artificial intelligence (GenAI) is reshaping how organizations create content, draft analyses and support decision-making at a scale that was, until recently, the domain of dedicated knowledge workers (Goodfellow et al., 2020; Kumar et al., 2025; Pillai et al., 2022). Firms across industries have integrated GenAI into customer service, internal documentation and frontline operations, drawn by the promise of agility and efficiency (Ooi et al., 2025). The same deployment speed, however, has exposed a parallel set of ethical, governance and operational failures that the existing literature on AI adoption has not yet fully addressed.
Three recent cases bring the nature of these failures into focus. First, Commonwealth Bank of Australia announced the redundancy of 45 customer service roles, citing an AI voice bot that was reducing call volumes by roughly 2,000 calls per week (ABC News, 2025). Within weeks, internal data and union evidence showed that call volumes had in fact risen, that staff were being asked to work overtime and that team leaders were being pulled onto the phones, which, in turn, forced the bank to reverse the decision, apologize to the affected employees and concede that the redundancy decision had not adequately considered the relevant business factors (ABC News, 2025). What failed in this case was not the AI system as such but the organizational governance around it, which had acted on inflated efficiency claims before those claims had been verified against operational reality. Second, Anthropic agreed to pay US$1.5bn to settle a class action brought by authors whose books had been downloaded from pirate repositories and used to train its large language models, in what became the largest publicly reported copyright recovery on record (NPR, 2025). The court had earlier ruled that training on lawfully acquired works qualified as fair use, suggesting that the settlement concerned the ethics and legality of the data sources rather than model training as such. Third, courts in Australia have begun to penalize legal practitioners who submit AI-generated citations without independent verification, with the Victorian Legal Services Board imposing conditions on one solicitor's practicing certificate after a Family Court case was found to contain fabricated case references (The Guardian, 2025). Read alongside one another, the three cases share a common signature, that is, the ethical breach in each is not produced by the AI system in isolation but by the organizational decision to deploy it without the accountability, fairness and verification controls that responsible deployment would have required.
The rapid surge in AI deployment has drawn parallel attention to a set of normative and regulatory frameworks designed to set boundaries around how AI systems should be built and used (European Commission, 2024; Floridi and Cowls, 2019; Jobin et al., 2019; OECD AI Principles, 2019). These frameworks offer important guidance at the policy and principle levels, yet they were not developed for empirical organizational analysis and most are difficult to operationalize as testable constructs at the firm level. Closer to the empirical task is the fairness, accountability, transparency, accuracy and autonomy (FATAA) framework, which translates ethical principles into measurable dimensions of organizational AI practice (Kieslich et al., 2022; Rana et al., 2024; Shin and Park, 2019). FATAA is well suited to the present study for two reasons. First, each of its five dimensions can be operationalized through validated measurement items grounded in prior work on algorithmic ethics. Second, the framework spans both the normative and operational sides of AI governance, thereby allowing the relationship between ethical principles and organizational outcomes to be tested rather than asserted.
Within this growing body of work on ethical AI, three gaps remain that the present study is positioned to address. First, much of the existing literature treats ethical principles either as a single composite construct or examines individual principles without testing their relative salience for adoption decisions (Casali, 2011). Closely related work that does engage FATAA dimensions has examined ethical considerations alongside GenAI adoption and performance (Rana et al., 2024), yet the empirical question of which specific dimensions exert the strongest influence on adoption, and which carry weaker or null effects, has not been fully answered in a way that holds up under both explanatory and predictive analysis. Second, prior empirical work has tended to rely on either structural equation modeling (SEM) for explanatory inference or machine learning for prediction, but rarely on both within the same study (Shmueli, 2010), leaving open the question of whether the explanatory and predictive stories converge on the same set of ethical drivers. Third, the role of ethical leadership in AI-enabled environments has been theorized largely as a moderator that strengthens technology-performance relations (Al Halbusi et al., 2023; Bedi et al., 2016), yet the evidence base for this assumption in GenAI contexts is thin.
The proposed model addresses these gaps by linking the FATAA principles to organizational performance through a single, sequential mechanism. The five FATAA dimensions are positioned as antecedents of GenAI usage, on two complementary theoretical grounds. Institutional theory predicts that organizations adopt practices that meet the legitimacy expectations of their regulatory and normative environments, which, in turn, allows ethical principles to operate as institutional pressures rather than as abstract ideals on adoption decisions (Scott, 2005). Behavioral reasoning theory adds an individual-level microfoundation, that is, the way users translate ethical principles into reasons for or against adoption through cognitive evaluation (Westaby, 2005). The model further positions GenAI usage as the proximal driver of organizational performance, consistent with the view that (AI) capability functions as a strategic resource (in data-rich environments) (Barney, 1991). Ethical leadership is introduced as a contextual moderator at both the adoption and performance stages, on the reasoning that leader-level ethical signaling should strengthen both the translation of ethical principles into usage and the realization of performance gains from AI use (Al Halbusi et al., 2023). This sequential structure of antecedents, mediator, outcome and moderator allows the study to test not only whether ethical principles influence adoption, but also whether the effect flows through to performance and whether leadership shapes that flow.
Three research questions emerge directly from this structure:
Which FATAA principles drive the use of GenAI in organizations and how do their effects compare in magnitude?
What is the effect of GenAI usage on organizational performance?
Does ethical leadership moderate the relationship between GenAI usage and organizational performance?
To address these questions, the study combines two analytical traditions that are typically used in isolation. Partial least squares (PLS)-SEM is used for theory-driven explanatory inference and three machine-learning techniques (decision trees, logistic regression and random forests) are used for predictive analysis and feature-level importance. The combination is not parallel but integrative, whereby the machine-learning results are read against the PLS-SEM results to assess whether the explanatory and predictive accounts converge on the same set of ethical drivers, with the chord diagram serving as the integrative bridge between the two.
The study makes three contributions. First, it empirically shows that the FATAA dimensions are not equally consequential for GenAI adoption, thereby calling into question the prevailing assumption that ethical principles can be treated as a uniform construct. Second, the hybrid PLS-SEM and machine-learning design surfaces both explanatory effects and feature-level importance, an integration that single-method studies cannot deliver. Third, the study specifies ethical leadership as a moderator at both the adoption and performance stages and finds neither conditioning effect supported, which yields a dual-stage boundary condition on leadership that prior FATAA-based work, with its single-stage capability moderator, does not provide, with governance-by-design advanced as a candidate interpretation of this boundary condition. The combined effect of these three contributions is to shift the conversation on ethical AI in organizations from principle to practice, that is, from what organizations should do to what actually shapes adoption and performance in the field.
2. Literature review
This section develops the conceptual model that links ethical AI principles to the adoption of GenAI and to organizational performance. The model rests on the FATAA principles of fairness, accountability, transparency, accuracy and autonomy, treated not as a list of isolated normative commitments but as five interdependent dimensions of an integrated governance scheme (Floridi and Cowls, 2019; Jobin et al., 2019). Within this scheme, the dimensions interact rather than operate in isolation. The legitimacy, trust and operational reliability they generate are produced jointly rather than as separate effects (European Commission, 2024; OECD AI Principles, 2019).
Fairness concerns the absence of bias and the equitable treatment of stakeholders by algorithmic decisions (Akter et al., 2023). Accountability concerns the assignment of responsibility for AI outputs and the traceability of decisions back to identifiable actors (Dwivedi et al., 2021). Transparency concerns the interpretability of AI systems and the reduction of information asymmetry between systems and users (Kieslich et al., 2022). Accuracy concerns the reliability of AI outputs and their fitness for the task at hand (Ooi et al., 2025). Autonomy concerns the extent to which AI systems can act and decide without continuous human input (Mökander and Floridi, 2022). Each principle has a direct counterpart in current generative systems. Fairness maps onto bias carried in training data and model outputs (Wörsdörfer, 2024), accountability onto the traceability of machine-generated content (Stahl and Eke, 2024), transparency onto the opacity of large foundation models (European Commission, 2024), accuracy onto factual reliability where a system can produce fluent yet incorrect output (Chae and Yoon, 2025) and autonomy onto the rise of agentic systems that pursue goals with reduced human oversight (Acharya et al., 2025). The five dimensions are interdependent rather than independent, which, in turn, implies that an increase in one can come at the cost of another, for instance, when greater autonomy reduces accountability or when higher accuracy comes at the cost of transparency (Stahl and Eke, 2024). FATAA is thus best understood as a framework that balances ethical fidelity and operational delivery (Floridi and Cowls, 2019). Recent work on ethics-by-design reinforces this reading, treating fairness, transparency and accountability as properties to be engineered into AI systems from the outset rather than applied after deployment (Brey and Dainow, 2024).
Institutional theory provides the first lens for the framework. Organizations operate within environments shaped by regulatory and normative pressures and thus adopt practices that meet the legitimacy expectations of those environments (Scott, 2005). Rising societal and regulatory concerns about AI have, in turn, made the FATAA principles into legitimacy-bearing institutional pressures that organizations are expected to be seen to honor. The institutional reading is, however, only a partial account of adoption, since organizations may imitate normative pressures ceremonially without implementing them in operational practice (Scott, 2005).
Behavioral reasoning theory provides the second lens. The theory holds that users evaluate a technology through the reasons they generate for and against its adoption (Westaby, 2005), with the reasons functioning as the cognitive bridge between values and behavioral intention. The mechanism has been applied in technology-adoption contexts where users weigh competing considerations of benefit and cost or risk (Claudy et al., 2015; Sahu et al., 2020). In the present context, the FATAA dimensions act as the content of those reasons. Fairness and transparency reduce perceived ethical risk and increase trust in the system (Glikson and Woolley, 2020; Kumar et al., 2025). Accountability operates as a reason that strengthens confidence in the governance scheme around the system. Autonomy, by contrast, operates as a double-sided reason, since the same property that enables efficiency also raises concerns about loss of control and ethical responsibility (Ray, 2023). The FATAA dimensions enter the adoption decision not only as institutional pressures but also as cognitive evaluative inputs at the individual user level.
Read alongside one another, the two perspectives produce a coherent account (Hollebeek et al., 2025): institutional theory explains why organizations adopt FATAA-aligned practices in response to external environmental pressures (Scott, 2005) and behavioral reasoning theory explains how those principles are translated into adoption decisions at the individual level (Westaby, 2005). This dual-lens structure represents a point of departure from prior FATAA-based work. Where earlier accounts treat institutional pressures as the primary driver of GenAI adoption, the present framework positions behavioral reasoning theory as the primary explanatory lens, with the FATAA principles entering as individual-level reasons-for adoption, and institutional theory as a complementary interpretive lens that situates those reasons within their regulatory and normative context. The combined account aligns directly with the empirical model, in which the FATAA dimensions are tested as antecedents of GenAI usage while GenAI usage is tested as the proximal driver of organizational performance.
Fairness is expected to drive GenAI adoption on three grounds. First, perceived fairness reduces resistance to AI systems in organizational contexts by strengthening stakeholder trust, which is itself determined by the extent to which algorithmic decisions are perceived as fair and justified (Akter et al., 2023; Mehrabi et al., 2021). Second, fairness is increasingly scrutinized by regulators (Wörsdörfer, 2024), which, in turn, makes fairness a compliance requirement rather than a discretionary feature. Third, although bias mitigation can come at the cost of model accuracy in some settings (Rahimzadeh et al., 2023), the trade-off does not undermine the adoption argument, since organizations that fail on fairness face reputational and regulatory exposure that outweighs the marginal accuracy gains. Accordingly, we hypothesize:
Fairness positively influences the usage of GenAI in organizations.
Accountability is also expected to drive GenAI adoption, although through a different mechanism. Clarifying who is responsible for AI outputs and how those outputs can be traced back to specific actors reduces the perceived risk of deploying GenAI systems in consequential settings (Dwivedi et al., 2021; Stahl and Eke, 2024). The expected effect, however, is likely to be smaller than that of fairness for two reasons. First, accountability mechanisms are often designed primarily to satisfy regulatory requirements rather than to facilitate organizational adoption decisions, which means their salience for adoption is partly indirect. Second, the over-formalization of accountability can generate bureaucratic friction that limits operational flexibility (Hu et al., 2021), which sets a ceiling on the size of the positive effect that can be expected. As such, we hypothesize:
Accountability positively influences the usage of GenAI in organizations.
Transparency is expected to promote GenAI adoption by reducing uncertainty about how the system arrives at its outputs and by increasing users' understanding of the algorithmic process (European Commission, 2024; Kieslich et al., 2022), implying that increased understanding strengthens trust and the system's perceived reliability. The expected effect, however, is bounded for two reasons. First, users in organizational settings may prioritize performance outcomes over interpretability, which, in turn, might limit the marginal value of additional transparency once a baseline level is met. Second, excessive transparency may lead to information overload, reducing practical usability (Casal and Kessler, 2023). The net effect is expected to be positive but modest. Consequently, we hypothesize:
Transparency positively influences the usage of GenAI in organizations.
Accuracy is expected to drive GenAI adoption directly, since the perceived usefulness of an AI system is closely tied to the reliability of its outputs (Chae and Yoon, 2025). Decision-making, operational performance and risk management in organizational settings all depend on accurate outputs (Ooi et al., 2025), which, in turn, makes accuracy a high-salience consideration in adoption decisions. The expected effect is, however, moderated by an internal tension within the model. Higher accuracy often requires more complex models that are harder to explain, which can come at the cost of transparency and users' ability to contest outputs (Lipton, 2018). Bias-mitigation requirements can also pull against pure accuracy optimization, since fairness constraints sometimes require the model to depart from the empirical distribution of the training data (Barocas et al., 2018; Mehrabi et al., 2021). Despite these tensions, accuracy is expected to remain the most directly observable driver of adoption, since organizations evaluate AI primarily by what the system produces rather than by how it produces it. Hence, we hypothesize:
Accuracy positively influences the usage of GenAI in organizations.
Autonomy is expected to drive GenAI adoption by allowing organizations to automate processes that would otherwise require human input at every step (Mökander and Floridi, 2022). Greater autonomy increases efficiency and scalability, particularly for routine and repetitive tasks (Brynjolfsson et al., 2025). The expected effect, however, is constrained by the parallel concerns about loss of control, the allocation of ethical responsibility and the risk exposure that comes with delegating consequential decisions to autonomous systems (Zerilli et al., 2019). The net effect is expected to be positive but smaller than the effects of fairness and accuracy. In this vein, we hypothesize:
Autonomy positively influences the usage of GenAI in organizations.
GenAI usage is expected to translate directly into organizational performance through three mechanisms (Kumar et al., 2025; Mikalef and Gupta, 2021). First, GenAI capabilities function as valuable, less-imitable resources that contribute to competitive advantage (Barney, 1991). Second, GenAI enables organizations to process and analyze large data volumes at speed, which, in turn, improves decision quality and operational efficiency (Kumar et al., 2025). Third, the realization of these performance gains is contingent on dynamic capability conditions, in particular collaboration, governance design and alignment of GenAI use with organizational priorities (Mikalef and Gupta, 2021; Teece et al., 2016). Therefore, we hypothesize:
The usage of GenAI positively influences organizational performance.
Ethical leadership is expected to condition how strongly the FATAA principles translate into GenAI adoption, which positions leadership as a contextual influence at the antecedent stage and not only at the outcome stage. Behavioral reasoning theory holds that the reasons users generate for or against a technology gain or lose force according to the social context in which the evaluation occurs (Westaby, 2005), which implies that the salience of ethical principles as adoption reasons is shaped by the signals leaders send about what the organization values. Ethical leaders model, communicate and reward ethical conduct (Brown and Treviño, 2006), which supplies employees with a frame that elevates the relevance of FATAA when they reason about whether to adopt GenAI. Social learning theory supplies the mechanism, as employees acquire and prioritize ethical considerations by observing credible leaders rather than through formal rules alone (Bandura, 1986; Bedi et al., 2016). This contextual reading distinguishes the present model from prior FATAA-based work that positions organizational innovativeness as the moderator of the usage–performance link alone (Rana et al., 2024), given that ethical leadership is a relational and normative condition rather than a capability and given that the present model tests its conditioning effect at both the adoption stage and the performance stage rather than at the performance stage only. The expectation, in turn, is that the association between each FATAA principle and GenAI usage strengthens as ethical leadership rises. Accordingly, we hypothesize:
Ethical leadership positively moderates the relationship between (a) fairness, (b) accountability, (c) transparency, (d) accuracy and (e) autonomy and the usage of GenAI in organizations, such that each relationship strengthens as ethical leadership increases.
Beyond the adoption stage, ethical leadership is also expected to moderate the relationship between GenAI usage and organizational performance, since leader-level ethical signaling shapes how AI systems are used and how their outputs are scrutinized in operational practice (Al Halbusi et al., 2023; Brown and Treviño, 2006). The way leaders frame and enforce ethical principles affects employee trust, compliance and the cultural alignment that sustains responsible AI use, with meta-analytic evidence supporting ethical leadership's influence on a range of organizational outcomes including ethical behavior, citizenship behavior and performance (Bedi et al., 2016). The strength of this moderation effect is, however, expected to vary across organizational contexts, levels of technological maturity and the formality of governance arrangements (Islam and Greenwood, 2024). Thus, we hypothesize:
Ethical leadership positively moderates the relationship between GenAI usage and organizational performance.
In combination, the eight hypotheses provide the empirical specification through which the differential salience of the FATAA dimensions, the GenAI–performance relationship and the conditioning role of ethical leadership at both the adoption and performance stages are tested. The empirical model and its estimation are described in the next section.
3. Methodology
3.1 Design
The study employs a quantitative, survey-based design that combines PLS-SEM with three machine-learning techniques (Hair et al., 2022), with the methodological flow presented in Figure 1. The combination is motivated by the complementary strengths of the two approaches.
PLS-SEM is well suited to testing latent-variable models with mediation and moderation paths under the conventional assumption of linear relationships (Hair et al., 2022). Machine-learning techniques relax the linearity assumption, capture interactions and non-linear patterns that PLS-SEM cannot represent directly, and yield feature-importance scores that rank predictors by their contribution to predictive accuracy (Richter and Tudoran, 2024). The two analytical traditions are integrated through the chord diagram visualization, which allows the explanatory results from PLS-SEM and the feature-level results from machine learning to be inspected against each other in a single visual representation.
3.2 Instrumentation
A structured survey instrument was developed to operationalize the constructs in the proposed model. Accountability, accuracy, ethical leadership and transparency were measured using items adapted from validated prior scales. Fairness and autonomy require item development beyond the available scales, as the existing literature offers limited, validated instruments for these dimensions in GenAI contexts. Specifically, the new fairness items were derived from theoretical work on distributive justice and stakeholder theory, while the new autonomy items were derived from work on AI-enabled decision-making capacity.
All items were reviewed by domain experts for face and content validity, and four items were clarified and trimmed based on expert feedback to improve clarity and reduce redundancy (Lim et al., 2026). The items and their properties and sources are presented in our later evaluation of the measurement model.
One construct deserves explicit definitional care. GenAI usage is operationalized through three employee-reported indicators adapted from validated adoption and usage instruments (Lin et al., 2018; Pillai et al., 2022; Rana et al., 2024), which capture the organization's proposal to use GenAI in the near future, its inclination to increase the use of GenAI and its capacity to use GenAI. The construct thus represents reported organizational engagement with GenAI, spanning use intention, expansion inclination and readiness, rather than system-logged frequency, intensity or task coverage of use, in line with the position that the appropriate conceptualization of usage is the one that fits the theoretical context in which usage is embedded (Burton-Jones and Straub, 2006) and with the long-standing reliance on self-reported use measures in information systems theory testing (Straub et al., 1995). Terminology is applied consistently on this basis throughout the manuscript, as GenAI usage refers to this construct, the adoption stage refers to the stage of the model at which the ethical principles relate to GenAI usage, and the performance stage refers to the stage at which GenAI usage relates to organizational performance.
3.3 Pilot study
A pilot study with 50 respondents was conducted to assess the reliability and validity of the measurement instrument before full data collection. Most constructs met the conventional Cronbach's alpha threshold of 0.70 in the initial pilot, but the fairness construct returned a low value (α = 0.465), prompting further development of the construct, in line with iterative refinement principles for empirical construct validation (Lim et al., 2026).
Three additional fairness items were added to the survey, drawn from established literature on AI ethics and organizational justice theory. The new items were designed to capture procedural justice (consistency of process and equality of treatment) and outcome parity (predictive validity of outputs across stakeholder groups), thus providing the construct with more comprehensive coverage in relation to fairness dimensions.
Three additional items were also added to the autonomy construct, designed to reflect varying levels of AI-enabled decision-making capacity rather than only the binary presence or absence of autonomy.
Table 1 reports the initial pilot Cronbach's alpha values, where fairness returned a low value with three items, motivating the addition of three further fairness items. After these additions and the corresponding three additional autonomy items, the full-sample measurement model showed acceptable reliability for fairness and autonomy, as reported in Table 3. Composite reliability and average variance extracted (AVE) were also assessed in line with PLS-SEM guidelines (Hair et al., 2022), and both indices supported the measurement model's reliability and convergent validity.
3.4 Main study
3.4.1 Data collection
Data were collected through Prolific, an online research platform commonly used in business research (e.g. Kumar et al., 2025). Prolific maintains a large pool of pre-screened participants and allows researchers to apply stratified sampling on demographic and professional characteristics. The present study used stratified sampling on industry and employment type to ensure coverage across organizational contexts.
Two participation criteria were applied to ensure that respondents could speak knowledgeably about GenAI use in organizational settings. First, respondents had to use workplace technology at least three times a week. Second, respondents had to have used at least one of these GenAI products: Bard, Bing AI, ChatGPT, Claude, GitHub Copilot or Llama. The two criteria, when combined, define an AI-literate professional sample.
The two criteria also establish respondents as competent informants for the constructs they rate, as all constructs are measured as individual-level perceptions with the respondent as the consistent referent, in line with the individual-level reasoning mechanism of behavioral reasoning theory, and respondents report on the GenAI systems they use at work, the supervisors they interact with and the performance consequences they observe in their daily roles. Perceptual measures of organizational performance are an established practice in this respect, supported by evidence of convergence between subjective and objective performance indicators (Dess and Robinson, 1984; Richard et al., 2009; Wall et al., 2004). Middle-level employees, who form the largest share of the sample, additionally occupy the organizational positions where GenAI deployment is enacted in daily work, which positions them as appropriate informants for the system properties and usage practices under study. The design remains an individual-level one for all constructs, which means the discussion and conclusion draw the organizational-level reading only as far as the perceptual evidence allows.
The final analytic dataset comprised 301 valid responses, after filtering for response consistency, reaction speed and internal logical coherence. One early respondent had missing values on items added later in the pilot iteration (FR4–FR6 and AT5–AT7), yielding n = 300 for those item-level analyses. The sample size approximates the conventional PLS-SEM rule of thumb of about 10 responses per indicator, since the model contains 33 indicators and the achieved sample yields a 1:9 ratio, and substantially exceeds the alternative rule of 10 times the largest number of structural arrows pointing into a single construct (Hair et al., 2022; Lim, 2025b). The sample also shows demographic diversity across age, gender, organizational size and industrial sector, which, in turn, supports the generalizability of the findings across the populations represented.
3.4.2 Data analysis
This study introduces a hybrid analytical method that integrates PLS-SEM and machine learning techniques with chord diagram visualization.
3.4.2.1 PLS-SEM
PLS-SEM was conducted in SmartPLS using a two-step process (Hair et al., 2022). First, the measurement model was assessed via Cronbach's alpha and composite reliability for internal consistency, factor loadings and AVE for convergent validity and the heterotrait–monotrait (HTMT) ratio for discriminant validity. To examine multicollinearity, we examined variance inflation factors (VIFs). Second, the structural model was assessed via path coefficients and their significance, which were estimated via bootstrapping (5,000 resamples). Goodness of the model's fit and predictions was examined using R2 and Stone–Geisser's Q2. PLS-SEM is especially well-suited for examining complex models with many latent variables and mediation relationships. However, PLS-SEM assumes a latent linear structure and prioritizes explanatory power over predictive accuracy; thus, it may not accurately represent non-linear interactions or complex decision-making processes characteristic of GenAI adoption contexts.
3.4.2.2 Machine learning
Machine learning techniques are introduced here to address two known limitations of PLS-SEM, specifically its reliance on linear assumptions and its emphasis on explanation rather than prediction (Hair et al., 2022; Richter and Tudoran, 2024). Three complementary algorithms are applied: decision trees (Quinlan, 1986), logistic regression (Peng et al., 2002) and random forests (Breiman, 2001). Decision trees offer interpretable, rule-based partitioning of the predictor space, thereby making the role of each variable visible at every split. Random forests aggregate many such trees, reduce the overfitting risk that any single tree carries and generate feature-importance scores that rank predictors by their contribution to model accuracy. Logistic regression provides a probabilistic comparison point against which the tree-based models can be benchmarked. Read alongside the PLS-SEM results, the three techniques add a predictive and feature-ranking layer to the explanatory inference produced by the structural model, which, in turn, allows the study to assess whether the explanatory and predictive accounts converge on the same set of ethical drivers.
Class imbalance in the survey responses is addressed through two complementary resampling strategies. The synthetic minority oversampling technique (SMOTE) generates synthetic minority-class observations by interpolating between existing instances of the minority class and their nearest neighbors (Chawla et al., 2002). Random undersampling, in turn, draws a balanced subset of the majority class to prevent the model from defaulting to the dominant class. The two strategies are applied in parallel, with the original distribution retained as a third comparison condition. Both resampling strategies are applied to the training partition only, after the stratified split, which ensures that the validation and test partitions retain the original class distribution and remain untouched by synthetic or resampled observations. Model performance is evaluated through accuracy, precision, recall and F1-score for the classification tasks (Sokolova and Lapalme, 2009) and through root mean square error (RMSE) and normalized root mean square error (NRMSE) for the regression-style tasks (Willmott and Matsuura, 2005). Robustness of the single-split results is assessed through repeated stratified cross-validation (10 repetitions of 5 folds) with resampling re-applied inside each training fold, with agreement beyond chance quantified through Cohen's kappa (Cohen, 1960). The quartile discretization decision is, in turn, checked through a random-forest regression estimated directly on the continuous composite outcomes.
To contextualize the absolute level of predictive accuracy, the machine-learning models are compared against a majority-class baseline, which assigns every test case to the most frequent response category in the training set and which serves as the floor below which a classification model would amount to a heuristic guess. The GenAI usage and organizational performance Likert composite scores were discretized into four ordinal classes by quartile-based binning, with cutoffs at the 25th, 50th and 75th percentiles. The dataset was subsequently split into training (60%), validation (20%) and test (20%) subsets using a fixed random seed, and classification metrics (precision, recall and F1-score) were computed using weighted averaging to account for class imbalance.
3.4.2.3 Visualization of chord diagrams
Chord diagram visualization is used to display the predictive structure produced by the random-forest analysis in a single integrated representation. For each predictor, the contribution to the target variable is computed from the random forest's feature-importance scores, standardized across predictors and rendered as the thickness of the corresponding chord between the predictor and the target. The diagrams are constructed at two levels of granularity, the construct level (Figure 8) and the item level (Figure 9), which, in turn, allows the predictive structure to be inspected at the level at which adoption decisions are made and at the level at which the underlying survey items operate.
4. Findings
4.1 Descriptive statistics
The summary statistics provide a profile of the key variables in the analysis. The highest item-level means appear in transparency (TR2: M = 4.28, SD = 0.883; TR1: M = 4.26, SD = 0.904) and ethical leadership (EL3: M = 4.23, SD = 0.878). The fairness items, such as FR4 (M = 4.09, SD = 1.005), cluster at the upper end of the scale. These patterns suggest that respondents view their organizations as relatively open and ethically led. They also view algorithmic processes as generally fair. The lowest item-level mean (M = 3.43, SD = 0.989) sits at the modifiability-of-system-configuration item, which is consistent with a more cautious view of how much users can or should reshape the underlying behavior of GenAI systems. The full item-level descriptive statistics are reported in Table 3.
VIFs were computed to assess multicollinearity among the predictors. All VIFs are well below the conventional threshold of 10 (Shrestha, 2020), indicating that multicollinearity does not threaten the structural model estimates.
Table 2 summarizes the demographic profile of respondents. The sample is balanced enough to support generalizable inference. Respondents skew toward female participants, the 25–29 age cohort and bachelor's-level education, with cross-regional coverage spanning Africa, Europe, and smaller shares from the Americas, Asia and Oceania rather than confining the findings to a single context. Most respondents occupy middle-level roles with one to six years of professional experience, while a clear majority report direct experience with AI technologies, which supports the study's focus on AI-literate adoption settings. Industry coverage is led by technology and education, while organizational size is balanced across small, medium and large firms.
Figure 2 shows the distribution of responses (raw frequencies) across all measurement items, separated by whether respondents have direct prior experience with AI technologies. Given the larger number of respondents with prior AI experience (n = 228) relative to those without (n = 73), the bar heights partly reflect group-size imbalance, so the figure is read as a descriptive distribution rather than as direct evidence of differential endorsement. With this caveat in mind, two suggestive patterns are visible. First, the response distributions among respondents with prior AI experience appear concentrated at the upper end of the scale, particularly on the transparency items (TR1, TR2, TR3) and ethical leadership items (EL1, EL3). Second, respondents without prior AI experience exhibit wider response distributions, particularly on accountability and fairness items, which is consistent with greater variability in perceptions of these dimensions in the absence of direct experience with the underlying technology. This observation is most pronounced on the autonomy items (AT2, AT3) and the GenAI usage items (GU1, GU2), which suggests that direct experience with AI calibrates how respondents interpret claims about system independence and organizational use. The most plausible reading is that prior experience produces a learning effect, in which respondents become more attuned to the practical and ethical implications of GenAI as they accumulate exposure. The implication for practice is that training and onboarding programs that build AI literacy are likely to shape how the FATAA dimensions are perceived in organizational settings, which, in turn, has consequences for how those dimensions translate into adoption decisions.
Figure 3 displays the pairwise relationships among the FATAA dimensions across higher and lower GenAI usage groups, defined by a threshold split on the GenAI usage composite score, where respondents with a composite score ≥3.5 are classified as higher GenAI users and the remainder as lower GenAI users. The diagonal shows the marginal distribution of each dimension, the upper triangle presents bivariate density contours and the lower triangle presents pairwise scatter plots. The visible separation between the two groups across most pairs of dimensions indicates that higher self-reported GenAI usage is associated with a shifted joint distribution of ethical perceptions, not only the marginals.
The marginal distributions along the diagonal show that skewness and dispersion differ between higher and lower GenAI usage groups, particularly for autonomy and accuracy. The bivariate scatter and density plots show systematic departures from a linear pattern across most pairs of dimensions, which, in turn, motivates the use of machine-learning techniques in addition to PLS-SEM, since the latter is constrained by the linearity assumption.
The non-linear pattern of relationships visible in Figure 3 supports the analytical strategy adopted in this study, which combines linear structural modeling with machine-learning techniques that can capture the non-linear structure that linear models cannot.
4.2 Partial least squares structural equation modeling (PLS-SEM)
This section reports the PLS-SEM evaluation, beginning with the measurement model before proceeding to the structural model.
4.2.1 Measurement model
Reliability was assessed using Cronbach's alpha (α) and composite reliability (ρ), with the conventional 0.70 threshold applied to both indices (Lim, 2025b; Taber, 2018). As shown in Table 3, though seven of the eight constructs meet the α threshold, all eight constructs meet the ρ threshold. Accountability returns α = 0.633, which falls just below the 0.70 reference point but exceeds the 0.60 threshold sometimes adopted in exploratory work. The construct is retained on the strength of its composite reliability (ρ = 0.803) and its AVE (0.578), as well as on the conceptual ground that the three accountability items capture the distinct but complementary sub-dimensions of external monitoring (AB1), third-party scrutiny (AB2) and intervention or modifiability (AB3), in line with extant guidance on scale development and validation (Lim et al., 2026). The intervention item also has direct conceptual grounding, as the account of accountability in AI as answerability treats the interrogation of a system and the limitation of its power as conditions of accountability rather than as separate traits (Novelli et al., 2024), which places the capacity to modify a system's configuration inside the accountability construct rather than outside it. A composite-level robustness check further shows that the structural pattern reproduces almost exactly at the composite level, with the accountability coefficient at 0.019 against the PLS-SEM estimate of 0.021, and that removing the intervention item shifts the accountability coefficient only from 0.019 to 0.023, both non-significant, while shifting no other coefficient by more than 0.002, which confirms that the accountability findings are not driven by this item. The accountability findings reported below are read with appropriate caution, and the limitations section returns to this point.
Convergent validity was assessed using factor loadings and the AVE, against the conventional thresholds of 0.50 for AVE and 0.708 for factor loadings (Cheung and Wang, 2017; Lim, 2025b). As shown in Table 3, all of the constructs meet the AVE threshold. A small number of individual loadings (FR2 = 0.670, AB3 = 0.655, EL2 = 0.617) fall just below the 0.708 reference point yet remain within the 0.40–0.70 band for which the established guidance recommends retention unless removal raises composite reliability or AVE above the relevant threshold (Hair et al., 2022). As each affected construct already satisfies the AVE and reliability thresholds, the three indicators are retained on the grounds of content validity and the supporting AVE evidence (Lim et al., 2026).
Factor loadings range from 0.617 to 0.910 across the eight constructs, indicating generally strong relationships between indicators and their respective latent variables. The construct-level ranges are: accountability (0.655–0.809), accuracy (0.849–0.906), autonomy (0.739–0.818), ethical leadership (0.617–0.837), fairness (0.670–0.876), GenAI usage (0.848–0.892), organizational performance (0.873–0.910) and transparency (0.813–0.833). The few loadings that fall just below the 0.708 reference point are retained on the grounds of content validity and the supporting AVE evidence (Lim et al., 2026). The measurement model as a whole meets the conventional thresholds for reliability and convergent validity.
The HTMT ratio was used to test discriminant validity. Table 4 shows that all HTMT values were less than 0.85, indicating that the constructs were distinct enough (Henseler et al., 2015). This means the constructs examine distinct theoretical ideas with little overlap.
The study assessed common method bias (CMB) using a variance-explained methodology, given reliance on self-reported measurements. Table 5 shows that the first unrotated factor explains 29.591% of the total variance, well below the 50% criterion set by Podsakoff et al. (2003). After rotation, the largest factor accounts for 14.445% of the variance. The pattern confirms that CMB is not a serious concern in this study.
4.2.2 Structural model
The structural model was estimated through bootstrapping with 5,000 resamples to test the eight hypotheses. Table 6 reports the path coefficients (β), t-values and 95% confidence intervals for the main (direct) effects (H1–H6) in Panel A and the mediation (indirect) effects through GenAI usage in Panel B, while Table 7 reports the moderation tests (H7 and H8). Of the five FATAA dimensions, accuracy (β = 0.217, t = 3.132, p < 0.01) and fairness (β = 0.208, t = 3.179, p < 0.01) emerge as significant positive predictors of GenAI usage, supporting H1 and H4. Autonomy exhibits a positive association with GenAI usage (β = 0.116, t = 1.658), though the 95% confidence interval includes zero ([−0.021, 0.253]), which provides only limited support for H5. Accountability (β = 0.021, t = 0.269, n.s.) and transparency (β = 0.065, t = 1.093, n.s.) are not statistically significant, which means H2 and H3 are not supported in this sample. GenAI usage, in turn, has a strong positive effect on organizational performance (β = 0.527, t = 10.236, p < 0.01), supporting H6. The full results are reported in Table 6, Panel A.
The sequential mechanism implied by the model, in which the ethical principles reach organizational performance through GenAI usage, was tested formally through the specific indirect effects of each principle on organizational performance through GenAI usage, estimated from the same 5,000-resample bootstrap with all interaction terms entered simultaneously and evaluated through percentile confidence intervals, in line with established guidance on mediation analysis in PLS-SEM (Nitzl et al., 2016) and on judging mediation through the significance of the bootstrapped indirect effect (Zhao et al., 2010). As reported in Table 6, Panel B, accuracy (β = 0.104, t = 2.394, p = 0.017, 95% CI [0.022, 0.194]) and fairness (β = 0.097, t = 2.424, p = 0.015, 95% CI [0.021, 0.176]) transmit significant positive indirect effects to organizational performance through GenAI usage, whereas the indirect effects of accountability (β = 0.015, p = 0.606), autonomy (β = 0.054, p = 0.171) and transparency (β = 0.024, p = 0.469) are not significant. The mediation evidence thus mirrors the direct-path pattern, as the two principles that predict GenAI usage are also the only principles whose effects carry through usage to organizational performance, which completes the differential-salience account at the level of the full sequential mechanism.
Table 7 reports the moderation tests for ethical leadership at both stages of the model, estimated with all interaction terms entered simultaneously. At the adoption stage, none of the five interactions between ethical leadership and the FATAA principles is statistically significant, with all interaction coefficients small and all t-values below 1.35, which means H7 is not supported in any of its five components. At the performance stage, the interaction between ethical leadership and GenAI usage is also non-significant (β = 0.059, t = 1.145, p = 0.252), which means H8 is not supported. The combined evidence is a consistent pattern of non-significance across both stages, as ethical leadership conditions neither the translation of ethical principles into adoption nor the conversion of adoption into performance. The discussion section returns to this double null in interpretive depth. Two further checks bound what the double null can and cannot establish. A sensitivity analysis shows that, at the achieved sample size with α = 0.05 and statistical power of 0.80, the design was able to detect interaction effects as small as f2 ≈ 0.026, a value just above the conventional small-effect threshold of 0.02 (Cohen, 1988), whereas the observed interaction effect sizes range from f2 = 0.001 to f2 = 0.010 across the six interaction terms and the bootstrapped 95% confidence interval for the performance-stage interaction ([−0.043, 0.154]) caps the conditioning effect at a small magnitude even at its upper bound. The pattern indicates that the double null is unlikely to reflect inadequate power against small-to-medium conditioning effects, although interactions of the still smaller magnitude typical of field research, where the median observed moderator effect is f2 = 0.002 (Aguinis et al., 2005), would require substantially larger samples to detect and cannot be ruled out.
The PLS-SEM model explains 22% of the variance in GenAI usage (R2 = 0.22) and 36.6% of the variance in organizational performance (R2 = 0.366). The R2 for GenAI usage corresponds to a moderate effect size, while the R2 for organizational performance corresponds to a large effect size by the conventional Cohen (1988) thresholds.
Predictive relevance was assessed through the Stone–Geisser Q2 statistic, computed in SmartPLS using a blindfolding procedure (Shmueli et al., 2019). The Q2 values are 0.151 for GenAI usage and 0.278 for organizational performance, which exceed the zero threshold for predictive relevance and indicate that the model has acceptable predictive power for both endogenous constructs.
4.3 Machine learning
4.3.1 Data preprocessing
The dataset was divided into training, validation and test subsets in a stratified 60:20:20 split with a fixed random seed for supervised learning evaluation. The three-way partition follows standard supervised-learning protocol, in which the validation subset is reserved for model selection and hyperparameter tuning while the test subset is kept fully held out for final evaluation. As the models were estimated with default configurations and no hyperparameter tuning, the validation subset was not required for model selection, and the three-way split was retained so that the test subset remained untouched during training. Final performance is reported on the held-out test subset.
4.3.2 Sampling technique
The training data exhibits class imbalance across the response categories of both GenAI usage and organizational performance, which can bias machine-learning models toward the dominant class if left unaddressed. Two complementary resampling strategies were applied to address this. SMOTE generates synthetic minority-class observations by interpolating between existing minority-class instances and their nearest neighbors (Chawla et al., 2002). Random undersampling, in turn, draws a balanced subset of majority-class observations by randomly removing instances from the dominant class until the class sizes are equalized (Chen et al., 2014; Riskiyadi, 2024; Shang et al., 2021). Models were trained and evaluated under three conditions: the original distribution, the SMOTE-oversampled distribution and the random-undersampled distribution. The performance differences across the three conditions are reported in the next subsections.
4.3.3 Machine learning models
Three machine-learning models were estimated for each of the two outcomes of interest. For GenAI usage as the outcome, the individual FATAA items served as the predictors. For organizational performance as the outcome, GenAI usage, ethical leadership and the GenAI usage by ethical leadership interaction term served as the predictors. The three models (decision tree, logistic regression and random forest) were estimated under each of the three sampling conditions, yielding nine model fits per outcome. Feature-importance scores from the random-forest models were retained for the variable-importance analysis.
4.3.4 Evaluation of models
Model performance is reported under each of the three sampling conditions for both outcomes, with GenAI usage and organizational performance in the next sub-sections.
4.3.4.1 Influence of ethical principles on GenAI usage
Table 8 shows that all three classifiers perform within a narrow band, with no model clearly dominating across sampling conditions. Under the original distribution, the random forest achieves the highest accuracy (0.33), followed by the logistic regression and decision tree (both at 0.30). Under undersampling, the random forest reaches its highest accuracy (0.36, F1 = 0.36), whereas the decision tree achieves its highest accuracy (0.36) under oversampling. Logistic regression performs most consistently across the three conditions but does not exceed 0.30. These results are based on a four-class classification scheme derived from quartile-based discretization of the outcome variable. Error metrics (NRMSE) indicate that the random forest under oversampling achieves the lowest prediction error (NRMSE = 0.39). Figure 4 shows substantial misclassification across classes for GenAI usage. The random forest identifies Q2 most accurately (12 of 21, 57%) while Q1 (12%), Q3 (18%) and Q4 (33%) show weaker class-specific prediction. Adjacent-class confusion is the dominant error pattern, with Q1 most frequently misclassified as Q2.
To contextualize the absolute level of predictive performance, the machine-learning models are compared against a majority-class baseline, that is, a heuristic that assigns every test case to the most frequent response category in the training set. In the present sample, this baseline yields an accuracy of approximately 33.9% for GenAI usage. The best-performing model is the random forest under undersampling at 0.36, with the random forest under the original distribution at 0.33. The remaining model and sampling combinations perform at or near the baseline. The narrow performance band suggests that the FATAA item set carries limited predictive signal for GenAI usage when treated as a four-class classification target. Two factors plausibly limit the absolute accuracy ceiling. First, GenAI usage, as operationalized here, is shaped by individual, organizational and contextual influences that go beyond the FATAA dimensions captured here, which sets a natural ceiling on what the predictor set can explain. Second, perceived class boundaries are not sharp in self-reported survey data, as reflected in the confusion matrices, which show systematic overlap between adjacent response categories. The machine-learning results are therefore interpreted alongside the PLS-SEM findings as complementary evidence on variable importance and non-linear pattern detection, rather than as a stand-alone predictive system. Robustness checks reinforce this reading while attaching uncertainty to it. Under repeated stratified cross-validation (10 repetitions of 5 folds, with resampling re-applied inside each training fold), the best-performing configuration for GenAI usage reaches a mean accuracy of 0.373 (SD = 0.052) with a mean Cohen's kappa of 0.138, against a majority-class proportion of 0.340. A random-forest regression estimated directly on the continuous usage composite, in turn, yields a cross-validated R2 of 0.019, which confirms that the limited item-level predictive signal is a property of the data rather than an artifact of the quartile discretization or of any single train-test split. The gap between the in-sample explanatory variance of the structural model and the out-of-sample predictive accuracy reported here is itself informative, as explanation and prediction are distinct scientific goals whose divergence carries diagnostic value (Shmueli, 2010; Shmueli and Koppius, 2011). The divergence indicates that the significant average associations identified by PLS-SEM do not translate into an individual-level prediction rule, which bounds the practical reading of the effects.
4.3.4.2 GenAI usage and organizational performance under ethical leadership
Table 9 shows that logistic regression performs best on the original distribution (accuracy = 0.46, F1 = 0.41), while the decision tree and random forest both achieve 0.39. Oversampling slightly reduces performance for the decision tree and random forest, while logistic regression remains stable at 0.44. Under undersampling, the three models perform within a narrow band (0.36–0.41), with logistic regression again leading. Error metrics (RMSE and NRMSE) indicate that logistic regression under oversampling achieves the lowest prediction error (NRMSE = 0.40). These results are based on a four-class classification scheme derived from quartile-based discretization of the outcome variable, with GenAI usage, ethical leadership and their interaction term as predictors. Figure 5 shows considerable class overlap for organizational performance, with both adjacent and non-adjacent misclassifications. Q4 (55%) and Q2 (50%) are predicted more accurately than the other classes, while Q3 (20%) shows the weakest performance. A notable pattern is the misclassification of Q1 as Q3, indicating that errors are not confined to neighboring classes. Under the same repeated cross-validation protocol, the best-performing configuration for organizational performance reaches a mean accuracy of 0.496 (SD = 0.041) with a mean Cohen's kappa of 0.253, against a majority-class proportion of 0.402, which places the performance-stage models more clearly above chance than the usage-stage models and locates the predictive signal primarily in GenAI usage itself.
Table 10, Panel A, and Figure 6 report the feature-importance scores from the random-forest model for GenAI usage as the outcome. The accountability item AB3 (0.066) is the single most influential predictor, followed by the accountability items AB1 (0.060) and AB2 (0.055), the transparency item TR3 (0.055), the fairness item FR6 (0.051) and the accuracy items AC1 (0.049) and AC3 (0.049). Accountability items occupy three of the top four ranks, with transparency, fairness and accuracy items completing the top tier, indicating a multi-faceted item-level structure rather than a single dominant principle. The impurity-based ranking should not, however, be read as evidence of substantive importance. Accountability's zero-order correlation with GenAI usage is near zero (r = 0.011), which means the bivariate evidence and the structural estimate agree that the construct carries no meaningful linear association with usage. Impurity-based importance is, in turn, known to be biased when predictors differ in their scale characteristics and to be distorted in the presence of correlated predictors (Strobl et al., 2007, 2008). A permutation-importance re-estimation on the held-out test partition, which is less susceptible to these biases, places the three accountability items at effectively zero (0.017, 0.005 and −0.009) while ranking an accuracy item (AC1 = 0.061) and a fairness item (FR4 = 0.045) at the top, which brings the predictive account into convergence with the structural account on fairness and accuracy. The chord diagrams that follow are read on this basis as descriptive visualizations of the random-forest impurity structure rather than as rankings of substantive importance.
Table 10, Panel B, and Figure 7 report the feature-importance scores for organizational performance as the outcome. The interaction term GenAI usage × ethical leadership emerges as the most influential predictor (0.370), with GenAI usage at 0.324 and ethical leadership at 0.306. The interaction-term importance score is not, however, accompanied by a significant moderation effect in the PLS-SEM analysis. To demonstrate that this divergence reflects shared rather than unique variance, a hierarchical regression was estimated in which the product term was added to a model already containing GenAI usage and ethical leadership. The product term raises the explained variance by only ΔR2 = 0.003 (from 0.348 to 0.351, non-significant), and the product term itself correlates near zero with organizational performance (r = −0.058), while GenAI usage and ethical leadership correlate with performance at r = 0.563 and r = 0.335. The high feature-importance score reflects the variance the interaction term shares with its constituent predictors, since random-forest importance is computed without partialling out that shared variance, whereas the PLS-SEM moderation estimate isolates the unique interaction effect and finds it non-significant. A permutation-importance re-estimation for the performance outcome reinforces this reading, as GenAI usage retains a clear positive score (0.081) while ethical leadership (0.014) and the interaction term (−0.013) contribute essentially nothing out of sample, which is fully consistent with the non-significant moderation estimate and with the dominance of the usage main effect.
4.4 Chord diagram visualizations
Chord diagrams visualize the relative contribution of each predictor to GenAI adoption and to organizational performance. Each chord's thickness reflects the magnitude of the predictor's contribution as derived from the random forest's feature-importance scores, which, in turn, gives a single visual representation of the predictive structure that the structural and machine-learning analyses surface separately.
Figure 8 shows the construct-level view, in which the FATAA dimensions and ethical leadership are plotted against GenAI usage and organizational performance.
Figure 9 disaggregates the same relationships at the item level, mapping each measurement statement onto the target variable rather than aggregating to the construct level. Read alongside Figure 8, Figure 9 surfaces the predictive structure of the model at two complementary levels of granularity, which, in turn, identifies both the construct-level and item-level drivers of GenAI adoption and the corresponding effects on organizational performance.
5. Discussion
The empirical results identify fairness and accuracy as the ethical principles most strongly associated with GenAI usage in organizational settings, which, in turn, indicates that organizations weigh both moral legitimacy and functional reliability in their adoption decisions. The pattern is consistent with prior research linking fairness to stakeholder trust and perceived equity in algorithmic outputs and those linking accuracy to the perceived usefulness and decision-making effectiveness of AI systems (Rana et al., 2024; Shin and Park, 2019; Song et al., 2022). Read through the behavioral reasoning lens, fairness and accuracy operate as cognitive rationales for adoption that reduce perceived risk and increase perceived benefit (Kumar et al., 2025).
The non-significant effects of accountability and transparency on GenAI usage, alongside the marginal effect of autonomy, suggest a disconnect between ethical intent and practical adoption levers. Prior work emphasizes these principles in normative governance frameworks, but the empirical evidence here implies that, in the studied sample, they operate more as compliance and legitimacy structures than as direct drivers of adoption decisions (Dwivedi et al., 2021). Rigid accountability requirements and high transparency standards can also create operational friction that limits adoption, particularly in fast-moving deployment contexts (Hu et al., 2021).
The strong positive link between GenAI usage and organizational performance is consistent with the resource-based view (Barney, 1991) and dynamic capabilities theory (Teece et al., 2016), in which such (GenAI) capabilities are read as strategic resources that enable organizations to identify opportunities, optimize processes and redeploy resources across changing environments (Ciasullo et al., 2026). The result is also consistent with prior empirical work that links GenAI use to improved decision-making and operational outcomes, particularly in data-rich settings (Kumar et al., 2025; Pillai et al., 2022).
The non-significant moderation by ethical leadership suggests that the leadership–performance relationship in AI-enabled environments is more contextual than the original moderation hypothesis implied. While ethical leadership remains a recognized driver of organizational culture and ethical behavior (Al Halbusi et al., 2023), the present results indicate that, once GenAI is in use, post-adoption performance is dominated by system-level attributes rather than amplified by leader-level signaling. This points toward a more contingent view of leadership, in which the effect of ethical leadership on performance depends on its integration with technical maturity and operational governance rather than operating as an independent moderator.
The machine-learning analyses initially appear to surface a divergent picture from the PLS-SEM results, given that the impurity-based random-forest scores rank the accountability items AB3, AB1 and AB2 in three of the top four positions despite the construct's non-significance in the structural model. The divergence dissolves once the importance metric is corrected, as the permutation-importance re-estimation reported with the findings places the accountability items at effectively zero and ranks accuracy and fairness items at the top, which aligns the predictive account with the structural account and with accountability's near-zero bivariate correlation with GenAI usage. The convergence of the two analytical traditions on fairness and accuracy, reached only after the known biases of impurity-based importance are corrected (Strobl et al., 2007, 2008), is precisely the kind of methodological insight the hybrid design is built to deliver, as a single-method study would have either over-read the impurity ranking or never surfaced it.
Model performance results further show that resampling does not consistently improve classification performance. For GenAI usage, the random forest performs best under undersampling rather than under oversampling, while logistic regression performs most consistently across the three conditions. For organizational performance, oversampling improves logistic regression's F1-score and NRMSE modestly, although accuracy is slightly lower than under the original distribution, and oversampling does not improve the tree-based models. The pattern indicates that the effectiveness of resampling strategies depends on the model type and the outcome rather than reflecting a general result for class-imbalance correction.
The confusion matrices show systematic misclassification between adjacent response categories, suggesting that the survey-based predictors do not clearly distinguish between GenAI usage levels. This pattern suggests that future research may benefit from approaches that better capture the ordinal nature of the outcome.
5.1 Theoretical implications
The empirical findings apply behavioral reasoning theory to the GenAI adoption setting and refine it in one specific respect. The FATAA dimensions are modeled as reasons-for adoption, and the results indicate that these reasons do not carry equal weight. Fairness and accuracy operate as pragmatic rationales that lower cognitive resistance and confer moral legitimacy on the technology, whereas accountability, autonomy and transparency are weaker in this sample. The differential salience identified here aligns the behavioral reasoning account with prior work on stakeholder trust and decision-making quality and recasts those moral dimensions as behavioral inputs to adoption rather than only as background normative commitments. The contribution is best read as a refinement of how reasons-for enter algorithmic-adoption decisions rather than as a modification of behavioral reasoning theory as a whole, since the present design captures the reasons-for component and does not measure reasons-against (Westaby, 2005). Refinement of this kind, which specifies which classes of reasons hold predictive weight, is a recognized form of theoretical contribution (Colquitt and Zapata-Phelan, 2007; Corley and Gioia, 2011; Whetten, 1989), and the differential salience itself stands as a noteworthy finding in the sense of Lim's (2026b) typology, that is, a more granular reading of how ethical principles enter adoption decisions than the universalist framing of AI ethics permits.
The non-significant effects for accountability and transparency, alongside the limited support for autonomy, provide an empirical boundary condition for the universalist framing of AI ethics, in which ethical principles are assumed to operate uniformly and to improve organizational outcomes once embedded in governance frameworks. A more defensible reading is that fairness and accuracy function as the ethical considerations most directly evaluated at the point of adoption, whereas accountability and transparency are governance-oriented principles whose influence is more indirect, working through implementation and oversight rather than through the initial adoption decision. This interpretation follows from the differential-salience logic developed in the framework rather than from a retrospective appeal to institutional pressures and, more importantly, it avoids treating institutional theory as a device that can account for both significant and non-significant results after the fact. Against Lim's (2026b) typology, the result is counterintuitive relative to the universalist prior, which holds that institutionalized ethical principles should operate uniformly across organizational settings.
The strong positive association between GenAI usage and organizational performance locates the performance payoff in reported usage rather than in the endorsement of ethical language. The specific indirect effects reinforce this reading, as fairness and accuracy are the only principles whose effects carry through usage to performance, which positions usage as the proximal channel through which ethical principles reach outcomes and which is now established through formal mediation evidence rather than asserted from the pattern of direct paths. The theoretical implication is that the considerations relating ethical principles to usage operate at a different level from the usage through which performance is realized, which means the two are better theorized as complementary rather than substitutable mechanisms.
The consistent non-significance of ethical leadership as a moderator at both the adoption stage and the performance stage invites a more careful theoretical reading than a single-stage framing would supply. The moderation hypotheses drew on social learning theory and upper-echelons reasoning (Bandura, 1986; Bedi et al., 2016), under which leader-level ethical signaling should strengthen both the translation of ethical principles into adoption and the conversion of adoption into performance. The data support neither conditioning path in this cross-sectional sample, even though ethical leadership relates directly to GenAI usage and, more strongly, to organizational performance as a main effect. The combined pattern points to a boundary condition that can be read as governance-by-design rather than governance-by-person (Brey and Dainow, 2024). Where ethical considerations are already embedded in system properties, documented controls and institutionalized oversight, the marginal signal a leader adds at the moment of evaluation is small, since the ethical frame employees draw on is sourced from the design of the technology and its governance rather than from leader modeling alone. Ethical leadership, on this reading, operates additively as an antecedent of adoption and performance rather than interactively as an amplifier of either link, which positions leadership upstream in governance formation and oversight rather than at the point where principles convert into adoption or adoption converts into performance. A moderated-mediation specification, in which ethical leadership shapes the governance design and operational controls that channel GenAI use into performance, is the natural next step, though it awaits a longitudinal or multi-stage design before it can be advanced as a positive claim. Treating this double null as informative rather than as a failed test is consistent with methodological guidance that non-significant findings carry theoretical value when they bound the reach of an established explanation (Cortina and Folger, 1998; Edwards, 2010) and, against Lim's (2026b) typology, the result is counterintuitive relative to the social-learning prediction. Nonetheless, alternative explanations for the double null also deserve explicit statement rather than silent dismissal. Insufficient statistical power, measurement limitations in the leadership construct, the individual level at which all constructs are measured, construct misspecification and heterogeneity across sectors and regulatory contexts could each produce non-significance without any governance mechanism at work. The sensitivity analysis reported with the structural results addresses the first of these directly, as the design held adequate power against small-to-medium conditioning effects and the observed interaction effect sizes fall below even the small-effect threshold, with the remaining alternatives acknowledged as open. Governance-by-design is thus advanced as the candidate interpretation most consistent with the observed pattern and, specifically, with the presence of significant main effects alongside uniformly trivial interaction effect sizes.
A further theoretical implication arises from the divergence between the structural and predictive results for accountability. The construct is not significant as a linear antecedent in the structural model (β = 0.021, n.s.) yet emerges as the dominant item-level predictor in the random-forest analysis, with the three accountability items occupying three of the top four feature-importance ranks (AB3 = 0.066, AB1 = 0.060, AB2 = 0.055). The divergence is not a contradiction but a feature of combining two analytical lenses, since structural model coefficients capture average linear effects after partialling out shared variance, whereas feature-importance scores capture non-linear and conditional contributions that the linear specification cannot detect. To examine whether the pattern reflects a conditional effect rather than a standalone one, a supplementary analysis using construct composite scores tested accountability for a quadratic term and for interactions with the other FATAA dimensions, with significance assessed through 5,000 bootstrap resamples. Accountability shows no quadratic effect on adoption, yet it interacts significantly with transparency (β = 0.181, t = 2.73) and with autonomy (β = 0.137, t = 2.26), such that its association with usage is negative when transparency is low and positive when transparency is high. This exploratory evidence is consistent with accountability operating in conjunction with other governance principles rather than as an independent linear driver, which reconciles its null structural path with its high item-level importance. The implication for theory is that the relative importance of ethical principles for adoption depends on the analytical lens applied, which, in turn, suggests that single-method studies risk under-reporting principles whose effects operate through conditional channels, and which supports the case for hybrid empirical designs in AI ethics research. Against Lim's (2026b) typology, this divergence is paradoxical in the precise sense the typology articulates, that is, the co-existence of two opposing analytical readings of the same construct.
Last but not least, it is worthwhile to clarify that the closest prior study, Rana et al. (2024), examined ethical considerations, GenAI adoption and organizational performance through institutional theory, with organizational innovativeness moderating the usage–performance link at a single stage. The present study departs from that account on four grounds rather than extending it incrementally. First, behavioral reasoning theory serves as the primary lens, which reframes the FATAA principles as adoption reasons of unequal weight rather than as collectively important pressures, and the results identify fairness and accuracy as the salient reasons while accountability and transparency are not. Second, ethical leadership enters as a relational and normative moderator at both the adoption stage and the performance stage, which produces a dual-stage conditional model where the prior account specified a single capability-based moderator at the performance stage only. Third, the consistent non-significance of leadership at both stages is theorized as a governance-by-design boundary condition rather than set aside as a null result, which adds an explanatory mechanism the prior account does not contain. Fourth, the design pairs explanatory PLS-SEM with predictive machine learning across a cross-sector sample of AI-literate professionals, which triangulates the differential-salience finding in a way a single-method study cannot. These points of departure are summarized in Table 11. The contrast also functions as an informative non-replication, as the uniform antecedent support reported for 384 managers in Indian IT and ITeS firms does not extend to cross-sector AI-literate professionals, with context-bounded results of this kind constituting the raw material of cumulative theory building in replication scholarship (Bettis et al., 2016; Köhler and Cortina, 2021; Tsang and Kwan, 1999).
5.2 Practical implications
The findings translate into four practical implications for organizations adopting GenAI, each grounded directly in the empirical results rather than in normative argument.
First, fairness and accuracy are the ethical principles most strongly associated with GenAI usage, which, in turn, places these two principles at the top of the operational governance agenda. In practice, this means investing in fairness audits and bias-detection routines that operate continuously across the AI lifecycle, with explicit metrics for demographic parity, equal opportunity across stakeholder groups and outcome consistency over time. The corresponding accuracy-validation protocols should test outputs against documented benchmarks before deployment decisions are made, with specific provision for held-out evaluation, edge-case stress testing and ongoing drift monitoring once the system is in production. These are operational controls, and the strength of their empirical associations justifies the investment they require.
Second, accountability and transparency did not predict adoption in this sample, with autonomy showing only marginal predictive evidence. This observation calls for a more measured operational stance toward these three principles. The empirical reading is not that these principles are unimportant, but that their effect on adoption is mediated by how they are designed and embedded in governance systems. The implication for practice is that organizations should treat accountability, autonomy and transparency as enabling conditions, that is, as features that have to be operationalized through audit structures, decision boundaries for human oversight and explainability artifacts, rather than as direct levers that can be pulled on their own. Concrete governance patterns include responsibility matrices that assign decision rights and intervention authority for AI outputs, model documentation conventions (model cards, data sheets) that make system behavior inspectable to internal and external stakeholders, and escalation protocols that route ambiguous or high-stakes decisions to human review.
Third, the strong positive link between GenAI usage and organizational performance supports framing GenAI as a strategic asset rather than as an ancillary technical project. The implication for practice is to align GenAI deployment with the organization's core operational and competitive priorities and to invest in the surrounding infrastructure (data quality, model monitoring, integration with existing workflows) that allows the technology to deliver against those priorities at scale.
Fourth, the data do not support the assumption that ethical leadership amplifies GenAI–performance relations, which, in turn, sets a different practical agenda for leadership in AI-enabled environments. Rather than relying on individual leadership signaling to convert GenAI use into performance, organizations should invest in formal governance structures that institutionalize ethical norms and operational controls, so that ethical standards are protected against turnover, role changes and shifts in leadership style. This is a system design implication, not a behavioral one and aligns with the empirical finding rather than working around it.
6. Conclusion
6.1 Key takeaways
The empirical question that drives this study is which ethical principles, if any, are associated with the usage of GenAI inside organizations and whether their effects flow through to organizational performance once GenAI is in use. To address this question, the FATAA framework was tested in a hybrid design that combines SEM with three machine-learning techniques on survey data from 301 AI-literate professionals.
Three findings stand out. Fairness and accuracy are the ethical principles most strongly associated with GenAI usage. Accountability, autonomy and transparency show weaker or non-significant associations. GenAI usage, in turn, is strongly associated with organizational performance, with the bootstrapped indirect effects confirming that only fairness and accuracy carry through usage to performance. Ethical leadership does not condition either the adoption stage or the performance stage, which calls for a more careful empirical specification of where leadership operates in AI-enabled environments.
The implication for theory is that ethical principles are not equally consequential for adoption, thereby calling into question the universalist framing that has dominated the AI ethics literature. The implication for practice is that fairness and accuracy require concrete operational controls. Accountability, autonomy and transparency require governance design that converts principle into operational consequence. Both implications are grounded in empirical findings rather than in normative claims and both invite further empirical scrutiny in longitudinal and cross-sectoral directions that the limitations section outlines.
6.2 Key limitations and future directions
Although this study contributes empirical evidence on how ethical principles shape GenAI adoption, four limitations open directions for future research.
First, the cross-sectional design limits the strength of causal inference and prevents observation of how the relationships among ethical principles, GenAI usage and performance evolve as AI deployments mature. Longitudinal designs that track organizations from initial governance design through long-term operation would, in turn, allow the temporal dynamics of ethical AI adoption to be observed directly. The analysis also pools respondents across sectors, which, in turn, may obscure variation that is likely to be present in heavily regulated domains such as defense, finance and healthcare. Sector-disaggregated replications, ideally combined with industry-specific instruments, would help establish how the FATAA dimensions operate across different regulatory environments.
Second, the study relies on quantitative self-report data, which captures perceptions of fairness, accountability and other ethical dimensions but does not capture the lived organizational dynamics through which ethical decisions actually unfold. The reliance on self-report extends to the two outcome-side constructs, as GenAI usage is captured through reported use intention, expansion inclination and capacity rather than through system-logged frequency or intensity of use, while organizational performance is captured through respondents' perceptions rather than through archival indicators, with all constructs measured at the individual level from single informants. Designs that pair matched informants across organizational levels with behavioral telemetry would allow the usage–performance link to be established at the organizational level of analysis and with logged rather than reported usage (Burton-Jones and Straub, 2006; Straub et al., 1995). Beyond measurement, qualitative work, in particular ethnographic case studies, expert interviews and grounded theory designs (Lim, 2025a), would, upon reaching saturation (Lim, 2026a), effectively surface the contextual reasoning that survey instruments cannot reach and provide a complementary evidence base alongside the structural and predictive analyses presented herein.
Third, the analysis examines respondents' perceptions of FATAA, but does not assess the effectiveness of specific technical explainability tools that are increasingly central to the operational delivery of these principles. Future work that integrates Local Interpretable Model-agnostic Explanations (Ribeiro et al., 2016) and SHapley Additive exPlanations (Lundberg and Lee, 2017) into the empirical design would, in turn, allow the relationship between technical transparency and perceived fairness to be tested directly, rather than inferred.
Fourth, ethical norms vary across cultural and institutional contexts, which, in turn, limits the generalizability of the present findings beyond the populations represented in the sample. Cross-cultural comparative designs would establish whether the relative salience of fairness and accuracy holds across regulatory regimes and whether the null moderation by ethical leadership reflects a feature of AI-literate professional populations or a more general property of GenAI–performance relations.
Declaration of Generative AI statement
During the preparation of this work, the authors used Anthropic Claude, OpenAI GPT and Microsoft Editor to check for argumentative coherence and improve readability, including expression, tone and style of writing. OpenAI GPT was also used to illustrate Figure 1 under the supervision of the authors. After using these tools/services, the authors reviewed and edited the content as needed and take full responsibility for the content of the publication.










