This study examines how chatbot capability failures lead to customers' incivility in AI-enabled service delivery through customers' emotional responses. It also investigates whether chatbot-expressed empathy, as a recovery-oriented capability, weakens or intensifies the relationship between customers' negative emotions and incivility.
Drawing on social response theory, we analyzed 32,586 text conversations from a university library chatbot using a framework that combines natural language processing and supervised machine learning to extract the constructs of interest. We then estimated a moderated mediation model using PROCESS to examine the proposed relationships.
The study identified chatbot capability failures and empathy, customer digital incivility and customer emotional profiles. The findings showed that technical and functional chatbot capability failures elicited negative customer emotions, and that these negative emotions fully mediated the relationship between chatbot capability failures and customer incivility behaviors. Although we anticipated that chatbot empathy, as a recovery-oriented capability, would mitigate the effect of negative emotions on incivility, the findings revealed the opposite: empathy strengthened rather than attenuated this effect. We termed this paradox the “empathy backfire,” indicating that empathic expressions alone were insufficient and could even intensify customer incivility when they were not matched by adequate chatbot capability.
The study advances social response theory by revealing a counterintuitive empathy backfire effect in which chatbot empathy intensifies, rather than mitigates, customer incivility behaviors. Although empathy often helps mitigate the negative effects of service failure on customer behavior in human interactions, managers should not assume that empathetic language embedded in chatbot scripts will produce comparable benefits.
1. Introduction
AI-enabled service delivery agents such as chatbots are increasingly central to frontline service delivery, enabling firms to provide scalable, personalized and always-available customer support across digital channels (Castelo et al., 2023; Crolic et al., 2022; Sheehan et al., 2020; Van Pinxteren et al., 2023; Zhang et al., 2023). Despite their growing prominence, chatbot capability failures remain common. Chatbot capability failure is defined as a chatbot's inability to meet customer expectations by providing inaccurate information, misinterpreting customer questions or failing to resolve customer problems (Li et al., 2025). When chatbots misunderstand customer needs, provide irrelevant responses or fail to resolve problems, customers may respond with frustration, anger and incivility behavior (Brendel et al., 2023). These responses have important implications for service managers because customer incivility behaviors can intensify service recovery challenges, prolong interactions, complicate data management and spill over into broader evaluations of the firm (Brendel et al., 2023; Castelo et al., 2023; Huang and Dootson, 2022).
Customer–chatbot interactions are not only transactional exchanges but also social interactions that can shape broader customer-firm relationships (Brendel et al., 2023). Consistent with social response theory (Nass and Moon, 2000), customers often respond to chatbots as if they were social actors, applying human expectations of responsiveness, reciprocity and emotional sensitivity (Castelo et al., 2023). Consequently, chatbot capability failure may be understood not only as a technical and functional breakdown but also as a social failure that violates customers' expectations of appropriate social conduct, thereby increasing the likelihood of negative emotional responses and customer incivility behaviors. Building on prior research (Lages et al., 2023), we define customers' incivility toward chatbots as customer behavior that violates norms of respectful conduct during human-chatbot interactions. Although prior research has advanced understanding of chatbot adoption, service quality and customer evaluations of services, three important gaps remain.
First, existing research offers limited insight into the factors that trigger customers' incivility within AI-enabled service delivery (Han et al., 2023). Second, while customer incivility is well documented in human service settings, much less is known about how it manifests in AI-enabled service delivery (Brendel et al., 2023). Third, empathy is widely viewed as beneficial in human service interactions; however, its role in chatbot capability failure remains unclear. More specifically, there is a paucity of knowledge on whether chatbot’s empathetic responses mitigate or intensify the progression from negative emotion to customer incivility behaviors (Brendel et al., 2023; Han et al., 2023). We define chatbot empathy as the chatbot's expressed capacity to recognize and respond to a customer's emotional state and service-related difficulty through supportive, understanding and affective language.
These gaps are compounded by methodological limitations in the literature. Most existing evidence relies on scenario-based experiments rather than naturally occurring interactions (e.g. Huang and Dootson, 2022). Such approaches offer limited insight into how emotional and behavioral dynamics unfold in real customer–chatbot interactions. To address these gaps, this study raises two significant research questions: RQ1: To what extent do chatbot capability failures increase customers' incivility toward chatbots through customers' negative emotions? RQ2: To what extent does chatbot empathy moderate the relationship between customers' negative emotions and customers' incivility toward chatbots? We examine these questions using 32,586 real-world text-based conversations between customers and chatbots in a service failure setting to provide a more realistic account of emotional and behavioral dynamics in AI-enabled service delivery. In Phase 1, we follow Fisher et al. (2021) and identify key constructs in customer–chatbot interactions, customers' emotional expressions and forms of customer incivility behaviors. In Phase 2, we transform text data into numeric values and test whether chatbot empathy, as recovery-oriented capability, moderates the relationship between customers' negative emotions and customer incivility behaviors.
This research makes three contributions to social response theory within AI-enabled service delivery research. First, it advances research on AI-enabled service delivery by unfolding how chatbot failure gives rise to negative emotional responses that shape subsequent customer behavior. Second, it extends customers' incivility research into chatbot interactions by showing that norm-violating behavior can be directed toward chatbots, thereby broadening customer incivility behaviors beyond human-to-human service settings. Third, it extends work on the paradox of humanization in AI interactions (Brendel et al., 2023) and clarifies the role of chatbot empathy in AI-enabled service delivery failure by examining whether chatbot empathetic responses, as recovery-oriented capability, mitigate or intensify the progression from negative emotion to incivility. Methodologically, the study also responds to calls for research based on naturally occurring service data by drawing on a large corpus of real-world chatbot conversations rather than hypothetical scenarios. These contributions deepen understanding of customers' responses to chatbot capability failure and subsequent emotional escalation in AI-enabled service delivery. We offer practical insight into the negative consequences of chatbot emotional expression and emotionally responsive chatbot systems.
2. Literature review
2.1 Customers' incivility in AI-enabled service delivery
Service failures disrupt customers' expectations of reliable and responsive service delivery. When customers perceive that a chatbot has failed to understand, respond to or resolve their problem, the resulting negative emotions may escalate into incivility behaviors directed toward the chatbot. Customer incivility in both traditional service encounters and AI-enabled service delivery shares a common behavioral core, namely, rude, disrespectful or norm-violating customer behavior. However, the two contexts differ in theoretically important ways.
In traditional service encounters, incivility is directed at a human frontline employee and is therefore embedded in interpersonal norms, emotional labor and reciprocal social expectations between customers and employees. In AI-enabled service delivery, by contrast, the immediate target may be a non-human service interface, such as a chatbot, which changes the social meaning of the behavior (Crolic et al., 2022). In AI-enabled service delivery, the presence of a human target is reduced as customers interact through a digital service medium, and blame may be attributed ambiguously to the chatbot, the firm or the underlying service system (Ryoo et al., 2024). At the same time, human-like chatbot cues can activate interpersonal expectations of responsiveness, understanding and empathy, meaning that capability failures may still be interpreted through social scripts despite the non-human nature of the chatbot (Crolic et al., 2022).
In traditional service encounters, Grönroos (1984) distinguished between the technical and functional competence of frontline service employees. Technical competence refers to employees' ability to deliver the core service outcome accurately and effectively, whereas functional competence refers to the manner in which the service is delivered. For example, designing a loan package that suits a customer's financial circumstances reflects an employee's technical competence, while communicating that package clearly reflects the employee's functional competence. Customers interacting with chatbots may expect similar capabilities from these systems. When such expectations are violated, customers are more likely to engage in incivility toward chatbots. In AI-enabled service delivery, customers' incivility is often triggered by customer-expressed negative emotions aroused from chatbot failures (Castillo et al., 2021). Customer emotions in AI-enabled service delivery have received some attention (see Table 1). However, findings across studies remain inconsistent (i.e. Luo et al., 2019; Chin et al., 2020). Table 1 presents the current state of the literature on customers' emotions and incivility behaviors in AI-enabled service delivery. The studies are categorized into three streams of research.
Literature review table
| Study | Theory | Methodology | Findings |
|---|---|---|---|
| Luo et al. (2019) | Emotion management theory | Experimental design | Appraisals and post-recovery emotions sequentially mediated the relationship between emotion regulation and consumer word-of-mouth |
| Chin et al. (2020) | Not specified | Experimental design | Empathy emerged as the most effective strategy to mitigate aggressive behavior |
| Sheehan et al. (2020) | Not specified | Experimental design | Unresolved chatbot errors decreased adoption and perceived humanness |
| Castillo et al. (2021) | Not specified | Interview | Customer resource loss determined customer coping strategies in AI-based service failure and customers passed the blame to technology for service failure |
| Choi et al. (2021) | Not specified | Experimental design | Warmth robots increased customer dissatisfaction during failures. Humanoids could effectively recover trust with sincere apologies or explanations |
| Seeger and Heinzl (2021) | Theory of anthropomorphism | Experimental design | Interaction with a human-like chatbot compared to a machine-like chatbot considerably decreased customer trust |
| Crolic et al. (2022) | Functionalist theory of emotion | Text analysis, experimental design | Too much chatbot humanization undermined service evaluation |
| Filieri et al. (2022) | Not specified | Text analysis | Customers’ emotions in chatbot interaction included positive emotions such as joy, surprise, interest and excitement. Robots malfunction decreased satisfaction |
| Huang and Dootson (2022) | Theory of stress and coping | Experimental design | High customer participation increased emotion-focused coping (i.e. frustration, aggression) when the availability of a human assistant was disclosed early on |
| Pantano and Scarpi (2022) | Multiple intelligences, social interaction theory | Survey after interaction with AI | Visual spatial intelligence affected positive emotions, no effect on negative emotions. Social intelligence affected positive and negative emotions. Verbal intelligence did not affect any type of emotions. Processing speed only affected negative emotions |
| Brendel et al. (2023) | Frustration–aggression theory | Experimental design | Perceived humanness directly increased the frustration with the chatbot when it produced errors. Perceived humanness increased service satisfaction which in turn reduced frustration. Perceived humanness influences the nature of aggression when users become frustrated |
| Herhausen et al. (2023) | Theory of arousal | Experimental design, text analysis | High- versus low-arousal emotions reduced gratitude. Active listening and empathy in the firm response de-escalated high arousal emotions and increased gratitude. For low-arousal emotions, there were diminishing effects for active listening while the effect of empathy varied across studies |
| Liu et al. (2023) | Implicit personality theory | Experiment | Using humorous emojis by chatbots increased consumers' reuse intention through consumers' perceived intelligence |
| Zhang et al. (2023) | Lay belief and emotional competence | Experimental design | Chatbot apology led to lower customer satisfaction than symbolic recovery from human employees due to chatbots lack of emotional competence |
| Chen et al. (2025) | Benign violation theory, relief theory | Experiment | Chatbot humor and informal language increased customer perceived failure. Chatbot failures were misunderstanding, lack of competence, personalization, and assurance |
| Liang et al. (2024) | – | Field and lab experiment | Chatbot gender mattered dealing with angry customers. Male chatbots were suitable for using apology, while female chatbots were suitable when using appreciation strategy |
| Ozuem et al. (2024) | Frustration–aggression theory | Qualitative design, interview | Customers' frustration and aggression affected both customer loyalty and technology adoption |
| Tang et al. (2024) | Social affordance | Experiment | Customer anger decreased customer satisfaction and chatbot empathy mitigated such effect |
| Zhang et al. (2024) | Stress-and-coping theory | Semi-structured interviews | Chatbot failure capabilities were misunderstanding, failing to solve problems, requesting sensitive information, faking humanization. Customers emotions were anger, frustration, betrayal and defeat |
| This study | Social response theory | Text analysis | This study identified customer incivility behaviors and customer emotions in chatbot interactions, identified chatbot capability failure and recovery-oriented capability and revealed a paradoxical effect of chatbot empathy in chatbot failure contexts |
| Study | Theory | Methodology | Findings |
|---|---|---|---|
| Emotion management theory | Experimental design | Appraisals and post-recovery emotions sequentially mediated the relationship between emotion regulation and consumer word-of-mouth | |
| Not specified | Experimental design | Empathy emerged as the most effective strategy to mitigate aggressive behavior | |
| Not specified | Experimental design | Unresolved chatbot errors decreased adoption and perceived humanness | |
| Not specified | Interview | Customer resource loss determined customer coping strategies in AI-based service failure and customers passed the blame to technology for service failure | |
| Not specified | Experimental design | Warmth robots increased customer dissatisfaction during failures. Humanoids could effectively recover trust with sincere apologies or explanations | |
| Theory of anthropomorphism | Experimental design | Interaction with a human-like chatbot compared to a machine-like chatbot considerably decreased customer trust | |
| Functionalist theory of emotion | Text analysis, experimental design | Too much chatbot humanization undermined service evaluation | |
| Not specified | Text analysis | Customers’ emotions in chatbot interaction included positive emotions such as joy, surprise, interest and excitement. Robots malfunction decreased satisfaction | |
| Theory of stress and coping | Experimental design | High customer participation increased emotion-focused coping (i.e. frustration, aggression) when the availability of a human assistant was disclosed early on | |
| Multiple intelligences, social interaction theory | Survey after interaction with AI | Visual spatial intelligence affected positive emotions, no effect on negative emotions. Social intelligence affected positive and negative emotions. Verbal intelligence did not affect any type of emotions. Processing speed only affected negative emotions | |
| Frustration–aggression theory | Experimental design | Perceived humanness directly increased the frustration with the chatbot when it produced errors. Perceived humanness increased service satisfaction which in turn reduced frustration. Perceived humanness influences the nature of aggression when users become frustrated | |
| Theory of arousal | Experimental design, text analysis | High- versus low-arousal emotions reduced gratitude. Active listening and empathy in the firm response de-escalated high arousal emotions and increased gratitude. For low-arousal emotions, there were diminishing effects for active listening while the effect of empathy varied across studies | |
| Implicit personality theory | Experiment | Using humorous emojis by chatbots increased consumers' reuse intention through consumers' perceived intelligence | |
| Lay belief and emotional competence | Experimental design | Chatbot apology led to lower customer satisfaction than symbolic recovery from human employees due to chatbots lack of emotional competence | |
| Benign violation theory, relief theory | Experiment | Chatbot humor and informal language increased customer perceived failure. Chatbot failures were misunderstanding, lack of competence, personalization, and assurance | |
| – | Field and lab experiment | Chatbot gender mattered dealing with angry customers. Male chatbots were suitable for using apology, while female chatbots were suitable when using appreciation strategy | |
| Frustration–aggression theory | Qualitative design, interview | Customers' frustration and aggression affected both customer loyalty and technology adoption | |
| Social affordance | Experiment | Customer anger decreased customer satisfaction and chatbot empathy mitigated such effect | |
| Stress-and-coping theory | Semi-structured interviews | Chatbot failure capabilities were misunderstanding, failing to solve problems, requesting sensitive information, faking humanization. Customers emotions were anger, frustration, betrayal and defeat | |
| This study | Social response theory | Text analysis | This study identified customer incivility behaviors and customer emotions in chatbot interactions, identified chatbot capability failure and recovery-oriented capability and revealed a paradoxical effect of chatbot empathy in chatbot failure contexts |
The first stream examines how chatbot anthropomorphism shapes customers' emotional and behavioral responses. This research suggests that anthropomorphic cues can create both beneficial and problematic effects. On the one hand, human-like chatbots may enhance customer engagement, trust, positive emotions and favorable behavioral responses by making the interaction feel more socially responsive and relational (Sheehan et al., 2020; Seeger and Heinzl, 2021). On the other hand, anthropomorphism can raise customers' expectations of social understanding, emotional sensitivity and competent service delivery, which may intensify frustration when the chatbot fails to meet these expectations. Consistent with this view, Crolic et al. (2022) showed that customers respond more negatively to chatbot failures when they attribute responsibility to the chatbot, experiencing stronger emotions such as anger and dissatisfaction. Similarly, Castillo et al. (2021) suggested that the absence of chatbot empathy can reduce customers' perceived value, while Choi et al. (2021) showed that customer satisfaction with humanoid chatbots depends on whether they convey warmth. Overall, this stream indicates that anthropomorphism is not inherently beneficial; rather, its effects depend on whether human-like chatbot cues are supported by appropriate emotional and relational capabilities (Pantano and Scarpi, 2022).
The second stream of research examines customer emotions in chatbot failure contexts (Herhausen et al., 2023; Huang and Dootson, 2022; Zhang et al., 2024). These studies mainly show that when chatbots fail to meet customer expectations, their emotional expressions can intensify customers' negative emotions (Table 1). Although this research advances understanding of customer–chatbot interactions, it focuses mainly on emotional responses to chatbot failure (Ozuem et al., 2024) rather than on how customer emotions emerge and change during the interaction. Related work by Filieri et al. (2022) examined customer emotions in robot interactions and found that positive emotions, such as joy, surprise and excitement, can emerge during these interactions. However, findings from robot interactions may not transfer directly to chatbot interactions because robots have a tangible physical presence that chatbots lack. This physical presence may elicit customer emotional expressions that differ from those generated in chatbot interactions (Filieri et al., 2022). Consequently, there is still limited understanding of how customers' emotions vary in type and intensity during chatbot interactions. This gap is important because negative emotions triggered by expectancy violations may be expressed through incivility behaviors, including verbal abuse and mockery (Saeed et al., 2024). While prior research has advanced knowledge of customer emotions in AI-enabled service delivery, including robots and chatbots, little is known about the dynamics of customers' emotions during chatbot interactions and how these dynamics may relate to customers' incivility.
The third stream of research has explored how linguistic cues in AI-enabled service delivery ease customers' negative emotions. Studies demonstrate that humor and emojis can reduce customer anger and frustration during chatbot interactions (Chen et al., 2025; Liu et al., 2023). For example, Tang et al. (2024) highlighted the role of empathetic language in attenuating negative customer emotions, while Brendel et al. (2023) found the opposite. Moreover, symbolic service recovery strategies such as apology and appreciation are more effective depending on the perceived gender of the chatbot. For instance, Liang et al. (2024) found that male chatbots are more effective when they use apology strategies, while female chatbots are more effective when they use appreciation strategies. Despite these valuable insights, existing research predominantly relies on qualitative interviews and experimental methods. A significant gap remains in understanding how customers express incivility in response to chatbot capability failure, and whether chatbot emotional expressions can effectively mitigate such behavior.
2.2 Social response theory
Social response theory provides a strong foundation for understanding how customers respond when chatbots fail to deliver the requested service. Social response theory argues that people often apply social rules, expectations and interaction scripts to computers and other digital platforms when those technologies display cues associated with human communication (Moon, 2000). Importantly, these responses do not require users to believe that technology is actually human. Rather, social heuristics are activated because the interaction resembles interpersonal exchange (Filieri et al., 2022). Foundational work shows that individuals often respond to technological agents with patterns typically observed in human interaction, including politeness, reciprocity and emotional reaction (Nass and Moon, 2000). More recent work has extended this logic, showing that social responses to AI-enabled agents are especially likely when technology communicates through language, turn-taking and other human-like cues (Han et al., 2023).
This perspective is highly relevant in AI-enabled service delivery because chatbots are designed to interact through conversational language and often occupy the role of a frontline service employee (Brendel et al., 2023; Han et al., 2023). As a result, customers may evaluate chatbot encounters not only in terms of technical and functional performance but also in terms of social appropriateness (Herhausen et al., 2023). When a chatbot fails to understand user intent, breaks conversational flow, offers irrelevant information or does not resolve the problem, the interaction can be interpreted as a technical, functional and social failure (Castillo et al., 2021). Prior research shows that consumers evaluate chatbots more negatively than equivalent human-provided service, and that chatbot failures can trigger frustration and anger (Chin and Yi, 2019).
Social response theory also helps explain the role of chatbot empathy. Empathic language, acknowledgment and relational cues can signal responsiveness and fairness, lowering the risk that negative emotions escalate into incivility (Tang et al., 2024). However, human-like and empathic cues do not always generate positive outcomes (Brendel et al., 2023). In AI-enabled service recovery, they can improve user reactions (Chin et al., 2020), but they can also intensify frustration when the chatbot cannot resolve the problem (Zhang et al., 2024). Thus, social response theory explains both the pathway from chatbot capability failure to incivility and the conditional role of chatbot empathy in shaping that escalation.
3. Hypothesis development
In traditional service interactions, customers evaluate both the technical quality of the service they receive and the functional quality of how that service is delivered (Grönroos, 1984). Service delivery unfolds through real-time interactions with frontline employees; therefore, these evaluations are inherently relational and emotional (Bitner et al., 1990). Customers therefore assess not only whether their problem is resolved, but also whether the chatbot appears understanding, responsive and appropriate in the interaction (Siahtiri et al., 2024). Frontline employees use cognitive and emotional capacities to interpret customer needs and respond to both the problem and the emotional tone of the interaction (Bitner et al., 1990).
These service expectations are also relevant in customer–chatbot interactions. As chatbots increasingly assume frontline service roles, customers are likely to evaluate them using similar technical and functional criteria. In other words, customers expect chatbots not only to provide accurate and effective solutions, but also to respond in ways that appear attentive and appropriate to the interaction (Castelo et al., 2023). When chatbots fail in their service capability, for example by providing inaccurate information, misinterpreting customer questions or failing to resolve the issue, they fall short of these expectations (Castillo et al., 2021; Chen et al., 2025; Zhang et al., 2024). Such chatbot capability failures interrupt customers' progress toward problem resolution and make the chatbot appear unresponsive or incapable of addressing their needs. This disruption is likely to evoke negative emotions such as frustration, anger and disappointment during interactions (Castillo et al., 2021; Chen et al., 2025). These negative emotions are important because they provide the mechanism through which chatbot capability failure is translated into behavioral responses. When customers feel frustrated or angry, they may externalize these emotions through hostile, disrespectful or aggressive language directed at the chatbot. Such behavior reflects customer incivility and a violation of norms of respectful conduct in interactions (Fullerton and Punj, 2004). Chatbot capability failure may lead to incivility because it generates negative emotions that are expressed behaviorally, not merely because the service outcome is poor. Thus, negative emotions are expected to mediate this relationship (Figure 1). Therefore:
The diagram illustrates a research model depicting the relationships between chatbot capability failure, customers' negative emotions, chatbot empathy, and customer incivility behaviors. The model starts with chatbot capability failure, which leads to customers' negative emotions. These negative emotions are then influenced by chatbot empathy, which in turn affects customer incivility behaviors. The diagram uses dashed lines to demonstrate mediation effects, indicating that customers' negative emotions mediate the relationship between chatbot capability failure and customer incivility behaviors. The model suggests that when chatbots fail to meet service expectations, it evokes negative emotions in customers, which can lead to incivility behaviors if not mitigated by chatbot empathy.Proposed research model. Source: Authors' own work
The diagram illustrates a research model depicting the relationships between chatbot capability failure, customers' negative emotions, chatbot empathy, and customer incivility behaviors. The model starts with chatbot capability failure, which leads to customers' negative emotions. These negative emotions are then influenced by chatbot empathy, which in turn affects customer incivility behaviors. The diagram uses dashed lines to demonstrate mediation effects, indicating that customers' negative emotions mediate the relationship between chatbot capability failure and customer incivility behaviors. The model suggests that when chatbots fail to meet service expectations, it evokes negative emotions in customers, which can lead to incivility behaviors if not mitigated by chatbot empathy.Proposed research model. Source: Authors' own work
Customers' negative emotions mediate the relationship between chatbot capability failure and customer incivility behaviors.
Customers' negative emotions and customer incivility are closely connected in service settings, where feelings such as frustration, anger and disappointment often increase the likelihood of rude, disrespectful or aggressive behavior (Lages et al., 2023; Sheng et al., 2025). In customer–chatbot interactions, chatbot failures can similarly evoke negative emotions, which customers may then externalize through uncivil responses toward the chatbot (Crolic et al., 2022). However, whether negative emotions escalate into incivility may depend on how the chatbot responds during the interaction and attempts to recover from the failure. Thus, chatbot empathy may serve as an important recovery-oriented capability that shapes the link between negative emotions and customers' incivility.
Chatbot empathy demonstrates how a chatbot recognizes, acknowledges and responds appropriately to the customer's emotional state (Brendel et al., 2023). Although chatbot users perceive such empathy as programmed rather than genuine, it may still send social cues in the interaction, and customers may interpret it as a signal of responsiveness, understanding and relational attentiveness (Nass and Moon, 2000). When a chatbot acknowledges a customer's frustration, expresses understanding or uses supportive language, the customer may feel heard and validated because they perceive the chatbot as a frontline employee (Juquelier et al., 2025; Kim et al., 2004). Such cues can help reduce the likelihood that negative emotions are translated into hostile or disrespectful behavior. In contrast, when empathic cues are absent or weak, customers may perceive the chatbot as indifferent or socially unresponsive, making negative emotions more likely to intensify and translate into incivility. Although human-like cues can sometimes heighten frustration when AI performance falls short, empathic acknowledgment during chatbot failure is more likely to buffer escalation because it restores a sense of interpersonal responsiveness. Thus, chatbot empathy is expected to weaken the extent to which negative emotion is externalized in the form of uncivil behavior (Figure 1). Therefore:
Chatbot empathy moderates the relationship between customers' negative emotions and customer incivility behaviors, such that the effect of customers' negative emotions on customer incivility behaviors decreases as chatbot empathy increases.
4. Methodology
4.1 Research design
We employed a two-phase methodology. Following Fisher et al. (2021), Phase 1 aimed to deepen understanding of how customers' incivility and emotional reactions emerge during customer–chatbot interactions in service failure contexts, to identify key constructs, their relative importance during interactions and to provide insights into this undertheorized domain (Fisher et al., 2021; Zeithaml et al., 2020). This phase feeds into Phase 2 to help operationalize identified constructs and transform text data into numeric data to test hypotheses.
4.2 Phase 1: data and data processing
We utilized a human-chatbot conversation transcript dataset that has been made available by the University of California. The chatbot, named ANTswers, was built and launched online in 2014 as an AI library assistant to augment library services and resolve student queries. The ANTswers application has been accumulating data since 2014. We extracted 4,438 conversations comprising 32,586 posts from 2014 to 2021, which included users' queries and answers generated by the chatbot. The collected conversation transcripts were processed to clean the textual content related to chatbot capability failure and to format for analysis in line with the study's objectives. This process employed an AI framework implemented in Python, combining natural language processing, machine learning and supervised classification to examine the customer–chatbot interaction phenomenon. Figure 2 illustrates the framework and the steps used to analyze the data.
The diagram illustrates an AI framework designed to analyze and structure chatbot-customer interactions. The process begins with data collection from chatbot-customer service interactions, followed by data processing. The AI framework consists of three main components: interaction extraction, deep emotion modeling, and misbehavior post classification. Interaction extraction involves extracting chatbot and customer interactions. Deep emotion modeling focuses on creating emotion intensity profiles. Misbehavior post classification identifies types of misbehavior such as teasing, disrespect, and offensive language. The framework utilizes natural language processing and machine learning techniques. The output of the framework includes identifying chatbot failures, customer incivility, chatbot emotions, and customer emotions.The proposed AI framework to transform chatbot–customer conversations into a structured representation of interactions. Source: Authors' own work
The diagram illustrates an AI framework designed to analyze and structure chatbot-customer interactions. The process begins with data collection from chatbot-customer service interactions, followed by data processing. The AI framework consists of three main components: interaction extraction, deep emotion modeling, and misbehavior post classification. Interaction extraction involves extracting chatbot and customer interactions. Deep emotion modeling focuses on creating emotion intensity profiles. Misbehavior post classification identifies types of misbehavior such as teasing, disrespect, and offensive language. The framework utilizes natural language processing and machine learning techniques. The output of the framework includes identifying chatbot failures, customer incivility, chatbot emotions, and customer emotions.The proposed AI framework to transform chatbot–customer conversations into a structured representation of interactions. Source: Authors' own work
To train the classification model for detecting incivility in customer–chatbot interactions, we developed a balanced training dataset of 300 annotated samples [100 samples from each category]. We first used unsupervised topic modeling (Jelodar et al., 2019) to identify key constructs, then manually checked and categorized these terms to develop a domain-specific dictionary based on Chin and Yi (2019). The dictionary was used to machine-annotate samples, with all machine-generated labels manually reviewed and validated. This balanced design ensured adequate class representation and reduced classification bias from imbalanced data (Chawla et al., 2002). Agreement between machine-assigned and human-validated labels was 93%, indicating high annotation consistency. This hybrid human-in-the-loop approach enabled efficient dataset construction while maintaining reliable label quality.
Second, we applied a sentence-embedding approach to extract semantic features from the content by converting each post into a numerical representation. Sentence embeddings captured the contextual meaning of entire sentences rather than isolated words; therefore, they provided a richer representation of each message (Conneau et al., 2018). These representations were then used as inputs to train the classification model (Delashmit and Manry, 2005). To mitigate overfitting, we employed 10-fold cross-validation during the training process (Wong and Yeh, 2020). The developed classification model yielded a validation accuracy of 94.23% and an F1 score of 84.3% in correctly classifying a post into three incivility behavior classes (Bhagat and Bakariya, 2025). This approach resulted in identifying chatbot capability failure, customers' emotions and customer incivility behaviors during customer–chatbot interactions. Afterward, we used descriptive analyses, group comparisons and one-way analysis of variance (ANOVA) to examine the interconnection among the identified phenomena.
4.3 Customer–chatbot interaction analyses
4.3.1 Extracting chatbot capability failure
To identify chatbot capability failures, we built on the prior research on employee–customer interactions in physical service environments (Grönroos, 1984) and AI-enabled service delivery, which highlights the frontline role of the chatbots and customers' expectations of human-like capabilities (Castillo et al., 2021). Thus, we examined two forms of chatbot capability failures known as technical and functional capability failures. Human-to-human service recovery often relies on interpersonal strategies such as apology, attentiveness and responsiveness (Gelbrich and Roschk, 2011); therefore, we examined whether the chatbot used any recovery-oriented capability that mirrors these interpersonal recovery efforts in AI-enabled service delivery. Accordingly, chatbot capability failures were operationalized as observable conversational indicators reflecting the chatbot's inability to resolve customer requests. These statements included requesting customers to repeat or rephrase questions (i.e. “Can you rephrase your question, please?”), inability statements (i.e. “I'm not trained”) and redirection to external services without resolution (i.e. “All of our services are available at Ask a Librarian.”). To assess the chatbot's recovery-oriented capability, we identified observable conversational indicators of empathic responses, including statements such as “Sorry you feel that way!” Table 2 presents the operationalization of chatbot capability failures and recovery-oriented capability.
Operationalization of constructs
| Variables | Representation of chatbot failures | Description/Used templates |
|---|---|---|
| Chatbot capability failure is the frequency of | ||
| Technical capability failure | Questions asked | Number of questions asked by the chatbot |
| Rephrasing behaviors | “Can you rephrase your question, please?” | |
| Repetitions | Same consecutive responses | |
| Functional capability failure | Scripted behaviors | “All of our services are available at Ask a Librarian.” |
| Referring to human | “That's a question I can't answer. Try asking a Librarian.” | |
| Recovery-oriented capability is the frequency of | ||
| Chatbot empathy | Empathetic statements | “Sorry you feel that way!” |
| Customer incivility behavior is frequency of | ||
| Disrespect language | Terms and expressions to indicate frustration and degrading the ability of the chatbot | “You are worthless.” |
| Offensive language | Use of profanity terms/obscene jokes | “what the … !” |
| Teasing language | Expressions of mockery and unrelated questions | “Will you marry me?” |
| Customer emotions are the average of | ||
| Positive emotions | Joy, Trust, Anticipation, Surprise | “Great--found what I was looking for. Thx!” |
| Negative emotions | Sadness, Anger, Disgust, Fear | “This is a terrible service!” |
| Variables | Representation of chatbot failures | Description/Used templates |
|---|---|---|
| Chatbot capability failure is the frequency of | ||
| Technical capability failure | Questions asked | Number of questions asked by the chatbot |
| Rephrasing behaviors | “Can you rephrase your question, please?” | |
| Repetitions | Same consecutive responses | |
| Functional capability failure | Scripted behaviors | “All of our services are available at Ask a Librarian.” |
| Referring to human | “That's a question I can't answer. Try asking a Librarian.” | |
| Recovery-oriented capability is the frequency of | ||
| Chatbot empathy | Empathetic statements | “Sorry you feel that way!” |
| Customer incivility behavior is frequency of | ||
| Disrespect language | Terms and expressions to indicate frustration and degrading the ability of the chatbot | “You are worthless.” |
| Offensive language | Use of profanity terms/obscene jokes | “what the … !” |
| Teasing language | Expressions of mockery and unrelated questions | “Will you marry me?” |
| Customer emotions are the average of | ||
| Positive emotions | Joy, Trust, Anticipation, Surprise | “Great--found what I was looking for. Thx!” |
| Negative emotions | Sadness, Anger, Disgust, Fear | “This is a terrible service!” |
4.3.2 Extracting customer emotions
Emotions are embedded in conversations through varied verbal expressions, examples of which are presented in Table 2. We employed the computational framework proposed by Adikari et al. (2021) to identify customer-expressed emotions using an emotion vocabulary grounded in Plutchik's (1982) eight-emotion model. Extending the valence aware dictionary and sentiment reasoner (VADER) sentiment analysis package, the framework also accounts for negations (“not happy,” “never sad”) and intensifiers (“so angry,” “very disappointed”) associated with emotional expressions. The extracted emotions were subsequently visualized using Plutchik's (1982) wheel of emotions via a Python-based visualization tool, producing emotion profiles with intensity scores ranging from 0 to 1, where 1 denoted the highest level of emotional intensity. This approach offers a more nuanced understanding of customer emotions than conventional categorical classifications and allows us to trace emotional trajectories and intensity patterns throughout the interaction.
4.3.3 Extracting customer incivility behaviors
Behaviorally, customers engaged in various categories of incivility behavior in online interactions. Most of these incivility behaviors related to verbal abuse, such as insulting, threatening and using offensive language (Chin et al., 2020; Chin and Yi, 2019). Following the dictionary suggested by Chin and Yi (2019), we explored customer incivility behaviors in customer–chatbot interactions. These behaviors are presented in Table 2.
4.4 Findings of phase 1
4.4.1 Chatbot capability failure analysis
Our analysis identified two chatbot capabilities that may fail during service delivery, namely technical capability and functional capability. These failures occurred when the chatbot misidentified intent, failed to extract key information, provided irrelevant answers or repeatedly asked customers to rephrase their questions. Functional failures concerned how the service was delivered across turns. They appeared as rigid dialog management, loss of context, slow or poorly timed responses and weak escalation that burdened customers rather than resolving their issues. In practice, technical failures produced wrong or incomplete answers, whereas functional failures created awkward, effortful interactions that felt unhelpful or uncaring. The analysis also showed that chatbots sometimes used recovery-oriented responses after failures to repair technical or functional failures by apologizing, acknowledging frustration or expressing sympathy. We conceptualized this capability as empathy.
4.4.2 Customer emotion analysis
The analysis revealed that customers responded to chatbots emotionally, as reflected in the words they used. Customers expressed different emotional patterns in response to chatbot capabilities. Customers reacted to chatbot capabilities with positive (joy, anticipation, surprise, trust) and negative emotions (anger, disgust, sadness and fear). Figure 3 shows the average emotions expressed by customers.
A radar chart displays the average intensity of various customer emotions across all conversations. The chart features eight axes, each representing a different emotion: joy, trust, fear, surprise, sadness, disgust, anger, and anticipation. The intensity of each emotion is plotted on a scale from 0 to 1. Joy has the highest intensity at 0.35, followed by anger at 0.34. Anticipation and disgust both have an intensity of 0.10. Trust is at 0.07, fear at 0.02, surprise at 0.03, and sadness at 0.09. The chart visually represents these emotions with colored areas extending from the center, indicating their relative intensities. All values are approximated.Average customer emotion intensity across all conversations. Source: Authors' own work
A radar chart displays the average intensity of various customer emotions across all conversations. The chart features eight axes, each representing a different emotion: joy, trust, fear, surprise, sadness, disgust, anger, and anticipation. The intensity of each emotion is plotted on a scale from 0 to 1. Joy has the highest intensity at 0.35, followed by anger at 0.34. Anticipation and disgust both have an intensity of 0.10. Trust is at 0.07, fear at 0.02, surprise at 0.03, and sadness at 0.09. The chart visually represents these emotions with colored areas extending from the center, indicating their relative intensities. All values are approximated.Average customer emotion intensity across all conversations. Source: Authors' own work
The analysis revealed that among all emotions, anger and joy emerged as the most prominent emotions, while trust and fear were the least prominent emotions in responding to chatbot capability failures. Findings showed that customers expressed joy (0.35) and anger (0.34) with a similar intensity. These two extreme emotions indicated the variation in emotions in customer service experiences. Anger is also a particularly prevalent emotion in customer service settings, with estimates suggesting that up to 20% of call center interactions involve hostile or angry customers. Further, we found that customers' emotional responses varied across individuals. The profiles of six individual customers were selected randomly, and their emotion trajectories and intensity were presented in Figure 4. Emotion trajectories represented customers' emotional variations over time, and the intensity profiles showed an aggregation of emotions during their service interactions with the chatbot.
The image contains three sets of graphs, each representing the emotional trajectory and profile of three different customers. Each set includes a line graph on the left and a radar chart on the right. The line graphs display emotion intensity over time for various emotions, including anger, anticipation, disgust, fear, joy, sadness, surprise, and trust. The radar charts show the emotion profiles, indicating the intensity of each emotion for the respective customer. For Customer 1, the line graph shows peaks in joy, anger, and trust around specific time units, while the radar chart indicates higher intensities in joy and trust. For Customer 2, the line graph shows multiple peaks in anger and smaller peaks in joy and trust, with the radar chart showing a prominent intensity in anger. For Customer 3, the line graph shows a significant peak in anger and smaller peaks in joy and disgust, with the radar chart indicating higher intensities in anger and disgust.
The image contains three sets of graphs, each representing the emotional trajectories and intensity profiles of three different customers during their interactions with a chatbot. Each set includes a line graph and a radar chart. The line graphs display the intensity of various emotions over time units, with different colors representing different emotions: anger, anticipation, disgust, fear, joy, sadness, surprise, and trust. The radar charts show the overall intensity of these emotions for each customer. For Customer 4, the line graph shows peaks in anger and sadness, with the radar chart indicating high anger and low sadness. For Customer 5, the line graph shows a peak in joy and trust, with the radar chart indicating moderate joy and trust. For Customer 6, the line graph shows peaks in anger, joy, and trust, with the radar chart indicating high anger, moderate joy, and low trust. The graphs illustrate the variation in emotional responses among different customers.Examples of emotional trajectory. Source: Authors' own work
The image contains three sets of graphs, each representing the emotional trajectory and profile of three different customers. Each set includes a line graph on the left and a radar chart on the right. The line graphs display emotion intensity over time for various emotions, including anger, anticipation, disgust, fear, joy, sadness, surprise, and trust. The radar charts show the emotion profiles, indicating the intensity of each emotion for the respective customer. For Customer 1, the line graph shows peaks in joy, anger, and trust around specific time units, while the radar chart indicates higher intensities in joy and trust. For Customer 2, the line graph shows multiple peaks in anger and smaller peaks in joy and trust, with the radar chart showing a prominent intensity in anger. For Customer 3, the line graph shows a significant peak in anger and smaller peaks in joy and disgust, with the radar chart indicating higher intensities in anger and disgust.
The image contains three sets of graphs, each representing the emotional trajectories and intensity profiles of three different customers during their interactions with a chatbot. Each set includes a line graph and a radar chart. The line graphs display the intensity of various emotions over time units, with different colors representing different emotions: anger, anticipation, disgust, fear, joy, sadness, surprise, and trust. The radar charts show the overall intensity of these emotions for each customer. For Customer 4, the line graph shows peaks in anger and sadness, with the radar chart indicating high anger and low sadness. For Customer 5, the line graph shows a peak in joy and trust, with the radar chart indicating moderate joy and trust. For Customer 6, the line graph shows peaks in anger, joy, and trust, with the radar chart indicating high anger, moderate joy, and low trust. The graphs illustrate the variation in emotional responses among different customers.Examples of emotional trajectory. Source: Authors' own work
As shown in Figure 4, some customers expressed high levels of joy, while others experienced intense anger, which can lead to more aggressive behavior toward the chatbot. For instance, Customer 1 began the service interaction with a sense of joy, transitioned to sadness and anger and ultimately concluded the conversation with an expression of trust. These positive emotions appeared to be due to the chatbot's service success in providing satisfactory responses to the customers' queries. In contrast, Customer 2 experienced a different emotional progression, starting with anger, transitioning to joy and trust, but ultimately reverting to the same level of anger at the conclusion of the interaction. This trajectory might reflect either the satisfactory solution delivered by the chatbot or the intensity of the customer's initial negative emotion, which may have been too strong for the chatbot to effectively mitigate during the interaction. Customer 6 presented another unique pattern, beginning the service interaction with anger and transitioning into trust and joy, concluding the interaction with positive emotions. A review of the conversations showed that in such instances, the chatbot provided accurate and satisfying answers to the customers' queries.
4.4.3 Customer incivility behaviors analysis
Figure 5 shows the classification model results for customer incivility behaviors. Disrespectful behavior was most frequent (40.07%), including dismissive remarks, aggressive demands and derogatory language that undermined the chatbot's authority, such as “You're useless” or “Get me a real human.” Offensive expressions accounted for 34.46% and included explicit hostility, profanity and obscene jokes. Teasing accounted for 25.47% and involved mocking programmed responses, testing the chatbot's boundaries with sarcasm, or confusing it with nonsensical queries, such as “Are you male or female?”
The horizontal bar graph compares three categories of customer incivility behaviors: Teasing, Offensive, and Disrespect. The x-axis represents the percentage of posts, ranging from 0% to 45%. The y-axis lists the three categories. The bars are colored differently: orange for Teasing, blue for Offensive, and green for Disrespect. Teasing accounts for 25.47% of the posts, Offensive expressions account for 34.46%, and Disrespectful behavior accounts for 40.07%. The graph highlights that Disrespectful behavior is the most frequent, followed by Offensive expressions and then Teasing. All values are approximated.Classification of customer incivility behaviors. Source: Authors' own work
The horizontal bar graph compares three categories of customer incivility behaviors: Teasing, Offensive, and Disrespect. The x-axis represents the percentage of posts, ranging from 0% to 45%. The y-axis lists the three categories. The bars are colored differently: orange for Teasing, blue for Offensive, and green for Disrespect. Teasing accounts for 25.47% of the posts, Offensive expressions account for 34.46%, and Disrespectful behavior accounts for 40.07%. The graph highlights that Disrespectful behavior is the most frequent, followed by Offensive expressions and then Teasing. All values are approximated.Classification of customer incivility behaviors. Source: Authors' own work
As summarized in Figure 6, one-way ANOVA results showed significant differences in anger and joy across incivility behavior groups (p < 0.05). Post-hoc t-tests indicated that customers with higher anger and lower joy were more likely to display incivility, particularly disrespectful and offensive language. These findings showed that incivility was closely linked to emotional arousal arising from unmet needs during failed chatbot interactions, with customers using such behaviors to express frustration, release negative affect or assert dominance.
The image contains three radar charts side by side, each representing different emotions expressed in various incivility behavior categories. Each chart has eight axes labeled with different emotions: joy, trust, anticipation, anger, fear, surprise, disgust, and sadness. The values on each axis range from 0 to 1. The first radar chart shows high anger with a value of 0.83 and low anticipation with a value of 0.03. The second radar chart indicates significant anger with a value of 0.59, moderate joy with a value of 0.24, and some sadness with a value of 0.15. The third radar chart displays high joy with a value of 0.64, moderate anger with a value of 0.17, and slight sadness with a value of 0.12. The charts illustrate how different emotions vary across incivility behavior categories, highlighting that higher anger and lower joy are associated with more incivility, particularly disrespectful and offensive language. All values are approximated.Emotions expressed in different incivility behaviors categories. Source: Authors' own work
The image contains three radar charts side by side, each representing different emotions expressed in various incivility behavior categories. Each chart has eight axes labeled with different emotions: joy, trust, anticipation, anger, fear, surprise, disgust, and sadness. The values on each axis range from 0 to 1. The first radar chart shows high anger with a value of 0.83 and low anticipation with a value of 0.03. The second radar chart indicates significant anger with a value of 0.59, moderate joy with a value of 0.24, and some sadness with a value of 0.15. The third radar chart displays high joy with a value of 0.64, moderate anger with a value of 0.17, and slight sadness with a value of 0.12. The charts illustrate how different emotions vary across incivility behavior categories, highlighting that higher anger and lower joy are associated with more incivility, particularly disrespectful and offensive language. All values are approximated.Emotions expressed in different incivility behaviors categories. Source: Authors' own work
Disrespectful behavior accounted for the largest share of incivility incidents in chatbot interactions. Anger intensity was significantly higher for disrespectful behavior (p < 0.05, β = 0.83) than for offensive language (p < 0.05, β = 0.59) and teasing (p < 0.05, β = 0.17). These findings suggested that when chatbots provided irrelevant or repetitive responses, failed to handle customer requests or de-escalate tense interactions, customers experienced stronger negative emotions that might be expressed as sarcasm, scolding, dismissive remarks or other disrespectful behavior.
Offensive language was another form of customers' incivility. It was associated with moderate anger (p < 0.05, β = 0.59) and lower joy (p < 0.05, β = 0.24), suggesting that offensive expressions might be driven by both frustration and amusement. This finding aligns with research showing that some users perceive chatbots as non-human entities and therefore test their limits or use inappropriate language without the social repercussions typically associated with human interactions (De Angeli and Brahnam, 2008).
Teasing was the final form of incivility identified in the analysis. It included behaviors such as questioning the chatbot's gender or asking about love and relationships. Teasing was associated with both anger (p < 0.05, β = 0.17) and joy (p < 0.05, β = 0.64), although the difference between these effects was not significant (p > 0.05). This pattern showed that teasing may reflect playful mockery, boundary testing or a less confrontational expression of negative emotion during interactions.
5. Phase 2: hypothesis testing
5.1 Data preparation and analysis method
To test the hypotheses, we transformed unstructured text data into a structured set of variables suitable for quantitative analysis. Following Phase 1, the chatbot transcripts were converted into an interaction-level dataset, with each complete chatbot interaction treated as one observation. The categories identified during Phase 1 were then operationalized as numerical variables. Customer emotion was represented as a continuous value between 0 and 1, which was calculated as the intensity of the expression using the VADER sentiment tool (Adikari et al., 2021). Customers experienced a range of negative emotions; thus, we calculated the average level of negative emotion for each interaction (Table 2). Chatbot capability failure was operationalized as failures in the chatbot's technical and functional capabilities, as reflected in texts involving requests for repetition or rephrasing and the frequency of repeated questions (Table 2). Chatbot empathy, as recovery-oriented capability, was operationalized using expressions such as “Sorry you feel that way” and “I am sorry.” Customer incivility behaviors were operationalized by identifying posts indicating disrespect, offensive language and teasing. Chatbot capability failure, chatbot empathy and customer incivility behaviors were all modeled as count variables, representing the number of times each construct occurred within an interaction. The final dataset therefore contained one row per interaction and separate columns for each construct.
For data analysis, we employed PROCESS Macro Model 4 with 10,000 bootstrap resamples, and bias-corrected 95% confidence intervals to examine mediation effects (Hayes, 2013). This analysis method provides robust estimates of the indirect effects of chatbot failure on customer incivility behaviors within the chatbot environment. To test Hypothesis 2, we used Model 1 in PROCESS (Hayes, 2013) and employed floodlight analysis using the Johnson–Neyman technique to identify the turning point at which the effect of the focal predictor on the outcome became significant across values of the moderator (Johnson and Neyman, 1936).
5.2 Result of hypothesis testing
Hypothesis 1 stated that customers' negative emotions mediate the relationship between chatbot capability failure and customers' incivility. As shown in Table 3, chatbot capability failure had a significant positive effect on customer incivility behaviors (β = 0.0917, t = 21.124, p < 0.001, 95% CI [0.0832, 0.1003]). Moreover, customers' negative emotions exerted a significant positive effect on customer incivility behaviors (β = 0.593, t = 57.15, p < 0.001, 95% CI [0.572, 0.613]). The bootstrapped indirect effect was significant, with a 95% confidence interval ranging from 0.0351 to 0.075. After accounting for the mediator, the direct effect of chatbot capability failure on customer incivility behaviors became non-significant, indicating full mediation. These results suggested that greater chatbot capability failure increased customers' negative emotions, which in turn led to higher levels of customer incivility. Overall, these findings supported Hypothesis 1 and answered RQ1.
Hypothesis results
| Customer incivility behaviors | |||||
|---|---|---|---|---|---|
| Variables | b | t-value | p | LLCI | ULCI |
| Mediation effect | |||||
| Chatbot failure | 0.091 | 21.124 | 0.001 | 0.083 | 0.100 |
| Customer negative emotions | 0.593 | 57.15 | 0.001 | 0.572 | 0.613 |
| Direct effect after mediator | 0.003 | 1.08 | 0.27 | −0.002 | 0.009 |
| Indirect effect | 0.054 | 0.0351 | 0.075 | ||
| Moderation effect | |||||
| Customer negative emotions × Chatbot empathy | 0.155 | 11.27 | 0.00 | 0.128 | 0.181 |
| Customer incivility behaviors | |||||
|---|---|---|---|---|---|
| Variables | b | t-value | p | LLCI | ULCI |
| Mediation effect | |||||
| Chatbot failure | 0.091 | 21.124 | 0.001 | 0.083 | 0.100 |
| Customer negative emotions | 0.593 | 57.15 | 0.001 | 0.572 | 0.613 |
| Direct effect after mediator | 0.003 | 1.08 | 0.27 | −0.002 | 0.009 |
| Indirect effect | 0.054 | 0.0351 | 0.075 | ||
| Moderation effect | |||||
| Customer negative emotions × Chatbot empathy | 0.155 | 11.27 | 0.00 | 0.128 | 0.181 |
Note(s): 10,000 bootstrap samples, 95% confidence interval
Hypothesis 2 posited that the effect of customers' negative emotions on customer incivility behaviors decreases as chatbot empathy increases. The results of this analysis addressed RQ2. As shown in Table 3, chatbot empathy significantly moderated the relationship between customers' negative emotions and customer incivility behaviors (β = 0.155, t = 11.27, p < 0.001). The floodlight analysis presented in Figure 7 showed that the moderating effect of chatbot empathy was positive and significant across all observed values of chatbot empathy. Specifically, the positive relationship between customers' negative emotions and customer incivility behaviors became stronger as chatbot empathy increased. Contrary to our prediction, higher levels of chatbot empathy exacerbated rather than attenuated the effect of customers' negative emotions on incivility directed toward the chatbot. Thus, Hypothesis 2 was not supported.
A line graph illustrates the conditional effect of customer negative emotions on customer incivility as chatbot empathy increases. The x-axis represents chatbot empathy, ranging from 0.0 to 6.5. The y-axis represents the conditional effect of customer negative emotions on customer incivility, ranging from 0.0 to 0.3. The graph includes four lines: the conditional effect line in blue, the lower 95% confidence interval line in orange, the upper 95% confidence interval line in green, and a zero line in light blue. The data shows that as chatbot empathy increases, the conditional effect of customer negative emotions on customer incivility also increases. The lines for the lower and upper 95% confidence intervals indicate the range of this effect. All values are approximated.Moderation plot. Source: Authors' own work
A line graph illustrates the conditional effect of customer negative emotions on customer incivility as chatbot empathy increases. The x-axis represents chatbot empathy, ranging from 0.0 to 6.5. The y-axis represents the conditional effect of customer negative emotions on customer incivility, ranging from 0.0 to 0.3. The graph includes four lines: the conditional effect line in blue, the lower 95% confidence interval line in orange, the upper 95% confidence interval line in green, and a zero line in light blue. The data shows that as chatbot empathy increases, the conditional effect of customer negative emotions on customer incivility also increases. The lines for the lower and upper 95% confidence intervals indicate the range of this effect. All values are approximated.Moderation plot. Source: Authors' own work
5.3 Further analysis
We further employed PROCESS Model 14 to test the full model and assess whether chatbot empathy moderates the indirect effect of chatbot capability failure on customer incivility behaviors through customers' negative emotions. The results supported this mediated moderation model, with a significant bootstrapped indirect effect (95% CI [0.007, 0.033]). These findings provided additional support for the robustness of the proposed model and show that chatbot empathy exacerbates the indirect effect of chatbot capability failure on customers' incivility.
6. Discussion and contributions
This study sets out to explain customer–chatbot interactions; specifically, how chatbot capability failures escalate into customers' incivility in AI-enabled service delivery, and whether chatbot empathy can de-escalate that process. Phase 1 identified chatbot capability failure as technical and functional, and empathy as a recovery-oriented capability. It further revealed that customers experience different types of emotions with varied intensity throughout interactions with the chatbot, and that emotional responses to chatbot capability failure were dynamic rather than static. Moreover, customers' incivility took multiple forms, including disrespect, offensive language and teasing, each emerging distinctly across individuals and interactions.
Phase 2 explained the mechanism underlying customer–chatbot interactions in AI-enabled service delivery. The full mediation finding revealed that chatbot capability failure drove incivility primarily by generating negative emotions rather than through a direct effect. More importantly, rather than mitigating this process, chatbot empathy amplified the indirect effect of chatbot capability failure on incivility through negative emotion, producing an empathy backfire effect.
6.1 Theoretical contributions
This study offers several theoretical implications for research on AI-enabled service delivery (Brendel et al., 2023; Han et al., 2023), customers' incivility (Lages et al., 2023) and social response theory (Filieri et al., 2022; Nass and Moon, 2000). First, this study contributes to research on AI-enabled service delivery by explaining that customer–chatbot interactions are emotionally charged social exchanges rather than purely informational transactions. While prior work has established that chatbot capability failures can evoke negative reactions (Han et al., 2023), we extend this logic by explaining how chatbot capability failure generates emotional and behavioral responses consistent with interpersonal norm violations. Further, we advance this line of research by explaining that emotional responses to chatbot capability failure are processual rather than static. It appears in our dataset that some customers move from positive to negative emotions as the interaction deteriorates, while others begin with strong negative emotions that either recover or intensify depending on how the chatbot responds. This suggests that chatbot service interactions cannot be reduced to a single emotional state at a fixed point in time. Rather, customers' emotions shift from positive to negative throughout the interaction, and these sequences help explain why similar chatbot capability failures can produce markedly different behavioral outcomes.
Second, this study contributes to customers' incivility research by demonstrating that AI-enabled service agents, such as chatbots, can be targets of norm-violating behavior, thereby expanding the conceptual boundaries of incivility beyond traditional human-to-human service interactions. This study complements and confirms prior work (i.e. Saeed et al., 2024) showing that, although chatbots are technological artifacts, customers apply human-like expectations of responsiveness, understanding and relational appropriateness, while simultaneously expecting technical competence. When these expectations are violated, they may escalate into negative emotions and incivility, such as the use of disrespectful, offensive and teasing language. One potential reason customers engage in incivility behavior when interacting with chatbots is the likelihood of fewer social consequences.
Third, this study advances research on the humanization paradox in AI interactions (Brendel et al., 2023) by investigating chatbot empathetic responses, deployed as a recovery-oriented capability when a chatbot fails. Contrary to our initial theorization, chatbot empathy exacerbates the progression from customers' negative emotions to incivility. This finding supports and extends recent research showing that AI-human-like cues can backfire by creating unmet expectations (Brendel et al., 2023; Han et al., 2023; Juquelier et al., 2025; Zhang et al., 2024). Social response theory suggests that customers apply human social scripts to technologies that communicate in socially meaningful ways. Our study challenges this assumption within social response theory (Nass and Moon, 2000; Filieri et al., 2022) by showing that human-like cues do not always improve AI-enabled service delivery. Instead, empathetic cues may exacerbate negative feelings when not paired with effective problem resolution. This counterintuitive result suggests an empathy backfire effect. When chatbots offer apologies, acknowledgment or caring language but still fail to solve the issue, customers may interpret those responses as scripted, hollow or patronizing. Under such conditions, empathy may intensify the mismatch between relational tone and functional incompetence, increasing frustration and making uncivil responses more likely to happen. This counterintuitive result supports and extends recent research suggesting that humanness and emotional expression in AI are double-edged, producing positive outcomes when matched with competence but adverse outcomes when they generate expectations the agent cannot fulfill (Brendel et al., 2023; Han et al., 2023; Juquelier et al., 2025; Zhang et al., 2024).
Finally, this study advances phenomenon-based research (Fisher et al., 2021; Zeithaml et al., 2020) by grounding its findings in naturally occurring conversational data rather than scenario-based experiments. Instead of assuming which chatbot capabilities drive failure or recovery, this study identifies the capabilities that appear in real customer–chatbot interactions. These include technical and functional capabilities, as well as the chatbot's recovery-oriented capability, however rudimentary, when attempting to recover from chatbot capability failure. By analyzing real-world service interactions, the study shows how these capabilities shape customer emotional and behavioral responses and may help explain conflicting findings on chatbot emotional expression. This approach is theoretically meaningful because it captures the fluidity, intensity and heterogeneity of chatbot capabilities, customers' emotions and incivility behavior as they unfold in live encounters, offering a richer and more grounded account of digital service failure.
6.2 Managerial implications
Our findings offer several implications for managers designing and deploying chatbots in frontline service settings. First, service managers should not assume that adding empathetic language to chatbot scripts will automatically calm customers when a chatbot fails. The findings suggest that empathetic responses can backfire when they are not supported by competent problem resolution. Empathy is particularly risky when the chatbot apologizes, acknowledges frustration or expresses sympathy, but then continues to misunderstand the request, provides irrelevant answers, repeats scripted responses or fails to resolve the problem. Managers should therefore avoid treating empathy as a cosmetic layer that can compensate for weak system performance. Instead, empathy should be used selectively and paired with concrete recovery actions, such as accurate answers, clear next steps, meaningful clarification or rapid escalation to a human agent. In practice, chatbot empathic behavior does not compensate for a lack of functional competence.
Second, service firms should prioritize investment in core chatbot capabilities before expanding human-like or relational features. Our findings show that technical and functional failures trigger negative emotions and subsequent incivility, suggesting that emotional cues should not compensate for poor service performance. Features such as apologies, sympathy or emotional acknowledgment should be introduced only when chatbots can accurately interpret customer requests, maintain conversational context, provide relevant responses and escalate unresolved issues appropriately. This is because customers primarily expect their problems to be resolved, not merely acknowledged. Managers should therefore assess chatbot performance not only through containment rates or response speed, but also by examining whether chatbots reduce or create friction in the service delivery process.
Third, managers should recognize that even brief customer–chatbot interactions can evoke emotions that increase the risk of uncivil behavior. Addressing customers' incivility toward chatbots is critically important as it signals service process breakdowns, increases recovery complexity and cost, burdens escalation systems and may spill over into broader judgments of the firm and its services (Castelo et al., 2023; Herhausen et al., 2023). Rather than dismissing chatbot-directed incivility because the target is non-human, firms should treat it as a diagnostic signal of service system breakdown. Repeated disrespect, offensive language and teasing can reveal unresolved pain points, chatbot-induced frustration and vulnerable points in the service journey. Monitoring such behaviors can help firms identify failure hotspots, refine escalation rules, prioritize system redesign and detect early signs of service quality deterioration.
Fourth, service managers should differentiate among types of incivility when designing response protocols. Since disrespect, offensive language and teasing reflect distinct emotional profiles, a uniform response is unlikely to be effective. Disrespectful or offensive interactions may require immediate escalation, brief corrective messaging or reduced conversational loops. Teasing, by contrast, may signal experimentation rather than serious dissatisfaction and may require lighter intervention unless it recurs with failure cues. Managers should therefore build detection systems that distinguish escalated misconduct from playful provocation, enabling more precise and efficient service recovery.
Finally, these findings imply that chatbot governance should include explicit safeguards for emotional escalation. Firms should establish escalation thresholds based on emotional deterioration within the conversation. For example, repeated expressions of anger, repeated clarification cycles or repeated failed empathetic statements should trigger a handoff to a human employee. Such safeguards can reduce interaction length, limit further customer frustration and protect overall perceptions of service quality.
7. Limitations and future research directions
This study offers novel insights into the emotional dynamics and customer incivility behaviors in customer–chatbot interactions when the chatbot fails. However, the limitations of this study provide avenues for future research. First, the dataset used in this study was derived from interactions with a single chatbot application, ANTswers. While the large sample size and naturalistic setting enhance the ecological validity of our findings, the context may limit generalizability. Customers in academic or informational service settings may differ in emotional expectations and behavioral norms from those engaging with chatbots in retail, banking, healthcare or travel, for example. Future studies could replicate the findings across diverse service sectors and chatbot functionalities to assess the boundary conditions of the dynamics between emotion and customer incivility. Second, the emotion detection and classification of incivility relied on machine learning and natural language processing techniques, which, although robust, are inherently constrained by the quality of the training data and the interpretive limitations of current sentiment analysis algorithms. For instance, sarcasm, cultural nuances and indirect forms of aggression may not be fully captured. Future work could integrate multimodal data, such as vocal tone in voice bots and biometric cues, to improve emotion and behavior classification. Third, this study did not account for the perceived sincerity or contextual appropriateness of such responses. As our findings suggest a backfire effect of empathy, future research should examine how customers perceive chatbot empathy, distinguishing between mechanical expressions and adaptive emotional intelligence, and how this perception influences emotional regulation and behavior. Perhaps the logic of the uncanny valley effect and the closely related AI sycophancy can shed more light on this domain.
Future research could also examine whether future generations develop a higher tolerance for the uncanny valley due to lifelong exposure to AI chatbots. These are questions that require further research. Fourth, a longitudinal design could help confirm causal pathways and show emotional change over time. Future research could design real-time intervention studies that manipulate chatbot behaviors (e.g. varying levels of empathy, apology and escalation strategies) and track their effects on customers' emotion and incivility across multiple encounters. Finally, this study focused primarily on customer-facing outcomes. Yet, chatbot incivility also has upstream effects on service firms, including damage to brand equity, reduced customer trust and satisfaction with the service or increased costs due to human intervention (see Huang and Dootson, 2022). Future research could investigate how emotional interactions in AI-enabled service interactions affect broader firm-level outcomes and how service firms can develop resilient technological and organizational strategies to manage the emotional outcomes of chatbot failures in AI-empowered service encounters.

