This study aims to investigate how anthropomorphic design (appearance and behavior) and interactivity (viewer–streamer and viewer–viewer) among artificial intelligence (AI)-driven virtual streamers influence consumer purchase intention in tourism and hospitality e-commerce live streaming (THCLS), which is mediated by social presence, telepresence and continuous watching intention, while accounting for demographic and behavioral heterogeneities.
An integrated model was developed and tested using survey data from 747 consumers experienced in watching virtual streamers in the THCLS context. A hybrid approach was used, combining partial least squares structural equation modeling (PLS-SEM) for linear relationships, multigroup analysis (MGA) for subgroup differences and artificial neural networks (ANNs) for nonlinear effects.
Anthropomorphic cues and interactivity significantly enhance social presence and telepresence, which in turn positively influence continuous watching intention and purchase intention. MGA reveals meaningful moderating effects of gender, age and viewing frequency, indicating heterogeneous responses to anthropomorphism and interactivity. ANN analysis further reveals nonlinear relationships, identifying social presence as the most influential predictor of purchase intention beyond linear rankings.
Practitioners should prioritize anthropomorphic appearance and behavior alongside real-time interaction to strengthen social presence, the primary driver of purchase intention. The development of community-oriented features that foster viewer–viewer engagement and the tailoring of strategies based on age and viewing frequency can enhance effectiveness. Continuous updates to content and interaction design are essential to mitigate novelty decay and sustain conversion performance.
This research extends presence-based theories and the stimulus–organism–response framework to AI-powered THCLS, uncovering dynamic linear, heterogeneous and nonlinear mechanisms in consumer decision-making, thereby addressing gaps in the virtual streamer literature and offering a novel hybrid analytical lens for THCLS.
1. Introduction
The digitalization of the tourism and hospitality industry has fundamentally reshaped the ways in which destinations and tourism offerings are promoted, moving beyond traditional static online communication toward more interactive, personalized and real-time forms of engagement (Ben Saad, 2024; Hua et al., 2023; Liu et al., 2025; Shao and Huang, 2025). Within this evolving digital ecosystem, live streaming has become an increasingly influential marketing and communication approach, allowing tourism organizations to visually demonstrate experiential value and establish instant connections with prospective travelers (Liang et al., 2024; Wang et al., 2026). Through immersive and interactive encounters that simulate aspects of firsthand experience before actual purchase, live streaming helps consumers evaluate tourism products more confidently and alleviates uncertainty during travel-related decision-making (Liang et al., 2024). More recently, another technological transformation has further disrupted this domain: the emergence of virtual streamers – artificial intelligence (AI)-powered or human-operated digital avatars capable of replacing conventional human hosts in live-streaming environments (Foroudi et al., 2025; Shao and Huang, 2025).
The global AI-powered virtual streamer live-streaming market is projected to experience remarkable expansion, with its value anticipated to rise from US$5.85bn in 2025 to US$35.83bn by 2031, reflecting a compound annual growth rate (CAGR) of 35.3% (LP Information, 2025). Comparable upward trends have been observed in key markets such as the USA, China and Europe, indicating the accelerating commercialization and worldwide adoption of virtual streamer technologies (LP Information, 2025). In China, a global frontrunner in live-streaming innovation, the virtual human (digital human) industry reached a market scale of over 36 billion yuan in 2023 and is expected to maintain rapid growth in the foreseeable future, driven by expanding applications across sectors including tourism marketing and promotion (Statista, 2025). Notably, virtual streamers have achieved impressive outcomes in mainstream live-commerce contexts and have, in certain cases, surpassed human influencers in commercial effectiveness, highlighting their potential as an emerging communication tool for experience-intensive industries such as tourism and hospitality (Liang et al., 2024; Wang and Zhang, 2025).
Virtual streamers offer distinct operational advantages for hospitality businesses, including 24/7 availability, consistent brand representation and novelty appeal (Truong and Chen, 2025; Wang and Zhang, 2025; Zheng et al., 2026). Despite these advantages, the deployment of virtual streamers in tourism and hospitality presents a fundamental challenge – referred to as the “reality gap.” Unlike tangible retail products, tourism and hospitality offerings require a high degree of trust, emotional engagement, and experiential imagination. The artificial nature of virtual avatars raises concerns regarding their ability to evoke social presence (i.e. the sense of being with others) and telepresence (i.e. the sense of being immersed in a destination), both of which are essential for influencing booking intentions (Fu et al., 2024; Gao et al., 2023). Consequently, understanding how virtual streamers can effectively bridge this gap through design and interaction mechanisms has become a critical issue for both researchers and practitioners.
As shown in Table S3 of the supplementary materials, the literature suggests that this bridging process is driven primarily by two key mechanisms: anthropomorphism and interactivity (Cao et al., 2026; Chen et al., 2024; Fu et al., 2024; Hou et al., 2024). Anthropomorphism describes the process through which individuals perceive and assign human-like attributes to nonhuman entities, including both appearance-based cues and human-like behaviors. By creating a sense of familiarity and social connection, anthropomorphic features can enhance users’ perceived affinity, credibility and trust toward artificial agents (Cao et al., 2026; Chen et al., 2024; Zhang et al., 2025). Interactivity, manifested through both viewer–streamer and viewer–viewer communication, transforms passive viewing into an active and socially embedded experience (Fu et al., 2024; Hou et al., 2024). While prior studies have examined these factors in general e-commerce and live-streaming contexts (Cao et al., 2026; Chen et al., 2024; Hou et al., 2024; Wang and Zhang, 2025; Zhang et al., 2025), as well as in THCLS with human hosts (Liang et al., 2024), three specific gaps motivate the present study. First, few studies situate AI-driven virtual streamers within tourism and hospitality, where experiential and intangible products may fundamentally alter how virtual hosts operate psychologically. Second, anthropomorphism has been treated either as a unidimensional construct (Chen et al., 2024) or narrowly reduced to physical appearance cues (Zhang et al., 2025), leaving the independent roles of anthropomorphic appearance and anthropomorphic behavior unexamined. Third, viewer–streamer and viewer–viewer interactivity are rarely modeled as distinct parallel drivers alongside anthropomorphism (Fu et al., 2024; Hou et al., 2024), leaving their relative contributions to social presence, telepresence and ultimately behavioral intentions unquantified.
To address these gaps, this study develops an integrated research model to examine consumer engagement with virtual streamers in THCLS. Importantly, this study adopts a hybrid analytical framework that combines partial least squares structural equation modeling (PLS-SEM), multigroup analysis (MGA) and artificial neural networks (ANNs) to provide a more comprehensive understanding of consumer decision-making. Specifically, PLS-SEM is used to test the hypothesized causal relationships and mediating mechanisms of social presence and telepresence. MGA is used to capture heterogeneity across demographic and behavioral segments, including gender, age and viewing frequency. In addition, an ANN is incorporated to uncover nonlinear relationships and assess the relative importance of key predictors, thereby complementing the limitations of linear modeling approaches. Accordingly, this study pursues three primary objectives:
to examine how specific virtual attributes (anthropomorphic appearance and behavior) and interaction types (viewer–streamer and viewer–viewer) influence consumers’ perceptions of social presence and telepresence;
to investigate the mediating roles of these presence constructs in shaping continuous watching intention and purchase intention; and
to assess both linear and nonlinear relationships, as well as demographic heterogeneity, using a hybrid PLS–MGA–ANN approach, thereby offering actionable insights for practitioners deploying virtual streamers in THCLS.
This study makes three key theoretical contributions. First, it extends presence theory in AI-mediated tourism contexts by jointly examining social presence and telepresence as dual psychological mechanisms through which virtual streamer attributes influence consumer behavior. Second, it advances the stimulus–organism–response (S–O–R) framework by integrating anthropomorphism (appearance and behavior) and interactivity (viewer–streamer and viewer–viewer) as complementary stimulus dimensions, offering a more nuanced understanding of how virtual design and social interaction jointly shape consumer responses. Third, it contributes methodologically by introducing a hybrid PLS–MGA–ANN approach, which enables the simultaneous examination of linear relationships, subgroup heterogeneity and nonlinear effects, thereby providing a more comprehensive analytical perspective on consumer decision-making in emerging digital environments.
2. Hypothesis development
The hypotheses are developed within an integrated theoretical framework that combines the SOR paradigm (Mehrabian and Russell, 1974) with presence theory (Lombard and Ditton, 1997; Short et al., 1976). In this model, anthropomorphic design (appearance and behavior) and interactivity features (viewer–streamer and viewer–viewer) function as environmental stimuli (S) that shape the states of social presence and telepresence (O). These organism states, in turn, drive the behavioral responses of continuous watching intention and purchase intention (R). Gender, age and viewing frequency are posited as boundary conditions that moderate the strength of the proposed relationships. The subsections below elaborate each path with theoretical and empirical support, progressing logically from stimuli to organism states, behavioral outcomes and moderators.
2.1 The effect of anthropomorphic appearance
Anthropomorphic appearance pertains to the visual human likeness of virtual streamers, such as realistic facial features and body forms, which can make digital entities more relatable in THCLS (Chen et al., 2024). In virtual environments, human-like appearances reduce psychological distance, fostering a sense of connection and immersion (Cao et al., 2026). With respect to social presence, anthropomorphic designs evoke perceptions of human contact and warmth, as viewers attribute social qualities to visually familiar avatars (Truong and Chen, 2025). Similarly, for telepresence, lifelike appearances enhance the illusion of being transported to a tourism setting, such as a virtual hotel tour, by blurring digital boundaries (Chen et al., 2024; Fu et al., 2024). Prior research on live-streaming commerce has indicated that appearance-based anthropomorphic cues enhance viewers’ perceptions of presence, which in turn fosters stronger engagement (Gao et al., 2023; Sun et al., 2024). Accordingly, we formulate the following hypotheses:
Anthropomorphic appearance positively influences social presence.
Anthropomorphic appearance positively influences telepresence.
2.2 The effect of anthropomorphic behavior
While an anthropomorphic appearance supplies the visual foundation for presence, anthropomorphic behavior introduces the dynamic, responsive dimension of human likeness, completing the stimulus side of the SOR framework (Di Dalmazi et al., 2026). Anthropomorphic behavior involves attributing human-like mental states, such as intentions, desires and emotions, to virtual streamers (Cao et al., 2026). In THCLS, behaviors such as expressing enthusiasm for a destination or responding empathetically mimic human guides, enhancing relational dynamics (Fu et al., 2024). Drawing from social identity theory, such behaviors promote identification, increasing social presence through perceived personalness and sensitivity (Chen et al., 2024). With respect to telepresence, behavioral cues create immersive narratives, making viewers feel present in the streamed environment (Liang et al., 2024). Behavior-dominant anthropomorphism outperforms appearance alone in reducing distance and enhancing responses, particularly in experiential contexts such as tourism (Cao et al., 2026). Therefore, the following hypotheses are formulated:
Anthropomorphic behavior positively influences social presence.
Anthropomorphic behavior positively influences telepresence.
2.3 The effect of viewer–streamer interactivity
Beyond anthropomorphic stimuli, interactivity constitutes a second major stimulus pathway. Operating through direct dialog between the host and audience, viewer–streamer interactivity refers to real-time, two-way communication between viewers and the virtual streamer, such as Q&A sessions or personalized recommendations in THCLS (Hou et al., 2024; Huang and Jiao, 2026). This interactivity enriches media cues, aligning with media richness theory to facilitate immediate feedback and build rapport (Sreejesh et al., 2020). It enhances social presence by conveying human warmth and sensitivity through direct engagement, fostering parasocial relationships (Hou et al., 2024; Liang et al., 2024). Khoi and Le (2026) demonstrate that both viewer–streamer and viewer–viewer interactions increase users’ sense of presence and the overall technology-enabled consumption experience, with these effects contingent on the facilitating role of THCLS platforms. Research on THCLS has demonstrated that gratification drives product and gift-giving intentions via telepresence and trust (Fu et al., 2024). Hence, the following hypotheses are proposed:
Viewer–streamer interactivity positively influences social presence.
Viewer–streamer interactivity positively influences telepresence.
2.4 The effect of viewer–viewer interactivity
Complementing the host-driven pathway, viewer–viewer interactivity introduces peer-to-peer social dynamics that further amplify the organism’s states of presence. Viewer–viewer interactivity involves concurrent exchanges among viewers, such as chat discussions or shared reactions in THCLS sessions (Hou et al., 2024). This peer-to-peer element creates a communal atmosphere, amplifying social cues in digital spaces (Fu et al., 2024). Grounded in the SOR framework, it boosts social presence by enabling a sense of human contact and collective warmth (Liang et al., 2024). Interactivity functions as a foundational experiential stimulus that precedes and enhances higher-order perceptual states such as telepresence by fostering users’ active cognitive and affective involvement in an online environment (Mollen and Wilson, 2010). Prior studies highlight that viewer-viewer features, alongside technological esthetics, elicit cognitive absorption and presence, leading to stronger engagement in THCLS (Hou et al., 2024; Liang et al., 2024). Accordingly, the following hypotheses are formulated:
Viewer–viewer interactivity positively influences social presence.
Viewer–viewer interactivity positively influences telepresence.
2.5 The effect of social presence
With the stimuli fully specified, attention now shifts to the states of the organism. Social presence operates as the first key mediator that translates stimuli into behavioral responses (Di Dalmazi et al., 2026; Gao et al., 2025; Zhang et al., 2025). Social presence, the perceived sense of being with others in a mediated environment, is crucial in THCLS for evoking emotional connections (Gao et al., 2025; Truong and Chen, 2025). It influences continuous watching intention by making sessions feel personally engaging and warm, encouraging repeated viewership (Hou et al., 2024). With respect to purchase intention, social presence builds trust and empathy, prompting decisions such as booking trips (Liang et al., 2024). Consistent with Gao et al. (2023), social presence functions as a critical psychological mechanism through which virtual streamer characteristics translate into behavioral outcomes, as heightened social presence significantly increases consumers’ purchase intention in THCLS. Based on presence theory, greater social presence mediates positive outcomes in interactive media, as seen in virtual gifts and impulsive intentions (Fu et al., 2024; Liang et al., 2024). Hence, the following hypotheses are formulated:
Social presence positively influences continuous watching intention.
Social presence positively influences purchase intention.
2.6 The effect of telepresence
Telepresence serves as the complementary organism state, capturing spatial immersion that works alongside social presence. Telepresence creates the sensation of being immersed in a remote environment, vital for THCLS where viewers “visit” virtual destinations (Liang et al., 2024). Ying et al. (2022) reported that greater telepresence significantly enhanced consumers’ intentions to (re)visit destinations. This relationship operated through both cognitive and affective mechanisms: telepresence strengthened cognitive responses related to educational value while simultaneously eliciting affective reactions in terms of entertainment and aesthetic appreciation. Telepresence promotes continuous watching intention by sustaining engagement through vivid, transportive experiences (Fu et al., 2024). With respect to purchase intention, telepresence enhances perceived realism, driving conversions in e-commerce (Gao et al., 2023). Empirical support from SOR-based studies shows that telepresence mediates interactions with intentions via flow and trust in tourism contexts (Fu et al., 2024; Liang et al., 2024). Hence, the following hypotheses are formulated:
Telepresence positively influences continuous watching intention.
Telepresence positively influences purchase intention.
2.7 The effect of continuous watching intention
Finally, the response stage links the states of the organism to downstream behavior. Continuous watching intention functions as a critical transitional mechanism between presence and purchase. Continuous watching intention denotes viewers’ propensity to maintain ongoing and repeated engagement with live-streaming content over time (Lv et al., 2022). Continuous watching intention is widely recognized as a critical antecedent of purchase intention in THCLS contexts, as sustained viewing deepens viewers’ trust in both the streamer and the promoted product (Yu, 2026). As an important behavioral stage within THCLS, continuous watching intention reflects heightened enjoyment, perceived information value and engagement, which facilitate the accumulation of product-related information and emotional involvement, thereby increasing purchase likelihood (Liu et al., 2023). Moreover, prolonged exposure to THCLS strengthens satisfaction and affective attachment to virtual influencers, amplifying engagement and, in turn, stimulating purchase and impulsive buying intentions (Yu et al., 2025c). Collectively, these insights indicate that continued viewing functions as a critical pathway through which consumers’ live-streaming experiences are translated into subsequent purchase-related behaviors. Accordingly, the following hypothesis is proposed:
Continuous watching intention positively influences purchase intention.
2.8 The moderating effects of gender, age and viewing frequency
To account for heterogeneity within the SOR framework, the final set of hypotheses examines whether the strength of the above relationships varies systematically across consumer segments. Consumer responses to virtual streamers and THCLS environments may vary across demographic and behavioral segments (Lv et al., 2022; Yu et al., 2025a; Yu et al., 2025b). Previous studies have indicated that gender, age and viewing frequency influence technology perceptions, social interaction preferences and media engagement patterns (Laor, 2022; Lin et al., 2017; Yu et al., 2026b). Prior research has indicated that consumers respond differently to virtual streamers depending on individual characteristics and contextual cues, with the MGA revealing significant heterogeneity in how virtual streamer attributes shape social presence, telepresence and purchase intention across different audience segments (Yu et al., 2025a). Moreover, although overall gender differences may appear subtle, evidence from THCLS suggests that certain relational and trust-based mechanisms (e.g. trust transfer and parasocial relationship formation) exhibit differential strengths across genders, implying that gender can act as a meaningful boundary condition that moderates the effects among key constructs in virtual influencer–driven THCLS contexts (Yu et al., 2025b). Both gender and age systematically condition the strength of key structural relationships in interactive digital environments, with multiple paths varying significantly between male and female users as well as between younger and older digital natives (Zhang et al., 2021). Lv et al. (2022) reported that gender, age, and viewing experience significantly moderate multiple core paths in live-streaming models – such as interest–intention and desire–purchase relationships – with stronger effects observed among female, younger and less experienced viewers, thereby providing strong theoretical justification for positing broad moderating roles of gender, age and viewing experience. Thus, the following hypotheses are formulated:
Gender moderates all the hypothesized relationships.
Age moderates all the hypothesized relationships.
Viewing frequency moderates all the hypothesized relationships.
Hence, as presented in Figure 1, the theoretical framework is introduced.
The model groups anthropomorphic appearance and anthropomorphic behaviour under anthropomorphism. It groups viewer-streamer interactivity and viewer-viewer interactivity under interactivity. These 4 factors connect to social presence and telepresence through H 1 a, H 1 b, H 2 a, H 2 b, H 3 a, H 3 b, H 4 a, and H 4 b. Social presence connects to continuous watching intention through H 5 a and to purchase intention through H 5 b. Telepresence connects to continuous watching intention through H 6 a and to purchase intention through H 6 b. Continuous watching intention connects to purchase intention through H 7. M G A lists gender, H 8, age, H 9, and frequency, H 10.Research framework
Source: Developed by authors
The model groups anthropomorphic appearance and anthropomorphic behaviour under anthropomorphism. It groups viewer-streamer interactivity and viewer-viewer interactivity under interactivity. These 4 factors connect to social presence and telepresence through H 1 a, H 1 b, H 2 a, H 2 b, H 3 a, H 3 b, H 4 a, and H 4 b. Social presence connects to continuous watching intention through H 5 a and to purchase intention through H 5 b. Telepresence connects to continuous watching intention through H 6 a and to purchase intention through H 6 b. Continuous watching intention connects to purchase intention through H 7. M G A lists gender, H 8, age, H 9, and frequency, H 10.Research framework
Source: Developed by authors
3. Research methodology
3.1 Ethical considerations and informed consent
This study followed the ethical standards of the 1964 Declaration of Helsinki and received approval from the Institutional Review Board of [anonymous] university. All the procedures complied with the relevant national and institutional guidelines. Prior to participation, the respondents were provided with clear information about the study’s purpose, voluntary nature and data usage. They were assured of anonymity and confidentiality, with no personally identifiable information collected and all the responses analyzed in aggregate form. Participants also retained the right to withdraw at any time without consequence. Informed consent was obtained electronically before the survey commenced. The data were securely stored and used solely for academic research purposes. These procedures ensured the protection of participants’ rights, privacy and well-being throughout the study.
3.2 Measurement
To empirically validate the proposed integrated model of consumer intentions in THCLS, measurement items were adapted from the established literature and contextualized for the virtual streamer setting. Guided by a positivist research paradigm, this study adopted a quantitative, deductive approach to test theoretically derived hypotheses using empirical data. A cross-sectional survey design was used, as it enables systematic measurement of latent constructs and statistical examination of their interrelationships within a real-world live-streaming context. Given the highly experiential, interactive and context-dependent nature of THCLS, particular attention was given to ensuring that the survey captured respondents’ lived experiences rather than abstract evaluations. Accordingly, respondents were explicitly instructed to recall a recent and specific virtual streamer session (e.g. a destination showcase or hotel live tour) when they answered the questionnaire, thereby grounding their responses in concrete consumption situations. All the constructs were measured using a seven-point Likert scale, ranging from “strongly disagree” (1) to “strongly agree” (7).
The survey captured key dimensions of anthropomorphism, interactivity, presence and behavioral intentions. Anthropomorphic appearance items were adapted from Cao et al. (2026) and Dabiran et al. (2024), while anthropomorphic behavior was measured using scales from Cao et al. (2026) and Waytz et al. (2010). Viewer–streamer and viewer–viewer interactivity was operationalized on the basis of Hou et al. (2024) and Ou et al. (2014). Social presence items were adapted from Gao et al. (2023), Kim and Park (2024) and Liang et al. (2024), whereas telepresence was measured using scales from Fu et al. (2024), Gao et al. (2023), and Pelet et al. (2017). Continuous watching intention was adapted from Lv et al. (2022) and Ma et al. (2025) and purchase intention was based on Shen et al. (2022). Table S1 in the supplementary materials summarizes the constructs and their sources.
To ensure content validity, a multistep adaptation and validation process was conducted. First, the items were reworded to reflect the THCLS context, such as replacing generic “user” references with “viewer” and situating interactions within virtual streamer broadcasts. In addition, several items were contextually enriched with scenario-based cues (e.g. real-time interaction, destination presentation, and audience engagement) to better reflect the dynamic and immersive characteristics of THCLS environments. Moreover, three domain experts (two academics and one live-streaming operations manager) reviewed the instrument, leading to wording refinements for clarity. A back-translation procedure was then applied to ensure conceptual equivalence (Brislin, 1980). Finally, a pilot study was conducted with 50 experienced viewers. This included cognitive interviews (n = 10) and a structured survey (n = 40). During the interviews, participants were encouraged to relate each item to their own viewing experiences. They also evaluated the realism and relevance of the questions. Based on exploratory factor analysis and qualitative feedback, ambiguous or weak items were revised or removed, resulting in a clear and reliable measurement instrument.
3.3 Data collection and sampling
A survey was conducted from September to October 2025 to investigate viewers’ intentions toward THCLS, specifically targeting individuals with prior experience watching virtual streamers in the THCLS context. To strengthen contextual validity, only respondents with recent and direct exposure to virtual streamer–led sessions were included. This ensured that the responses were based on actual behavior rather than hypothetical scenarios. The survey was administered online through a widely recognized data collection platform, Credamo. To improve representativeness, random sampling was combined with demographic quotas to approximate the distribution of typical THCLS audiences.
The survey was structured into three main sections. Initially, it outlined the study’s objectives, assured participants of confidentiality and provided a brief overview of virtual streamers in THCLS. To enhance situational immersion, participants were given short contextual prompts and examples before they answered the main questions. These prompts helped them recall realistic viewing situations. Section 2 featured eligibility screening questions, requiring respondents to be at least 18 years old and to have watched a THCLS hosted by a virtual streamer. Additional data quality controls were implemented. These included attention-check questions and minimum response-time thresholds. Such measures help identify and eliminate inattentive or superficial responses. The respondents who satisfied the eligibility requirements and successfully completed the survey received a modest participation incentive. The survey collected respondents’ demographic profiles and usage-related information, including gender, age, education level, income status and viewing frequency.
The minimum sample size was calculated a priori in accordance with established guidelines for structural model complexity, expected effect magnitude and statistical power (Cohen, 1992; Soper, 2018). A power analysis with a medium effect size (f2 = 0.15), α = 0.05 and 95% power indicated a minimum required sample size of 129 valid responses. Out of the initial pool of 912 responses, 747 questionnaires were deemed valid after the screening rules were applied. The relatively large sample size enhances the robustness of the analysis. It also improves the generalizability of the findings within the context of the THCLS. As illustrated in Table S2 of the supplementary materials, the proportion of female participants (55.02%) exceeded that of male participants (44.98%). The largest age segment was 26–35 years (28.11%), followed by the 18–25 (23.43%) and 36–45 (21.02%) age groups. The largest share of respondents held a bachelor’s degree (25.97%), with a significant portion holding a master’s degree (19.54%). Income levels were distributed mainly between RMB 6,001–9,000 (26.24%) and RMB 9,001–12,000 (21.15%). With respect to usage habits, the most common viewing frequency was three–four times each week (27.71%), followed by one–two times each week (21.82%), indicating a highly engaged sample.
3.4 Common method variance
Given that the data were collected through self-reported questionnaires, this study adopted both procedural and statistical approaches to mitigate potential common method variance (CMV). Procedurally, the questionnaire was designed to be clear and concise, demographic items were placed at the end to minimize priming effects and respondents were assured of anonymity and confidentiality to encourage accurate reporting. In addition, item randomization was applied to reduce response bias and order effects. Statistically, CMV was assessed using Harman’s single-factor test and the marker variable technique. The first factor accounted for 34.185% of the total variance, remaining below the recommended 50% threshold (Hair et al., 2010; Podsakoff et al., 2003). Following Lindell and Whitney (2001) and Williams et al. (2010), an unrelated environmental concern construct was included as a marker variable. The results showed no significant associations between the marker variable and the focal constructs, suggesting that CMV was unlikely to pose a substantial concern.
4. Data analysis and results
In this study, a hybrid multistage approach combining PLS–SEM, MGA and ANN was applied. PLS-SEM was preferred over CB-SEM because the former is predictive and exploratory. Moreover, the research model is complex, involving multiple mediators and moderators and the data exhibited a nonnormal distribution – conditions under which PLS-SEM offers greater robustness and statistical power (Foroughi et al., 2025; Hair et al., 2016; Yu et al., 2026a). After measurement invariance was established, subgroup heterogeneity across gender, age and viewing frequency was examined (Lv et al., 2022). Finally, ANN models nonlinear interactions and ranks antecedent importance through sensitivity analyses, reconciling theoretical explanations with predictive power (Foroughi et al., 2025; Shang et al., 2025).
4.1 Measurement model
The measurement model was assessed following the recommended procedures for PLS-SEM (Hair et al., 2016). As shown in Table 1, all factor loadings ranged from 0.754 to 0.906, exceeding the recommended 0.70 level and indicating satisfactory indicator reliability. Internal consistency reliability was supported by Cronbach’s alpha values (0.798–0.865) and composite reliability (CR) values (0.869–0.908), all above the 0.70 criterion. Convergent validity was also confirmed, as all AVE values ranged from 0.624 to 0.712, surpassing the 0.50 threshold (Fornell and Larcker, 1981). Discriminant validity (see Table 2) was examined using the Fornell–Larcker criterion and HTMT ratio (Henseler et al., 2015). The square roots of AVE values exceeded the correlations among constructs, and all HTMT values were below 0.85, with the highest value being 0.838 between telepresence and viewer–viewer interactivity. Overall, the measurement model demonstrated adequate reliability, convergent validity and discriminant validity, supporting its suitability for subsequent structural model analysis.
Reliability and validity
| Construct | Items | Factor loading | Cronbach’s alpha | CR | AVE |
|---|---|---|---|---|---|
| Anthropomorphic appearance | AA1 | 0.831 | 0.827 | 0.885 | 0.659 |
| AA2 | 0.785 | ||||
| AA3 | 0.819 | ||||
| AA4 | 0.811 | ||||
| Anthropomorphic behavior | AB1 | 0.844 | 0.850 | 0.899 | 0.690 |
| AB2 | 0.829 | ||||
| AB3 | 0.836 | ||||
| AB4 | 0.813 | ||||
| Viewer-streamer interactivity | VSI1 | 0.831 | 0.836 | 0.891 | 0.671 |
| VSI2 | 0.836 | ||||
| VSI3 | 0.777 | ||||
| VSI4 | 0.830 | ||||
| Viewer-viewer interactivity | VVI1 | 0.848 | 0.852 | 0.900 | 0.692 |
| VVI2 | 0.810 | ||||
| VVI3 | 0.832 | ||||
| VVI4 | 0.837 | ||||
| Social presence | SP1 | 0.763 | 0.798 | 0.869 | 0.624 |
| SP2 | 0.757 | ||||
| SP3 | 0.871 | ||||
| SP4 | 0.763 | ||||
| Telepresence | TP1 | 0.842 | 0.865 | 0.908 | 0.712 |
| TP2 | 0.858 | ||||
| TP3 | 0.838 | ||||
| TP4 | 0.837 | ||||
| Continuous watching intention | CWI1 | 0.754 | 0.852 | 0.900 | 0.693 |
| CWI2 | 0.838 | ||||
| CWI3 | 0.884 | ||||
| CWI4 | 0.849 | ||||
| Purchase intention | PI1 | 0.906 | 0.838 | 0.892 | 0.675 |
| PI2 | 0.770 | ||||
| PI3 | 0.821 | ||||
| PI4 | 0.782 |
| Construct | Items | Factor loading | Cronbach’s alpha | ||
|---|---|---|---|---|---|
| Anthropomorphic appearance | AA1 | 0.831 | 0.827 | 0.885 | 0.659 |
| AA2 | 0.785 | ||||
| AA3 | 0.819 | ||||
| AA4 | 0.811 | ||||
| Anthropomorphic behavior | AB1 | 0.844 | 0.850 | 0.899 | 0.690 |
| AB2 | 0.829 | ||||
| AB3 | 0.836 | ||||
| AB4 | 0.813 | ||||
| Viewer-streamer interactivity | VSI1 | 0.831 | 0.836 | 0.891 | 0.671 |
| VSI2 | 0.836 | ||||
| VSI3 | 0.777 | ||||
| VSI4 | 0.830 | ||||
| Viewer-viewer interactivity | VVI1 | 0.848 | 0.852 | 0.900 | 0.692 |
| VVI2 | 0.810 | ||||
| VVI3 | 0.832 | ||||
| VVI4 | 0.837 | ||||
| Social presence | SP1 | 0.763 | 0.798 | 0.869 | 0.624 |
| SP2 | 0.757 | ||||
| SP3 | 0.871 | ||||
| SP4 | 0.763 | ||||
| Telepresence | TP1 | 0.842 | 0.865 | 0.908 | 0.712 |
| TP2 | 0.858 | ||||
| TP3 | 0.838 | ||||
| TP4 | 0.837 | ||||
| Continuous watching intention | CWI1 | 0.754 | 0.852 | 0.900 | 0.693 |
| CWI2 | 0.838 | ||||
| CWI3 | 0.884 | ||||
| CWI4 | 0.849 | ||||
| Purchase intention | PI1 | 0.906 | 0.838 | 0.892 | 0.675 |
| PI2 | 0.770 | ||||
| PI3 | 0.821 | ||||
| PI4 | 0.782 |
CR = composite reliability; AVE = average variance extracted
Discriminant validity
| Fornell–larcker criterion | AA | AB | VSI | VVI | SP | TP | CWI | PI |
|---|---|---|---|---|---|---|---|---|
| Anthropomorphic appearance (AA) | 0.812 | |||||||
| Anthropomorphic behavior (AB) | 0.407 | 0.831 | ||||||
| Viewer-streamer interactivity (VSI) | 0.545 | 0.409 | 0.819 | |||||
| Viewer-viewer interactivity (VVI) | 0.566 | 0.495 | 0.534 | 0.832 | ||||
| Social presence (SP) | 0.430 | 0.375 | 0.443 | 0.440 | 0.790 | |||
| Telepresence (TP) | 0.699 | 0.502 | 0.680 | 0.719 | 0.444 | 0.844 | ||
| Continuous watching intention (CWI) | 0.242 | 0.437 | 0.290 | 0.268 | 0.344 | 0.235 | 0.833 | |
| Purchase intention (PI) | 0.282 | 0.362 | 0.298 | 0.335 | 0.388 | 0.322 | 0.369 | 0.822 |
| Heterotrait–monotrait ratio | ||||||||
| Anthropomorphic appearance (AA) | ||||||||
| Anthropomorphic behavior (AB) | 0.486 | |||||||
| Viewer-streamer interactivity (VSI) | 0.652 | 0.483 | ||||||
| Viewer-viewer interactivity (VVI) | 0.672 | 0.580 | 0.630 | |||||
| Social presence (SP) | 0.528 | 0.453 | 0.538 | 0.530 | ||||
| Telepresence (TP) | 0.824 | 0.584 | 0.796 | 0.838 | 0.533 | |||
| Continuous watching intention (CWI) | 0.291 | 0.504 | 0.340 | 0.311 | 0.407 | 0.271 | ||
| Purchase intention (PI) | 0.335 | 0.425 | 0.352 | 0.393 | 0.461 | 0.375 | 0.437 |
| Fornell–larcker criterion | ||||||||
|---|---|---|---|---|---|---|---|---|
| Anthropomorphic appearance ( | 0.812 | |||||||
| Anthropomorphic behavior ( | 0.407 | 0.831 | ||||||
| Viewer-streamer interactivity ( | 0.545 | 0.409 | 0.819 | |||||
| Viewer-viewer interactivity ( | 0.566 | 0.495 | 0.534 | 0.832 | ||||
| Social presence ( | 0.430 | 0.375 | 0.443 | 0.440 | 0.790 | |||
| Telepresence ( | 0.699 | 0.502 | 0.680 | 0.719 | 0.444 | 0.844 | ||
| Continuous watching intention ( | 0.242 | 0.437 | 0.290 | 0.268 | 0.344 | 0.235 | 0.833 | |
| Purchase intention ( | 0.282 | 0.362 | 0.298 | 0.335 | 0.388 | 0.322 | 0.369 | 0.822 |
| Heterotrait–monotrait ratio | ||||||||
| Anthropomorphic appearance ( | ||||||||
| Anthropomorphic behavior ( | 0.486 | |||||||
| Viewer-streamer interactivity ( | 0.652 | 0.483 | ||||||
| Viewer-viewer interactivity ( | 0.672 | 0.580 | 0.630 | |||||
| Social presence ( | 0.528 | 0.453 | 0.538 | 0.530 | ||||
| Telepresence ( | 0.824 | 0.584 | 0.796 | 0.838 | 0.533 | |||
| Continuous watching intention ( | 0.291 | 0.504 | 0.340 | 0.311 | 0.407 | 0.271 | ||
| Purchase intention ( | 0.335 | 0.425 | 0.352 | 0.393 | 0.461 | 0.375 | 0.437 |
Values on diagonal, which are italic, are the square root of AVE. HTMT < 0.85 (Kline, 2023)
4.2 Structural model
The structural model was evaluated through a bootstrapping procedure with 5,000 resamples to assess the significance, effect sizes and robustness of the hypothesized relationships. As reported in Table 3 and Figure 2, path coefficients, standard deviations, t-values, f2 effect sizes, variance inflation factors (VIFs) and confidence intervals were evaluated. The model demonstrated satisfactory fit and explanatory power (SRMR = 0.048), indicating an acceptable overall model fit.
The model groups anthropomorphic appearance and anthropomorphic behaviour under anthropomorphism, and viewer-streamer interactivity and viewer-viewer interactivity under interactivity. Anthropomorphic appearance connects to social presence with 0.167 and to telepresence with 0.310. Anthropomorphic behaviour connects to social presence with 0.140 and to telepresence with 0.086. Viewer-streamer interactivity connects to social presence with 0.205 and to telepresence with 0.291. Viewer-viewer interactivity connects to social presence with 0.167 and to telepresence with 0.346. Social presence, R squared 0.289, connects to continuous watching intention with 0.299 and to purchase intention with 0.229. Telepresence, R squared 0.706, connects to continuous watching intention with 0.103 and to purchase intention with 0.161. Continuous watching intention, R squared 0.127, connects to purchase intention, R squared 0.234, with 0.253.Results of path analysis
Note(s): *p < 0.05, **p < 0.01 and ***p < 0.001
Source: Developed by authors
The model groups anthropomorphic appearance and anthropomorphic behaviour under anthropomorphism, and viewer-streamer interactivity and viewer-viewer interactivity under interactivity. Anthropomorphic appearance connects to social presence with 0.167 and to telepresence with 0.310. Anthropomorphic behaviour connects to social presence with 0.140 and to telepresence with 0.086. Viewer-streamer interactivity connects to social presence with 0.205 and to telepresence with 0.291. Viewer-viewer interactivity connects to social presence with 0.167 and to telepresence with 0.346. Social presence, R squared 0.289, connects to continuous watching intention with 0.299 and to purchase intention with 0.229. Telepresence, R squared 0.706, connects to continuous watching intention with 0.103 and to purchase intention with 0.161. Continuous watching intention, R squared 0.127, connects to purchase intention, R squared 0.234, with 0.253.Results of path analysis
Note(s): *p < 0.05, **p < 0.01 and ***p < 0.001
Source: Developed by authors
Results of structural model assessment
| H | Relationship | Coefficient | SD | t-value | f2 | Confidence interval | VIF | p-value | Supported |
|---|---|---|---|---|---|---|---|---|---|
| H1a | AA → SP | 0.167 | 0.043 | 3.866 | 0.023 | 0.080, 0.252 | 1.698 | 0.000*** | YES |
| H1b | AA → TP | 0.310 | 0.034 | 9.234 | 0.192 | 0.244, 0.375 | 1.698 | 0.000*** | YES |
| H2a | AB → SP | 0.140 | 0.038 | 3.671 | 0.020 | 0.064, 0.215 | 1.398 | 0.000*** | YES |
| H2b | AB → TP | 0.086 | 0.023 | 3.797 | 0.018 | 0.043, 0.129 | 1.398 | 0.000*** | YES |
| H3a | VSI → SP | 0.205 | 0.044 | 4.712 | 0.036 | 0.120, 0.291 | 1.629 | 0.000*** | YES |
| H3b | VSI → TP | 0.291 | 0.034 | 8.503 | 0.177 | 0.224, 0.358 | 1.629 | 0.000*** | YES |
| H4a | VVI → SP | 0.167 | 0.042 | 3.941 | 0.022 | 0.083, 0.250 | 1.804 | 0.000*** | YES |
| H4b | VVI → TP | 0.346 | 0.031 | 11.120 | 0.226 | 0.285, 0.406 | 1.804 | 0.000*** | YES |
| H5a | SP → CWI | 0.299 | 0.040 | 7.426 | 0.082 | 0.220, 0.376 | 1.246 | 0.000*** | YES |
| H5b | SP → PI | 0.229 | 0.041 | 5.635 | 0.051 | 0.146, 0.307 | 1.348 | 0.000*** | YES |
| H6a | TP → CWI | 0.103 | 0.039 | 2.654 | 0.010 | 0.027, 0.178 | 1.246 | 0.008** | YES |
| H6b | TP → PI | 0.161 | 0.038 | 4.255 | 0.027 | 0.086, 0.236 | 1.258 | 0.000*** | YES |
| H7 | CWI → PI | 0.253 | 0.039 | 6.548 | 0.073 | 0.178, 0.330 | 1.145 | 0.000*** | YES |
| Model fit | Q2 | R2 | |||||||
| SRMR | 0.048 | SP | 0.176 | 0.289 | |||||
| TP | 0.499 | 0.706 | |||||||
| CWI | 0.085 | 0.127 | |||||||
| PI | 0.154 | 0.234 |
| H | Relationship | Coefficient | SD | t-value | f2 | Confidence interval | p-value | Supported | |
|---|---|---|---|---|---|---|---|---|---|
| H1a | AA → SP | 0.167 | 0.043 | 3.866 | 0.023 | 0.080, 0.252 | 1.698 | 0.000 | |
| H1b | AA → TP | 0.310 | 0.034 | 9.234 | 0.192 | 0.244, 0.375 | 1.698 | 0.000 | |
| H2a | AB → SP | 0.140 | 0.038 | 3.671 | 0.020 | 0.064, 0.215 | 1.398 | 0.000 | |
| H2b | AB → TP | 0.086 | 0.023 | 3.797 | 0.018 | 0.043, 0.129 | 1.398 | 0.000 | |
| H3a | VSI → SP | 0.205 | 0.044 | 4.712 | 0.036 | 0.120, 0.291 | 1.629 | 0.000 | |
| H3b | VSI → TP | 0.291 | 0.034 | 8.503 | 0.177 | 0.224, 0.358 | 1.629 | 0.000 | |
| H4a | VVI → SP | 0.167 | 0.042 | 3.941 | 0.022 | 0.083, 0.250 | 1.804 | 0.000 | |
| H4b | VVI → TP | 0.346 | 0.031 | 11.120 | 0.226 | 0.285, 0.406 | 1.804 | 0.000 | |
| H5a | SP → CWI | 0.299 | 0.040 | 7.426 | 0.082 | 0.220, 0.376 | 1.246 | 0.000 | |
| H5b | SP → PI | 0.229 | 0.041 | 5.635 | 0.051 | 0.146, 0.307 | 1.348 | 0.000 | |
| H6a | TP → CWI | 0.103 | 0.039 | 2.654 | 0.010 | 0.027, 0.178 | 1.246 | 0.008 | |
| H6b | TP → PI | 0.161 | 0.038 | 4.255 | 0.027 | 0.086, 0.236 | 1.258 | 0.000 | |
| H7 | CWI → PI | 0.253 | 0.039 | 6.548 | 0.073 | 0.178, 0.330 | 1.145 | 0.000 | |
| Model fit | Q2 | R2 | |||||||
| 0.048 | 0.176 | 0.289 | |||||||
| 0.499 | 0.706 | ||||||||
| 0.085 | 0.127 | ||||||||
| 0.154 | 0.234 |
AA = anthropomorphic appearance; AB = anthropomorphic behavior; VSI = viewer-streamer interactivity; VVI = viewer-viewer interactivity; SP = social presence; TP = telepresence; CWI = continuous watching intention; PI = Purchase intention; SD = standard deviation; VIF = variance information factor. *p < 0.05, **p < 0.01 and ***p < 0.001
The findings offer robust empirical support for the proposed SOR framework and presence theory. Specifically, the significant positive effects of anthropomorphic appearance and behavior on both social presence and telepresence confirm that stimulus attributes (S) shape organism states (O), which in turn drive behavioral responses (R) to continuous watching and purchase intention. Anthropomorphic features significantly influenced perceptions of presence. Specifically, anthropomorphic appearance positively affected social presence (H1a: β = 0.167, t = 3.866, p < 0.001) and telepresence (H1b: β = 0.310, t = 9.234, p < 0.001). Similarly, anthropomorphic behavior had significant positive effects on social presence (H2a: β = 0.140, t = 3.671, p < 0.001) and telepresence (H2b: β = 0.086, t = 3.797, p < 0.001). Interactivity effects were also substantial. Viewer–streamer interactivity significantly enhanced social presence (H3a: β = 0.205, t = 4.712, p < 0.001) and telepresence (H3b: β = 0.291, t = 8.503, p < 0.001). Similarly, viewer–viewer interactivity positively influenced social presence (H4a: β = 0.167, t = 3.941, p < 0.001) and telepresence (H4b: β = 0.346, t = 11.120, p < 0.001), with the latter exhibiting a substantial effect size (f2 = 0.226). With respect to the downstream effects of presence, social presence significantly drove continuous watching intention (H5a: β = 0.299, t = 7.426, p < 0.001) and purchase intention (H5b: β = 0.229, t = 5.635, p < 0.001). Telepresence also positively predicted continuous watching intention (H6a: β = 0.103, t = 2.654, p < 0.01) and purchase intention (H6b: β = 0.161, t = 4.255, p < 0.001). Finally, continuous watching intention strongly positively affected purchase intention (H7: β = 0.253, t = 6.548, p < 0.001). Overall, the model accounted for a substantial proportion of the variance in the endogenous constructs, particularly telepresence (R2 = 0.706) and social presence (R2 = 0.289). All the Q2 values were positive (SP = 0.176; TP = 0.499; CWI = 0.085; PI = 0.154), confirming the model’s predictive relevance. Multicollinearity was not a concern, as all the VIF values ranged from 1.145 to 1.804, well below the conservative threshold of 3.3.
4.3 Multigroup analysis
To examine potential subgroup differences, MGA was conducted using Henseler’s nonparametric MGA and permutation-based tests (Yu et al., 2025a). Prior to performing MGA, measurement invariance across groups was established through configural and compositional invariance assessments, ensuring the comparability of path coefficients between subgroups (Hair et al., 2016; Henseler et al., 2016). As other demographic variables did not significantly affect continuance intention, MGA was restricted to three grouping criteria: gender (male, n = 336; female, n = 411), age (≤35 years, n = 385; > 35 years, n = 362) and viewing frequency (low frequency ≤ 2 times/week, n = 395; high frequency ≥ 3 times/week, n = 352). Two complementary nonparametric approaches—Henseler’s MGA and the permutation test – were used to assess the robustness of group differences (Henseler et al., 2016) (see Table 4). First, regarding gender, significant heterogeneity was observed in the relationship between anthropomorphic behavior and telepresence (H2b). The effect was significant for females (β = 0.127, p < 0.001) but nonsignificant for males (β = 0.031, p > 0.05), with a path difference of −0.096 (p < 0.05). This gender-specific pattern extends prior findings in social commerce (Lv et al., 2022; Zhang et al., 2021) by showing that female viewers in virtual streamer contexts are particularly responsive to behavioral anthropomorphism. Second, regarding age (younger ≤ 35 vs older > 35), significant differences emerged in the outcomes of presence and intentions. The path from telepresence to continuous watching intention (H6a) was significant for younger viewers (β = 0.194, p < 0.001) but not for older viewers (β = 0.002, p > 0.05), with a significant group difference of 0.192 (p < 0.05). Furthermore, the link between continuous watching and purchase intention (H7) was significantly stronger (p < 0.01) for younger viewers (β = 0.368, p < 0.001) than for older viewers (β = 0.157, p < 0.01), with a significant group difference of 0.211 (p < 0.01). These age-related differences align with and extend digital-native research (Laor, 2022; Yu et al., 2025a). Third, regarding viewing frequency (low vs high), the effect of anthropomorphic appearance on telepresence (H1b) was significantly stronger for low-frequency viewers (β = 0.408, p < 0.001) than for high-frequency viewers (β = 0.212, p < 0.001), with a significant group difference of 0.196 (p < 0.01). This pattern suggests that the immersive impact of visual anthropomorphic cues diminishes as viewers accumulate repeated exposure. In contrast, viewer–viewer interactivity exerted a significantly stronger influence on telepresence among high-frequency viewers (β = 0.441, p < 0.001) than among low-frequency viewers (β = 0.259, p < 0.001), with a significant group difference of −0.181 (p < 0.01), indicating that frequent viewers increasingly derive immersion from social and community-based interactions rather than visual novelty. This novelty-decay effect advances prior work on repeated exposure to THCLS (Yu, 2026).
Multigroup analysis results
| H | Relationship | Path coefficient (gender) | Group 1–Group 2 | Path coefficient (age) | Group 1–Group 2 | Path coefficient (frequency) | Group 1–Group 2 | MGA results | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Male 336 | Female 411 | Coefficient difference | p-value | Younger (≤35) 385 | Older (>35) 362 | Coefficient difference | p-value | Low frequency (≤2 times/week) 395 | High frequency (≥3 times/week) 352 | Coefficient difference | p-value | |||
| H1a | AA → SP | 0.246*** | 0.105 | 0.141 | 0.113 | 0.235*** | 0.101 | 0.134 | 0.119 | 0.212*** | 0.118 | 0.094 | 0.280 | No/No/No |
| H1b | AA → TP | 0.346*** | 0.289*** | 0.056 | 0.409 | 0.295*** | 0.317*** | −0.023 | 0.735 | 0.408*** | 0.212*** | 0.196 | 0.002** | No/No/Yes |
| H2a | AB → SP | 0.146* | 0.116* | 0.031 | 0.697 | 0.180** | 0.096 | 0.083 | 0.275 | 0.109 | 0.164** | −0.055 | 0.474 | No/No/No |
| H2b | AB → TP | 0.031 | 0.127*** | −0.096 | 0.035* | 0.100** | 0.080* | 0.020 | 0.662 | 0.085* | 0.091** | −0.006 | 0.893 | Yes/No/No |
| H3a | VSI → SP | 0.142* | 0.248*** | −0.106 | 0.230 | 0.144* | 0.273*** | −0.129 | 0.138 | 0.195** | 0.217*** | −0.022 | 0.799 | No/No/No |
| H3b | VSI → TP | 0.287*** | 0.290*** | −0.003 | 0.964 | 0.243*** | 0.341*** | −0.097 | 0.153 | 0.276*** | 0.294*** | −0.019 | 0.783 | No/No/No |
| H4a | VVI → SP | 0.111 | 0.209*** | −0.098 | 0.267 | 0.127* | 0.204** | −0.077 | 0.363 | 0.208** | 0.125* | 0.084 | 0.331 | No/No/No |
| H4b | VVI → TP | 0.335*** | 0.351*** | −0.017 | 0.785 | 0.372*** | 0.320*** | 0.052 | 0.407 | 0.259*** | 0.441*** | −0.181 | 0.003** | No/No/Yes |
| H5a | SP → CWI | 0.314*** | 0.270*** | 0.044 | 0.600 | 0.302*** | 0.306*** | −0.004 | 0.957 | 0.374*** | 0.230*** | 0.144 | 0.066 | No/No/No |
| H5b | SP → PI | 0.248*** | 0.226*** | 0.022 | 0.786 | 0.190*** | 0.260*** | −0.070 | 0.389 | 0.314*** | 0.164** | 0.150 | 0.065 | No/No/No |
| H6a | TP → CWI | 0.135* | 0.049 | 0.086 | 0.281 | 0.194*** | 0.002 | 0.192 | 0.014* | 0.062 | 0.134* | −0.072 | 0.341 | No/Yes/No |
| H6b | TP → PI | 0.084 | 0.185*** | −0.101 | 0.204 | 0.096 | 0.205*** | −0.109 | 0.158 | 0.107 | 0.202*** | −0.096 | 0.213 | No/No/No |
| H7 | CWI → PI | 0.199** | 0.208*** | −0.008 | 0.928 | 0.368*** | 0.157** | 0.211 | 0.007** | 0.236*** | 0.260*** | −0.024 | 0.753 | No/Yes/No |
| H | Relationship | Path coefficient (gender) | Group 1–Group 2 | Path coefficient (age) | Group 1–Group 2 | Path coefficient (frequency) | Group 1–Group 2 | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Male 336 | Female 411 | Coefficient difference | p-value | Younger (≤35) 385 | Older (>35) 362 | Coefficient difference | p-value | Low frequency (≤2 times/week) 395 | High frequency (≥3 times/week) 352 | Coefficient difference | p-value | |||
| H1a | AA → SP | 0.246 | 0.105 | 0.141 | 0.113 | 0.235 | 0.101 | 0.134 | 0.119 | 0.212 | 0.118 | 0.094 | 0.280 | No/No/No |
| H1b | AA → TP | 0.346 | 0.289 | 0.056 | 0.409 | 0.295 | 0.317 | −0.023 | 0.735 | 0.408 | 0.212 | 0.196 | 0.002 | No/No/Yes |
| H2a | AB → SP | 0.146 | 0.116 | 0.031 | 0.697 | 0.180 | 0.096 | 0.083 | 0.275 | 0.109 | 0.164 | −0.055 | 0.474 | No/No/No |
| H2b | AB → TP | 0.031 | 0.127 | −0.096 | 0.035 | 0.100 | 0.080 | 0.020 | 0.662 | 0.085 | 0.091 | −0.006 | 0.893 | Yes/No/No |
| H3a | VSI → SP | 0.142 | 0.248 | −0.106 | 0.230 | 0.144 | 0.273 | −0.129 | 0.138 | 0.195 | 0.217 | −0.022 | 0.799 | No/No/No |
| H3b | VSI → TP | 0.287 | 0.290 | −0.003 | 0.964 | 0.243 | 0.341 | −0.097 | 0.153 | 0.276 | 0.294 | −0.019 | 0.783 | No/No/No |
| H4a | VVI → SP | 0.111 | 0.209 | −0.098 | 0.267 | 0.127 | 0.204 | −0.077 | 0.363 | 0.208 | 0.125 | 0.084 | 0.331 | No/No/No |
| H4b | VVI → TP | 0.335 | 0.351 | −0.017 | 0.785 | 0.372 | 0.320 | 0.052 | 0.407 | 0.259 | 0.441 | −0.181 | 0.003 | No/No/Yes |
| H5a | SP → CWI | 0.314 | 0.270 | 0.044 | 0.600 | 0.302 | 0.306 | −0.004 | 0.957 | 0.374 | 0.230 | 0.144 | 0.066 | No/No/No |
| H5b | SP → PI | 0.248 | 0.226 | 0.022 | 0.786 | 0.190 | 0.260 | −0.070 | 0.389 | 0.314 | 0.164 | 0.150 | 0.065 | No/No/No |
| H6a | TP → CWI | 0.135 | 0.049 | 0.086 | 0.281 | 0.194 | 0.002 | 0.192 | 0.014 | 0.062 | 0.134 | −0.072 | 0.341 | No/Yes/No |
| H6b | TP → PI | 0.084 | 0.185 | −0.101 | 0.204 | 0.096 | 0.205 | −0.109 | 0.158 | 0.107 | 0.202 | −0.096 | 0.213 | No/No/No |
| H7 | CWI → PI | 0.199 | 0.208 | −0.008 | 0.928 | 0.368 | 0.157 | 0.211 | 0.007 | 0.236 | 0.260 | −0.024 | 0.753 | No/Yes/No |
A p-value below 0.05 or above 0.95 suggests significant differences (at the 5% level) in specific path coefficients across the Two groups for the Henseler’s MGA method. *p < 0.05, **p < 0.01 and ***p < 0.001
4.4 Importance–performance map analysis
To complement the PLS-SEM results, importance–performance map analysis (IPMA) was conducted to identify managerial priorities for purchase intention (Table 5 and Figure 3). Social presence emerged as the most important antecedent (Importance = 0.315), with moderate performance (54.747), highlighting it as a primary area for improvement. Continuous watching intention was highly important (0.256) and was associated with the highest level of performance (67.227). In contrast, the telepresence (0.169) and interactivity constructs were less important for directly driving purchases, although they remain essential for indirectly enhancing purchase intention by strengthening viewers’ immersive experiences and sustained engagement throughout the live-streaming process. This IPMA ranking directly corroborates the theoretical priority of social presence within the SOR organism state.
The horizontal axis shows importance, total effects, from about 0.048 to 0.327, and the vertical axis shows performance from 0 to 100. S P has the highest importance at about 0.314 with performance about 55. C W I has importance about 0.257 and the highest performance at about 67. T P has importance about 0.170 and performance about 59. V S I has importance about 0.113 and performance about 60. A A has importance about 0.106 and performance about 55. V V I has importance about 0.109 and performance about 51. A B has the lowest importance at about 0.061 and performance about 61. The legend identifies A A, A B, C W I, S P, T P, V S I, and V V I.Importance-performance map analysis
Note(s):AA = anthropomorphic appearance; AB = anthropomorphic behavior; VSI = viewer-streamer interactivity; VVI = viewer-viewer interactivity; SP = social presence; TP = telepresence; CWI = continuous watching intention; PI = purchase intention
Source: Developed by authors
The horizontal axis shows importance, total effects, from about 0.048 to 0.327, and the vertical axis shows performance from 0 to 100. S P has the highest importance at about 0.314 with performance about 55. C W I has importance about 0.257 and the highest performance at about 67. T P has importance about 0.170 and performance about 59. V S I has importance about 0.113 and performance about 60. A A has importance about 0.106 and performance about 55. V V I has importance about 0.109 and performance about 51. A B has the lowest importance at about 0.061 and performance about 61. The legend identifies A A, A B, C W I, S P, T P, V S I, and V V I.Importance-performance map analysis
Note(s):AA = anthropomorphic appearance; AB = anthropomorphic behavior; VSI = viewer-streamer interactivity; VVI = viewer-viewer interactivity; SP = social presence; TP = telepresence; CWI = continuous watching intention; PI = purchase intention
Source: Developed by authors
IPMA On purchase intention
| Variables | Importance index | Performance index |
|---|---|---|
| Anthropomorphic appearance (AA) | 0.106 | 55.249 |
| Anthropomorphic behavior (AB) | 0.060 | 60.931 |
| Viewer-streamer interactivity (VSI) | 0.113 | 59.548 |
| Viewer-viewer interactivity (VVI) | 0.109 | 51.253 |
| Social presence (SP) | 0.315 | 54.747 |
| Telepresence (TP) | 0.169 | 59.394 |
| Continuous watching intention (CWI) | 0.256 | 67.227 |
| Variables | Importance index | Performance index |
|---|---|---|
| Anthropomorphic appearance ( | 0.106 | 55.249 |
| Anthropomorphic behavior ( | 0.060 | 60.931 |
| Viewer-streamer interactivity ( | 0.113 | 59.548 |
| Viewer-viewer interactivity ( | 0.109 | 51.253 |
| Social presence ( | 0.315 | 54.747 |
| Telepresence ( | 0.169 | 59.394 |
| Continuous watching intention ( | 0.256 | 67.227 |
4.5 Artificial neural network analysis
To examine potential nonlinear relationships and enhance predictive assessment, an ANN analysis was performed using the significant predictors of purchase intention (i.e. social presence, telepresence and continuous watching intention) as input variables. A tenfold cross-validation procedure was used to evaluate model robustness. As reported in Table 6, the ANN model demonstrated satisfactory predictive performance, with an average training RMSE of 0.6232 (SD = 0.0156) and an average testing RMSE of 0.5710 (SD = 0.0500), indicating stable predictive accuracy across different data partitions. The sensitivity analysis (Table 7) assessed the relative contribution of each predictor based on normalized importance scores. Social presence emerged as the most influential predictor, with a normalized importance of 100.0%, followed by continuous watching intention (92.1%) and telepresence (76.2%). These findings suggest that social presence plays the most prominent role in explaining purchase intention within the ANN framework. A comparison between PLS-SEM and ANN results (Table 8) revealed differences in predictor rankings across the two analytical approaches. PLS-SEM identified continuous watching intention as the strongest predictor based on effect size (f2 = 0.073), followed by social presence (f2 = 0.051) and telepresence (f2 = 0.027). In contrast, ANN analysis ranked social presence as the most important predictor, followed by continuous watching intention and telepresence. The inconsistent rankings between the two methods indicate that nonlinear patterns may exist among the predictors, which may not be fully captured by the linear assumptions of PLS-SEM. Therefore, the combined PLS-SEM–ANN approach provides a more comprehensive understanding of the factors driving purchase intention.
RMSE Values (model A)
| Training | Training | Testing | Testing | |||
|---|---|---|---|---|---|---|
| n | SSE | RMSE | n | SSE | RMSE | Total samples |
| 668 | 259.987 | 0.6239 | 79 | 28.839 | 0.6042 | 747 |
| 664 | 266.538 | 0.6336 | 83 | 27.894 | 0.5797 | 747 |
| 660 | 267.372 | 0.6365 | 87 | 31.655 | 0.6032 | 747 |
| 668 | 266.268 | 0.6314 | 79 | 28.774 | 0.6035 | 747 |
| 674 | 284.889 | 0.6501 | 73 | 20.212 | 0.5262 | 747 |
| 670 | 258.625 | 0.6213 | 77 | 21.738 | 0.5313 | 747 |
| 669 | 233.03 | 0.5902 | 78 | 19.515 | 0.5002 | 747 |
| 668 | 247.681 | 0.6089 | 79 | 24.169 | 0.5531 | 747 |
| 665 | 251.897 | 0.6155 | 82 | 23.156 | 0.5314 | 747 |
| 673 | 259.704 | 0.6212 | 74 | 33.941 | 0.6772 | 747 |
| Mean | 259.599 | 0.6232 | Mean | 25.989 | 0.5710 | |
| SD | 13.042 | 0.0156 | SD | 4.686 | 0.0500 |
| Training | Training | Testing | Testing | |||
|---|---|---|---|---|---|---|
| n | n | Total samples | ||||
| 668 | 259.987 | 0.6239 | 79 | 28.839 | 0.6042 | 747 |
| 664 | 266.538 | 0.6336 | 83 | 27.894 | 0.5797 | 747 |
| 660 | 267.372 | 0.6365 | 87 | 31.655 | 0.6032 | 747 |
| 668 | 266.268 | 0.6314 | 79 | 28.774 | 0.6035 | 747 |
| 674 | 284.889 | 0.6501 | 73 | 20.212 | 0.5262 | 747 |
| 670 | 258.625 | 0.6213 | 77 | 21.738 | 0.5313 | 747 |
| 669 | 233.03 | 0.5902 | 78 | 19.515 | 0.5002 | 747 |
| 668 | 247.681 | 0.6089 | 79 | 24.169 | 0.5531 | 747 |
| 665 | 251.897 | 0.6155 | 82 | 23.156 | 0.5314 | 747 |
| 673 | 259.704 | 0.6212 | 74 | 33.941 | 0.6772 | 747 |
| Mean | 259.599 | 0.6232 | Mean | 25.989 | 0.5710 | |
| 13.042 | 0.0156 | 4.686 | 0.0500 |
SSE = sum square of errors; RMSE = root mean square of errors, n = sample size
Sensitivity analysis (model A)
| NN | SP | TP | CWI |
|---|---|---|---|
| ANN (1) | 0.360 | 0.290 | 0.350 |
| ANN (2) | 0.362 | 0.362 | 0.275 |
| ANN (3) | 0.338 | 0.312 | 0.350 |
| ANN (4) | 0.410 | 0.212 | 0.378 |
| ANN (5) | 0.332 | 0.308 | 0.360 |
| ANN (6) | 0.384 | 0.307 | 0.309 |
| ANN (7) | 0.369 | 0.220 | 0.411 |
| ANN (8) | 0.336 | 0.301 | 0.363 |
| ANN (9) | 0.420 | 0.263 | 0.317 |
| ANN (10) | 0.438 | 0.238 | 0.324 |
| Average importance | 0.375 | 0.281 | 0.344 |
| Normalized importance (%) | 100.0% | 76.2% | 92.1% |
| 0.360 | 0.290 | 0.350 | |
| 0.362 | 0.362 | 0.275 | |
| 0.338 | 0.312 | 0.350 | |
| 0.410 | 0.212 | 0.378 | |
| 0.332 | 0.308 | 0.360 | |
| 0.384 | 0.307 | 0.309 | |
| 0.369 | 0.220 | 0.411 | |
| 0.336 | 0.301 | 0.363 | |
| 0.420 | 0.263 | 0.317 | |
| 0.438 | 0.238 | 0.324 | |
| Average importance | 0.375 | 0.281 | 0.344 |
| Normalized importance (%) | 100.0% | 76.2% | 92.1% |
SP = social presence; TP = telepresence; CWI = continuous watching intention
Comparison between PLS-SEM and ANN results
| PLS-SEM path | Effect size (f2) | Normalised relative importance (%) | Ranking (PLS-SEM) [based on effect size] | Ranking (ANN) [based on normalized relative importance] | Remark |
|---|---|---|---|---|---|
| Model A (output: purchase intention) | |||||
| SP → PI | 0.051 | 100.0 | 2 | 1 | Not matched |
| TP → PI | 0.027 | 76.2 | 3 | 3 | Matched |
| CWI → PI | 0.073 | 92.1 | 1 | 2 | Not matched |
| PLS-SEM path | Effect size (f2) | Normalised relative importance (%) | Ranking (PLS-SEM) [based on effect size] | Ranking ( | Remark |
|---|---|---|---|---|---|
| Model A (output: purchase intention) | |||||
| SP → PI | 0.051 | 100.0 | 2 | 1 | Not matched |
| TP → PI | 0.027 | 76.2 | 3 | 3 | Matched |
| CWI → PI | 0.073 | 92.1 | 1 | 2 | Not matched |
5. Discussion and conclusions
5.1 Conclusions
This study enhances the understanding of virtual streamers in THCLS by explaining how anthropomorphism and interactivity shape consumer intentions. As AI-enabled live streaming becomes a key digital interface in tourism and hospitality, an integrated framework is developed and empirically tested to clarify the mediating roles of social presence, telepresence and continuous watching intention. Using a hybrid PLS–MGA–ANN approach, the study captures linear effects, subgroup heterogeneity and nonlinear predictive patterns. The results show that anthropomorphic appearance and behavior, together with viewer–streamer and viewer–viewer interactivity, enhance consumers’ immersive and social experiences. These effects are primarily transmitted through presence perceptions, which act as central psychological mechanisms linking AI-driven stimuli to engagement and purchase intentions. Social presence emerges as the most influential driver, underscoring the importance of perceived human connection for trust formation in experience-based THCLS contexts. Continuous watching intention further serves as a key transitional mechanism, as it translates immersive experiences into commercial outcomes. In addition, the findings reveal meaningful heterogeneity across gender, age and viewing frequency, indicating that the effectiveness of virtual streamers depends on user characteristics and accumulated experience. By extending presence-based theories to AI-powered THCLS, this study offers theoretical advancements and actionable insights for designing immersive and interactive THCLS experiences.
5.2 Theoretical implications
This study contributes significantly to the literature on THCLS, virtual streamers and technology-mediated consumer behavior. First, it advances presence theory by disentangling the roles of social presence and telepresence. Unlike prior research that conceptualizes presence as a unidimensional construct or examines social and spatial presence separately (Kim and Park, 2024; Mollen and Wilson, 2010), the findings demonstrate that social presence primarily drives behavioral intentions, whereas telepresence enhances experiential immersion. This distinction clarifies how relational and immersive cues jointly reduce psychological distance in AI-enabled, avatar-mediated consumption contexts (Fu et al., 2024; Liang et al., 2024; Truong and Chen, 2025). In addition, the study contributes to anthropomorphism research by decomposing the construct into appearance-based and behavior-based dimensions. In contrast to prior monolithic treatments that produced mixed findings (Kim and Park, 2024; Sun et al., 2024), the results show that an anthropomorphic appearance fosters telepresence, whereas anthropomorphic behavior more strongly enhances social presence. This extends social response theory and computer-as-social-actor (CASA) perspectives by demonstrating that distinct anthropomorphic cues activate different perceptual systems – sensory immersion versus social cognition – in experiential service settings such as tourism. Moreover, the research enriches the interactivity literature by highlighting the asymmetric roles of viewer interactions. While previous studies have emphasized streamer-centric interaction as the primary driver of engagement (Hou et al., 2024; Sreejesh et al., 2020), the findings reveal that viewer–viewer interactivity exerts a particularly strong influence on telepresence. This extends SOR theory by conceptualizing interactivity as a socially embedded mechanism that amplifies experiential realism through shared attention and communal participation (Liang et al., 2024; Mollen and Wilson, 2010). In addition, the inclusion of MGA introduces important boundary conditions related to gender, age and viewing frequency (Laor, 2022; Lv et al., 2022; Zhang et al., 2021). In particular, the observed novelty decay of anthropomorphic appearance among frequent viewers highlights the dynamic and experience-dependent nature of consumer responses to virtual streamers (Yu et al., 2025a, 2025b). Finally, the hybrid PLS–MGA–ANN approach advances theory development by integrating explanatory and predictive perspectives. While PLS-SEM captures theory-driven linear relationships, the ANN results reveal nonlinear effects – most notably the dominant predictive role of social presence – that linear models may underestimate (Hou et al., 2024; Wang and Zhang, 2025).
5.3 Practical implications
On the basis of the multianalytical findings, several strategic recommendations are derived for tourism stakeholders to optimize virtual streamer deployment in THCLS. First, social presence should be the primary managerial lever for driving purchases. ANN analysis identifies it as the dominant nonlinear predictor (100.0% importance), whereas IPMA highlights it as a critical improvement area. Developers must prioritize anthropomorphic behavioral capabilities – such as emotion-sensitive dialog and socially adaptive feedback – over graphical realism. Virtual streamers should use personalized greetings to simulate authentic interactions, reducing the uncertainty of intangible tourism products. In addition, THCLS must be managed as a social ecosystem. The strong effect of viewer–viewer interactivity on telepresence indicates that “being there” is cocreated. Platform operators should facilitate digital crowd dynamics through shared reactions, polls and tagging, transforming THCLS into socially immersive events rather than isolated viewing experiences. By fostering viewer–viewer interaction, firms can increase user engagement duration and platform stickiness, which are key drivers of revenue generation through higher booking probabilities and cross-selling opportunities. Moreover, lifecycle-based strategies are essential for mitigating “novelty wear-off.” This has direct financial implications, as sustaining user engagement over time reduces churn rates and enhances customer lifetime value, which is critical for long-term profitability in digital tourism platforms. The results of the MGA show that visual appearance has a weaker effect on immersion for high-frequency viewers. While high-quality visual design works as an entry stimulus, sustained engagement requires a shift toward community building and social interaction. In addition, audience segmentation is crucial. Female viewers respond more strongly to interactional responsiveness, while younger viewers show a greater propensity to convert continuous watching into purchases. Tailoring interaction styles to these segments enhances conversion efficiency. Finally, managers should leverage predictive analytics to identify “tipping points” where investments in social presence yield disproportionate gains. This enables more efficient investment decisions in AI development and content design, ensuring that expenditures on virtual streamer technologies translate into measurable economic returns for tourism destinations and hospitality firms.
5.4 Limitations and future research
Despite its theoretical and methodological contributions, this study has several limitations that suggest avenues for future research. First, the reliance on cross-sectional survey data restricts strong causal inference and limits insight into the dynamic evolution of user perceptions and behaviors. Although the hypothesized relationships are theoretically grounded and empirically supported, future studies could adopt longitudinal or panel designs to examine how anthropomorphism, interactivity, and presence perceptions develop over time. Such approaches would be particularly valuable for capturing novelty decay and habit formation effects associated with repeated exposure to virtual streamers. In addition, the sample is drawn from a single cultural context, which may limit generalizability. Cultural values shape social interaction norms, responses to anthropomorphism, and technology-mediated trust formation. Future research could conduct cross-cultural or multicountry comparisons to assess whether the relative effects of social presence, telepresence and interactivity differ across collectivist versus individualist cultures or across emerging and mature tourism and hospitality markets. Moreover, although this study distinguishes between anthropomorphic appearance and anthropomorphic behavior, these constructs were modeled at an aggregate level. Future research could decompose anthropomorphism into finer-grained dimensions, such as facial realism, voice naturalness, emotional expressiveness, perceived autonomy or moral agency. Experimental designs manipulating specific cues would allow clearer causal attribution and identification of optimal configurations across THCLS contexts. Furthermore, this study focuses on purchase intention rather than actual behavior. Given the well-documented intention–behavior gap in experiential services, future studies could incorporate behavioral data (e.g. click-through rates, booking records or dwell time) to validate whether the proposed psychological mechanisms translate into economic outcomes. In addition, the moderating analysis was limited to gender, age and viewing frequency. Future research could examine additional individual-level moderators, such as technology readiness, parasocial tendency, trust propensity or travel involvement. Moreover, while the hybrid PLS–MGA–ANN approach enhances predictive power, ANN models lack interpretability. Future studies could apply explainable AI techniques to unpack nonlinear effects. Finally, future research should examine hybrid human–virtual co-streaming formats to better understand trust, responsibility and authenticity in mixed-agent live-streaming environments.
Ethics statement
The author sought and gained ethical approval from the institution’s Research Ethical Board and the study complied with ethical standards. There was no number attached to the approval.
Data availability
The original contributions presented in the study are included in the article/supplementary material, further inquiries can be directed to the corresponding author.
Informed consent
The researcher sought and gained consent of the participants to take part in the study. Out of the 912 sampled participants, all 912 accepted and voluntarily participated in the study after the researcher assured them of anonymity and that their responses were solely for academic purposes.
Author contribution
Teng Yu: design of the work, data collection, analysis and interpretation, drafting the article and final approval. Zhaokang Zeng and Teng Liu: data interpretation, drafting the article, critical revision of the article. Chengliang Wang: data analysis, interpretation, editing and formatting. All authors contributed to the article and approved the submitted version.
References
Further reading
Supplementary material
The supplementary material for this article can be found online.

