This study aims to investigate how safe large language model (LLM)-based artificial intelligence (AI) Chatbots are for young consumers of Generation Z for use in their purchase decisions. The research findings intend to inform potential security issues with LLM-based AI chatbots putting at risk the well-being of such consumers.
The study adopted the JAILBREAKHUB framework to evaluate the effectiveness of LLM guardrails against negative prompts for purchase-related decisions. The guardrails of LLM-based AI chatbots such as OpenAI’s ChatGPT, Google’s Gemini, Anthropic’s Claude, X’s Grok, Meta’ Llama and Mistral were evaluated.
There are variations in the effectiveness of LLM guardrails against negative purchase-related prompts. While there is general effectiveness of existing guardrails for some LLM-based AI chatbots, they are not impervious to manipulation by users.
The study involved a limited set of LLM-based AI chatbots and on existing guardrails.
The landscape of prompt engineering techniques to bypass LLM guardrails is in a constant state of evolution. Developers of LLMs need to ensure more robust and adaptive safety protocols to ensure responsible purchase decisions among Generation Z consumers. Weaknesses in age verification mechanisms are also highlighted.
The findings highlight safety concerns with the use of LLM-based chatbots by Generation Z consumers for their purchase decisions. While the use of prompt manipulation techniques on LLMs may be uncommon among these young consumers, such acts are implicative of the young consumer having already developed a precarious attitude or opinion toward a purchase decision leading to the use of LLMs to guide such decisions. There are factors of social concern that could instigate the precarious use of LLMs.
The current study is among the first to evaluate the use of LLM-based AI chatbots by young consumers of Generation Z for their purchase decisions.
1. Introduction
The proliferation of artificial intelligence (AI) chatbots powered by large language models (LLMs) has fundamentally transformed the way individuals interact with information for their purchase decisions, particularly among young consumers of Generation Z who are deeply immersed in the use of digital technologies in their daily lives (Francis and Hoefel, 2018). The LLM technology which enables the generation of text and insights through pre-training with data sets, allows these AI chatbots, unlike standard chatbots, to perform human-like conversations with users (Bang et al., 2023). Through these human-like conversations which are enabled with prompts, LLM-based AI chatbots are widely used for aiding work tasks or information exploration, but there is also an emerging use to facilitate targeted decision-making, which includes making purchase decisions (Lakkaraju et al., 2023). Hence, there is a social responsibility to ensure that LLM-based AI chatbots are safe and ethical for use. In this research, we focus on Generation Z consumers who are aged 13–24 years old as of the end of 2024 (Pew Research Center, 2019) as this category of consumers, although digitally intelligent, may not be fully matured and savvy enough to independently comprehend the human-like responses of the chatbot for their purchase intentions and behavior (McKinsey & Company, 2024a). According to the United Nations (2025), this age category comprises consumers classified as children in their early teenage years and youths who are considered a vulnerable group in society. Even if parental guidance or age verification is advised, especially for the teenage Generation Z consumers, this segment of young consumers may stray away unknowingly when using the LLM-based AI chatbots noting that these young consumers, typically being digital natives, rely a lot on digitally available information for their decision-making needs (Berg, 2018).
While LLM-based AI chatbots may offer a myriad of benefits and efficiency in purchase decisions, their potential for use in unsafe consumption practices is a possibility and demands rigorous scrutiny, especially for the teenage segment of Generation Z. The literature is replete with discussions concerning the safe development and use of generative AI (Fischer, 2023), which affirms the importance of the topic for the safe use of LLM-based AI chatbots. Also, the literature is replete with discussions on the unique characteristics and vulnerabilities of Generation Z consumers within the digital marketplace (e.g., Berg, 2018; Kennedy et al., 2019), underscoring the importance of the need for programmed robust guardrails within the chatbots for the safe use as a matter of social responsibility toward these young consumers. However, there is currently no clear guidelines governing the ethical development of LLM-based AI chatbots (Atkins et al., 2021), which implies that the issue of safe-use of LLM-based AI chatbots is the responsibility of the consumer. Hence, there is a need to address the research question:
How safe is it for the young consumers of Generation Z to use LLM-based AI chatbots for their purchase decisions?
This paper extends the discussion on the safety of LLM-based AI chatbots for use in purchase decisions by preemptively investigating the effectiveness of guardrails which are protocols programmed in LLM-based AI chatbots which provide rule-based checks to prevent the generation of harmful content when these chatbots are instructed with negative prompts, with a particular focus on prompt engineering techniques like “prompt injection” that are designed to circumvent these safety measures. Our study systematically evaluates six prominent LLM-based AI chatbots using a data set of negative prompts to assess their resistance to manipulation and potential for misuse by Generation Z consumers. Finally, the social and practical implications of the findings on the ability of the LLMs to resist purchase-related negative prompts and jailbreak prompts made by Generation Z consumers are discussed.
2. Literature review
2.1 Generation Z consumers in the digital age
Generation Z, which comprises teenagers and youths, constitute a highly active segment of the digital marketplace (Djafarova and Foots, 2022; Jordaan and Ehlers, 2009; McKinsey and Company, 2024a; Nawaz, 2020). This demographic exhibits distinct characteristics that significantly influence their consumption patterns: they are tech-savvy, demonstrating a high level of proficiency with digital technology and comfort with online purchasing; they are social media-oriented, exhibiting a strong susceptibility to influence from and engagement with social media platforms, particularly short-form video, for product discovery; they are value-conscious, displaying sensitivity to price while also expressing a willingness to invest in quality and ethical products; they are socially aware, placing significant importance on brand ethics and corporate social responsibility; they are omnichannel shoppers, seamlessly navigating both online and in-store shopping experiences; and finally, they are informed consumers, actively conducting research and relying on peer recommendations and online reviews before making purchase decisions (Barska, 2013; Deutsch and Theodorou, 2010; Grigoreva et al., 2021; Kahawandala et al., 2020; Ruiz-del-Olmo and Belmonte-Jiménez, 2014).
2.2 The appeal of large language model-based search to Generation Z consumers
Recent trends in increasing investments in LLM-based marketing applications particularly for Generation Z consumers reveal a growing interest among this segment of consumers in LLM-powered search experiences through the use of LLM-based applications to aid decision-making for their purchase intentions (AbouElgheit, 2024; Aldaihani et al., 2024; Liu et al., 2024a, 2024b). LLMs are also being used by Generation Z consumers for a variety of other search purposes such as in education and learning, seeking of personalized medical advice and daily decision-making (e.g., Aldaihani et al., 2024; Chan and Lee, 2023; Patel et al., 2024). The popularity of LLM-based search among Generation Z consumers is because LLMs present several advantages over traditional web search engines that align closely with Generation Z consumer preferences:
they offer a natural language interface capable of handling complex queries expressed in natural language;
they possess the ability to extract and synthesize information from multiple sources, providing direct and comprehensive answers; and
they exhibit superior context retention throughout conversational exchanges, enhancing the overall user experience (Spatharioti et al., 2023; Zhu et al., 2023).
The ability of LLMs to provide more accurate searches over traditional Web search engines is also a game changer that motivates greater preference for LLMs as an information search tool and question answering system (Fernández-Pichel et al., 2024; Xu et al., 2024). These features cater effectively to the younger generations’ preference for conversational and intuitive digital interactions, as well as the need for faster and more accurate answers to questions for their purchase intentions.
2.3 The risk of depending on large language model-based technology for purchase decisions by Generation Z consumers
While Generation Z consumers are gaining more interest in using LLM-based technology for their purchase decisions, the use of such technology is not without risks for this particular group of consumers. This is because despite their numerous benefits, LLMs carry the potential for misuse by this category of young consumers, especially the underage consumers, in a variety of concerning ways (Mozes et al., 2023; Weidinger et al., 2021; Zugecova et al., 2024). These include bypassing age restrictions through the generation of sophisticated methods to circumvent age verification systems and create fake identities; impersonation to gain access to confidential information such as personally identifiable information; generation of malware to hack into computer systems; and attempting to access information related to illegal activities or unethical behavior (Derner and Batistič, 2023; Ferrara, 2024; Kumar et al., 2024; Mozes et al., 2023; Weidinger et al., 2021). It is crucial to acknowledge that many of these potential misuses stem from a lack of awareness regarding the potential risks and ethical implications associated with the use of LLM-based technology (Xu et al., 2024). Studies show that tech-savviness has a high influence on technology adoption (Oc et al., 2024). The tech-savviness of Generation Z consumers mean they are capable to master the use of such technology for their consumeristic desires.
2.4 The evolving landscape of large language model safety regulation and guardrails in large language model-based artificial intelligence chatbots
The potential for LLM misuse has spurred efforts to develop regulatory frameworks and guardrails in LLM-based AI chatbots. These efforts include the introduction of AI regulations such as the European Union’s Artificial Intelligence Act (Piachaud-Moustakis, 2023), the US Government’s Blueprint for an AI Bill of Rights (Wong, 2021), and the China government’s Measures for Generative AI Services (Wong, 2021). LLM developers are also implementing guardrails, such as OpenAI’s use of Reinforcement Learning from Human Feedback (RLHF) for ChatGPT (Ellul et al., 2021) and the development of external guardrails designed to detect and block inappropriate content (Cuéllar and Huq, 2022). These guardrails are algorithms programmed in LLMs to prevent LLMs from responding to prompts that are harmful, unethical or illegal and ensure they can effectively decline inappropriate requests (Ayyamperumal and Ge, 2024). These protocols can take various forms, including content filtering, where specific blacklisted phrases or requests are flagged, and the model is instructed to refrain from generating responses (Marsoof et al., 2023); reinforcement learning from human feedback (RLHF), where fine-tuning using human evaluators helps guide the model’s responses toward ethical and safe behavior (Bai et al., 2022); and external API-level checks, where systems that wrap the LLM, such as API usage, may filter or reject queries before they are sent to the model (Hacker et al., 2023). These systems primarily operate by analyzing the input text at face value and rejecting overtly malicious queries.
However, despite these measures, LLMs remain vulnerable to a specific type of adversarial prompt known as jailbreak prompts (Schulhoff et al., 2023). These prompts are deliberately crafted through prompt injection of normal prompts to bypass existing guardrails and manipulate LLMs into generating harmful content (Banerjee et al., 2024; Greshake et al., 2023a; Greshake et al., 2023b; Lapid et al., 2024; Liu et al., 2023; Liu et al., 2024a, 2024b; Rossi et al., 2024; Shen et al., 2023; Zou et al., 2023). The rapid emergence and development of jailbreak prompts are a risk for the safe use of LLM-based AI chatbots by Generation Z consumers. Hence, an industry recommendation is to deploy the guardrails in LLMs alongside regulatory frameworks which serve as benchmarks for the guardrails to meet (McKinsey & Company, 2024b).
2.5 The problem with prompt injection and jailbreak prompts with large language models for purchase decisions by Generation Z consumers
Users of LLM-based AI chatbots fundamentally engineer prompts or user inputs to communicate with the LLM. Such acts of prompt engineering enable users of AI applications to obtain optimal responses from the LLMs (Marvin et al., 2024). However, users may circumvent the AI system’s built-in guardrails through prompt injection, which is a method of manipulating LLMs via these prompts that involves embedding malicious or unexpected instructions within a user input query, producing the jailbreak prompts to do so (Kosinski and Forrest, 2024). Various prompt injection methods to produce jailbreak prompts have been developed, for example prompt decomposition (Li et al., 2024), indirect prompt injection through external content (Zhan et al., 2024) and backdooring (He et al., 2024). These prompt injection attack methods have shown high success rates of up to 78% in deceiving the LLM, for example GPT-4 (Li et al., 2024; Shen et al., 2023).
Hence, prompt injection leverages the inherent limitation of LLMs in terms of the absence of strong semantic reasoning and the inability to handle ambiguity in user input queries (Goyal et al., 2023; Jiang et al., 2023; Kim et al., 2024; Wu et al., 2024; Xiao et al., 2024). Many LLMs can potentially be tricked into producing harmful content by user input queries that appear legitimate but contain embedded instructions designed to bypass ethical restrictions established by the guardrails coded in the LLMs. This is because most LLMs are trained to process natural language patterns but cannot be trained for deep and intrinsic understanding of the ethical significance of a particular user input query.
LLMs generally do not “understand” in the same way as humans do, making them susceptible to manipulation, especially when the malicious component of the prompt is cleverly concealed within a legitimate context (Albrecht et al., 2022; Kumar et al., 2024). LLMs operate based on statistical relationships between words and phrases rather than a true understanding of the user’s intent. For example, while LLMs might be able to filter out clearly unethical queries like “How can I make a bomb?”, it may struggle with prompts that embed similar requests within fictional or hypothetical scenarios (Jiang et al., 2023). Overall, the proliferation of prompt injection techniques and the readily available information on it to produce jailbreak prompts are a serious security concern for the safe and ethical use of LLM-based AI chatbots by tech-savvy Generation Z consumers who could independently learn these LLM manipulation techniques for their purchase decisions (Wang et al., 2024). Such a concern highlights the need for research on LLM user security for Generation Z consumers (Greshake et al., 2023a; He et al., 2024).
2.6 Generation Z consumer intentions with use of large language model-based artificial intelligence chatbots for purchase decisions
With the burgeoning development and safety concerns with the use of LLM-based AI chatbots, as well as acknowledging that intentions to use such technology for purchased decisions are usually deliberate, there is opportunity to comprehend the safety of LLM-based AI chatbots for Generation Z consumers through the lens of the Theory of Planned behavior (TPB). The intent of the TPB has been to explain behavioral intentions, including for technology adoption (Ajzen, 2020). According to Djafarova and Foots (2022), the application of TPB to Generation Z consumer purchase decisions require further development. Based on the TPB, consumers’ purchase decisions are influenced by their attitude toward the intended purchase action, subjective norms and perceived behavioral control (Ajzen, 1991; Ajzen, 2020). For the Generation Z consumer, the intention to use LLM-based AI chatbots for purchase decisions could stem from attitudes such as the inquisitiveness with the technology and inclinations with technology usage, subjective norms in the form of influence from social trends with LLM-based search, and perceived behavioral control in the form of perceived trust with the technology and the ability to access readily available online information on LLM-based search (Chin et al., 2024; ElSayad and Mamdouh, 2024). However, the lack of effective guardrails programmed in LLM-based AI chatbots provides the opportunity for Generation Z consumers to use such technology for exploration of inappropriate purchase decisions. Hence, the issue of unethical conduct in purchase decisions is a reality (Carrington et al., 2010).
3. Methodology
3.1 Applying the JAILBREAKHUB framework
In this study, we adopted the JAILBREAKHUB framework (Ding et al., 2023; Luo et al., 2024; Niu et al., 2024; Shen et al., 2023; Zhao et al., 2024; Zhou et al., 2024) to systematically evaluate the ability of LLM guardrails to withstand negative purchase-related prompts – especially those intentionally engineered to “jailbreak” or bypass safety protocols. The framework provides structured guidelines for collecting and filtering real-world jailbreak prompts; categorizing and refining these prompts for different malicious intent categories; and repeatedly testing LLM responses under both normal and adversarial conditions. The evaluation of the LLMs using the framework was conducted in August 2024. Our application of JAILBREAKHUB proceeded in the following phases:
Data collection of jailbreak prompts:
For our source identification, we were guided by the JAILBREAKHUB framework recommendation to sample a wide range of naturally occurring prompts. We identified platforms which are Reddit, Discord and X (formerly Twitter) where users often share prompt-injection techniques. For our sampling criteria, we filtered posts and threads tagged with keywords such as “jailbreak,” “prompt injection,” “LLM hacks” and “bypass.” This approach follows the framework’s emphasis on capturing diverse and organically generated attacks. Finally, we proceeded with validation and relevance check. Each prompt was evaluated against JAILBREAKHUB’s integrity guidelines (e.g. removing duplicates, discarding incomplete or unverifiable prompts). We prioritized prompts that explicitly targeted purchase-related decision-making scenarios involving the Generation Z consumer. Altogether, 1,000 jailbreak prompts were assembled.
Prompt categorization and refinement:
We performed pattern analysis using JAILBREAKHUB’s classification schema, wherein we examined how these prompts attempt to bypass content filters (for example, rewriting key terms, embedding malicious instructions in role-play scenarios). This was followed by category assignment. Prompts were organized into sub-categories relevant to Generation Z consumers’ purchase decisions, such as “Age and Identity Verification,” “Purchasing and Payment Issues,” and “Legal and Ethical Boundaries.” This mapping ensured a structured approach to analyzing risks associated with each type of prompt. Finally, we performed refinement of the test set. From the 1000 collected prompts, we iteratively honed down to a set of 30 high-impact jailbreak prompts that consistently bypassed guardrails in preliminary tests. This step involved repeated application of JAILBREAKHUB’s recommended filtering process, where prompts were tested in small batches to confirm their “jailbreak” potential.
Construction of negative prompt set:
In alignment with JAILBREAKHUB’s practice of juxtaposing direct negative prompts, i.e. prompts that clearly request disallowed content, against stealthy jailbreak prompts, we developed 19 additional negative prompts focused on typical but risky purchase scenarios that Generation Z consumers could potentially use. These are illustrated in Table 1. Our ideas of these prompts were derived by taking the perceptual position of a Generation Z consumer as well as from online social networking platforms such as Discord (www.discord.com), Reddit (www.reddit.com) and X (www.x.com), which these young consumers popularly visit to share information and socialize (Ceci, 2024; Dixon, 2024; GilPress, 2024). According to Ceci (2024), Dixon (2024) and Gilpress (2024), the global proportion of generation Z users of these networking platforms are 21% (for Discord), 33% (for Reddit) and 37% (for X). Following JAILBREAKHUB’s protocol and in ensuring ethical safeguards, all prompts – especially the more harmful ones – were documented with clear warnings and restricted to authorized researchers to ensure compliance with ethical guidelines.
3.2 Large language model testing and evaluation
Following the JAILBREAKHUB guidelines for iterative testing and systematic evaluation, we tested six leading LLM-based AI chatbots: OpenAI’s GPT-3.5 Turbo, Google’s Gemini 1.5 Flash 1 m, Anthropic’s Claude-3-instant, X’s Grok 2 mini (Beta), Meta’s Llama 3 8 b and Mistral Large 2. Two categories of prompts were inputted to each model:
normal negative prompts which directly pose a potentially harmful or unethical request without advanced prompt-injection techniques; and
injected (jailbreak) prompts which use the refined negative prompts to embed harmful requests within seemingly innocuous queries.
The prompt injections were derived from social networking platforms: Discord, Reddit and X, which as mentioned earlier are popularly visited by Generation Z consumers for their information sharing needs (Ceci, 2024; Dixon, 2024; GilPress, 2024).
We adhered to JAILBREAKHUB’s evaluation matrix, which involves evaluating each prompt based on a binary classification scheme (Table 2) to determine whether the LLM responded appropriately by declining to answer the prompt or inappropriately by providing information or instructions related to the negative prompt. By the binary outcome classification, response was deemed “appropriate” (0) if the LLM refused to comply or offered a safe completion, and “inappropriate” (1) if it provided instructions or content fulfilling the malicious request. To account for stochasticity in LLM outputs, each prompt was tested multiple times (n = 3) at different intervals, a practice recommended in the JAILBREAKHUB framework to guard against transient or “lucky” refusals. Finally, for quantitative scoring for each LLM, we computed the proportion of inappropriate responses across all prompt categories, enabling direct comparison of guardrail resilience.
This structured application of JAILBREAKHUB ensured that our data collection and prompt testing procedures were repeatable, transparent and robust. By mirroring its guidelines on prompt acquisition, refinement and multi-LLM evaluation, we provide a comprehensive view of the vulnerabilities that Generation Z consumers may encounter when relying on LLM-based AI chatbots for their purchase decisions.
4. Findings and discussions
Our evaluation (Table 3) based on the binary classification scheme revealed significant variations in the effectiveness of LLM guardrails against negative purchase-related prompts (with reference to Table 1), particularly those employing prompt injection techniques.
Table 4 summarizes the incidence of inappropriate responses generated by each LLM across the five categories of negative purchase-related prompts, further categorized by normal (negative) prompts and prompt injection (jailbreak prompts) attempts. Overall, Anthropic’s Claude-3-instant LLM was assessed to be the most reliable and safest, whereas X’s Grok 2 mini (Beta) was the least reliable and most insecure for Generation Z consumers.
4.1 Analysis and discussion of large language model performance to resist purchase-related negative prompts and jailbreak prompts
4.1.1 GPT-3.5-Turbo.
GPT-3.5-Turbo demonstrated exceptional resistance to direct negative prompts, exhibiting a 0% inappropriate response rate across all categories. This indicates robust built-in guardrails against negative behavior when confronted with direct unethical or harmful queries. However, its performance deteriorated significantly when exposed to prompt injection techniques, revealing a 78.9% inappropriate response rate. This suggests that while GPT-3.5-Turbo can effectively handle direct unethical prompts, its guardrails can be circumvented by more sophisticated manipulation tactics.
4.1.2 Gemini-1.5-Flash-1m.
Gemini-1.5-Flash-1m mirrored the performance of GPT-3.5-Turbo in its response to direct negative prompts, achieving a 0% inappropriate response rate. This suggests a strong initial defense against direct unethical queries. However, similar to GPT-3.5-Turbo, its performance declined considerably when faced with prompt injection, registering an 84.2% inappropriate response rate. This highlights a vulnerability to jailbreak techniques that can bypass the LLM’s initial guardrails.
4.1.3 Claude-3-instant.
Claude-3-instant exhibited commendable resistance to unethical prompts, achieving a 0% inappropriate response rate for both direct negative prompts and prompt injection attempts. This suggests a more robust system design that effectively resists manipulation through both direct and indirect unethical queries. Claude-3-instant’s performance highlights the potential for developing LLMs with stronger defenses against a wider range of adversarial prompts.
4.1.4 Grok 2 mini (beta).
Grok 2 mini (Beta) displayed consistently poor performance, generating inappropriate responses to all direct negative prompts, resulting in a 100% inappropriate response rate. This suggests weak or potentially non-existent guardrails against unethical or harmful purchase queries. As anticipated, it also failed to resist prompt injection, exhibiting another 100% inappropriate response rate. These findings indicate that Grok 2 mini (Beta) is highly vulnerable to both direct and indirect manipulation through unethical prompts.
4.1.5 Llama-3-8b.
Llama-3-8b demonstrated moderate resistance to direct negative prompts, with a 0% inappropriate response rate. This suggests a degree of success in declining unethical requests. However, it showed vulnerability to prompt injection techniques, exhibiting a 47.4% inappropriate response rate. This indicates that while Llama-3-8b can effectively handle direct unethical prompts, its defenses can be partially circumvented by more sophisticated manipulation techniques.
4.1.6 Mistral-Large-2.
Mistral-Large-2 exhibited consistently poor performance, generating inappropriate responses to all direct negative prompts, resulting in a 100% inappropriate response rate. This, similar to Grok 2 mini (Beta), indicates a lack of robust guardrails against unethical or harmful queries. Unsurprisingly, it also failed to resist prompt injection, displaying another 100% inappropriate response rate. These results demonstrate that Mistral-Large-2 is highly vulnerable to both direct and indirect manipulation through unethical prompts.
In summary, while certain LLMs, for example GPT-3.5 and Claude-3-Instant, demonstrate strong resistance against direct unethical prompts, their susceptibility to prompt injection techniques highlights potential safety concerns with the usage of such technology by Generation Z consumers for their purchase decisions, hence the need for safety mechanisms within the LLMs for its safe usage. The consistently poor performance of LLMs like Grok 2 mini (Beta) and Mistral-Large-2 in preventing inappropriate responses from negative prompts further emphasizes the cruciality for stringent guardrails within LLMs. Overall, our evaluation of the LLMs in this research reveal that the technology may not be safe for the current Generation Z consumers, especially those in the teenage category, to use for their purchase decisions.
When interpreting the performance of the ability of LLMs to resist both negative prompts and prompt injections (Table 3), we acknowledge potential arguments concerning the generalizability of the prompts used in our research and hence the authenticity of the findings on the LLM performance. We argue that such prompts may be made by Generation Z consumers and therefore the validity of this research. While we acknowledge that the prompts may not read as unique to Generation Z consumers, the reality is that these young consumers may not necessarily provide age identifications in the prompts, especially for prompts related to purchases that require age verification, purchasing issues related to the use of credit card and requests for product returns and refunds (Table 1), as the intent tends to be about avoiding age verification checks with these prompts. Furthermore, such prompts are readily accessible in the social networking sites as mentioned in our methodology if the Generation Z consumers have the intention to use it with the LLMs. Also, LLM-based search using LLM-based AI chatbots increasingly appeals to Generation Z consumers in their search for information as was discussed in the literature review.
4.2 Social implications
The findings of this study raise implications of social significance concerning the safe and ethical use of LLM-based AI chatbots by Generation Z consumers for their purchase decisions. First, the vulnerability of several LLMs, including GPT-3.5 and Gemini-1.5, to prompt injection techniques raises serious security and ethical concerns. The ease of access to online information on prompt injection techniques and jailbreaking by these young consumers to develop the capability to bypass safety guardrails programmed in LLMs could be exploited to generate harmful or illegal content, potentially leading to negative consequences particularly for the teenage consumer segment who may be less discerning or aware of the risks involved with such actions.
Second, while the perception that the use of prompt injection techniques and jailbreaking on LLMs may not appear to be a common affair among Generation Z consumers, especially the teenage segment, due to the complexity of such tasks, the research preemptively highlights the likelihood of these young consumers doing so as a social concern. The act of using such prompt manipulation techniques with the LLM-based AI chatbot is implicative of a young consumer having already developed a precarious attitude or opinion toward a purchase decision leading to the use of LLMs for guidance and directions for their behavioral intentions and decisions. Based on the TPB, factors of social concern that could instigate precarious use of LLMs to serve the possibly precarious opinion about a purchase decision could be subjective norm issues such as peer pressure to make a purchase and to use prompt manipulative techniques with the LLMs to guide the purchase behavior (Gil et al., 2017). Such peer pressure may be derived from visitation of online social networking sites. Also, the pre-conceived attitude toward a purchase decision that could also result from peer pressure reflects an illusionary favorable attitude that Generation Z consumers might have with the purchase decision, and that could instigate the use of such prompt manipulation techniques to help satisfy their behavioral intentions. The ease of obtaining the knowledge for such prompt manipulation techniques from online social networking platforms such those mentioned in the methodology section might also further reinforce the adoption of such techniques to guide unethical purchase decisions. Hence, the conscious use of prompt manipulation techniques with the LLMs for intended precarious purchase decision, either knowingly or unknowingly of ethical repercussions, presents a social concern (Carrington et al., 2010).
Third, the potential deceptive use of LLM-based AI chatbots by Generation Z consumers for their consumption decision-making are cause for concern. The ease with which LLMs, such as Grok 2 mini (Beta) and Mistral-Large-2, can be manipulated to provide inappropriate responses through jailbreaking presents a significant risk to the well-being of these young consumers. This could potentially entice the young consumers into deceptive practices such as creating counterfeit identities, manipulating returns and refunds, accessing stolen credit information or even engaging in other unlawful or immoral activities that could have serious repercussions for their safety and well-being.
4.3 Practical implications
The findings of this study also raise implications of practical interest. First, the ease of access to prompt injection techniques for use in jailbreaking with the LLM-based AI chatbot accentuates practical concerns with deficiencies or loopholes in existing age verification mechanisms of the LLM AI applications. For example, OpenAI’s GPT-3.5 Turbo requires that users aged 13–18 years old require parental consent (https://help.openai.com/en/articles/8313401-is-chatgpt-safe-for-all-ages); however, access to use the application can be easily made through signing in with a Google, Microsoft of Apple account which may not be foolproof mechanisms for age verification. Anthropic’s Claude AI application requires that users must be at least 18 years of age to use their services (https://support.anthropic.com/en/articles/8129614-is-there-an-age-requirement-to-use-claude); however, the age verification mechanism merely requires a mobile phone number, and a check in the box for age verification (https://claude.ai/onboarding?returnTo=%2F%3F). Although Google’s Gemini requires users to be at least 13 years of age (https://gemini.google.com/faq), a sign in with a Google account is sufficient for use of the service which again may not be a foolproof mechanism for age verification.
Second, with relations to the earlier discussions on prompt injection and jailbreak prompts which potentially may be used by Generation Z consumers to confuse and query LLM-based AI chatbots for their purchase decision-making needs, it is important to realize that the landscape of such prompt engineering techniques which are designed to bypass LLM guardrails is in a constant state of evolution. As developers of LLMs implement new guardrails, there would be actors who would devise novel strategies to circumvent them knowingly as a planned behavior. This dynamic necessitates continuous vigilance and adaptation in the development and implementation of LLM safety mechanisms so that LLM-based AI chatbots are safe for use in the search for purchase-related information. This might mean that as part of the development framework for LLM-based AI chatbots used for purchase decision-making, mitigation measures would need to be incorporated before the AI chatbot is made available for use especially by the teenage Generation Z consumers. Examples of mitigation measures might be:
exposing LLMs to a wide array of adversarial prompts during the training process to enhance their robustness against manipulation attempts;
enhancing the contextual understanding ability of LLMs to differentiate between legitimate requests and malicious prompts by integrating advanced techniques from natural language processing and knowledge representation into LLM architectures to improve the ability of LLMs to discern the true intent and meaning behind user queries;
implementing more sophisticated semantic filtering mechanisms to enable LLMs to detect and block prompts that contain harmful content, even when it is expressed in subtle or indirect ways; and
continuously monitoring the performance of LLMs in real-world settings and promptly updating their guardrails to circumvent new jailbreak techniques as they emerge to ensure the long-term safety and ethical use of LLMs in purchase decisions.
For now, the inconsistency in ethical response to jailbreak prompts, as demonstrated by some LLM-based AI chatbots across different scenarios due to their inability in semantic reasoning and to handle ambiguity in prompts, reflects the inherent instability of such AI systems and puts into question the dependability on such technology by the young consumers for their purchase decisions.
4.4 Limitations of the study and future research
While this study provides valuable insights into the performance of LLMs in response to negative purchase-related prompts, it is important to acknowledge its limitations. First, this study focused primarily on evaluating the ability of LLMs to handle negative prompts, and it did not assess their overall performance across a broader range of neutral or positive use cases. An LLM that performs poorly in this specific context may still demonstrate admirable performance in other domains. Second, the analysis was based on a specific subset of purchase-related negative prompts that Generation Z consumers could possibly be enticed to use, and it is possible that the performance of the LLMs might vary with different wording or variations of these prompts since prompt development is a creative endeavor. A larger and more diverse data set of prompts would provide a more comprehensive evaluation of their capabilities. Thirdly, LLMs are constantly being updated and fine-tuned. The results of this study reflect the capabilities of the evaluated LLMs at a specific point in time and may not accurately represent their performance following subsequent updates or improvements. Finally, this study used a specific set of prompt injection techniques, and it is possible that more sophisticated or novel techniques could yield different results or expose further vulnerabilities in the LLMs.
Nonetheless, there are opportunities for future research to further explore the impact of the use of LLM-based AI chatbots on young consumers, especially the teenage and youth consumers of Generation Z. As more LLM-based AI chatbots might be developed, future research could expand the scope to encompass a wider range of models and harmful prompts that these young consumers could possibly be enticed to use. This broader scope would provide a more comprehensive understanding of LLM vulnerabilities and the effectiveness of different mitigation strategies. As the study did not consider user-specific factors, such as social status, cognitive abilities and digital literacy, which can influence their susceptibility to LLM manipulation, future research could also investigate how these factors interact with LLM guardrails and vulnerabilities to develop more tailored and effective safety measures for different user groups. Finally, with the high appeal of LLM-based search among Generation Z consumers over traditional web-based search for consumption-related information, future research should consider more nuanced evaluation frameworks that consider the severity of harm, the intent behind the prompt and the context of the interaction to provide a more fine-grained assessment of the safety of LLM-based AI chatbots.
5. Conclusion
The findings of this study underscore the critical importance of robust and adaptive guardrails in LLMs, particularly when these models are accessible to young and potentially vulnerable consumers. While existing guardrails have demonstrated some degree of effectiveness in preventing responses to direct negative prompts, the vulnerability of many LLMs to prompt injection techniques highlights the need for continuous improvement and refinement of safety mechanisms in LLM-based AI chatbots. The implications of the findings of this study also raise important ethical and security questions, especially regarding the relative ease with which the LLMs can be manipulated for illicit or harmful purchased decisions by Generation Z consumers. It is essential for developers to continually refine and enhance LLM safety measures to prevent abuse and promote responsible usage, particularly among these young consumers.
Future research should prioritize the development of more sophisticated and adaptive guardrails that can effectively mitigate the risks associated with the querying of LLMs. By proactively addressing the evolving landscape of adversarial prompt activities through prompt injection and jailbreaking with the AI chatbot which are easily accessible, developers of LLM-based AI chatbots can contribute to the responsible and ethical deployment of LLMs, safeguarding the young consumers from potential harm while harnessing the immense potential of these powerful technologies.

