The potential for large language models (LLMs) to improve behavioral science research has generated significant discussion. But, the specific role that LLMs should serve in behavioral research, especially in terms of simulating human participants, remains an open research question. The purpose of this work is to engage with this open question and address a critical gap in the literature stemming from the lack of a practical framework for realistically using artificial research participants.
Google Scholar was systematically searched for modern, peer-reviewed literature. Additional articles were found by both backward and forward citation searching the relevant articles. Exclusion criteria were articles that were not directly related to artificial intelligence (AI) and/or research participants, and articles written in a language other than English. This approach resulted in 26 citations that comprehensively capture current perspectives.
This study proposes two novel stances: that artificial research participants can complement human participants during data collection, and replace human participants during pilot testing. This framework engages with the open question of artificial research participants usage while addressing a framework gap in the literature.
This workadvances discourse LLMs potentially transforming behavioral science by establishing a framework differentiating the use of artificial research participants in data collection versus pilot testing. This study reinforces this framework with clear implementation guidelines that maximize the strengths of AI while respecting the human element and the methodological integrity of behavioral research.
Introduction
The rapid development of artificial intelligence (AI) large language models (LLMs) is influencing various sectors of modern society. It is no surprise that there is growing attention toward applying LLMs to behavioral research. Behavioral science relies tremendously on human participants, so one potential application of LLMs is simulating human participants. A consensus among researchers is that using LLMs to simulate human participants would have a substantial impact on behavioral science. What remains unknown, and a point of contention, is the nature of this impact. How and if LLMs should be used to support behavioral research, particularly in terms of simulating human participants, is an open research question.
Methodology
To address this open question, we conducted a systematic search on Google Scholar for recent, peer-reviewed literature. Searches using keywords and phrases such as “artificial intelligence ethics,” “artificial intelligence human subjects” and “artificial intelligence participants” were conducted with a preference for articles published since 2020. Articles were primarily found through these systematic searches. Additional articles were found by backward citation searching (i.e. reviewing the References sections of relevant articles) as well as forward citation searching (i.e. using Google Scholar’s “cited by” option to find newer works that cited relevant articles after they were published). We evaluated identified articles through their abstracts before examining their contents in full. Exclusion criteria were only articles that primarily focused on a different topic than AI and/or research participants, and articles that were written in a language other than English. The resulting broad, global scope included theoretical and experimental academic journal articles, as well as a few other types of sources (e.g. conference proceedings). Together, these techniques yielded a final count of 26, mostly recent citations (see Table 1) comprising various current views and findings about using LLMs in behavioral research.
Citation summary
| Author(s) and year | Title | Country |
|---|---|---|
| Achiam et al. (2023) | “GTP-4 technical report” | United States* |
| Almeida et al. (2024) | “Exploring the psychology of LLMs’ moral and legal reasoning” | Brazil, Germany |
| Argyle et al. (2023) | “Out of one, many: Using language models to simulate human samples” | United States |
| Azizi et al. (2021) | “Can synthetic data be a proxy for real clinical trial data? A validation study” | Canada |
| Baack et al. (2025) | “Towards best practices for open datasets for LLM training” | Germany* |
| Bender et al. (2021) | “On the dangers of stochastic parrots: Can language models be too big” | United States |
| Binz and Schulz (2023) | “Using cognitive psychology to understand GPT-3” | Germany |
| Crockett and Messeri (2023) | “Should large language models replace human participants?” | United States |
| Dillion et al. (2023) | “Can AI language models replace human participants? | United States, India |
| Doyle et al. (2024) | “Training simulated participants for role portrayal and feedback practices in communication skills training: a BEME scoping review: BEME Guide No. 86” | Ireland, Australia, United States, Canada, United Kingdom |
| Duan et al. (2024) | “MacBehaviour: an R package for behavioural experimentation on large language models” | China |
| Grimm (2010) | “Social desirability bias” | United states |
| Grossmann et al. (2023) | “AI and the transformation of social science research” | Canada, United States |
| Gunn et al. (2009) | “A taxonomy of video games and AI” | United Kingdom |
| Hämäläinen et al. (2023) | “Evaluating large language models in generating synthetic HCI research data: a case study” | Finland |
| Harding et al. (2023) | “AI language models cannot replace human research participants” | United States, Germany |
| Harrer (2023) | “Attention is not all you need: the complicated case of ethically using large language models in healthcare and medicine” | Australia |
| Horton (2023) | “Large language models as simulated economic agents: What can we learn from homo silicus?” | United States |
| Kokosi and Harron (2022) | “Synthetic data in medical research” | United Kingdom |
| Kuru (2024) | “Lawfulness of the mass processing of publicly accessible online data to train large language models” | Netherlands |
| Liu (2024) | “Cultural bias in large language models: a comprehensive analysis and mitigation strategies” | China |
| Park et al. (2024) | “Diminished diversity-of-thought in a standard large language model” | United States, United Kingdom |
| Nov et al. (2023) | “Putting ChatGPT’s medical advice to the (turing) test: Survey study” | United States |
| Salali et al. (2024) | “Flourishing cultural diversity in AI” | United Kingdom |
| Schramowski et al. (2022) | “Large pre-trained language models contain human-like biases of what is right and wrong to do” | Germany |
| Sun et al. (2024) | “Random silicon sampling: Simulating human sub-population opinion using a large language model based on group-level demographic information” | South Korea, Qatar |
| Westfall (2023) | “New research shows ChatGPT reigns supreme in AI tool sector” | United States |
| Author(s) and year | Title | Country |
|---|---|---|
| “GTP-4 technical report” | United States | |
| “Exploring the psychology of LLMs’ moral and legal reasoning” | Brazil, Germany | |
| “Out of one, many: Using language models to simulate human samples” | United States | |
| “Can synthetic data be a proxy for real clinical trial data? A validation study” | Canada | |
| “Towards best practices for open datasets for | Germany | |
| “On the dangers of stochastic parrots: Can language models be too big” | United States | |
| “Using cognitive psychology to understand GPT-3” | Germany | |
| “Should large language models replace human participants?” | United States | |
| “Can | United States, India | |
| “Training simulated participants for role portrayal and feedback practices in communication skills training: a | Ireland, Australia, United States, Canada, United Kingdom | |
| “MacBehaviour: an R package for behavioural experimentation on large language models” | China | |
| “Social desirability bias” | United states | |
| “AI and the transformation of social science research” | Canada, United States | |
| “A taxonomy of video games and AI” | United Kingdom | |
| “Evaluating large language models in generating synthetic | Finland | |
| “AI language models cannot replace human research participants” | United States, Germany | |
| “Attention is not all you need: the complicated case of ethically using large language models in healthcare and medicine” | Australia | |
| “Large language models as simulated economic agents: What can we learn from homo silicus?” | United States | |
| “Synthetic data in medical research” | United Kingdom | |
| “Lawfulness of the mass processing of publicly accessible online data to train large language models” | Netherlands | |
| “Cultural bias in large language models: a comprehensive analysis and mitigation strategies” | China | |
| “Diminished diversity-of-thought in a standard large language model” | United States, United Kingdom | |
| “Putting ChatGPT’s medical advice to the (turing) test: Survey study” | United States | |
| “Flourishing cultural diversity in AI” | United Kingdom | |
| “Large pre-trained language models contain human-like biases of what is right and wrong to do” | Germany | |
| “Random silicon sampling: Simulating human sub-population opinion using a large language model based on group-level demographic information” | South Korea, Qatar | |
| “New research shows ChatGPT reigns supreme in | United States |
This is a structured summary of the analyzed citations. The summary includes author(s), publication year, title and country of origin. For country of origin, works with authors from different countries have countries of origin ordered by authorship. Works with 30+ authors are denoted with “*” and assigned the first author’s country of origin
Results
While some perspectives have raised concerns about LLMs simulating human participants (Crockett and Messeri, 2023; Harding et al., 2023; Park et al., 2024), a wave of work has also proposed the usefulness of LLMs replacing and/or complementing human participants (Dillion et al., 2023; Grossmann et al., 2023). We refer to this application of LLMs as “artificial research participants” and specify “research” as “artificial participants” is already used in the computing and entertainment literature to describe AI characters in games (for a review, see Gunn et al., 2009). We also distinguish artificial research participants more conceptually from two other types of artificial data: synthetic data and simulated participants. Synthetic data are artificially generated data that are constructed through machine learning, and emulate existing human data (Kokosi and Harron, 2022). This type of data is often used in the medical field to avoid the cost, consent and privacy concerns associated with using patient data from health records. Privacy concerns are not sufficiently addressed through anonymization, as anonymized patient data have been subject to the successful reidentification of patients. Overall, the use of synthetic data circumvents the above concerns and also produces results that are similar to those from real patient datasets (Azizi et al., 2021; Kokosi and Harron, 2022). Simulated participants are also used in the medical field to address similar concerns, but consist of humans in defined roles rather than computer-generated models of existing data. Additionally, simulated participants can be used for both research and educational purposes. An example of the latter is that health professionals are regularly trained by working with simulated participants who teach them how to communicate and interact in a range of simulated scenarios (for a review, see Doyle et al., 2024).
Our definition for artificial research participants is distinct from both of the above. We define artificial research participants as represented exclusively by LLM outputs instead of machine learning more broadly or human actors. One could consider artificial research participants a subset of synthetic data because LLM outputs are being used to represent human outputs. However, artificial research participants have unique properties due to being exclusively tied to LLMs. Also, the primary use case for artificial research participants is advancing behavioral science research. Below, we provide clear guidelines for when to use them.
Artificial research participants framework
We acknowledge that human participants are, and will continue to be, of highest importance in behavioral research. We do not argue for the replacement of human data collection, which is a core element of behavioral research. Dillion et al. (2023) proposed using LLMs as corroborating evidence after human data is collected to check for robustness and replicability. Harding et al. (2023) responded with the statement that using LLM outputs as comparison points proves that LLMs have not replaced human participants in psychology research, as suggested by Dillion et al.’s (2023) titled work, “Can AI language models replace human participants?” The current work adds nuance to this discussion of LLMs replacing versus complementing human participants, while also addressing a major gap in the literature: there is no existing framework for how exactly to use artificial research participants. Our framework is a two-part proposal that considers two fundamentally different stages of research.
Our first proposal is that artificial research participants can complement human participants during the data collection phase, which is a phase where a methodical understanding of human behavior and variability is essential. Our second proposal is that artificial research participants can replace human participants during the pilot testing phase, which is a phase that prioritizes quickly identifying basic issues and outcomes. Our proposals take advantage of the human-like, instant responses of LLMs by optimizing pilot testing and supporting the subsequent human data collection. We support our proposals with three main points. First, LLMs provide human-like responses, which underlies the overall validity of using artificial research participants. Second, LLMs are uniquely fit for pilot testing, which supports our divergent proposals of using artificial research participants to support human participants in data collection and replace human participants in pilot testing. Third, LLMs have a unique potential to represent diversity well, which supports considerations of artificial research participants relative to human participants.
Human-like responses
Our first main point is that LLMs have passed a crucial threshold of humanness: they provide convincingly human responses. GPT-3.5 produced outputs that are similar to human outputs, and provided human-like responses in terms of both meeting academic and professional benchmarks and exhibiting human failures in logical reasoning (Achiam et al., 2023; Nov et al., 2023). Regarding behavioral research specifically, various LLMs provide these convincing human-like responses across numerous research domains (Almeida et al., 2024; Binz and Schulz, 2023; Dillion et al., 2023; Grossmann et al., 2023; Hämäläinen et al., 2023; Horton, 2023; Schramowski et al., 2022). However, the finding that LLMs produce outputs indistinguishable from human outputs is not universal. While GPT-3.5 was able to produce moral judgments that were strongly correlated with human moral judgments overall in one study, the LLM did not align with human judgments in particular moral scenarios where moral tradeoffs were more ambiguous (Dillion et al., 2023). More questions for GPT-3.5 in another study about morality and other nuanced topics, such as political orientation and economic preference, elicited an intriguing “correct answer effect” that comprised little to no variation in responses as the LLM attempted to repeatedly provide the supposed “correct answer” to various ambiguous questions (Park et al., 2024). The LLM’s repeated “correct answers” contrasted to divergent human responses to the corresponding questions in that study. It seems that a prerequisite for using artificial research participants is that they should only be applied to domain-specific research situations where their validity has been established. For example, behavioral researchers should be especially cautious about using them in moral judgments research until there is a greater understanding about the morality paradigm boundary conditions that elicit human-like responses instead of the correct answer effect.
Overall, the capability of LLMs to provide convincing, human-like responses underscores the use of artificial research participants to begin with. However, the true nature of these human-like responses is an open question. The declaration that LLMs are “stochastic parrots” imitating human language without an intelligent sense of understanding their own outputs (Bender et al., 2021) has generated a philosophical debate examining our understanding of LLMs. This lies outside of the scope of our work and we do not attempt to make grand claims about the true nature or essence of LLMs. Rather, we are precisely highlighting emerging evidence demonstrating that LLMs consistently provide human-like responses across a number of research domains, and that these responses are often indistinguishable from human responses. Passing this significant threshold suggests that psychology researchers can place confidence in artificial research participants responding in ways that are consistently representative of human responses. Future work should evaluate the ability of LLMs to provide human-like responses in additional research domains, and should do so in a variety of LLMs other than GPT models. ChatGPT reigns supreme relative to other LLMs and AI tools (Westfall, 2023), and it is crucial to avoid centralization to increase nuance and representation in LLM responses.
Within research that only requires text responses (Dillion et al., 2023) and is in appropriate research domains, we posit that artificial research participants can supplement the data collection of human participants as a validation tool for conducting systematic verifications of results. While not defensible on their own as replacements for human participants, the use of artificial research participants to assess the replicability and robustness of human participant sample results is an existing stance (Dillion et al., 2023; Harding et al., 2023) that we support as our first proposal. Our second proposal, which is a more novel suggestion, is that artificial research participants can replace human participants during pilot testing.
Pilot testing
Our second proposal is that LLMs can replace human participants during the pilot testing phase due to their efficiency: they quickly provide human-like responses and do so without needing incentives or being compromised by fatigue. As such, artificial research participants should be effective for fast error detection and insights into initial reactions. Dillion et al. (2023) first alluded to artificial research participants supporting pilot testing, while Harding et al. (2023) countered their proposal with the concern that LLMs are exemplary test takers so they should not be able to mimic human comprehension difficulties. Our rebuttal to Harding et al. (2023) is that LLMs respond similarly to humans, biases included. LLM training data sets include the linguistics of unfiltered and biased behaviors that manifest as implicit, human-like biases (Schramowski et al., 2022). As such, GPT-4 commits some logical reasoning errors and exhibits human-level performance on both professional and academic test benchmarks (Achiam et al., 2023). This suggests that, at the very least, LLMs can provide representative initial information about study usability and practicality.
Pilot testing has different fundamental goals than data collection. The priorities during pilot testing are understanding a study’s clarity, basic errors and first reactions. Based on their massive training on naturalistic expressions, LLMs are uniquely fit for these tasks. They can quickly identify confusing instructions and questions, provide human-like responses that provide a basic sense of usability, do not need incentives and are not compromised by fatigue (Dillion et al., 2023). Avoiding conditions leading to inherent mental fatigue, such as mood and health conditions, is of high importance during pilot testing specifically given the typical small sample sizes of this phase. Such conditions can have exacerbated impacts in small samples, which can provide misleading initial conclusions about test usability and responses. Meanwhile, behavioral effects of these conditions can be averaged out in the larger sample sizes of the much more comprehensive data collection phase.
Even though every LLM response represents massive amounts of human thought, each artificial research participant ultimately serves as one participant (Dillion et al., 2023). As such, we recommend using a handful of LLMs to represent multiple artificial research participants during pilot testing, which is a phase that typically just needs a few participants to flag confusing questions and provide initial impressions. This is, of course, limited to research that only requires text responses (Dillion et al., 2023) as well as research in the appropriate domains. While artificial research participants can replace human participants to optimize this stage of research, a replacement approach is inappropriate for the data collection stage where sufficient sample size and an authentic sense of human variability are important elements. Thus, we maintain that artificial research participants play a supplementary, sanity check role during data collection. During both the pilot testing and data collection phases, although, a sense of diversity in responses is also important.
Representativeness
Our third point is that artificial research participants can have representative research value. Foundational work on behavioral research has highlighted that psychology studies are typically based on convenience samples of university students that exhibit WEIRD (i.e. Western, educated, industrialized, rich and Democratic) qualities overall. LLMs are at risk of perpetuating this WEIRD bias because they are primarily based on, and used by, Western cultures (Dillion et al., 2023; Salali et al., 2024). There is an extension of this argument positing that LLMS are HYPER-WEIRD (i.e. hegemonic, young, publicly expressed, Western, educated, industrialized, rich and Democratic). That is, on top of exhibiting WEIRD qualities, LLMs overrepresent hegemonic viewpoints, young users and publicly expressed statements that can differ from private, controversial thoughts (Crockett and Messeri, 2023). We agree with all of the above. LLMs are certainly at risk of perpetuating typical sampling biases given their WEIRD data sets and feedback loops. LLMs can further be considered HYPER-WEIRD. Despite this, our position is that LLMs still have great potential for representativeness relative to university student convenience samples, so artificial research participants should still be considered even in diversity contexts.
We begin with the novel argument that university student convenience samples are also HYPER-WEIRD. These standard WEIRD samples also exhibit HYPER qualities: they overrepresent hegemonic viewpoints, are young and are not immune to the phenomenon of expressing statements that can differ from private, controversial thoughts. Social desirability bias, or the tendency to express socially desirable responses instead of responses that reflect true feelings and thoughts, is well-documented in behavioral research settings (Grimm, 2010). Next, we extend our position by claiming that university student convenience samples could potentially be more HYPER-WEIRD than LLMs. One example of this is the “young” characteristic of HYPER-WEIRD. Traditional university students are 18–21 years old. Traditional internet users, while still young (Crockett and Messeri, 2023), likely exist at a much less constrained age range than 18–21. Another example of this is the “educated” characteristic of HYPER-WEIRD. Every single student in a university sample is by definition educated. By contrast, not every single internet user is by definition educated, and such users are represented in LLM training data sets. It is worth noting that this is purely a logical argument, and we are cautious to state outright that university student convenience samples are more HYPER-WEIRD than LLMs given that LLMs being WEIRD and HYPER-WEIRD is part of an ongoing debate. Additional research and discussion is needed to better understand LLMs before making strong claims related to this topic.
Apart from reframing HYPER-WEIRD to include, and potentially apply more to, university student convenience samples, we also argue that LLMs have distinct characteristics that makes them relatively more representative as artificial research participants. LLMs are unique in that they can achieve diversity through two paths. First, users can prompt LLMs to effectively represent subpopulations. LLMs instructed to simulate diverse datasets provide similar outputs to diverse human data sets that capture the complex, nuanced relationships that exist between attitudes and sociocultural contexts (Argyle et al., 2023; Sun et al., 2024). Second, LLMs can become inherently more representative through future version updates containing more diverse training data. Diversification through ingrained training datasets can produce higher quality data and is part of a current discussion on best practices toward LLM training (Baack et al., 2025). Overall, while we agree that LLMs currently perpetuate HYPER-WEIRD biases (Crockett and Messeri, 2023; Salali et al., 2024), carefully prompting dynamic and diverse outputs (Argyle et al., 2023; Sun et al., 2024) is an effective strategy for currently accessing the representative research value of LLMs. Advocating for more inclusive training data sets in LLM training (Liu, 2024) might be an effective strategy for unlocking the potential for inherent representativeness of LLMs in the future. Finally, urging for increased transparency about LLM datasets is important, as transparency naturally encourages accountability for best practices (Baack et al., 2025). Public knowledge about these datasets can strengthen arguments for more diverse training data as well.
Discussion
The rising popularity of LLMs is sparking important discussions about how to properly integrate them into various areas of society. We conducted a systematic search on Google Scholar using keyword searches, backward citation and forward citation to understand modern research on how emerging discourse, stemming from the consensus that LLMs can transform behavioral science as we know it, has created an open research question about how LLMs could be used to simulate human participants, and whether they should be employed to do so at all. We propose two use cases for artificial research participants in text-based, appropriate research domains: supplementing results from human participants during data collection, and replacing human participants during pilot testing. Together, these two proposals form a framework that addresses a critical gap in the literature: the lack of a clearly defined framework for how artificial research participants would realistically be used. Our novel framework also acknowledges the fundamentally different goals of these stages: data collection attempts to contribute to a comprehensive understanding about human nature and variability, while pilot testing beforehand attempts to conduct basic study design assessment by flagging errors and providing first impressions. We support these proposals with three major points: LLMs provide human-like responses often indistinguishable from human responses, LLMs are particularly well-suited for pilot testing and LLMs carry a unique potential for representativeness.
Our framework for artificial research participants must be prefaced with a careful attitude of caution. Caution must be exercised as the use of LLMs in behavioral research contains inherent ethical dilemmas. One inherent dilemma is that of transparency. While an organization’s LLM technical report may provide certain descriptions of training data (Achiam et al., 2023), we are generally not fully aware of the training data being used to train LLMs by OpenAI, Google and others in terms of the full data sets. This issue of transparency about LLM datasets extends to LLM outputs. Some outputs can be inaccurate as a result of either a “hallucination” (i.e. novel, incorrect information) or a reliance on sources that contain inaccuracies (Kuru, 2024). It is the responsibility of the user to judge the authenticity of an LLM output, and there is little transparency about how LLMs arrive at any particular output. That is, for each output, it is unclear how an LLM is selecting which sources from the vast amount of information available to it, or if it used real sources in its output at all. The lack of transparency in particular LLM outputs can ultimately make users prone to accepting inaccuracies as truths. This has the potential to become a growing source of misinformation that contaminates academic and technical writing, as well as healthcare and medicine sectors, at scale (Harrer, 2023).
Another inherent ethical dilemma is that of consent. LLMs are trained on enormous amounts of publicly available sources without the consent of the countless individuals associated with these sources, which creates a complex question of ethics. On one hand, the information is publicly available, so one could argue that these sources cannot be connected to an interest in privacy and do not need explicit consent to be used. On the other hand, one could also argue that the individuals who created these sources could not have known that this information would eventually be subject to LLM training, which is a result of both the modern nature of LLMs as well as the fact that privacy policies (e.g. by Google) regarding LLM training tend to be discreet (Kuru, 2024).
Another problem with artificial research participants, which is a general issue rather than an ethical dilemma, lies in the dominance of ChatGPT relative to other LLMs. Recent analytics showed that in one year, relative to other AI tools beyond exclusively LLMs, ChatGPT alone captured approximately 60% of the market share and accumulated billions of visits per month (Westfall, 2023). ChatGPT reigning supreme has trickled down to a dominance of ChatGPT, relative to other LLMs, in various areas including in the research studies cited within this very article. It is essential to diversify LLM use beyond ChatGPT to provide more nuanced, representative research (and artificial research participants). Until that point, the centralization of LLM use around ChatGPT remains a limitation of the current work as well as LLM use and discussion more broadly. Overall, these inherent issues of transparency, consent and GPT dominance in using LLMs in behavioral research are important to be aware of before proceeding with a potential framework for their use.
Conclusion
Ultimately, our framework is about a young and quickly evolving topic. There is much more work to be done before arriving at any conclusion about LLM integrations in behavioral research. Thus, we encourage three avenues for progress from this point. The first avenue is a continuation of the lively theoretical dialogue about various use cases for LLMs as well as their ethical considerations (e.g. Crockett and Messeri, 2023; Dillion et al., 2023; Grossman et al., 2023; Harding et al., 2023) to establish more foundations for careful, if any, implementation. The second avenue is a continuation of experimental research using artificial participants (Almeida et al., 2024; Binz and Schulz, 2023; Dillion et al., 2023; Grossmann et al., 2023; Hämäläinen et al., 2023; Horton, 2023; Schramowski et al., 2022) to better understand the appropriate research domains and the consistency of LLMs’ human-like responses. That is, more research should be conducted on how artificial participants behave in additional research domains, and specifically how researchers can prompt human-like responses versus the correct answer effect (Park et al., 2024) in more ambiguous research topics such as morality. The third avenue is advocacy for improved LLM transparency and practices to both increase the representativeness of training data sets and address the inherent LLM issues that affect their usage in behavioral science (i.e. lack of public training data and lack of consent). While questions remain to be answered about the validity, ethics and best use cases of artificial research participants, exploring our artificial research participant framework and these three related avenues of research is a helpful step toward establishing responsible guidelines for implementation.

