Purpose

This study aims to examine the artificial intelligence (AI)-witness effect, a recently observed phenomenon demonstrating impaired recognition performance for suspects following an AI-facial recall task, in which eyewitnesses describe a suspect’s face using a simulated AI interface and may then be shown an AI-generated facial composite of the suspect. In addition, this study examines whether the accuracy of the facial composite presented to eyewitnesses moderates the AI-witness effect and tests the interference and processing theoretical accounts for how AI-facial composites might affect recognition performance.

Design/methodology/approach

Mock witnesses viewed a suspect’s face and were randomly assigned to either complete an AI-facial recall task or a no AI-facial recall task control condition. Facial composite accuracy (accurate, inaccurate, no composite control) of the suspect was manipulated. Finally, participants identified the suspect from a showup. Logistic regression and signal detection analyses were used to assess memory strength and response criteria.

Findings

Results showed that when mock witnesses were presented with the accurate AI-facial composite of the suspect, engaging in the AI-facial recall task was associated with impaired recognition performance, worse memory strength and a less strict, more liberal response criterion, compared with the control condition.

Research limitations/implications

The findings are based on a controlled experimental paradigm and require replication across different samples and more ecologically valid settings. Future research should further disentangle the interference and processing accounts.

Practical implications

The results highlight the need for careful evaluation of AI-facial composite systems prior to their widespread use in criminal investigations. Findings are best interpreted as evidence of memory interference under controlled conditions, and further research is needed before conclusions can be drawn about real-world forensic practice.

Originality/value

This study provides one of the first tests of the AI-witness effect, extending established research on verbal overshadowing and facial composite construction to the context of generative AI-facial composite systems.

Eyewitness identifications play a central role in criminal investigations and cases, yet decades of research demonstrate that eyewitness memory is malleable and susceptible to distortion (Wells, 2020). One well-documented phenomenon is that engaging in recall of a perpetrator, such as describing a face or constructing a facial composite, can impair subsequent recognition performance. This effect has been widely studied in the context of the verbal overshadowing effect (e.g. Schooler and Engstler-Schooler, 1990) and related work on facial composite construction (e.g. Tredoux et al., 2021). Additionally, evidence from meta-analyses and reviews suggests that such impairments are reliable across a variety of paradigms and procedures (Meissner and Brigham, 2001; Sporer et al., 2020). These findings raise concerns about investigative practices that require witnesses to recall or reconstruct faces prior to identification tasks.

Recent technological developments have introduced artificial intelligence (AI)-generated facial composite systems, which allow eyewitnesses to generate highly realistic images of suspects by describing and constructing the face using AI-generative models. Compared to traditional composite methods (e.g. featural systems such as Identi-Kit or holistic systems such as EvoFIT), AI systems, almost instantaneously, produce more photorealistic facial composites, and witnesses can tweak their facial composites quickly and efficiently through more input with the AI system. Featural systems require a witness to select and arrange individual facial features, such as the eyes, nose, mouth, and hairline, from a database of options to build a likeness piece by piece. Holistic systems instead present witnesses with whole-face images generated by an evolutionary algorithm; witnesses repeatedly select the faces that most resemble the suspect, and these selections are “bred” across generations until a satisfactory composite is produced. AI-generative facial composite systems represent a different approach: a witness provides a verbal description of the suspect’s face, and a generative model (e.g. a text-to-image model) converts that description directly into a photorealistic image, which the witness can then request to revise through additional prompts (e.g. Sádaba-Campo and Gómez-Moreno, 2025). Although AI-generative systems can produce more photorealistic composites and do so more quickly than traditional methods, they have not been subjected to the same standard of evaluation used to assess featural and holistic systems, which examine, among other things, the likeness of a composite to the target and the rate at which composites support correct identifications (Frowd, 2021). The accuracy, consistency, and potential biases of AI-generative facial composite systems therefore remain largely untested against these benchmarks. Indeed, generative AI models have been shown to encode demographic and featural biases that can affect the realism and characteristics of the faces they produce (Sádaba-Campo and Gómez-Moreno, 2025), raising the possibility that such biases could contribute to interference effects during eyewitness recognition. Although there is evidence that police are using or considering using these tools in investigative contexts (e.g. INTERPOL, 2023; Wilkins, 2025), their cognitive consequences for eyewitness recognition performance remain largely unknown. Initial evidence suggests that engaging with AI-generated facial composites may impair recognition performance, a phenomenon referred to as the AI-witness effect (Mota et al., 2026). Specifically, the researchers observed that creating facial composites using AI appeared to reduce memory strength for the suspect without consistently affecting response bias, similar to impaired recognition performance observed in traditional recall paradigms (Bacharach and Baker, 2024; Wells et al., 2015).

Two main theoretical accounts have been proposed to explain why describing or reconstructing a face impairs later recognition. The first is an interference account, which posits that describing or constructing a face could generate a secondary representation that competes with or distorts the original memory trace (Loftus et al., 1978; Meissner et al., 2008; Sporer, 2007; Windschitl, 1996). According to this account, inaccuracies in a description or composite may function as postevent misinformation, leading witnesses to rely on the altered representation of the suspect’s face during recognition. Empirical work examining this account has produced mixed findings, with some studies reporting weak or inconsistent relationships between description accuracy and recognition performance (McQuiston-Surrett and Topp, 2008; Meissner et al., 2008; Sporer, 1996; Topp‐Manriquez et al., 2016).

The second is a processing account, which explains that impaired recognition performance could be due to a mismatch in cognitive processing styles throughout the encoding, retrieval and recognition phases. Face recognition is widely thought to rely on holistic processing (McIntyre et al., 2016; Meinhardt-Injac et al., 2013; Schooler, 2002; Tredoux et al., 2021), whereas describing or constructing a face is thought to require featural processing (Bruce and Young, 1986; Jacoby, 1991). According to this view, engaging in featural processing during recall might disrupt or override, the holistic processes required for accurate recognition. Consistent with this account, research has shown that conditions encouraging compatible processing strategies between encoding and retrieval processes (e.g., holistic encoding-holistic retrieval or featural encoding-featural retrieval) yield better recognition performance than mismatched strategies (e.g., featural encoding-holistic retrieval or holistic encoding-featural retrieval; Baker and Reysen, 2021).

The recent work of Mota et al. (2026), which observed impairing effects of AI-generated facial composites on recognition performance and dubbed the effect the AI-witness effect, attempted to examine the mechanisms underlying the impaired recognition performance. Most relevant is their Experiment 3: in the experiment, the researchers attempted to isolate the role of composite presentation on recognition performance. The AI-facial recall and -facial composite presentation components of the AI-facial composite task were separated, allowing for an initial test of whether simply viewing a facial composite influenced recognition. The findings demonstrated that mock witnesses who were presented with an AI-generated facial composite showed reduced memory strength compared to those who were not presented with a composite, regardless of whether they had previously described the suspect. In other words, exposure to the facial composite alone was sufficient to impair recognition performance. Additionally, AI-facial composite presentation influenced mock witnesses’ response tendencies during recognition, producing a stricter response criterion among mock witnesses who had not engaged in recall of the suspect’s face, while having less of an impact on response criterion among mock witnesses who described the suspect’s face. These findings suggest that AI-facial composite presentation might introduce a competing or altered representation of the suspect’s face, consistent with the interference account.

Extending Mota et al.’s (2026) research, an important next step is to examine what characteristics of the composite could drive any interference potentially impairing recognition performance. One possible factor is the accuracy of the facial composite itself. AI-generated facial composites are often highly realistic and may closely approximate a witness’s memory of a suspect, which could strengthen their influence on subsequent recognition decisions. Based on an interference account, more accurate composites might exert stronger effects by reinforcing the original memory trace. Another possibility could be that more accurate composites could create a confusable representation of the suspect’s face that competes with the originally encoded memory of the suspect. On the other hand, more inaccurate composites might impair recognition by introducing misleading features that distort memory. Thus, understanding how the accuracy of the AI-facial composite affects recognition performance, we think, is an essential next step for refining theoretical accounts of the AI-witness effect and for evaluating the practical utility of AI-facial composite systems in investigative contexts.

In addition to examining the direct effects of composite accuracy, it is important to consider how variability in witnesses’ descriptions of the suspect’s face relates to recognition performance. Prior research has produced mixed evidence regarding whether description quality predicts subsequent recognition accuracy, with some studies suggesting weak or no relationships (e.g. Brown and Lloyd-Jones, 2002; Fallshore and Schooler, 1995; Finger and Pezdek, 1999; Kitagami et al., 2002; Meissner et al., 2001; Schooler and Engstler-Schooler, 1990; Wickham and Swift, 2006). Relevant to the current study is whether the accuracy of the facial composite might serve as a moderator between description quality and recognition performance. While there are a few early studies which have observed that accurate facial composites are related to accurate recognition performance (e.g. Frowd et al., 2008), we are not aware of any work which has examined whether the accuracy of facial composites might help to explain whether description quality is related to recognition performance. We think it’s possible that more accurate facial composites could strengthen the influence of accurate descriptions on recognition, whereas inaccurate facial composites could weaken, or even reverse, the relationship by reinforcing erroneous features. Accordingly, examining whether facial composite accuracy moderates the relationship between description quality measures (e.g. correct descriptors, incorrect descriptors, subjective descriptors, number of descriptors, and description word count) and recognition performance provided a novel test of the interference account.

Our study was designed to provide a test of the interference account by experimentally manipulating the accuracy of AI-generated facial composites. Specifically, we examined whether exposure to accurate versus inaccurate facial composites influenced eyewitness recognition performance. Participants, in the study, viewed a mock crime video, ostensibly making them mock witnesses to the crime. Then, mock witnesses were assigned to a 2 (AI-facial recall task: yes, no) × 3 (composite accuracy: no composite, inaccurate, accurate) × 2 (showup: target present, target absent) between-subjects design. Recognition performance was assessed using signal detection measures: our main variable of interest was mock witnesses’ memory strength, but we were also interested in observing their response criteria.

Consistent with prior work (e.g. Mota et al., 2026; Schooler and Engstler-Schooler, 1990; Tredoux et al., 2021), we expected that participants who engaged in the AI-facial recall task would demonstrate reduced memory strength relative to those in the no AI-facial recall control condition, replicating the AI-witness effect. Importantly, we tested whether this impairment varied as a function of facial composite accuracy. Based on the interference account, both accurate and inaccurate composites could impair recognition, but potentially through different mechanisms: inaccurate composites could introduce misleading features, whereas accurate composites could create representations of the face that compete with the original memory. In contrast, if impaired recognition performance is due to shifts in processing styles, consistent with the processing account, then engaging in the AI-facial recall task could reduce memory strength regardless of the accuracy of the facial composite.

The study was also designed to examine whether the accuracy of the facial composite moderates the relationship between description quality and recognition performance. We explored whether various measures of description quality, including correct, incorrect, subjective descriptors, number of descriptors, and overall description quantity/word count, were related to recognition performance and depended on the accuracy of the AI-facial composite. Based on the limited research on this issue, we are unsure exactly how our description quality measures might interact with the accuracy of the facial composite on recognition performance. However, if the interference account holds, then the relationship between description quality and recognition performance should depend on the accuracy of the resulting facial composite: a possible interaction effect could be that inaccurate facial composites could intensify any negative effects of inaccurate descriptions (e.g. Topp‐Manriquez et al., 2016), while accurate facial composites could either strengthen (e.g. Frowd et al., 2008), alter, or compete with (e.g. Wells et al., 2005) the effects of correct descriptions.

The design of the study was a 2 (AI-facial recall task: yes, no) × 3 (accuracy of AI-facial composite presented: no composite, inaccurate, accurate) × 2 (showup: target present, target absent) between-subjects factorial design. The dependent variable was the probability of a “yes” response during the showup task. Participants were randomly assigned to conditions.

Participants were recruited from Prolific (N = 727, Mage = 41.81, %female = 56.5, %male = 42.0, %other = 1.5). In line with reported Prolific sample characteristics (Palan and Schitter, 2018), participants were racially/ethnically diverse (%Non-Hispanic White/Caucasian = 65.9, %Black/African American = 16.0, %Hispanic or Latina = 9.1, %Asian = 6.3, %Other = 2.7), educated (%four-year college degree = 35.1, %some college = 19.4, %master’s or professional degree = 14.7, %complete high school = 13.0, %other = 17.6), and employed (%full time = 59.5, %part time = 17.9, %unemployed = 7.3, %other = 15.1). Initially, 830 participants were obtained, and 103 were removed from the sample due to incomplete data or failing the attention checks. Participants were compensated $1.20 for their participation. The study was programmed using Qualtrics, and all stimuli were presented and all responses were recorded using computers. Participants completed the study online on their own devices.

Burglary video.

Prior to watching a crime video, participants were told to pay attention to the video with the following instructions: “You are going to view a brief video. Pay attention. You will be asked about the video later.” Participants were presented with these instructions for 1 min, after which they were automatically directed to the crime video stimuli. Participants viewed the home burglary video footage created and employed by Baker and Reysen (2020). The video was approximately 30 s long. In the video, a female burglar was shown breaking into a locked door, entering the home and stealing various items from a cabinet and purse. The burglar is shown briefly looking toward the camera. The face of the burglar is visible for approximately 5 s.

AI-facial recall.

After watching the video, participants were randomly assigned to either describe the suspect’s face using a simulated AI interface (AI-facial recall) or complete a no-AI-facial recall control task.

AI-facial recall: ChatGPT simulation.

Participants assigned to the AI-facial recall condition used a ChatGPT simulation to describe the face of the burglar shown in the video. Participants in the AI-facial recall condition were prompted to describe the burglar’s face using ChatGPT. Similar to previous studies’ instructions (e.g. Schooler and Engstler-Schooler, 1990; Smith and Flowe, 2015; Wilson et al., 2018), participants were instructed, “Now, describe the face you saw in the video. Describe the face in such a way that your description would aid someone else in attempting to identify the person. Focus your description on the facial features. Describe the shape and size of the eyes, eyebrows, nose, ears, mouth, chin, etc. Type your description of the face using ChatGPT.” The ChatGPT simulation was created using Qualtrics and resembled the format of ChatGPT’s website. As done with ChatGPT, participants typed their answers into a textbox. Their description was not used to generate a facial composite.

No AI-facial recall control: Tetris.

Participants in the no AI-facial recall condition played Tetris rather than describe the burglar’s face using ChatGPT. Tetris, as well as other non-verbal tasks, has been used as a no-description control task in previous research (e.g. Baker and Reysen, 2020), and implementation of the Tetris game was a practical task for both our online sample. After 5 min passed, participants were directed to stop the Tetris game and continue with the study. The Tetris (2024) game used is open source.

AI-facial composite accuracy.

When presented with an AI-facial composite, some participants were either presented with the accurate, inaccurate, or no facial composite at all (control group). Participants presented with either the accurate or inaccurate facial composite were told, “An AI-generated image of the suspect was created. Briefly view the image.” They were shown an AI-generated facial composite of the suspect. ChatGPT generated the facial composite image based on an accurate and inaccurate description of the suspect provided by the PI. To examine how facial composite accuracy affects recognition performance and to test the interference account of the AI-witness effect, the accuracy of the AI-facial composite was experimentally crossed with AI-facial recall. Thus, we needed a single, consistent accurate or inaccurate composite image to expose to participants in those conditions, including those who did not engage in the AI-facial recall. Manipulating AI-facial recall and the accuracy of the AI-facial composite allowed us to provide a test for whether any interference or shifts in processing affected recognition performance. Additionally, using AI-generated composite images produced from known accurate or inaccurate descriptions aligns with the intended real-world application of this research. As AI-driven tools become more integrated into investigative practices, AI-generated facial composites based on detailed, potentially reliable, eyewitness reports could become a more realistic issue. By using a facial composite created using an accurate or inaccurate researcher-produced description and knowing that the facial composite accurately or inaccurately resembled the suspect, we think we were able to provide a clear test of whether facial composite accuracy had an influence. Both composite images were generated using the image-generation feature of ChatGPT (OpenAI, 2025), with each image produced from a single prompt entered by the first author and without iterative refinement or selection among multiple candidate images, so that the accurate and inaccurate composites differed only in the description from which they were generated (see Figure 1 for images of the accurate and inaccurate AI-generated facial composites used in the study).

Figure 1
Two front-facing portraits present women with neutral expressions, with differences in hairstyle, facial features, and visible skin texture.The two portraits present women facing directly towards the viewer against plain backgrounds. The woman on the left has long straight hair parted near the centre and falling past her shoulders. Her eyebrows are defined, her eyes are open and level, and her lips are closed with a neutral expression. The woman on the right has shorter straight hair parted near the centre and ending around the jaw and neck. Her eyebrows are lighter, her eyes are open and level, and her lips are closed with a neutral expression. Fine lines and uneven skin texture are visible around the forehead, eyes, cheeks, and lower face.

The accurate and inaccurate AI-generated facial composites

Note(s): The figure depicts both the accurate (left) and inaccurate (right) composite images, which were generated using the image-generation feature of ChatGPT. Each image was produced from a single prompt entered by the researcher without iterative refinement or selection among multiple candidate images. The accurate and inaccurate composites differed only in the accurate and inaccurate descriptions from which they were generated

Source: ChatGPT

Figure 1
Two front-facing portraits present women with neutral expressions, with differences in hairstyle, facial features, and visible skin texture.The two portraits present women facing directly towards the viewer against plain backgrounds. The woman on the left has long straight hair parted near the centre and falling past her shoulders. Her eyebrows are defined, her eyes are open and level, and her lips are closed with a neutral expression. The woman on the right has shorter straight hair parted near the centre and ending around the jaw and neck. Her eyebrows are lighter, her eyes are open and level, and her lips are closed with a neutral expression. Fine lines and uneven skin texture are visible around the forehead, eyes, cheeks, and lower face.

The accurate and inaccurate AI-generated facial composites

Note(s): The figure depicts both the accurate (left) and inaccurate (right) composite images, which were generated using the image-generation feature of ChatGPT. Each image was produced from a single prompt entered by the researcher without iterative refinement or selection among multiple candidate images. The accurate and inaccurate composites differed only in the accurate and inaccurate descriptions from which they were generated

Source: ChatGPT

Close Figure 1
Accurate AI-facial composite presentation.

Participants shown the accurate facial composite were presented with the same accurate AI-generated facial composite image. The accurate facial composite was created using the following prompt in ChatGPT, “Create a facial composite based on the following description: female, white/Caucasian, twenties to thirties, light brown to blonde hair color, straight hair, shoulder length, rounded face shape, thick eyebrows, full lips, round chin.” The description consisted of descriptors that were considered “correct descriptors” in the subsequent description quality coding rubric. No descriptors were considered incorrect. See coding information for the description quality measures in the Results section.

Inaccurate AI-facial composite presentation.

Participants shown the inaccurate facial composite were presented with the same inaccurate AI-generated facial composite image. The inaccurate facial composite was created using the following prompt in ChatGPT, “Create a facial composite based on the following description: female, white/Caucasian, twenties to thirties, brown hair color, straight hair, length above the should, oval face shape, thin eyebrows, thin lips, round chin.” The description consisted of “incorrect descriptors,” according to our description quality rubric, that accounted for half of the provided description. The facial composite needed to resemble the target enough so as to, theoretically, provide some interference with the target’s face, so half of the descriptors in the description were correct (e.g. female, white/Caucasian).

No AI-facial composite presentation control.

Participants not presented with a facial composite, were not provided with any instructions regarding a facial composite. Participants were directed to a blank screen and then prompted to complete the recognition task.

Showup.

After the AI-facial recall and AI-facial composite presentation phases, participants were required to identify the burglar from a single photo, or a showup task. Before seeing the showup, participants were instructed, “Now you will be asked to identify the person who you believe committed the burglary in the video that you saw previously. You will be shown 1 photo.” This instruction is similar to a “neutral” or “standard” instruction (Mickes et al., 2017; Smith et al., 2018). When presented with the photo, participants were asked: “Is the person shown in the photo the same person you saw commit the burglary in the video?” Participants attempted to identify the burglar from either a photo of the burglar, the target present condition, or a photo of a foil, the target absent condition. The showups were those used by Bacharach and Baker (2024). The foil’s face matched the physical description of the target’s face. Participants could either respond “yes” or “no” to the showup. Participants were not given the option to say, “I don’t know.” In the cognitive literature, this type of identification is commonly referred to as a 1-alternative forced choice (AFC) task (Clark, 2003). In the forensic literature, the 1-AFC task is referred to as a showup task. Police administration of showups is common in the field, as creating a showup with a single person is easier than creating a six-person lineup. Additionally, the use of showups in studies allows researchers to assess assumptions of signal detection theory in eyewitness contexts and develop better theories of how variables impact eyewitness decisions (Smith et al., 2018).

All participants provided informed consent and then watched the crime video. Immediately after watching the video, participants were randomly assigned to one of the AI-facial recall tasks for approximately 5 min: participants either used the ChatGPT simulation to describe the face of the burglar or engaged in the no AI-facial recall control task and played a game of Tetris. Next, participants were randomly assigned to one of the AI-facial composite accuracy conditions: participants were either presented with the accurate or inaccurate facial composite of the suspect, or they were not presented with a facial composite, the control condition. Following the AI-facial recall and composite tasks, participants attempted to identify the burglar from a showup. Participants were either shown a target present or target absent photo of the burglar and chose whether they believed the photo depicted the burglar or not. Finally, all participants completed a demographic questionnaire, were debriefed, and compensated for their participation. The duration of the study was approximately 8.5 min from start to finish (Mmin = 8.39, SD = 6.79).

Recognition performance was examined using logistic regression. Logistic regression analyses were used to examine the effects of AI-facial recall, AI-facial composite accuracy, and showup on “yes” responses. Instead of computing hits and false alarms as has been done in traditional forensic face recognition studies (e.g. Baker and Reysen, 2020), the eyewitness decision that each participant makes (“yes” or “no” in a 1AFC or showup task) was the dependent variable (e.g., Bacharach and Baker, 2024; Sauerland et al., 2008; Wells et al., 2015; Wixted and Mickes, 2014; Wright et al., 2009). Logistic regression analysis allowed us to assess for main effects and an interaction between the factors. The interaction with showup indicated a memory strength effect, while main effects for “yes” indicated response criteria effects. See DeCarlo (1998, 2010) for excellent descriptions of how regression can be used to estimate signal detection measures such as d’ and C when using “yes” responses as the dependent variable. We also computed d’ and c values.

Using G*Power (version 3.1.9.7, Faul et al., 2009), we performed an a priori power analysis for a sample size estimation based on data from Mota et al. (2026, Exp 1). The researchers used “yes” responses as their dependent variable and used odds ratios (ORs) to describe their effect sizes. Their reported ORs for interaction or memory strength effects ranged from 2.163 to 10.739 and are considered to be relatively large. With a significance criterion of OR = 5.00, α = 0.05 and power = 0.95, the minimum number of subjects per experimental condition yielded n = 47, indicating that the total sample size needed to detect an effect in the current study was n = 564 (i.e. 47 subjects × 12 experimental conditions). Based on the parameters of the power analysis, the obtained sample size of N = 727, ranging from 56 to 69 subjects per condition, was adequate to test our hypotheses.

Memory strength: AI-facial recall × AI-facial composite accuracy × showup interaction

A 2 × 3 × 2 logistic regression consisting of the factors, AI-facial recall task (coded as no AI-facial recall control = 0, AI-facial recall = 1), accuracy of AI-facial composite presented (coded as no facial composite = 0, accurate facial composite = 1, inaccurate facial composite = 2), and showup (coded as target absent = 0, target present = 1), and the interactions between the factors, was performed on “yes” responses (coded as no = 0, yes = 1) (see Table 1 for regression output for the full model). Following Osborne’s (2017) recommendation, ORs below 1, and the corresponding 95% confidence interval (95% CI), are reported as their reciprocal (1/OR) so that all reported ORs in this section are greater than 1; this transformation does not change the direction or significance of any effect, only the framing of the comparison. ORs and 95% CIs are also reported for non-significant effects to convey the precision of the estimates. Results of the full model, -LL = 780.690, revealed an interaction between the three factors on “yes” responses, Wald(1) = 9.901, p = 0.007, indicating differences in memory strength between mock witnesses in the AI-facial recall and no AI-facial recall conditions depending on facial composite accuracy. The three-way interaction was broken down in a separate analysis by examining the interaction between AI-facial recall × showup for each AI-facial composite accuracy condition (see Table 2). For mock witnesses shown the accurate composite, those in the no AI-facial recall control condition had better memory strength than those in the AI-facial recall condition, d’no AI-recall = 1.704, d’AI-recall = 0.759, Wald(1) = 6.492, p = 0.011, OR = 5.933, 95% CI [1.508, 23.340]. There was no difference in memory strength between the AI-facial recall and no AI-facial recall conditions for the no composite, d’no AI-recall = 1.368, d’AI-recall = 1.247, Wald(1) = 0.462, p = 0.497, OR = 1.530, 95% CI [0.449, 5.210], and inaccurate composite condition, d’no AI-recall = 1.948, d’AI-recall = 1.865, Wald(1) = 3.502, p = 0.061, OR = 3.534, 95% CI [0.943, 13.240] (see Figure 2).

Table 1

Full model: AI-facial recall × AI-facial composite × showup effects and interactions

“Yes” responseBSEWaldpOR95% CI
AI-facial recall       
No AI-recall0 (base)
AI-recall0.1440.5000.8200.7741.1540.4333.074
AI-facial composite accuracy2.7980.247
No composite0 (base)
Accurate composite−0.8820.6202.0220.1550.4140.1231.396
Inaccurate composite0.1260.4860.0680.7951.1350.4382.941
Showup
Target absent0 (base)
Target present2.2950.43527.800<0.0019.9234.22823.287
AI-facial recall × showup
AI-recall × target present−0.4250.6250.4620.4970.6540.1922.227
AI-facial composite accuracy × showup2.3840.304
Accurate composite × target present0.7230.7290.9850.3212.0600.4948.592
Inaccurate composite × target present−0.3950.6100.4190.5170.6740.2042.227
AI-facial recall × AI-facial composite accuracy*8.0850.018*
AI-recall × accurate composite1.5720.7774.0910.0434.8171.05022.104
AI-recall × inaccurate composite−0.6920.7440.8640.3530.5010.1162.153
AI-facial recall × AI-facial composite accuracy × showup*9.9010.007*
AI-recall × accurate composite × target present−1.3560.9382.0890.1480.2580.0411.620
AI-recall × inaccurate composite × target present1.6890.9193.3680.0665.4050.89232.763
Constant−1.7750.34226.939<0.0010.169
Note(s):

“Base” refers to the level of each variable coded as 0. *Interactions were broken down in subsequent analyses and are presented in Tables 2 and 3 

Source(s): Authors’ own work
Table 2

Memory strength: AI-facial recall × showup by AI-facial composite

“Yes” responseBSEWaldpOR95% CI
No composite       
AI-facial recall
No AI-recall0 (base)
AI-recall0.1440.5000.8200.7741.1540.4333.074
Showup
Target absent0 (base)
Target present2.2950.43527.800<0.0019.9234.22823.287
AI-facial recall × showup
AI-recall × target present−0.4250.6250.4620.4970.6540.1922.227
Constant−1.7750.34226.939<0.0010.169
Accurate composite
AI-facial recall
No AI-recall0 (base)
AI-recall1.7160.5958.3060.0045.5611.73117.861
Showup
Target absent0 (base)
Target present3.0180.58426.681<0.00120.4466.50664.254
AI-facial recall × showup
AI-recall × target present−1.7810.6996.4920.0110.1690.0430.663
Constant−2.6570.51726.382<0.0010.070
Inaccurate composite
AI-facial recall
No AI-recall0 (base)
AI-recall−0.5490.5520.9890.3200.5780.1961.704
Showup
Target absent0 (base)
Target present1.9000.42719.756<0.0016.6862.89315.453
AI-facial recall × showup
AI-recall × target present1.2620.6743.5080.0613.5340.94312.240
Constant−1.6490.34522.797<0.0010.192
Note(s):

The table presents three separate follow-up regressions, one within each level of AI-facial composite accuracy, each testing the AI-facial recall × showup interaction

Source(s): Authors’ own work
Figure 2
A grouped bar graph compares target-present hits and target-absent false alarms across AI-facial recall conditions for no, accurate, and inaccurate composites.The grouped bar graph plots percentage of yes responses on the vertical axis from 0 to 100 per cent at intervals of 20 per cent. The legend identifies Target Present as Hits and Target Absent as False Alarms. Three composite conditions each contain No A I Recall and A I Recall groups. For No Composite with No A I Recall, hits are 62.7 per cent and false alarms are 14.5 per cent, with d prime equals 1.37 and c equals 0.35. For No Composite with A I Recall, hits are 55.9 per cent and false alarms are 16.4 per cent, with d prime equals 1.23 and c equals 0.37. For Accurate Composite with No A I Recall, hits are 58.9 per cent and false alarms are 6.6 per cent, with d prime equals 1.70 and c equals 0.62. For Accurate Composite with A I Recall, hits are 57.4 per cent and false alarms are 28.1 per cent, with d prime equals 0.76 and c equals 0.20. A bracket connects the Accurate Composite groups. For Inaccurate Composite with No A I Recall, hits are 83.9 per cent and false alarms are 16.1 per cent, with d prime equals 1.95 and c equals 0.02. For Inaccurate Composite with A I Recall, hits are 72.4 per cent and false alarms are 10.0 per cent, with d prime equals 1.87 and c equals 0.35. Asterisks precede d prime for both Accurate Composite groups and precede c for the No A I Recall and A I Recall groups with 2 and one asterisks, respectively.

AI-facial recall × AI-facial composite accuracy × showup on “yes” responses

Note(s): The figure displays the probability of a “yes” response in the showup conditions, or the hit and false alarm rates, for AI-facial recall conditions across the AI-facial composite accuracy conditions. Results showed an interaction between AI-facial recall, AI-facial composite accuracy and showup, indicating a memory strength effect: for mock witnesses shown the accurate composite, those in the no AI-facial recall control condition had better memory strength than those in the AI-facial recall condition, *p < 0.05. There was no difference in memory strength between the AI-facial recall and no AI-facial recall conditions for the no composite and inaccurate composite conditions. Results also showed an interaction between AI-facial recall and accuracy of AI-facial composite, indicating a response criteria effect: for mock witnesses shown the accurate composite, those in the no AI-facial recall control condition had a stricter response criterion than those in the AI-facial recall condition, **p < 0.01. There was no difference in response criteria between the AI-facial recall and no AI-facial recall conditions for the no composite and inaccurate composite conditions

Source: Authors’ own work

Figure 2
A grouped bar graph compares target-present hits and target-absent false alarms across AI-facial recall conditions for no, accurate, and inaccurate composites.The grouped bar graph plots percentage of yes responses on the vertical axis from 0 to 100 per cent at intervals of 20 per cent. The legend identifies Target Present as Hits and Target Absent as False Alarms. Three composite conditions each contain No A I Recall and A I Recall groups. For No Composite with No A I Recall, hits are 62.7 per cent and false alarms are 14.5 per cent, with d prime equals 1.37 and c equals 0.35. For No Composite with A I Recall, hits are 55.9 per cent and false alarms are 16.4 per cent, with d prime equals 1.23 and c equals 0.37. For Accurate Composite with No A I Recall, hits are 58.9 per cent and false alarms are 6.6 per cent, with d prime equals 1.70 and c equals 0.62. For Accurate Composite with A I Recall, hits are 57.4 per cent and false alarms are 28.1 per cent, with d prime equals 0.76 and c equals 0.20. A bracket connects the Accurate Composite groups. For Inaccurate Composite with No A I Recall, hits are 83.9 per cent and false alarms are 16.1 per cent, with d prime equals 1.95 and c equals 0.02. For Inaccurate Composite with A I Recall, hits are 72.4 per cent and false alarms are 10.0 per cent, with d prime equals 1.87 and c equals 0.35. Asterisks precede d prime for both Accurate Composite groups and precede c for the No A I Recall and A I Recall groups with 2 and one asterisks, respectively.

AI-facial recall × AI-facial composite accuracy × showup on “yes” responses

Note(s): The figure displays the probability of a “yes” response in the showup conditions, or the hit and false alarm rates, for AI-facial recall conditions across the AI-facial composite accuracy conditions. Results showed an interaction between AI-facial recall, AI-facial composite accuracy and showup, indicating a memory strength effect: for mock witnesses shown the accurate composite, those in the no AI-facial recall control condition had better memory strength than those in the AI-facial recall condition, *p < 0.05. There was no difference in memory strength between the AI-facial recall and no AI-facial recall conditions for the no composite and inaccurate composite conditions. Results also showed an interaction between AI-facial recall and accuracy of AI-facial composite, indicating a response criteria effect: for mock witnesses shown the accurate composite, those in the no AI-facial recall control condition had a stricter response criterion than those in the AI-facial recall condition, **p < 0.01. There was no difference in response criteria between the AI-facial recall and no AI-facial recall conditions for the no composite and inaccurate composite conditions

Source: Authors’ own work

Close Figure 2

Results of the full model also showed an interaction between AI-facial recall and accuracy of AI-facial composite on “yes” responses, Wald(2) = 8.085, p = 0.018, indicating differences in response criteria between mock witnesses in the no AI-facial recall and AI-facial recall conditions depending on facial composite accuracy. The interaction was broken down in a separate analysis by examining AI-facial recall for each AI-facial composite accuracy condition (see Table 3). For mock witnesses shown the accurate composite, those in the no AI-facial recall control condition had a stricter response criterion than those in the AI-facial recall condition, cno AI-recall = 0.624, cAI-recall = 0.204, Wald(1) = 3.901, p = 0.048, OR = 1.699, 95% CI [1.004, 2.874]. There was no difference in response criteria between the AI-facial recall and no AI-facial recall conditions for the no composite, cno AI-recall = 0.352, cAI-recall = 0.371, Wald(1) < 0.001, p = 0.984, OR = 1.005, 95% CI [0.596, 1.697], and inaccurate composite condition, cno AI-recall = 0.020, cAI-recall = 0.350, Wald(1) = 0.447, p = 0.504, OR = 1.193, 95% CI [0.712, 1.998]. Results also revealed a main effect of showup on “yes” responses, Wald(1) = 27.800, p <0.001, OR = 9.923, 95% CI [4.228, 23.287], indicating a response criterion effect. As expected, there was a shift in response criterion associated with showup such that the odds of “yes” responses were higher for mock witnesses shown the target present showup (60.4%hits) compared to mock witnesses shown the target absent showup (15.1%false alarms; see Figure 2).

Table 3

Response criteria: AI-facial recall × AI-facial composite interaction effect

“Yes” responseBSEWaldpOR95% CI
No composite       
AI-facial recall
No AI-recall0 (base)
AI-recall0.0050.2670.0000.9841.0050.5961.697
Constant−0.5440.1838.8120.0030.580
Accurate composite
AI-facial recall
No AI-recall0 (base)
AI-recall0.5300.2683.9010.0481.6991.0042.874
Constant−0.7710.19915.043<0.0010.463
Inaccurate composite
AI-facial recall
No AI-recall0 (base)
AI-recall0.1760.2630.4470.5041.1930.7121.998
Constant−0.5530.1858.9440.0030.575
Note(s):

The table presents three separate follow-up regressions, one within each level of AI-facial composite accuracy, each testing the effect of AI-facial recall

Source(s): Authors’ own work

We computed five description quality measures to get a sense of the quality of mock witnesses’ descriptions: (1) correct descriptors, (2) incorrect descriptors, (3) subjective descriptors, (4) number of descriptors and (5) description quantity. Correct descriptors included adjectives that correctly described the facial features of the burglar (e.g. female, White/Caucasian, blonde hair, etc.), and incorrect descriptors included adjectives that did not correctly describe the burglar (e.g. wore glasses, had freckles, tattoos, etc.). Subjective descriptors were odd adjectives that were neither correct nor incorrect, or just not helpful to produce the AI-facial composite (e.g. looked sneaky, had an ax to grind with somebody, etc.). The number of descriptors was the total number of adjectives/phrases used to describe the suspect and description quantity was the total word count of the description. Our measures of description quality were calculated according to previous studies that employed similar accuracy measures (e.g. Baker and Reysen, 2020; Meissner et al., 2001). Three independent raters scored each description provided by mock witnesses who completed the AI-facial recall task (n = 359). Intraclass correlation coefficients were calculated using a two-way mixed-effects model, absolute agreement and average measurement, to determine the inter-rater reliability of our description quality measures. Results revealed there was excellent to good agreement among the three raters for correct descriptors, ICC = 0.808, 95% CI [0.771, 0.840], F(357, 714) = 5.230, p < 0.001, number of descriptors, ICC = 0.771, 95% CI [0.580, 0.861], F(357, 714) = 6.163, p < 0.001 and description quantity, ICC = 0.937, 95% CI [0.925, 0.948], F(355, 710) = 16.060, p < 0.001. Results revealed decent agreement for the incorrect descriptors, ICC = 0.533, 95% CI [0.252, 0.691], F(357, 714) = 2.911, p < 0.001. Results revealed poor agreement for the subjective descriptors (M = 2.62, SD = 1.60), ICC = 0.296, 95% CI [0.075, 0.460], F(357, 714) = 1.719, p < 0.001, so it was not included in the analysis (Koo and Li, 2016; ten Hove, 2024). Subsequent analyses were performed using the average correct descriptors (M = 3.12, SD = 1.37), incorrect descriptors (M = 1.74, SD = 1.22), number of descriptors (M = 7.11, SD = 2.81) and description quantity (M = 39.34, SD = 21.66) measures across the three raters.

To examine whether any relationship between our description quality measures and recognition performance depended on the accuracy of the facial composite, we performed a logistic regression consisting of our four description quality measures (higher values indicated more correct descriptors, incorrect descriptors, number of descriptors, and a greater quantity), the accuracy of AI-facial composite presented (coded as no composite = 0, accurate composite = 1, inaccurate composite = 2) and their possible interactions, was performed on correct identifications (coded as incorrect ID [misses/false alarms] = 0, correct ID [correct rejections/hits] = 1). Note that we did not include the showup interaction IV or the “yes” response DV. Instead, and consistent with other researchers (Mota et al., 2026; Russ et al., 2018; Sporer, 1993), we used a simplified model by computing correct identifications across target present and absent showups. Results, -LL = 396.793, revealed no main effects, Walds ranged = 1.461–3.341, ps ranged = 0.068–0.333, ORs ranged = 1.017–2.464, or interactions between our description quality measures and facial composite accuracy factor, Walds ranged = 3.235–6.123, p-values ranged = 0.047-0.198, ORs ranged = 1.034–2.464. The full model showed an interaction between description quantity and composite accuracy, Wald(2) = 6.123, p = 0.047; however, when the interaction was broken down by composite accuracy, there were no significant effects, Walds ranged = 0.001–3.056, ps ranged = 0.080–0.974, ORs ranged = 1.000–1.001.

Mock witnesses who completed the AI-facial recall task and described the suspect’s face were not, on average, any better or worse at correctly identifying the suspect than mock witnesses who completed the no AI-facial recall control task. However, this depended on which facial composite, if any, mock witnesses were subsequently shown. When mock witnesses were shown the accurate AI-generated facial composite, those who had completed the AI-facial recall task showed substantially worse memory strength and a less strict, more liberal response criterion than those who had completed the no AI-facial recall control task. By contrast, completing the AI-facial recall task made little difference to memory strength or response criteria when mock witnesses were shown the inaccurate composite or no composite at all. In other words, the AI-witness effect, impaired recognition performance following the AI-facial recall task, emerged specifically when mock witnesses were exposed to a highly accurate AI-generated likeness of the suspect. Additionally, we found no evidence that the accuracy of the facial composite moderates the relationship between description quality and recognition performance.

With the present study, we investigated whether engaging in the AI-facial recall task, describing the suspect’s face using a simulated AI interface, and the accuracy of the facial composite subsequently presented to mock witnesses affected eyewitness recognition performance. Findings provided evidence for an AI-witness effect, such that the AI-facial recall task impaired subsequent recognition memory of the suspect, but only when mock witnesses were also shown an accurate AI-generated facial composite of the suspect. Specifically, when participants were shown the accurate facial composite, memory strength was reduced and the response criterion was less strict (more liberal) for participants who had completed the AI-facial recall task compared to those who had completed the no AI-facial recall control task. There was no difference in memory strength or response criteria between the AI-facial recall conditions for the inaccurate composite or no-composite conditions. Findings regarding description quality and recognition performance revealed no consistent relationship between our description quality measures and recognition performance, nor moderation by facial composite accuracy.

Findings are consistent with the large body of research demonstrating that recalling or reconstructing a face can impair later recognition (Schooler and Engstler-Schooler, 1990; McIntyre et al., 2016; Meissner and Brigham, 2001; Wilson et al., 2018). Meta-analytic and review work has shown that both verbal descriptions and facial composite reconstruction can disrupt recognition memory (Sporer et al., 2020; Tredoux et al., 2021), and the present results extend this research to AI-facial composite systems, which are increasingly realistic and interactive. In particular, Sporer et al.’s (2020) review suggested that exposure to a facial composite, even one constructed by someone other than the witness, can itself impair a witness’s later recognition of the target, indicating that this kind of impairment does not require the witness to have generated the composite themselves. Similarly, McIntyre et al. (2016) found that viewing a forensic facial composite could impair subsequent recognition of the depicted face, particularly under conditions that encouraged holistic processing of the composite. The present findings build on this work by showing that the degree of impairment associated with composite exposure depended on how closely the composite resembled the target: only the highly accurate AI-generated composite, and not the inaccurate one, was associated with reduced memory strength and a more liberal response criterion among mock witnesses who had also completed the AI-facial recall task. This pattern is consistent with the broader literature on composite-viewing effects, while extending it to the highly realistic images that AI-generative systems can now produce. Importantly, the finding that accurate facial composites produced the greatest impairment is particularly informative. The current results suggest that when composites are highly realistic and closely resemble the target, as is increasingly the case with AI systems, the accurate facial composites might have stronger detrimental effects on recognition memory. This explanation is consistent with research on misinformation and memory distortion, which demonstrates that postevent information can be detrimental to eyewitness performance (Loftus et al., 1978; Topp‐Manriquez et al., 2016).

The present findings have important implications for theory development regarding the AI-witness effect. Two accounts were considered: interference and processing accounts. First, the results provide support for an interference account. The finding that memory strength was most affected for mock witnesses presented with the accurate facial composite condition suggests that highly similar or plausible representations of the suspect may have competed with the original memory of the suspect. Rather than reinforcing memory, accurate AI-generated facial composites may create a competing representation that is difficult to distinguish from the original encoding. This interpretation is consistent with research on postevent misinformation, which shows that newly introduced information can alter or overwrite memory representations (e.g. Loftus et al., 1978), as well as recent extensions of interference accounts of verbal overshadowing (e.g. Goldman and Baker, 2026). It is also consistent with evidence that face recognition retains image-specific information rather than relying solely on a more abstract, identity-level representation (Dunn et al., 2019); exposure to a specific AI-generated image of the suspect’s face may therefore be encoded as a distinct image-specific representation that competes with mock witnesses’ memory of the originally encoded image, particularly when that AI-generated image closely resembles the target.

Our study’s results showing a lack of relationship between description quality and recognition performance, however, challenge the interference account’s notion that the quality of the description should be related to recognition performance. Our null findings align with prior work showing weak or inconsistent links between description quality and recognition performance (Brown and Lloyd-Jones, 2002; Fallshore and Schooler, 1995; Kitagami et al., 2002; Meissner et al., 2001; Wickham and Swift, 2006). Instead, the findings suggest that the accuracy of the generated AI-facial composite itself, rather than the underlying description of the face, plays a more central role in influencing recognition.

Second, findings are less informative for the processing account, which suggests that engaging in featural processing during description should impair recognition (Schooler, 2002; Baker and Reysen, 2021), and perhaps regardless of composite accuracy (Wells et al., 2005). Instead, the moderating role of facial composite accuracy suggests that the content of the representation could be critical, not just the shifts in processing strategies during recall, composite presentation, and recognition. This finding is also consistent with Mota et al. (2026) research, which observed impaired recognition performance for mock witnesses presented with an AI-generated facial composite. Thus, the AI-witness effect appears to involve both representational competition and, potentially, processing shifts, indicating the need for a theoretical framework incorporating both interference and processing accounts.

The present findings have potential implications for the use of AI-facial composite systems in forensic and investigative contexts. Law enforcement agencies are increasingly adopting AI technologies to generate suspect images (INTERPOL, 2023; Wilkins, 2025) and are often under the assumption that quickly constructed and more realistic composites should improve investigations. However, the current results challenge this assumption. First, we observed that mock witnesses who completed the AI-facial recall task and were subsequently shown an accurate AI-generated composite of the suspect had impaired recognition performance, consistent with the recently observed AI-witness effect (Mota et al., 2026). Additionally, the finding that accurate AI-generated composites might impair eyewitness recognition performance raises concerns about their use prior to identification procedures. If witnesses are exposed to AI-generated composites before making an identification, their memory for the suspect may be weakened, potentially increasing the risk of misidentification. Second, results suggest that AI-generated composites may function similarly to postevent information, introducing new representations that alter or compete with memory. As reflected in the title of our paper, an important question is whether AI may, in some respects, be too good. Specifically, when AI-generated facial composites are highly realistic and closely resemble the suspect, does this level of realism introduce postevent information that interferes with or alters the witness’s original memory of the suspect? This raises legal and procedural concerns regarding the admissibility and reliability of eyewitness identifications following exposure to AI-generated images (Wells, 2020).

We acknowledge several limitations of the present study and point to avenues for future research. First, the AI-facial recall task was simulated using a ChatGPT interface rather than a fully interactive AI image-generation system (e.g. EagleAI, 2022; Ekatpure et al., 2025). While this approach allowed for experimental control in our study, future research should examine the effects of fully interactive AI systems, where witnesses iteratively refine composites, as would be plausible if a real eyewitness to a crime was asked to create a facial composite. Relatedly, participants in our AI-facial recall condition typed their description of the suspect’s face directly into the simulated AI interface, whereas witnesses to crimes are not typically asked to type a description: their recall of the crime and suspect’s description is usually given verbally to an officer, such as during a cognitive interview. The distinction between written and spoken recall is potentially important, as researchers have found that spoken eyewitness accounts contain more detail than written accounts of the same event (Sauerland and Sporer, 2011). Because our AI-facial recall task required typed input, our findings may not fully capture how a spoken description, later entered or transcribed into an AI system by a practitioner, would affect other aspects of recall, facial composite construction, and eyewitness recognition performance.

Second, we recognize that for our AI-facial composite accuracy manipulation, the operationalizations of “accurate” and “inaccurate” composites were based on predefined descriptions. And we used these accurate and inaccurate AI-facial composites produced from the predefined descriptions to provide a test for the interference and processing accounts for how AI-facial composites might affect recognition performance. However, future work could examine how a broader range of AI-facial composite accuracy manipulations (e.g. graded similarity between composite and target) or measures (e.g. assessing the similarity of produced facial composites to the target) might influence subsequent recognition performance. Additionally, because every mock witness in the accurate- and inaccurate-composite conditions viewed the same accurate or inaccurate image, respectively, our research cannot distinguish effects that are general to accurate versus inaccurate AI-generated composites from effects that are specific to these particular images of this particular target; see Lewis (2024) for a description of the stimulus-as-fixed-effect problem. We therefore cannot be certain whether the pattern of impairment observed for the accurate composite would generalize to other faces, other AI-generated images of the same face or composites produced using different generative models or prompts, particularly given evidence that generative approaches to facial-composite production can vary considerably in their output (Sádaba-Campo and Gómez-Moreno, 2025). Future work could consider this limitation and sample multiple target identities and multiple AI-generated composites per identity, and should model both participants and stimuli as random effects, to establish whether the present findings reflect a general property of accurate AI-generated composites or are specific to the particular images used here.

A third limitation concerns the timing of our AI-facial recall, AI-composite, and recognition procedures. Mock witnesses in our study moved directly from the AI-facial recall and composite presentation phases to the showup task, with no delay. This timeline does not reflect forensic practice, in which witnesses typically provide a description, and may increasingly be shown an AI-generated composite, well before any identification procedure takes place, often after a delay of hours, days or longer. It is possible that the impairment observed among mock witnesses who completed the AI-facial recall task and viewed the accurate composite is temporary and would diminish, or even reverse, given a more realistic delay, analogous to the “release” from verbal overshadowing that has been reported when a delay or an intervening task is introduced between description and identification (Finger and Pezdek, 1999). If the AI-witness effect observed here is short-lived, its forensic relevance would be correspondingly limited; however, if it persists or grows over time, the practical concerns we raise would be strengthened. Determining whether, and for how long, this interference persists is therefore an important next avenue of future research.

A fourth limitation concerns the online administration of the study. Participants were recruited on Prolific and completed the study on their own devices, and we did not record or restrict which kind of device (e.g. desktop, laptop, tablet or mobile phone) they used. As a result, we acknowledge that factors such as screen size, image resolution, and typing interface across devices could have affected how clearly mock witnesses viewed the burglary video and the AI-facial composite images, as well as how they typed their descriptions during the AI-facial recall task. Future research could consider precautions such as recording participants’ device type and screen size, or restrict participation to a specific device type, to help control for such types of error.

Finally, future research should both replicate and extend the AI-witness effect, given the novelty of this finding. Establishing the effect’s reliability across various samples, methodologies and additional contexts is essential before strong theoretical or applied conclusions can be drawn. Additionally, research should continue to refine theoretical accounts of the AI-witness effect by designing studies that can more directly disentangle interference and processing mechanisms. For example, manipulating delays between composite exposure and identification (Goldman and Baker, 2026; Wilson et al., 2018) could help test predictions derived from interference accounts, whereas varying the AI-facial recall task to encourage different processing strategies (e.g. featural versus holistic; Frowd et al., 2008; McIntyre et al., 2016; Skelton et al., 2015) could provide insight into processing accounts.

The study provides evidence that engaging in AI-facial composite systems can impair eyewitness recognition performance, particularly when the facial composite is accurate and resembles the suspect. These findings extend existing research on verbal overshadowing and composite construction to emerging AI-facial composite systems and highlight important theoretical and applied considerations. Our findings should be interpreted as evidence of memory interference under controlled conditions (e.g. we used a simulated AI interface, a single target identity, and an immediate recall-composite presentation-recognition paradigm) rather than as a direct demonstration of how AI-facial composite systems would affect eyewitnesses in real investigations. As AI-facial composite systems continue to be developed and integrated into forensic practice, it is important to develop a better understanding of how AI-facial composite systems might influence the reliability of eyewitness evidence.

  • Highly accurate AI-generated facial composites can reduce eyewitnesses’ memory strength and shift their response criterion when they are next asked to identify a suspect, particularly when witnesses have also described the suspect’s face beforehand.

  • Before AI-facial composite systems are adopted for forensic practice, they should be evaluated against the same standardized protocols applied to traditional composite systems (Frowd, 2021), including assessment of their accuracy, consistency and potential demographic or featural biases.

  • Agencies using AI-facial composite tools should document when and how a composite was generated and whether and when it was shown to a witness, given the potential implications of this exposure for the reliability and legal admissibility of any subsequent identification.

  • Because these current study’s findings were obtained under controlled, time-compressed laboratory conditions with a single target identity, they should inform pilot evaluation and policy discussion rather than be treated as a definitive basis for operational forensic practice. Replication across samples, target stimuli, AI systems, and more realistic timeframes is needed before firm guidelines are established.

Ywomie A. Mota is based at the Department of Psychology,Coastal Carolina University, Conway, South Carolina, USA.

Melissa A. Baker is based at the Department of Psychology,Coastal Carolina University, Conway, South Carolina, USA.

Ywomie A. Mota, B.S., is Research Assistant, Psychology, Law, Emotions, and Attitudes (PLEA) Lab, Department of Psychology, Coastal Carolina University, yamota@coastal.edu. Melissa A. Baker, Ph.D., is Assistant Professor, Department of Psychology, Coastal Carolina University, mbaker4@coastal.edu, https://orcid.org/0000-0003-3365-8709. Data collection was completed at the authors’ institution, Coastal Carolina University. The authors have declared no conflicts of interest. Portions of this study were presented at the American Psychology-Law Society 2026 conference. All procedures performed in studies involving human participants were in accordance with the ethical standards of Coastal Carolina University’s institutional review board and with the 1964 Helsinki declaration and its later amendments or comparable ethical standards. The data that support the findings of this study are openly available in Mendeley Data: “AI-Facial Composite Accuracy”, Mendeley Data, V1, doi: 10.17632/cz3977bzzj.1 Correspondence concerning this article should be addressed to Melissa A. Baker at Department of Psychology, Smith Science 217A, 642 Century Circle, Conway, SC 29526.

The authors thank the research assistants of the PLEA Lab who powered through this project from start to finish: Avery Blinco, Frankie Cullen, Gianna Esposito, Kayla Janssen, Grace Kohls, Lauren Leatherman, Gabriela Pereira, Hannah Saluga, Norah Sullivan, Za’Mya Thompson and Ryan Wolyniec. These RAs assisted with data collection and coding. Their hard work and attention to detail transformed a mountain of mock witness descriptions into usable data, and this project would not have been possible without their contributions.

Bacharach
,
V.R.
and
Baker
,
M.A.
(
2024
), “
Verbal overshadowing and decision criterion effects on recognition memory for faces
”,
Journal of Cognitive Psychology
, Vol.
36
No.
8
, pp.
881
-
897
, doi: .
Baker
,
M.A.
and
Reysen
,
M.B.
(
2020
), “
The influence of recall instruction type and length on the verbal overshadowing effect
”,
American Journal of Forensic Psychology
, Vol.
38
No.
3
, pp.
1
-
29
.
Baker
,
M.A.
and
Reysen
,
M.B.
(
2021
), “
Using intentional and incidental encoding instructions to test the transfer inappropriate processing shift account of verbal overshadowing
”,
Journal of Cognitive Psychology
, Vol.
33
No.
5
, pp.
533
-
548
, doi: .
Brown
,
C.
and
Lloyd-Jones
,
T.J.
(
2002
), “
Verbal overshadowing in a multiple face presentation paradigm: effects of description instruction
”,
Applied Cognitive Psychology
, Vol.
16
No.
8
, pp.
873
-
885
, doi: .
Bruce
,
V.
and
Young
,
A.
(
1986
), “
Understanding face recognition
”,
British Journal of Psychology
, Vol.
77
No.
3
, pp.
305
-
327
, doi: .
Clark
,
S.E.
(
2003
), “
A memory and decision model for eyewitness identification
”,
Applied Cognitive Psychology
, Vol.
17
No.
6
, pp.
629
-
654
, doi: .
DeCarlo
,
L.T.
(
1998
), “
Signal detection theory and generalized linear models
”,
Psychological Methods
, Vol.
3
No.
2
, p.
186
, doi: .
DeCarlo
,
L.T.
(
2010
), “
On the statistical and theoretical basis of signal detection theory and extensions: unequal variance, random coefficient, and mixture models
”,
Journal of Mathematical Psychology
, Vol.
54
No.
3
, pp.
304
-
313
, doi: .
Dunn
,
J.D.
,
Ritchie
,
K.L.
,
Kemp
,
R.I.
and
White
,
D.
(
2019
), “
Familiarity does not inhibit image-specific encoding of faces
”,
Journal of Experimental Psychology: Human Perception and Performance
, Vol.
45
No.
7
, pp.
841
-
854
, doi: .
EagleAI
(
2022
), “
Forensic sketch AI-rtist
”,
lablab.ai.
,
available at:
Link to Forensic sketch AI-rtistLink to the cited article (
accessed
4 September 2026).
Ekatpure
,
M.J.N.
,
Sonali
,
E.
,
Mayur
,
K.
,
Tabasum
,
M.
and
Geeta
,
P.
(
2025
), “
Understanding VisionSketch: automated forensic sketch generation with artificial intelligence
”,
International Journal on Advanced Computer Engineering and Communication Technology
, Vol.
14
No.
1
, pp.
735
-
737
, doi: .
Fallshore
,
M.
and
Schooler
,
J.W.
(
1995
), “
The verbal vulnerability of perceptual expertise
”,
Journal of Experimental Psychology: Learning, Memory, and Cognition
, Vol.
21
No.
6
, pp.
1608
-
1623
, doi: .
Faul
,
F.
,
Erdfelder
,
E.
,
Buchner
,
A.
and
Lang
,
A.-G.
(
2009
), “
Statistical power analyses using G*power 3.1: tests for correlation and regression analyses
”,
Behavior Research Methods
, Vol.
41
No.
4
, pp.
1149
-
1160
, doi: .
Finger
,
K.
and
Pezdek
,
K.
(
1999
), “
The effect of cognitive interview on face identification accuracy: release from verbal overshadowing
”,
Journal of Applied Psychology
, Vol.
84
No.
3
, pp.
340
-
348
, doi: .
Frowd
,
C.D.
(
2021
), “Forensic facial composites”, in
Smith
,
A. M.
,
Toglia
,
M. P.
, &
Lampinen
,
J. M.
(Eds),
Methods, Measures, and Theories in Eyewitness Identification Tasks
,
Routledge
,
London
, pp.
34
-
64
, doi: .
Frowd
,
C.D.
,
Bruce
,
V.
,
Smith
,
A.J.
and
Hancock
,
P.J.B.
(
2008
), “
Improving the quality of facial composites using a holistic cognitive interview
”,
Journal of Experimental Psychology: Applied
, Vol.
14
No.
3
, pp.
276
-
287
, doi: .
Goldman
,
K.
and
Baker
,
M.A.
(
2026
), “
A test of a ‘Neo’ retrieval-based interference account of verbal overshadowing: effects of suspect description accuracy and post-description delay of suspect identifications on eyewitness recognition performance
”,
Journal of Cognitive Psychology
, Vol.
38
No.
2
, pp.
114
-
132
, doi: .
INTERPOL
(
2023
), “
ChatGPT impacts on law enforcement
”,
Background Paper
,
available at:
Link to ChatGPT impacts on law enforcementLink to the cited article (
accessed
4 September 2026).
Jacoby
,
L.L.
(
1991
), “
A process dissociation framework: separating automatic from intentional uses of memory
”,
Journal of Memory and Language
, Vol.
30
No.
5
, pp.
513
-
541
, doi: .
Kitagami
,
S.
,
Sato
,
W.
and
Yoshikawa
,
S.
(
2002
), “
The influence of test-set similarity in verbal overshadowing
”,
Applied Cognitive Psychology
, Vol.
16
No.
8
, pp.
963
-
972
, doi: .
Koo
,
T.K.
and
Li
,
M.Y.
(
2016
), “
A guideline of selecting and reporting intraclass correlation coefficients for reliability research
”,
Journal of Chiropractic Medicine
, Vol.
15
No.
2
, pp.
155
-
163
, doi: .
Lewis
,
M.B.
(
2024
), “
Fixing the stimulus-as-a-fixed-effect fallacy in forensically valid face-composite research
”,
Journal of Applied Research in Memory and Cognition
, Vol.
13
No.
2
, pp.
306
-
314
, doi: .
Loftus
,
E.F.
,
Miller
,
D.G.
and
Burns
,
H.J.
(
1978
), “
Semantic integration of verbal information in a visual memory
”,
Journal of Experimental Psychology: Human Learning and Memory
, Vol.
4
No.
1
, pp.
19
-
31
, doi: .
McIntyre
,
A.H.
,
Hancock
,
P.J.B.
,
Frowd
,
C.D.
and
Langton
,
S.R.H.
(
2016
), “
Holistic face processing can inhibit recognition of forensic facial composites
”,
Law and Human Behavior
, Vol.
40
No.
2
, pp.
128
-
135
, doi: .
McQuiston-Surrett
,
D.
and
Topp
,
L.D.
(
2008
), “
Externalizing visual images: examining the accuracy of facial descriptions vs composites as a function of the own-race bias
”,
Experimental Psychology
, Vol.
55
No.
3
, pp.
195
-
202
, doi: .
Meinhardt-Injac
,
B.
,
Persike
,
M.
and
Meinhardt
,
G.
(
2013
), “
Holistic face processing is induced by shape and texture
”,
Perception
, Vol.
42
No.
7
, pp.
716
-
732
, doi: .
Meissner
,
C.A.
,
Brigham
,
J.C.
and
Kelley
,
C.M.
(
2001
), “
The influence of retrieval processes in verbal overshadowing
”,
Memory & Cognition
, Vol.
29
No.
1
, pp.
176
-
186
, doi: .
Meissner
,
C.A.
,
Sporer
,
S.L.
and
Susa
,
K.J.
(
2008
), “
A theoretical review and meta-analysis of the description-identification relationship in memory for faces
”,
European Journal of Cognitive Psychology
, Vol.
20
No.
3
, pp.
414
-
455
, doi: .
Meissner
,
C.A.
and
Brigham
,
J.C.
(
2001
), “
A meta-analysis of the verbal overshadowing effect in face identification
”,
Applied Cognitive Psychology
, Vol.
15
No.
6
, pp.
603
-
616
, doi: .
Mickes
,
L.
,
Seale‐Carlisle
,
T.M.
,
Wetmore
,
S.A.
,
Gronlund
,
S.D.
,
Clark
,
S.E.
,
Carlson
,
C.A.
and
Wixted
,
J.T.
(
2017
), “
ROC s in eyewitness identification: instructions versus confidence ratings
”,
Applied Cognitive Psychology
, Vol.
31
No.
5
, pp.
467
-
477
, doi: .
Mota
,
Y.A.
,
Stum
,
M.D.
and
Baker
,
M.A.
(
2026
, manuscript under review), “
The AI-witness effect: AI-generated facial composites of suspects impair eyewitness recognition performance
”.
OpenAI
(
2025
), “
ChatGPT (version 5.0)
”,
available at:
Link to ChatGPT (version 5.0)Link to the cited article (
accessed
on 4 September 2026).
Osborne
,
J.W.
(
2017
),
Regression & Linear Modeling: Best Practices and Modern Methods
,
SAGE Publications, Inc
,
Thousand Oaks, CA
, doi: .
Palan
,
S.
and
Schitter
,
C.
(
2018
), “
Prolific.ac – a subject pool for online experiments
”,
Journal of Behavioral and Experimental Finance
, Vol.
17
, pp.
22
-
27
, doi: .
Russ
,
A.J.
,
Sauerland
,
M.
,
Lee
,
C.E.
and
Bindemann
,
M.
(
2018
), “
Individual differences in eyewitness accuracy across multiple lineups of faces
”,
Cognitive Research: principles and Implications
, Vol.
3
, p.
30
, doi: .
Sádaba-Campo
,
N.
and
Gómez-Moreno
,
H.
(
2025
), “
Exploration of generative neural networks for police facial sketches
”,
Big Data and Cognitive Computing
, Vol.
9
No.
2
, p.
42
, doi: .
Sauerland
,
M.
and
Sporer
,
S.L.
(
2011
), “
Written vs. spoken eyewitness accounts: does modality of testing matter?
”,
Behavioral Sciences & the Law
, Vol.
29
No.
6
, pp.
846
-
857
, doi: .
Sauerland
,
M.
,
Holub
,
F.E.
and
Sporer
,
S.L.
(
2008
), “
Person descriptions and person identifications: verbal overshadowing or recognition criterion shift?
”,
European Journal of Cognitive Psychology
, Vol.
20
No.
3
, pp.
497
-
560
, doi: .
Schooler
,
J.W.
(
2002
), “
Verbalization produces a transfer inappropriate processing shift
”,
Applied Cognitive Psychology
, Vol.
16
No.
8
, pp.
989
-
997
, doi: .
Schooler
,
J.W.
and
Engstler-Schooler
,
T.Y.
(
1990
), “
Verbal overshadowing of visual memories: some things are better left unsaid
”,
Cognitive Psychology
, Vol.
22
No.
1
, pp.
36
-
71
, doi: .
Skelton
,
F.C.
,
Frowd
,
C.D.
and
Speers
,
K.E.
(
2015
), “
The benefit of context for facial-composite construction
”,
Journal of Forensic Practice
, Vol.
17
No.
4
, pp.
281
-
290
, doi: .
Smith
,
A.M.
,
Wells
,
G.L.
,
Lindsay
,
R.C.L.
and
Myerson
,
T.
(
2018
), “
Eyewitness identification performance on showups improves with an additional-opportunities instruction: evidence for present-absent criteria discrepancy
”,
Law and Human Behavior
, Vol.
42
No.
3
, p.
215
, doi: .
Smith
,
H.M.J.
and
Flowe
,
H.D.
(
2015
), “
ROC analysis of the verbal overshadowing effect: testing the effect of verbalisation on memory sensitivity
”,
Applied Cognitive Psychology
, Vol.
29
No.
2
, pp.
159
-
168
, doi: .
Sporer
,
S.L.
(
1993
), “
Eyewitness identification accuracy, confidence, and decision times in simultaneous and sequential lineups
”,
Journal of Applied Psychology
, Vol.
78
No.
1
, pp.
22
-
33
.
Sporer
,
S.L.
(
2007
), “
Person descriptions as retrieval cues: do they really help?
”,
Psychology, Crime & Law
, Vol.
13
No.
6
, pp.
591
-
609
, doi: .
Sporer
,
S.L.
,
Tredoux
,
C.G.
,
Vredeveldt
,
A.
,
Kempen
,
K.
and
Nortje
,
A.
(
2020
), “
Does exposure to facial composites damage eyewitness memory? A comprehensive review
”,
Applied Cognitive Psychology
, Vol.
34
No.
5
, pp.
1166
-
1179
, doi: .
Sporer
,
S.L.
(
1996
), “Psychological aspects of person descriptions”, in
Sporer
,
S.L.
,
Malpass
,
R.S.
&
Köhnken
,
G.
(Eds),
Psychological Issues in Eyewitness Identification
,
Erlbaum
,
Mahwah, NJ
, pp.
53
-
86
.
ten Hove
,
D.
,
Jorgensen
,
T.D.
and
van der Ark
,
L.A.
(
2024
), “
Updated guidelines on selecting an intraclass correlation coefficient for interrater reliability, with applications to incomplete observational designs
”,
Psychological Methods
, Vol.
29
No.
5
, pp.
967
-
979
, doi: .
Topp‐Manriquez
,
L.D.
,
McQuiston
,
D.
and
Malpass
,
R.S.
(
2016
), “
Facial composites and the misinformation effect: how composites distort memory
”,
Legal and Criminological Psychology
, Vol.
21
No.
2
, pp.
372
-
389
, doi: .
Tredoux
,
C.G.
,
Sporer
,
S.L.
,
Vredeveldt
,
A.
,
Kempen
,
K.
and
Nortje
,
A.
(
2021
), “
Does constructing a facial composite affect eyewitness memory? A research synthesis and meta-analysis
”,
Journal of Experimental Criminology
, Vol.
17
No.
4
, pp.
713
-
741
, doi: .
Wells
,
G.L.
(
2020
), “
Psychological science on eyewitness identification and its impact on police practices and policies
”,
American Psychologist
, Vol.
75
No.
9
, pp.
1316
-
1329
, doi: .
Wells
,
G.L.
,
Charman
,
S.D.
and
Olson
,
E.A.
(
2005
), “
Building face composites can harm lineup identification performance
”,
Journal of Experimental Psychology: Applied
, Vol.
11
No.
3
, pp.
147
-
156
, doi: .
Wells
,
G.L.
,
Smith
,
A.M.
and
Smalarz
,
L.
(
2015
), “
ROC analysis of lineups obscures information that is critical for both theoretical understanding and applied purposes
”,
Journal of Applied Research in Memory and Cognition
, Vol.
4
No.
4
, pp.
324
-
328
, doi: .
Wickham
,
L.H.V.
and
Swift
,
H.
(
2006
), “
Articulatory suppression attenuates the verbal overshadowing effect: a role for verbal encoding in face identification
”,
Applied Cognitive Psychology
, Vol.
20
No.
2
, pp.
157
-
169
, doi: .
Wilkins
,
J.
(
2025
), “
Police admit they’re using ChatGPT to generate ‘sketches’ of suspects
”,
Futurism
,
available at:
Link to Police admit they’re using ChatGPT to generate ‘sketches’ of suspectsLink to the cited article (
accessed
4 September 2026).
Wilson
,
B.M.
,
Seale-Carlisle
,
T.M.
and
Mickes
,
L.
(
2018
), “
The effects of verbal descriptions on performance in lineups and showups
”,
Journal of Experimental Psychology: General
, Vol.
147
No.
1
, pp.
113
-
124
, doi: .
Windschitl
,
P.D.
(
1996
), “
Memory for faces: evidence of retrieval-based impairment
”,
Journal of Experimental Psychology: Learning, Memory, and Cognition
, Vol.
22
No.
5
, pp.
1101
-
1122
, doi: .
Wixted
,
J.T.
and
Mickes
,
L.
(
2014
), “
A signal-detection-based diagnostic-feature-detection model of eyewitness identification
”,
Psychological Review
, Vol.
121
No.
2
, pp.
262
-
276
, doi: .
Wright
,
D.B.
,
Horry
,
R.
and
Skagerberg
,
E.M.
(
2009
), “
Functions for traditional and multilevel approaches to signal detection theory
”,
Behavior Research Methods
, Vol.
41
No.
2
, pp.
257
-
267
, doi: .
Published by Emerald Publishing Limited. This article is published under the Creative Commons Attribution (CC BY 4.0) licence. Anyone may reproduce, distribute, translate and create derivative works of this article (for both commercial and non-commercial purposes), subject to full attribution to the original publication and authors. The full terms of this licence may be seen at Link to the terms of the CC BY 4.0 licenceLink to the terms of the CC BY 4.0 licence.

or Create an Account

Close subscription notice
Close access options