Studies included in the analysis, their research approaches, methods and data
| Study | Study approach | Methodology | Data |
|---|---|---|---|
| Binns et al. (2018) | An experimentation into people's perceptions of justice in algorithmic decision making under different scenarios and explanation styles | Three studies: (1) A scenario-based in-lab user study. Data collected via expert interviews (2) A between-subjects study of five scenarios, with 65 participants in each scenario (3) A within-subjects study with 65 participants in loan and insurance cases | In-lab study: 19 participants from a small town in the UK The two online studies: 390 British participants over 18 years old recruited via Prolific Academic |
| Brennen (2020) | An exploration into what stakeholders want from XAI | Expert interviews | Company founders, investors, potential end users, and academia members (n = 40). Presumably from the USA due to the authors' universities |
| Broekens et al. (2010) | An examination of the usefulness and naturalness of types of explanations for the actions of agents based on belief desire intention (BDI) | A between-subjects study of three conditions with 10 participants in each scenario | Study subjects were a “balanced mix of family, friends, colleagues, and students of the first two authors,” with an average education level between bachelor's and master's. (n = 30). Presumably from the Netherlands due to the authors' university |
| Bussone et al. (2015) | An exploration into how explanations are related to domain experts' trust and reliance on clinical decision support systems | An exploratory between-group user study employing two different versions of a CDSS prototype (comprehensive vs selective version). Using the think-aloud method, participants' decision making and trust when exploring the prototype was analyzed | Eight participants (seven primary care practitioners and one nurse), recruited through ads in medical network groups and forums, medical schools, and local primary care offices in the UK |
| Cai et al. (2019) | An investigation into what are the key types of information that medical experts want and need when introduced to a diagnostic AI assistant | A three-phase lab study with pathologists. Participants were interviewed before, during, and after presenting DNN predictions for prostate cancer diagnosis to explore the types of information they needed from the AI assistant | 21 pathologists participated in the study, recruited from a pool of remote contractors assisting Google Health with pathology projects. Presumably from the USA due to the authors' university |
| Chazette and Schneider (2020) | An exploration into what users see as the advantages and disadvantages of embedded explanations in software systems and what is their current level of transparency | An online questionnaire using LimeSurvey with 16 questions (11 multiple choice, five open-ended) on software skills and use, explanation needs and frequency, and the presentation of explanations | Snowball sampling with the help of the personal networks of the authors. Target population included adult end users of all ages, with different occupations. In total, 107 respondents completed the survey, of which 84% were from Brazil and 16% from Germany |
| Cheng et al. (2019) | An investigation of what kind of design principles would help non-expert stakeholders to understand how decision-making algorithms work | An online between-subjects study of five conditions, with each participant randomly assigned to a scenario, completed via Amazon MTurk | Participants of the study were recruited from Amazon MTurk. To qualify for the study participants needed to reside in the US, be aged 18 or above, and have a HIT approval rate of 90% or above. (n = 199) |
| Cirqueira et al. (2020) | A demonstration of the usage of scenario-based requirement elicitation for XAI in a fraud detection context | A problem-centered expert interview study to validate two fraud detection scenarios that could be adopted to identify expert requirements for adequate explanations | Three banking fraud specialists from one bank in Austria participated in the study, but the recruitment process was not provided |
| Cramer et al. (2008) | Examines the influence of transparency on users' trust and acceptance of content-based art recommendation systems | A between-subjects user study of three conditions with 22 participants in the first condition and 19 in the second and third | Participants were volunteers from the researchers' personal and professional networks, were relatively well educated, and had a good knowledge using computers. The participants' country of origin was not stated, presumably the Netherlands based on the authors' university |
| Dodge et al. (2019) | An exploration into how four types of programmatically created explanations affect people's fairness judgment of ML systems | An online survey with four explanation styles; each participant presented with six fairness judgment cases | 160 people from Amazon MTurk, with criteria that the participant must live in the US and have completed at least 1,000 tasks in MTurk with at least a 98% approval rate |
| Ehsan et al. (2019) | An investigation of how to train a neural rationale generator to create rationale styles and how people perceive them | Both between- and within-subjects user study with participants split into two equal groups with two identical experimental conditions, differing only by type of candidate rationale. The first group evaluated a focused-view rationale, whereas the second group had complete-view rationales. Participants were asked to view five videos with a set of rationales each and to rate each rationale based on four different statements | 128 participants were recruited through TurkPrime: 93% of the participants lived in the US, while the 7% that were left were from India |
| Eiband et al. (2018) | An explorative quest to advance existing UI guidelines for increased transparency and to improve users' mental models, with the particular case of Freeletics Bodyweight Application | A stage-based participatory process consisting of different phases: (1) semi-structured interviews of app users about their current mental models, (2) card sorting to find which components of the app users thought relevant for the perceived transparency of the app, (3) user testing of the prototype versions of the new UI, (4) evaluation of the prototypes with users | 14 active users of the app were recruited for the interviews in a park (presumably in Germany), and 11 users for the card sorting, a mixture of long- and short-term users of the app. The number of participants in the prototype testing was not stated. In the evaluation stage, 15 users participated. The participants in all stages were presumably from Germany, based on the authors' university |
| Eslami et al. (2018) | An investigation of how revealing users' parts of the algorithmic process affects their perceptions of online advertising and its platforms | A lab study in which users viewed the actual personalized ads and explanations the advertisers had given to them, followed by what an advertising algorithm had inferred about them. In the last phase, users created themselves advertising in an ad creation interface and wrote their own desired explanations for an ad of a product of interest | 32 participants from San Francisco, United States, and the surrounding area. Participants were picked from a larger group of interested people by non-probability modified quota sampling to balance five characteristics with the proportions of the US's population: gender, age, education, race/ethnicity and socioeconomic status |
| Lim and Dey (2009) | An experimentation into what kind of information demands and under which circumstances users have them using four real-world applications | Two experiments: (1) A between-subjects survey study in which participants were shown three to four scenarios of one application followed by two to five instances of the scenarios. Participants were asked to describe their feelings about the application and what kind of information needs they would have (2) A between-subjects study for intelligibility types, which was formulated based on the results of experiment one. Participants were assigned to a survey with only one intelligibility type. Participants were asked to rate their satisfaction with the application using a seven-point Likert scale and questions about the usefulness of explanations | (1) 250 participants in the first experiment recruited from Amazon Mechanical Turk. (2) 610 participants in the second experiment, recruited also from MTurk. Participants were distributed evenly across the 12 conditions The geographical distribution of the participants was not provided |
| Lim et al. (2009) | An examination into what kind of explanation types are the most effective to describe the workings of context-aware intelligent systems | Two experiments. In both, participants were allowed to explore the system, after which their understanding of the system was tested. (1) A between-subjects study with three conditions to explore the effectiveness of question types. (2) Same procedure as in the first, but combined with two additional conditions (“What If” and “How To”) | (1) 53 participants in the first experiment, divided between the three conditions, and (2) 158 participants in the second experiment, divided (not evenly) among the five conditions. Recruitment procedure and country of the participants were not stated |
| Ngo et al. (2020) | An examination of users' mental models in using recommender systems, namely, Netflix | A semi-structured interview study focusing on participants' experience with Netflix. Participants were asked about the workings and data processing of Netflix and asked to draw their own image of Netflix | 10 interviewees with advanced experience with Netflix. Recruitment procedure and country of origin not stated |
| Oh et al. (2018) | An investigation to understand the user experience in art co-creation with AI | A between-subjects study with four conditions and a treatment condition. Participants performed a series of drawing tasks with a think-aloud method and were interviewed afterwards about their experience with the tool. Users' experience was also quantitatively measured with a survey | 30 participants were recruited through an announcement in Seoul National University's online community website (thus, they are presumed to be South Korean) |
| Putnam and Conati (2019) | A quest to understand whether and when it is necessary for an intelligent tutoring system to explain its underlying user modeling techniques to students | An user experiment in which participants studied the materials provided, did a pre-test based on the context of materials, and used an adaptive constraint satisfaction problem (CSP) applet to solve two CSPs, followed by a post-test questionnaire and interview | Nine participants (university students) recruited from an introductory AI course at a university in North America |
| Schrills and Franke (2020) | An investigation of how prototypical visualization approaches aimed at increasing the explainability of ML systems affect users' perceived trustworthiness and observability of the system | An online within-subjects study with three conditions that presented different information visualizations. Users' agreement with the classification was measured after each stimulus | 83 participants were recruited via e-mail, social networks, and at the local university. Geographic distribution of the participants was not provided |
| van der Waa et al. (2020) | An investigation of what properties make a confidence measure desirable and why, and how an interpretable confidence measure (ICM) is interpreted by users | Two user experiments: (1) An interview study with domain experts to evaluate “transparency of the case-based reasoning approach underlying an ICM compared to other confidence measures” (2) Quantitative online survey with users to evaluate users' interests and preferences regarding the explanations provided by a decision-support system (autonomous car) regarding its confidence in its advice | (1) “Several domain experts” participated in the study. Recruitment process and country of origin not stated (2) 40 participants recruited via Amazon Mechanical Turk |
| Wang et al. (2019) | A quest for designing a conceptual framework for building human-centered, decision-theory-driven XAI, based on which an explainable clinical diagnostic tool for intensive care phenotyping was designed in co-creation with clinicians | A co-design lab study with clinicians. Participants were asked to use the diagnostics dashboard and diagnose patient cases using it. Sessions were recorded, and participants were instructed to think aloud during their diagnostic process | 14 medical professionals recruited from a local hospital. Country and background were not stated |
| Weitz et al. (2019a) | An exploration of how incorporating virtual agents into XAI designs affects the perceived trust of users | A between-subjects user study. Participants interacted with a graphical user interface and were split into four test groups with different types of visualizations and audio explanations. Participants were asked to rate their impressions and trust in the system | 60 participants. Recruitment process and background of the participants were not provided |
| Weitz et al. (2019b) | An examination of how using virtual agents in explanations affects the trustworthiness of autonomous intelligent systems | A between-subjects user study with two conditions. The first group received explanations from a virtual agent while the other received only visual explanations. Users' perceived trust of the system was measured afterwards with a questionnaire | 30 participants. Recruitment process and the participants' country of origin were not stated |
| Xie et al. (2019) | An exploration into what medical professionals consider as explainable when interacting with data for diagnosis and treatment purposes | An interview study consisting of questions revolving around the professionals' working practices, challenges, and experience using computer-based systems to facilitate medical work | Sample consisted of six medical professionals from California, US, recruited via online participant call |
| Yin et al. (2019) | An examination of whether people's trust in a model varies depending on the model's stated accuracy on “held-out data” and on its observed accuracy in practice | Three experiments: (1) A between-subjects study with 10 treatments. Users were randomly assigned to one of five accuracy levels and asked to make predictions about the outcomes of 40 speed dating events (2) A between-subjects study with two sub-experiments and two conditions with different levels of observed accuracy. Users were again asked to predict the outcome of 40 speed dates (3) A between-subjects study with six conditions varying along the stated accuracy and observed accuracy, again with the same 40 prediction tasks | There were 1,994 participants in the first experiment, 757 participants in the second, and 1,042 participants in the third. All participants were from the United States and recruited via Amazon MTurk |
| Study | Study approach | Methodology | Data |
|---|---|---|---|
| An experimentation into people's perceptions of justice in algorithmic decision making under different scenarios and explanation styles | Three studies: (1) A scenario-based in-lab user study. Data collected via expert interviews | In-lab study: 19 participants from a small town in the UK | |
| An exploration into what stakeholders want from XAI | Expert interviews | Company founders, investors, potential end users, and academia members ( | |
| An examination of the usefulness and naturalness of types of explanations for the actions of agents based on belief desire intention (BDI) | A between-subjects study of three conditions with 10 participants in each scenario | Study subjects were a “balanced mix of family, friends, colleagues, and students of the first two authors,” with an average education level between bachelor's and master's. ( | |
| An exploration into how explanations are related to domain experts' trust and reliance on clinical decision support systems | An exploratory between-group user study employing two different versions of a CDSS prototype (comprehensive vs selective version). Using the think-aloud method, participants' decision making and trust when exploring the prototype was analyzed | Eight participants (seven primary care practitioners and one nurse), recruited through ads in medical network groups and forums, medical schools, and local primary care offices in the UK | |
| An investigation into what are the key types of information that medical experts want and need when introduced to a diagnostic AI assistant | A three-phase lab study with pathologists. Participants were interviewed before, during, and after presenting DNN predictions for prostate cancer diagnosis to explore the types of information they needed from the AI assistant | 21 pathologists participated in the study, recruited from a pool of remote contractors assisting Google Health with pathology projects. Presumably from the USA due to the authors' university | |
| An exploration into what users see as the advantages and disadvantages of embedded explanations in software systems and what is their current level of transparency | An online questionnaire using LimeSurvey with 16 questions (11 multiple choice, five open-ended) on software skills and use, explanation needs and frequency, and the presentation of explanations | Snowball sampling with the help of the personal networks of the authors. Target population included adult end users of all ages, with different occupations. In total, 107 respondents completed the survey, of which 84% were from Brazil and 16% from Germany | |
| An investigation of what kind of design principles would help non-expert stakeholders to understand how decision-making algorithms work | An online between-subjects study of five conditions, with each participant randomly assigned to a scenario, completed via Amazon MTurk | Participants of the study were recruited from Amazon MTurk. To qualify for the study participants needed to reside in the US, be aged 18 or above, and have a HIT approval rate of 90% or above. ( | |
| A demonstration of the usage of scenario-based requirement elicitation for XAI in a fraud detection context | A problem-centered expert interview study to validate two fraud detection scenarios that could be adopted to identify expert requirements for adequate explanations | Three banking fraud specialists from one bank in Austria participated in the study, but the recruitment process was not provided | |
| Examines the influence of transparency on users' trust and acceptance of content-based art recommendation systems | A between-subjects user study of three conditions with 22 participants in the first condition and 19 in the second and third | Participants were volunteers from the researchers' personal and professional networks, were relatively well educated, and had a good knowledge using computers. The participants' country of origin was not stated, presumably the Netherlands based on the authors' university | |
| An exploration into how four types of programmatically created explanations affect people's fairness judgment of ML systems | An online survey with four explanation styles; each participant presented with six fairness judgment cases | 160 people from Amazon MTurk, with criteria that the participant must live in the US and have completed at least 1,000 tasks in MTurk with at least a 98% approval rate | |
| An investigation of how to train a neural rationale generator to create rationale styles and how people perceive them | Both between- and within-subjects user study with participants split into two equal groups with two identical experimental conditions, differing only by type of candidate rationale. The first group evaluated a focused-view rationale, whereas the second group had complete-view rationales. Participants were asked to view five videos with a set of rationales each and to rate each rationale based on four different statements | 128 participants were recruited through TurkPrime: 93% of the participants lived in the US, while the 7% that were left were from India | |
| An explorative quest to advance existing UI guidelines for increased transparency and to improve users' mental models, with the particular case of Freeletics Bodyweight Application | A stage-based participatory process consisting of different phases: (1) semi-structured interviews of app users about their current mental models, (2) card sorting to find which components of the app users thought relevant for the perceived transparency of the app, (3) user testing of the prototype versions of the new UI, (4) evaluation of the prototypes with users | 14 active users of the app were recruited for the interviews in a park (presumably in Germany), and 11 users for the card sorting, a mixture of long- and short-term users of the app. The number of participants in the prototype testing was not stated. In the evaluation stage, 15 users participated. The participants in all stages were presumably from Germany, based on the authors' university | |
| An investigation of how revealing users' parts of the algorithmic process affects their perceptions of online advertising and its platforms | A lab study in which users viewed the actual personalized ads and explanations the advertisers had given to them, followed by what an advertising algorithm had inferred about them. In the last phase, users created themselves advertising in an ad creation interface and wrote their own desired explanations for an ad of a product of interest | 32 participants from San Francisco, United States, and the surrounding area. Participants were picked from a larger group of interested people by non-probability modified quota sampling to balance five characteristics with the proportions of the US's population: gender, age, education, race/ethnicity and socioeconomic status | |
| An experimentation into what kind of information demands and under which circumstances users have them using four real-world applications | Two experiments: (1) A between-subjects survey study in which participants were shown three to four scenarios of one application followed by two to five instances of the scenarios. Participants were asked to describe their feelings about the application and what kind of information needs they would have | (1) 250 participants in the first experiment recruited from Amazon Mechanical Turk. (2) 610 participants in the second experiment, recruited also from MTurk. Participants were distributed evenly across the 12 conditions | |
| An examination into what kind of explanation types are the most effective to describe the workings of context-aware intelligent systems | Two experiments. In both, participants were allowed to explore the system, after which their understanding of the system was tested. (1) A between-subjects study with three conditions to explore the effectiveness of question types. | (1) 53 participants in the first experiment, divided between the three conditions, and (2) 158 participants in the second experiment, divided (not evenly) among the five conditions. Recruitment procedure and country of the participants were not stated | |
| An examination of users' mental models in using recommender systems, namely, Netflix | A semi-structured interview study focusing on participants' experience with Netflix. Participants were asked about the workings and data processing of Netflix and asked to draw their own image of Netflix | 10 interviewees with advanced experience with Netflix. Recruitment procedure and country of origin not stated | |
| An investigation to understand the user experience in art co-creation with AI | A between-subjects study with four conditions and a treatment condition. Participants performed a series of drawing tasks with a think-aloud method and were interviewed afterwards about their experience with the tool. Users' experience was also quantitatively measured with a survey | 30 participants were recruited through an announcement in Seoul National University's online community website (thus, they are presumed to be South Korean) | |
| A quest to understand whether and when it is necessary for an intelligent tutoring system to explain its underlying user modeling techniques to students | An user experiment in which participants studied the materials provided, did a pre-test based on the context of materials, and used an adaptive constraint satisfaction problem (CSP) applet to solve two CSPs, followed by a post-test questionnaire and interview | Nine participants (university students) recruited from an introductory AI course at a university in North America | |
| An investigation of how prototypical visualization approaches aimed at increasing the explainability of ML systems affect users' perceived trustworthiness and observability of the system | An online within-subjects study with three conditions that presented different information visualizations. Users' agreement with the classification was measured after each stimulus | 83 participants were recruited via e-mail, social networks, and at the local university. Geographic distribution of the participants was not provided | |
| An investigation of what properties make a confidence measure desirable and why, and how an interpretable confidence measure (ICM) is interpreted by users | Two user experiments: (1) An interview study with domain experts to evaluate “transparency of the case-based reasoning approach underlying an ICM compared to other confidence measures” | (1) “Several domain experts” participated in the study. Recruitment process and country of origin not stated | |
| A quest for designing a conceptual framework for building human-centered, decision-theory-driven XAI, based on which an explainable clinical diagnostic tool for intensive care phenotyping was designed in co-creation with clinicians | A co-design lab study with clinicians. Participants were asked to use the diagnostics dashboard and diagnose patient cases using it. Sessions were recorded, and participants were instructed to think aloud during their diagnostic process | 14 medical professionals recruited from a local hospital. Country and background were not stated | |
| An exploration of how incorporating virtual agents into XAI designs affects the perceived trust of users | A between-subjects user study. Participants interacted with a graphical user interface and were split into four test groups with different types of visualizations and audio explanations. Participants were asked to rate their impressions and trust in the system | 60 participants. Recruitment process and background of the participants were not provided | |
| An examination of how using virtual agents in explanations affects the trustworthiness of autonomous intelligent systems | A between-subjects user study with two conditions. The first group received explanations from a virtual agent while the other received only visual explanations. Users' perceived trust of the system was measured afterwards with a questionnaire | 30 participants. Recruitment process and the participants' country of origin were not stated | |
| An exploration into what medical professionals consider as explainable when interacting with data for diagnosis and treatment purposes | An interview study consisting of questions revolving around the professionals' working practices, challenges, and experience using computer-based systems to facilitate medical work | Sample consisted of six medical professionals from California, US, recruited via online participant call | |
| An examination of whether people's trust in a model varies depending on the model's stated accuracy on “held-out data” and on its observed accuracy in practice | Three experiments: (1) A between-subjects study with 10 treatments. Users were randomly assigned to one of five accuracy levels and asked to make predictions about the outcomes of 40 speed dating events | There were 1,994 participants in the first experiment, 757 participants in the second, and 1,042 participants in the third. All participants were from the United States and recruited via Amazon MTurk |
Sharing content requires targeting cookies to be enabled. Please update your cookie preferences to use this feature.