Table A1

Studies included in the analysis, their research approaches, methods and data

StudyStudy approachMethodologyData
Binns et al. (2018) An experimentation into people's perceptions of justice in algorithmic decision making under different scenarios and explanation stylesThree studies: (1) A scenario-based in-lab user study. Data collected via expert interviews
(2) A between-subjects study of five scenarios, with 65 participants in each scenario
(3) A within-subjects study with 65 participants in loan and insurance cases
In-lab study: 19 participants from a small town in the UK
The two online studies: 390 British participants over 18 years old recruited via Prolific Academic
Brennen (2020) An exploration into what stakeholders want from XAIExpert interviewsCompany founders, investors, potential end users, and academia members (n = 40). Presumably from the USA due to the authors' universities
Broekens et al. (2010) An examination of the usefulness and naturalness of types of explanations for the actions of agents based on belief desire intention (BDI)A between-subjects study of three conditions with 10 participants in each scenarioStudy subjects were a “balanced mix of family, friends, colleagues, and students of the first two authors,” with an average education level between bachelor's and master's. (n = 30). Presumably from the Netherlands due to the authors' university
Bussone et al. (2015) An exploration into how explanations are related to domain experts' trust and reliance on clinical decision support systemsAn exploratory between-group user study employing two different versions of a CDSS prototype (comprehensive vs selective version). Using the think-aloud method, participants' decision making and trust when exploring the prototype was analyzedEight participants (seven primary care practitioners and one nurse), recruited through ads in medical network groups and forums, medical schools, and local primary care offices in the UK
Cai et al. (2019) An investigation into what are the key types of information that medical experts want and need when introduced to a diagnostic AI assistantA three-phase lab study with pathologists. Participants were interviewed before, during, and after presenting DNN predictions for prostate cancer diagnosis to explore the types of information they needed from the AI assistant21 pathologists participated in the study, recruited from a pool of remote contractors assisting Google Health with pathology projects. Presumably from the USA due to the authors' university
Chazette and Schneider (2020) An exploration into what users see as the advantages and disadvantages of embedded explanations in software systems and what is their current level of transparencyAn online questionnaire using LimeSurvey with 16 questions (11 multiple choice, five open-ended) on software skills and use, explanation needs and frequency, and the presentation of explanationsSnowball sampling with the help of the personal networks of the authors. Target population included adult end users of all ages, with different occupations. In total, 107 respondents completed the survey, of which 84% were from Brazil and 16% from Germany
Cheng et al. (2019) An investigation of what kind of design principles would help non-expert stakeholders to understand how decision-making algorithms workAn online between-subjects study of five conditions, with each participant randomly assigned to a scenario, completed via Amazon MTurkParticipants of the study were recruited from Amazon MTurk. To qualify for the study participants needed to reside in the US, be aged 18 or above, and have a HIT approval rate of 90% or above. (n = 199)
Cirqueira et al. (2020) A demonstration of the usage of scenario-based requirement elicitation for XAI in a fraud detection contextA problem-centered expert interview study to validate two fraud detection scenarios that could be adopted to identify expert requirements for adequate explanationsThree banking fraud specialists from one bank in Austria participated in the study, but the recruitment process was not provided
Cramer et al. (2008) Examines the influence of transparency on users' trust and acceptance of content-based art recommendation systemsA between-subjects user study of three conditions with 22 participants in the first condition and 19 in the second and thirdParticipants were volunteers from the researchers' personal and professional networks, were relatively well educated, and had a good knowledge using computers. The participants' country of origin was not stated, presumably the Netherlands based on the authors' university
Dodge et al. (2019) An exploration into how four types of programmatically created explanations affect people's fairness judgment of ML systemsAn online survey with four explanation styles; each participant presented with six fairness judgment cases160 people from Amazon MTurk, with criteria that the participant must live in the US and have completed at least 1,000 tasks in MTurk with at least a 98% approval rate
Ehsan et al. (2019) An investigation of how to train a neural rationale generator to create rationale styles and how people perceive themBoth between- and within-subjects user study with participants split into two equal groups with two identical experimental conditions, differing only by type of candidate rationale. The first group evaluated a focused-view rationale, whereas the second group had complete-view rationales. Participants were asked to view five videos with a set of rationales each and to rate each rationale based on four different statements128 participants were recruited through TurkPrime: 93% of the participants lived in the US, while the 7% that were left were from India
Eiband et al. (2018) An explorative quest to advance existing UI guidelines for increased transparency and to improve users' mental models, with the particular case of Freeletics Bodyweight ApplicationA stage-based participatory process consisting of different phases: (1) semi-structured interviews of app users about their current mental models, (2) card sorting to find which components of the app users thought relevant for the perceived transparency of the app, (3) user testing of the prototype versions of the new UI, (4) evaluation of the prototypes with users14 active users of the app were recruited for the interviews in a park (presumably in Germany), and 11 users for the card sorting, a mixture of long- and short-term users of the app. The number of participants in the prototype testing was not stated. In the evaluation stage, 15 users participated. The participants in all stages were presumably from Germany, based on the authors' university
Eslami et al. (2018) An investigation of how revealing users' parts of the algorithmic process affects their perceptions of online advertising and its platformsA lab study in which users viewed the actual personalized ads and explanations the advertisers had given to them, followed by what an advertising algorithm had inferred about them. In the last phase, users created themselves advertising in an ad creation interface and wrote their own desired explanations for an ad of a product of interest32 participants from San Francisco, United States, and the surrounding area. Participants were picked from a larger group of interested people by non-probability modified quota sampling to balance five characteristics with the proportions of the US's population: gender, age, education, race/ethnicity and socioeconomic status
Lim and Dey (2009) An experimentation into what kind of information demands and under which circumstances users have them using four real-world applicationsTwo experiments: (1) A between-subjects survey study in which participants were shown three to four scenarios of one application followed by two to five instances of the scenarios. Participants were asked to describe their feelings about the application and what kind of information needs they would have
(2) A between-subjects study for intelligibility types, which was formulated based on the results of experiment one. Participants were assigned to a survey with only one intelligibility type. Participants were asked to rate their satisfaction with the application using a seven-point Likert scale and questions about the usefulness of explanations
(1) 250 participants in the first experiment recruited from Amazon Mechanical Turk. (2) 610 participants in the second experiment, recruited also from MTurk. Participants were distributed evenly across the 12 conditions
The geographical distribution of the participants was not provided
Lim et al. (2009) An examination into what kind of explanation types are the most effective to describe the workings of context-aware intelligent systemsTwo experiments. In both, participants were allowed to explore the system, after which their understanding of the system was tested. (1) A between-subjects study with three conditions to explore the effectiveness of question types.
(2) Same procedure as in the first, but combined with two additional conditions (“What If” and “How To”)
(1) 53 participants in the first experiment, divided between the three conditions, and (2) 158 participants in the second experiment, divided (not evenly) among the five conditions. Recruitment procedure and country of the participants were not stated
Ngo et al. (2020) An examination of users' mental models in using recommender systems, namely, NetflixA semi-structured interview study focusing on participants' experience with Netflix. Participants were asked about the workings and data processing of Netflix and asked to draw their own image of Netflix10 interviewees with advanced experience with Netflix. Recruitment procedure and country of origin not stated
Oh et al. (2018) An investigation to understand the user experience in art co-creation with AIA between-subjects study with four conditions and a treatment condition. Participants performed a series of drawing tasks with a think-aloud method and were interviewed afterwards about their experience with the tool. Users' experience was also quantitatively measured with a survey30 participants were recruited through an announcement in Seoul National University's online community website (thus, they are presumed to be South Korean)
Putnam and Conati (2019) A quest to understand whether and when it is necessary for an intelligent tutoring system to explain its underlying user modeling techniques to studentsAn user experiment in which participants studied the materials provided, did a pre-test based on the context of materials, and used an adaptive constraint satisfaction problem (CSP) applet to solve two CSPs, followed by a post-test questionnaire and interviewNine participants (university students) recruited from an introductory AI course at a university in North America
Schrills and Franke (2020) An investigation of how prototypical visualization approaches aimed at increasing the explainability of ML systems affect users' perceived trustworthiness and observability of the systemAn online within-subjects study with three conditions that presented different information visualizations. Users' agreement with the classification was measured after each stimulus83 participants were recruited via e-mail, social networks, and at the local university. Geographic distribution of the participants was not provided
van der Waa et al. (2020) An investigation of what properties make a confidence measure desirable and why, and how an interpretable confidence measure (ICM) is interpreted by usersTwo user experiments: (1) An interview study with domain experts to evaluate “transparency of the case-based reasoning approach underlying an ICM compared to other confidence measures”
(2) Quantitative online survey with users to evaluate users' interests and preferences regarding the explanations provided by a decision-support system (autonomous car) regarding its confidence in its advice
(1) “Several domain experts” participated in the study. Recruitment process and country of origin not stated
(2) 40 participants recruited via Amazon Mechanical Turk
Wang et al. (2019) A quest for designing a conceptual framework for building human-centered, decision-theory-driven XAI, based on which an explainable clinical diagnostic tool for intensive care phenotyping was designed in co-creation with cliniciansA co-design lab study with clinicians. Participants were asked to use the diagnostics dashboard and diagnose patient cases using it. Sessions were recorded, and participants were instructed to think aloud during their diagnostic process14 medical professionals recruited from a local hospital. Country and background were not stated
Weitz et al. (2019a) An exploration of how incorporating virtual agents into XAI designs affects the perceived trust of usersA between-subjects user study. Participants interacted with a graphical user interface and were split into four test groups with different types of visualizations and audio explanations. Participants were asked to rate their impressions and trust in the system60 participants. Recruitment process and background of the participants were not provided
Weitz et al. (2019b) An examination of how using virtual agents in explanations affects the trustworthiness of autonomous intelligent systemsA between-subjects user study with two conditions. The first group received explanations from a virtual agent while the other received only visual explanations. Users' perceived trust of the system was measured afterwards with a questionnaire30 participants. Recruitment process and the participants' country of origin were not stated
Xie et al. (2019) An exploration into what medical professionals consider as explainable when interacting with data for diagnosis and treatment purposesAn interview study consisting of questions revolving around the professionals' working practices, challenges, and experience using computer-based systems to facilitate medical workSample consisted of six medical professionals from California, US, recruited via online participant call
Yin et al. (2019) An examination of whether people's trust in a model varies depending on the model's stated accuracy on “held-out data” and on its observed accuracy in practiceThree experiments: (1) A between-subjects study with 10 treatments. Users were randomly assigned to one of five accuracy levels and asked to make predictions about the outcomes of 40 speed dating events
(2) A between-subjects study with two sub-experiments and two conditions with different levels of observed accuracy. Users were again asked to predict the outcome of 40 speed dates
(3) A between-subjects study with six conditions varying along the stated accuracy and observed accuracy, again with the same 40 prediction tasks
There were 1,994 participants in the first experiment, 757 participants in the second, and 1,042 participants in the third. All participants were from the United States and recruited via Amazon MTurk

or Create an Account

Close subscription notice
Close access options