The What Works in Character Education project (WWCE; Berkowitz & Bier, 2007) was a descriptive summary of significant findings in the empirical literature on character education for the period 1945 to 2004. It employed a broad definition of character education to include related positive youth development programs such as socioemotional learning. Drawing on studies carried out in K–12 education settings, it encompassed research on published character development curricula as well as research on particular instructional practices such as moral discussion and cooperative learning. The current meta-analysis revisits WWCE by applying meta-analytic techniques to the same set of studies collected as part of the WWCE review. While the current analysis did not include any studies not covered by WWCE, it did modify its approach to examining 5 systematic reviews for which Berkowitz and Bier (2007) looked to the overall conclusions. In contrast, the current analysis expanded these reviews to individually examine each of their respective studies. It complements the already substantial referencing and discussion of the WWCE that exists in the character education literature. Random effects meta-analysis of the original WWCE studies examined 836 comparisons across 64 studies involving 96,930 participants and yielded a small but statistically significant average effect size, g = 0.33, 95% CI = [0.21, 0.45]. This finding is consistent with prior evidence on the effect size associated with various educational interventions (Lipsey & Wilson, 1993) and reaffirms that character education, as conceptualized and analyzed in the original WWCE, is effective in positively influencing character outcomes. The metaanalytic techniques also yielded new evidence of substantial heterogeneity of effects across studies, suggesting the potential value of focusing more narrowly, in future meta-analyses, on the character education programs that appear most promising. In the meta-analysis reported here, there were also indications of possible bias in some studies’ reporting of results, but overall conclusions regarding effectiveness remaine valid even after correcting for possible bias
Character education has been a topic of renewed interest for many years (Berkowitz & Bier, 2007; Leming, 1993; Lickona, 1991; McClellan, 1999; Mulkey, 1997; Sojourner, 2012). Character education programs have garnered significant private and public support, with substantial funding from both corporations and government bodies (Berkowitz & Bier, 2007; Howard et al., 2004; Murphy, 2002; Sojourner, 2012).
The largest extant review of the character education literature was published by a project called What Works in Character Education (WWCE; Berkowitz & Bier, 2005, 2007). Though WWCE is now more than 15 years old, it has had considerable influence on perceptions of character education. WWCE is described in more detail below, though in short, it was a significant step toward establishing an empirical foundation for a rapidly growing educational movement that hoped to command the respect of scholars and the confidence of practitioners. WWCE brought a new level of credibility and guidance to character education by identifying effective programs and practices, including interventions that “worked” in the sense of often being linked to positive, character-related outcomes. However, WWCE was a purely descriptive analysis that generated no statistical summation of results from the individual studies reviewed and therefore could not offer any overall conclusions that were based on a comprehensive statistical analysis. Accordingly, we considered it worthwhile to reanalyze the studies reported in WWCE using meta-analytic techniques. As elucidated below, examining the data in this way would allow us to derive an overall conclusion about the effectiveness of interventions contained within the body of literature analyzed, determine the strength of evidence for coming to this conclusion, and investigate what specific variables contribute to this effectiveness. An updated literature search covering the years since WWCE and up to 2020 was also completed (Brown et al., 2021). The current study served as the first phase in that more comprehensive review. Given the important role the WWCE played in the history of efforts to evaluate the effectiveness of character education programs, we believed the present study deserved to be published independently of the larger effort, as a complement to the earlier WWCE publications.
As background to this study, we will first briefly define what is meant by character, and by character education in the context of the WWCE project. Second, we will provide a summary and critical review of the character education literature as well as research covering related approaches to positive youth development. Given the sheer size of the literature, this review will focus on systematic reviews covering positive youth interventions. Finally, we will summarize the methodology of Berkowitz and Bier (2007) as it provided the starting point for which the current research was conducted.
Defining Character and Character Education
Character has to do with the extent to which a person consistently thinks and acts in ways that are considered prosocial. Behaviors and attitudes thought to be indicative of character are typically ones that contribute to the functioning of a social community, or at least are considered socially desirable. This allows for a very broad array of behaviors to be treated as indicators of character, including intellectual virtues (like curiosity), performance virtues (like perseverance), and self-regulatory functioning. WWCE did not exclude studies that examined virtues other than moral/prosocial ones as long as a given study included some aspect of behavior that was considered socially valued. Character was conceptualized as a complex, dynamic, multidimensional psychological construct reflected in moral, intellectual, and self-regulatory functioning (Berkowitz et al., 2017 ; Lickona & Davidson, 2005; Linkins et al., 2014). Understood in this way, character is the whole set of psychological characteristics that enable one to operate as an effective member of society, flourish intellectually, and function as a moral agent.
Because of the complex multidimensionality of character, it is hardly surprising that definitions of character have varied among scholars and practitioners. Nor should it surprise us that there have been diverse conceptions of what constitutes “character education” (Lapsley, & Narváez, 2006; McGrath et al., 2021; Williams, 2000). Some character education approaches are pedagogical, focused on the direct teaching of character lessons or the use of pedagogical methods such as cooperative learning and project-based learning, while other character education initiatives focus on peer relationships and community involvement. Some approaches use fictional and/or nonfictional literature to teach character, while others emphasize mentorship. Still others focus on modifying school culture as a whole (Berkowitz & Bier, 2007). The original WWCE project also included programs that focused on the teaching of core values, prosocial behaviors, socioemotional reasoning, and conflict resolution skills (Berkowitz & Bier, 2007). As WWCE and the present study aimed to be inclusive of programs intended to promote development of character, character education was also defined as any program intended to encourage positive youth development (see Berkowitz & Bier, 2005, 2007; Berkowitz et al., 2017).
This ecumenical approach allows for a great deal of diversity in programs considered for inclusion in this review. In this context, “character education” encompasses any type of positive youth development program, rather than only including programs that explicitly labeled themselves as character education (Berkowitz & Bier, 2007). Programs focused on youth development can be considered character education if they target some aspect of social investment, including moral values, sociomoral reasoning, knowledge of ethical issues, moral emotional competencies, or moral identity; behavioral competencies such as conflict resolution skills; and/or characteristics that support prosocial behavior such as traditional character strengths (Berkowitz & Bier, 2007). Since most character education programs focus on the period from kindergarten through high school, the examination of the character education literature can be bounded to some degree by focusing on this period.
The What Works in Character Education Project
As one of the first systematic reviews of the research literature on character education programming, the WWCE project yielded broad, yet important, conclusions. The single most important finding was that character education can positively affect the character development of school students. Other major findings included evidence that character education can affect many aspects of character development; character education programs take a wide variety of forms; and character education programs tend to include multiple components (Berkowitz & Bier, 2005, 2007).
Despite its historic importance as the first large-scale review of the character education literature, limitations of the project must be recognized. One was the focus on the results of significance testing. As has been widely discussed in recent years, this is a problematic inferential strategy. Assuming some homogeneity in effects associated with character education, one implication of power analysis is that the likelihood of significant findings is partly a function of sample size. Given that some studies involved multiple schools with sometimes thousands of students, interventions of even trivial size can achieve significance though they have limited clinical value. This is particularly problematic in studies where the data analysis did not correct for dependencies among children experiencing character education in a shared classroom or school. In recent years, there has been growing recognition that violation of the independence of observations assumption associated with using standard t tests or analyses of variance to evaluate interventions administered to groups of individuals can substantially increase the probability of a Type I error, thereby suggesting a difference from a comparison treatment when that difference really reflects varying group effects for the students receiving the two treatments (e.g., Baldwin et al., 2005; Murray, 1998). Finally, the exclusion of articles with a balance of evidence suggesting no or negative effects for character education interventions meant the statistics reviewed were not a comprehensive survey of the character education literature. Accordingly, while the WWCE project was the most thorough compilation of research literature to that time, it addressed the question of whether the literature indicated character education could be efficacious, and if so under what conditions. It did not address the question of how efficacious that literature suggested character education programs to be.
Other Reviews of the Character Education Literature
Several other attempts have been made to provide insights into the effectiveness of character education programs and practices, each with their own strengths and weaknesses. Early such efforts focused on programs that emphasized moral development and values clarification interventions (Enright et al., 1983; Lockwood, 1978; Schlaefli et al., 1985), consistent with Kohlberg’s influence on the rebirth of interest in character education. As a group, they reported evidence of growth in student’s moral reasoning, but did not find much support for the hypothesis that those programs had an impact on secondary outcomes such as self-esteem, personal adjustment, drug use, or academic success. However, these earlier works noted various methodological concerns, including failure to use random assignment, lack of control groups, use of inappropriate statistics, and failure to control for potential confounds. The focus on morals and values also means they are of limited value as indicators of character education effectiveness in general.
At approximately the same time as WWCE, the U.S. Department of Education sponsored a large-scale effort to evaluate cross-program effectiveness, referred to as the Social and Character Development Research Consortium (2010). This involved original research simultaneously evaluating multiple character education programs with a common set of instruments, using a prospective design over the period 2004–2007, involving a large variety of schools and participants, evaluated using appropriate statistical methods such as correction for nonindependence among students. The project revealed little evidence of positive effects, and the investigators concluded that programs may even have been associated with a detrimental impact on some student outcomes. However, this research had several limitations. Only seven programs were examined, and it is unclear to what extent they represented effective programs in general. For example, only three of them were on Berkowitz and Bier’s (2007) list of programs for which there was existing evidence of effectiveness. Second, the substantial majority of teachers in the control groups reported integrating social and character development strategies into their classrooms.
More recent reviews of programs explicitly identifying as character education have employed meta-analytic methods and have reported modest positive effects on academic, behavior, social, and personality outcomes (Diggs & Akos, 2016; Jeynes, 2019). Jeynes (2019) found that longer programs were also associated with larger positive effects. Several meta-analyses have also been conducted focusing specifically on socioemotional learning programs (Durlak et al., 2011; Taylor et al., 2017). These have consistently generated positive effects for a number of outcomes, including socioemotional skills and academic performance. Overall, then, results across the literature that can be broadly conceived as character education have been inconsistent.
The Present Study
To summarize, prior attempts at reviewing the effectiveness of interventions that can be considered instances of character education have tended to use a restricted definition of character education or range of programs, or did not employ meta-analytic techniques to synthesize effects across the literature. The current study is intended to rectify these shortcomings and serve as a complement to Berkowitz and Bier (2007), augmenting their work by applying meta-analytic techniques to the WWCE literature base. Given the importance of WWCE in the story of the contemporary renewal of character education, this article focuses specifically on the effectiveness of character education in the studies that were considered for inclusion in WWCE, expanding the database beyond those that showed significant effects and changing the framing question from “What works?” to “Does it work?” The meta-analysis reported in this article examined the relationship between character education interventions and character outcomes across a range of programs, intervention targets, and outcome measures. As background to this effort, a more detailed description of the WWCE project’s methodology follows.
WWCE represented an extensive systematic review of the literature covering character education initiatives between the years 1945 and 2004 (Berkowitz & Bier, 2005, 2007). The WWCE project’s goal was to identify and review all scientifically sound research studies in character education that met their selection criteria. Given Berkowitz and Bier (2007)’s broad definition of character education, as described above, their review included a wide variety of programs and outcome measures. Various methods were used to find studies. The electronic search strategy for the WWCE project began with review of the ERIC, PsycINFO, Health-STAR, and Social Sciences Index databases for articles published between 1945 and 2004. Search terms included school, evaluation, intervention, program, and project as well as terms such as prosocial, virtues, values education, moral education, civics education, ethics, school culture, school climate, and social and emotional. Berkowitz and Bier (2007) also assembled an advisory panel of five individuals considered experts in various domains of youth interventions, including service learning, socioemotional learning, violence prevention, drug and alcohol prevention, and teacher/class- room effects on students. Members of this panel were each asked to supply lists of studies from their fields that fit the project’s inclusion criteria. Berkowitz and Bier (2007) additionally contacted the developers and/or evaluators of then-existing programs to request further outcome evaluation reports. This process generated 760 studies (see Figure 1).
These studies were then reviewed for Berkowitz and Bier (2007)’s inclusionary criteria. To be included, studies had to have evaluated a program that could be considered character education, that is, a program that focused on youth personal development, including the teaching of moral values, prosocial behaviors, socioemo- tional reasoning, and conflict resolution skills. Other inclusionary criteria required that programs of interest were implemented in a kindergarten to Grade 12 setting, did not focus on students with mental and/or developmental disabilities (though samples described as “at-risk” were included), and were offered to the general student population; and the research was quantitative in nature. These criteria reduced the pool to 114 articles, one of which was not retrievable for the present study. Five of these were meta-analyses or other systematic reviews of prior literature. Berkowitz and Bier (2007) reviewed each of these again for additional criteria suggesting acceptable scientific quality and sound research design. This was defined as including a comparison group, a pre-post design or other approach to establishing equivalence between the experimental and control groups, and inclusion of significance test results. This step reduced the pool to 73 studies and five systematic reviews. These were further reduced to 64 studies and five systematic reviews, to include only those studies where the bulk of the tests reflected significant positive outcomes, as the primary goal of WWCE was to identify attributes of effective programs rather than provide an overall judgment about the effectiveness of character education. This process resulted in 33 programs being deemed effective (Berkowitz & Bier, 2007).
Berkowitz and Bier (2007) focused on statistically significant evidence in support of character education programs from the studies featuring these 33 programs. For example, the Berkowitz and Bier (2007) reported that the most commonly reported implementation strategies among programs deemed to be effective were professional development, peer interactive teaching strategies, direct teaching strategies, family and community participation, modeling and mentoring, behavior management, and community service and service learning.
WWCE has been widely used in subsequent years to support the conduct of character education programs. The original article (Berkowitz & Bier, 2007) and spin-off publications (e.g., Berkowitz, 2011) have been cited hundreds of times, and provided the foundation for hundreds of trainings for faculty and administrators in character education. As noted, though, the project did not address the question of how well character education worked in the studies reviewed. The present study provides an addendum to this classic work, focusing on those studies drawn from the key 113 WWCE sources that provided sufficient information to generate estimates of the mean effect size between continuing education and comparison conditions. It therefore represents a meta-analysis specific to the WWCE endeavor.
Method
Search Strategy
Where WWCE focused on those studies of the character education literature for which the researchers reported significant results and found program effectiveness, the current research investigated whether that set of studies suggested character education generally is effective. The meta-analysis began with the previously described pool of 113 articles. As noted above, these were all the studies or reviews Berkowitz and Bier (2007) were able to identify up through 2004 that summarized one or more quantitative evaluations of a character education program for a general or “at-risk” population of students in a K–12 setting (see Figure 1). The first step in the present project involved examining the five reviews included in this set for studies not included in the 113. That process yielded an additional 117 references. Of these, 86 were unpublished manuscripts, three duplicated prior results, and seven did not meet the selection criteria. The remaining 21 new references were added to the 108 primary sources from WWCE, raising the total to 129 studies considered for the present study.
Two graduate students conducted a full text review of each ofthese articles to evaluate compliance with the following additional selection criteria appropriate to the present project. The study had to be in a peer-reviewed publication; non-peer-reviewed reports, theses, and dissertations were excluded due to potential concerns about research quality. Second, studies had to include sufficient statistical information for the purposes of computing or estimating at least one effect size based on the standardized mean difference between character education and control conditions. In cases of disagreement on whether a study met these criteria, a third graduate student evaluated the study. This process resulted in the exclusion of an additional 65 studies, leaving the final pool of studies to be included in the analysis at 64. Of these, 42 (66%) were studies used in the WWCE review in Berkowitz and Bier (2007). The additional 22 studies (bolded in Table 1) either did not meet their criteria for inclusion, in that significant positive outcomes did not sufficiently outweigh negative and nonsignificant outcomes, or were articles we identified from one of the five systematic reviews covered in Berkowitz and Bier (2007), but not themselves reviewed on the individual study level.
A summary of the studies may be found in Table 1. These studies involved a total of 96,930 participants. About 28% of studies studied a majority Black sample, participants in 38% of studies were predominantly White, and 9% of studies included predominantly Hispanic participants. The majority of studies (57%) took place in an urban environment, while 20% occurred in a suburban environment and 5% were in rural settings. Only 26 of the 64 studies discussed issues of risk among their participants. Of those, 40% identified their participants as being at-risk in some way. The average sample size was 709 (median = 280; range = 11–8,280 participants). The average number of schools per study was 12.7 (median = 7.5; range = 1–66 schools). In a majority of studies, experimental and control groups were housed in different schools. Forty-two percent of the 64 studies used random assignment at the individual, class, or school level; 50% of the studies did not, typically due to investigators using convenience samples from schools that had already been implementing the youth intervention program of interest. The remaining studies involved a mix of random and nonrandom assignment of schools or students, or did not describe the method of assignment sufficiently. For additional information about assignment and program characteristics, please refer to Table 2.
There were 36 different youth intervention programs featured across the 64 studies, including 29 established/branded programs, and seven site-specific programs. The most common target of intervention involved “productive” elements of character (43%) such as school adjustment, social skills, or problematic behaviors; followed by outcomes reflecting moral functioning such as moral reasoning, ethical sensibility, or racist attitudes. An additional 18% reflected more cognitive or intellectual contributors to character such as knowledge of various topics. A final 9% of outcomes did not fit into any of these three categories. Examples of such outcomes included expectations for peer/adult substance use, visits to the school nurse, and awareness of social norms.
Data Extraction
All data were initially extracted by graduate students, again with two graduate students reviewing and extracting data for each study. Discrepancies were resolved by one of the current authors, who double checked the data in question for accuracy. Since WWCE included all findings from an article, results for a study were aggregated across subgroups, outcome measures, interventions, and time points. In addition to outcome statistics, potential moderators were also extracted consistent with WWCE, including methods of program implementation and outcome targets.
Methods of program implementation included direct teaching methods, interactive teaching methods, behavioral management strategies, modeling or mentoring, family and community participation, community service or service learning, teacher professional development, and schoolwide organizational change. Elements of study design considered for the current analysis were the method of participant assignment (random assignment or other methods) and level of participant assignment (by individual, classroom, or school).
Finally, research quality was evaluated using criteria based on the Cochrane Handbook for Systematic Reviews of Interventions standards for identifying systematic bias (Higgins & Green, 2011). Nine potential contributors to poor quality were independently evaluated by two graduate students according to these guidelines. A detailed explanation of each of these can be found in Table 3. Raters evaluated each study on each of the nine contributors as demonstrating a high, low, or unclear risk to research quality, resulting in a total of 576 ratings. Of these, 220 (38.2%) were instances where risk of bias was unclear due to inadequate reporting of methodology in studies. There were 203 (35.2%) instances of a high likelihood of bias and only 153 (26.6%) of a low likelihood of bias.
General Analytic Strategy
All analyses for which computation or estimation of a standardized mean difference was possible were included. Meta-analytic results were generated using Comprehensive Meta-Analysis, Version 3.0 (Borenstein et al., 2013). The first step involved computing an estimate of Cohen’s d for each analysis. A few studies reported rates of dichotomous outcomes rather than dimensional outcomes. These were converted to an estimated standardized mean difference using the log odds ratio (Hasselblad & Hedges, 1995). There were 15 such analyses that were excluded because of a base rate of 0 in at least one group, that is, the targeted behavior never occurred in that group. These typically had to do with substance use in younger children. In instances where t or F values were used in a cluster randomized trial and were not analyzed using mixed models or a similar strategy that corrected for clustering, the significance test value was corrected per Hedges (2007) Equation 5. Except when computed from a t or F value, which had already been adjusted, standardized mean differences and standard errors were also corrected for clustering bias (Equations. 18.19 and 18.20; Hedges, 2009). Researchers rarely provided the intraclass correlation coefficient reflecting the impact of case clustering on the results, so in all cases this statistic was assumed to be .15. This value is consistent with prior studies examining group effects in educational settings (Hedges & Hedberg, 2007; Schochet, 2008; What Works Clearinghouse, 2014). A number of unusually large effects were noted, so the top and bottom 2.5% of d values were then winsorized (Martin & Roberts, 2010), and standard errors were recomputed based on winsorized values. Final d values and standard errors were then used to generate a mean Hedges’s g value for each study (Hedges, 2009).
Results
As space does not permit adequate explanation each of the below statistical terms or why a given analytic procedure was judged appropriate, and as doing so is out of the scope of the current research, the current authors refer readers to consult the primary sources as cited below. Given the diversity of study populations, interventions, and outcomes, the random effects model was considered appropriate (Borenstein et al., 2010). Statistical procedures for detecting and adjusting for publication bias such as Egger’s test (Egger, 1997), the funnel plot (Light & Pillemer, 1984), and the trim and fill method (Taylor & Tweedie, 1998) were also computed. Potential moderating variables were analyzed via meta-regression, as outlined in Borenstein et al. (2015). There were 836 comparisons for which standardized mean differences could be reported across the 64 studies. Of these, 83% were directly estimated from means, standard deviations, and group sizes.
On average, comparisons yielded a small, statistically significant effect size, g = 0.33, p < .001, 95% CI = [0.21, 0.45] with Hartung-Knapp-Sidik-Jonkman correction (IntHout et al., 2014) for heterogeneity in effects. These findings suggest that character education, broadly defined, is effective in positively influencing character outcomes relative to control conditions, in the kindergarten to grade 12 general education setting. However, this interpretation must be tempered by significant evidence of heterogeneity, Q(63) = 540.9, p < .001. This heterogeneity is amplified by both sampling error and by heterogeneity across the populations aggregated in this study. The I2 statistic suggested that 88% of variation in study effects was due to population variability rather than sampling error. The standard deviation of population effects was estimated at 𝛕 = 0.40, which suggests more than the typical level of variability across populations found in a random effects meta-analyses (Linden & Hönekopp, 2021). Not surprisingly, the prediction interval of population effects was quite large, [0.48, 1.14], indicating substantial variability in intervention effects across treatment populations, including the potential for medium-sized deleterious effects in some settings.
Possible Selection Bias
Qualitative examination of the funnel plot (see Figure 2) revealed asymmetry of effects suggestive of probable selection bias even with winsorizing of extreme values. Estimating the true effect size under the assumption that asymmetry in study effects was a result of selection bias generated a substantially reduced estimate of g = 0.07. However, this trim and fill method can be quite inaccurate under circumstances of heterogeneity among population effects (Peters et al., 2007; Terrin et al., 2003), which clearly applies in the present circumstances. It is also noteworthy that Egger’s test of the regression intercept did not reach significance, t(62) = 1.71, p = .09, though the test of the rank correlation was, z = 2.8, p < .01. Vevea and Hedges (1995) also developed a likelihood ratio test for the null hypothesis of no bias. Using the weightr package in R (Coburn & Vevea, 2019) across four levels of p values for the mean effect for each study, this test was significant, χ2( df = 3) = 44.99, p < .001, suggesting evidence of bias. Assuming moderate one-tailed selection bias (Vevea & Woods, 2005), the corrected effect was estimated to be g = 0.23. Taken together, the bulk of the evidence suggests some degree of selection bias is present, resulting in some overestimation of the mean effect.
Moderating Variables
Meta-regression analyses were conducted evaluating eight program elements (direct teaching, student activities, behavioral management, school organizational change, mentoring, community/family participation in the program, service learning, and teacher professional development) as potential moderators of effects. We similarly examined 11 types of outcome variables (personal competency, school attitudes, knowledge of character, risky behavior, antisocial behavior, school behavior, academic achievement, attendance, problem-solving, interpersonal competency, and sociomoral cognition) for variations in program impact. None of these analyses were significant. Because mentoring is of particular interest in this literature, we note that mentoring programs resulted in larger effects on average, mean g for mentoring programs = 0.38, 95% CI = [0.12, 0.64] versus 0.30 [0.20, 0.40] for those without, suggesting the possibility a larger sample of studies might have sufficient power to achieve significance.
Discussion
The primary intention of this study was to provide an addendum to the WWCE project, generating an estimate of the effectiveness of the character education studies that were part of that project. Across 64 studies, character education demonstrated an overall positive impact compared to control groups, with a mean effect size of g = .33, indicating that on average the mean for students receiving character education was about 1/3 a standard deviation better than the mean for students in a control condition. Traditionally, this is considered a small to moderate effect across various disciplines (Cohen, 1988). However, such effects are not considered trivial, and can still be quite important in large populations such as that represented by K–12 students. The most famous example of this phenomenon is Rosenthal’s (1990) demonstration that the effect of aspirin on prevention of heart attacks, which is now widespread practice, was far smaller than the effect reported here, but across the entire population has a profound effect on heart disease. It is noteworthy that the mean value reported here is the same as the mean value Lipsey and Wilson (1993) reported for all meta-analyses they found in educational research that compared interventions, .34.
However, it should be noted that estimates of the mean effect based on the assumption of selection bias generated smaller mean effects, with the procedure introduced by Vevea and Woods (2005) suggesting a mean effect of .23 even after winsorizing extreme effects. The existing information does not allow us to identify the cause(s) of potential selection bias, or even the extent to which these findings are due to heterogeneity in effects across treatment populations. There is a tendency to attribute such evidence to the “file drawer problem” (Rosenthal, 1979), that is, the tendency not to publish studies with null findings. This explanation would not seem particularly relevant in the present circumstances. Given the size of the samples and the amount of effort typically involved in conducting a school-based study, it is unlikely that researchers would forsake publication if possible. Large sample sizes are also likely to result in some significant findings assuming even a small effect, particularly given the tendency not to correct for clustered administration of treatment in the statistical analysis.
More likely, then, is the presence of some use of questionable research practices (Simmons et al., 2011), such as omitting non–significant analyses, adding covariates until an analysis becomes significant, and the aforementioned tendency to conduct tests without correction for clustered treatment. Even when corrected for selection bias, however, the mean effect is of a size that suggests the potential for practical value.
The heterogeneity of findings in this study also merits consideration. There was substantial variability in effects across studies. It is important to keep in mind that this finding does not imply ineffectiveness of treatment, only that we could not account for the variability in results based on the different educational approaches and targets we were able to compare to each other. None of the moderators we evaluated based on program or outcome variables examined in WWCE accounted for significant variation in effect sizes. A more likely interpretation has to do with variability in the effectiveness of character education programs across settings. This interpretation is consistent with finding a prediction interval from the present results of [0.48, 1.14]. To understand the meaning of this interval, it is helpful to compare it to the confidence interval reported earlier in conjunction with the mean effect, [0.22, 0.45]. The latter helps us gauge how precise 0.33 is as a mean effect estimate across all the student populations, character education programs, and treatment outcomes considered in this meta-analysis. The fact that both endpoints of the interval point to a small-to-moderate mean effect suggests reasonable precision in the overall estimate. In contrast, the prediction interval attempts to estimate the range of effects across settings. The fact that the interval includes negative values as large as 0.48 suggests the potential for some character education programs in some settings with some populations to have deleterious effects on some outcomes.
Dramatic variation in the efficacy of different character education programs could also have contributed to the null findings of the contemporaneous Social and Character Development Research Consortium (2010), if the programs identified were of relatively low effectiveness. Unfortunately, we are unable to draw firm conclusions about the causes of relatively negative findings with the available data. For example, one of the two programs with the poorest outcomes (the largest negative mean effect sizes) was only evaluated in one study, while the other was associated with a positive mean effect in the only other study examining that program. Clearly, it would be important to identify programs that are particularly promising and conduct multiple studies to generate reliable evidence of effectiveness across settings.
Ultimately, the findings of this study support the overall effectiveness of the character education programs incorporated in WWCE. The effect size reported here is in line with similar findings from other studies in education. However, though this article represents an important addendum to the WWCE project, that project is now more than 15 years old. Therefore what we present here cannot be considered a current profile of the character education literature. Even so, it points to goals for future character education research. Perhaps most important is the need to focus on programs of established effectiveness in order to more precisely determine the relative contribution of the different program components to overall program effectiveness.
Authors’ Notes
Keith Johnson https://orcid.org/0000-0001-5707-6223
Robert E. McGrath https://orcid.org/0000-0002-2589-5088
Mitch Brown https://orcid.org/0000-0001-6615-6081
Marvin W Berkowitz https://orcid.org/0000-0003-4335-6405
Keith Johnson is now at the Central Western Massachusetts VA Healthcare System. We are grateful to the following individuals who contributed to extracting the data for this study: Francesca Bates, Tamar Blanchard, Peter Hart, Melissa Jermann, Megan McGrath, Allison Parente, Shirley-Anne Paul, Alec Twibell, Norah Wallace, and Bina Westrich.
This work is part of a larger project done in collaboration with the Center for Character and Citizenship at the University of Missouri-St. Louis, and was made possible through the support of a grant from the John Templeton Foundation (grant #61178 to Melinda C. Bier and Marvin W Berkowitz). The opinions expressed in this publication are those of the authors and do not necessarily reflect the views of the John Templeton Foundation.
References marked with an asterisk indicate studies included in the meta-analysis.


