Skip to article sections
Purpose

The current literature on school teacher-created summative assessment lacks a clear consensus regarding its definition and key principles. The purpose of this research was therefore to arrive at a cohesive understanding of what constitutes effective summative assessment.

Design/methodology/approach

Conducting a systematic literature review of 95 studies, this research adhered to the Preferred Reporting Items for Systematic Reviews and Meta-Analyses guidelines. The objective was to identify the core principles governing effective teacher-created summative assessments.

Findings

The study identified five key principles defining effective summative assessment creation: validity, reliability, fairness, authenticity and flexibility.

Research limitations/implications

The expansiveness of education research is such that not all relevant studies may have been identified, particularly outside of mainstream databases. This study considered only the school environment, so contextual limitations will exist.

Originality/value

To the best of the authors’ knowledge, this study contributes original insights by proposing a holistic definition that can facilitate consensus-building in further research. The assimilation of core principles guided the development of quality indicators beneficial for teacher practice. The comprehensive definition, key principles and quality indicators offer a unique perspective on summative assessment discourse.

Assessment of learning, or summative assessment, can be broadly defined as a formal piece of work conducted at the end of a unit to give an absolute indication of student progress. It is an evaluation of the extent to which the outcomes of a unit of study have been achieved and counts towards calculating an overall achievement level (Brady and Kennedy, 2019). In some countries, this is also known as classroom assessment written for summative purposes (Rao and Banerjee, 2023). Summative assessment is a crucial issue in school education, not only in measuring the acquisition of knowledge and skills at the end of a unit of study but also in reporting these outcomes to students, parents and administrators (Cumming et al., 2019; Wyatt-Smith et al., 2017).

Over the past three decades, there has been a notable global emphasis on accountability in the field of primary and secondary education (DeLuca and Bellara, 2013). Various stakeholders, including governments, schools, industry and parents, have increasingly directed their attention towards assessing the performance of educational institutions (Donnelly and Wiltshire, 2014). Many countries now maintain national databases on education, providing education statistics and indicators, while international benchmarking has become more prevalent and influential in national education discussions (Donnelly and Wiltshire, 2014). In this context, evaluating student achievements has become the primary yardstick for measuring performance, necessitating an accurate and dependable approach to determining these accomplishments.

Despite broadly held agreement on its importance, the literature as a collective is vague about what constitutes effective summative assessment, and no universally agreed-upon definition of what effective teacher-created summative assessment looks like has been published to date. In response, this paper first conducts a systematic literature review (SLR) to assimilate existing research and propose such a definition. The key findings – a set of principles characterising effective summative assessment – contribute to a set of quality indicators that, when used by teachers to design assessment items, can assure stakeholders of the quality of the resultant data. This reliable data can then be trusted to inform educational decisions at a student, class, school, district and even national level.

Research indicates that teachers (particularly early career teachers) have a knowledge and skill deficit regarding the creation of summative assessment items (Coombs et al., 2020) and lack adequate theoretical understanding and confidence to create and implement effective assessment aligned to both the required curriculum and the specific needs of their class (DeLuca and Bellara, 2013). As such, a key premise of this paper is that an SLR is necessary to identify and assimilate principles of effective summative assessment and, consequently, develop a set of indicators for the quality assurance of such items.

While some systems determine a student’s achievement level solely based on the results of externally set summative assessment, for example, in the USA and England, countries such as Australia, New Zealand, Scotland and Wales primarily rely on the results of teacher-created assessment items. It is therefore essential to clearly define the scope of this review by defining what is meant by “teacher-created summative assessment item”, which may take many forms, including formal examination, project, essay, oral presentation, interview, performance, experiment report or creation of a new work based on what has been learned during the teaching period (Rao and Banerjee, 2023).

The underlying structure of a summative assessment item is that it contains two documents created by the teacher for students: a task description that outlines the assessment requirements, context and associated expectations, and a rubric or marking criteria that generally consist of a matrix listing criteria and descriptions of varying response quality. The process of creating a summative assessment item can be a complex task for which many early-career teachers feel unprepared (DeLuca and Bellara, 2013).

Beyond informing teachers and students of student learning after a unit of work, the results from the summative assessment are also used as a source of quantitative data for external organisations to determine the effectiveness of teaching and school performance (Gonski et al., 2018). While the stakeholders most impacted by the quality of summative assessment items are students, it cannot be denied that over the past few decades, there has been increasing value placed on the data emerging from these items by multiple key stakeholders, including government departments, schools and subject departments (Gonski et al., 2018). A teacher’s competence to create and implement effective summative assessment is thus essential both to garner confidence and mastery of this critical skill in their pedagogical practice and also for them to use these assessment instruments to optimise the educational outcomes of their students (Gonski et al., 2018; Craven et al., 2014). Further, given that assessment is used to support both decision-making and communication with key stakeholders, the quality of the decisions made and the resulting communication with stakeholders such as students, parents, potential employers and policymakers is dependent on the assessment written being effective (Wiliam, 2008).

In this paper, it is not the authors’ intention to quantify the extent to which an item of assessment practice is effective. Instead, the purpose is to elicit from the extant literature a set of principles that educators may use to assure quality and guide good practice when creating assessment items. In this way, the term “effective” is used in its most common form – to describe an item that does what it is designed to do. The literature reviewed to date mentions phrases such as “sound and useful assessment” (Levy-Vered and Nasser-Abu Alhija, 2015, p. 378) or “understanding principles of sound assessment by integrating practices, theories and philosophies of assessment to support teaching and learning” (DeLuca and Bellara, 2013, p. 356). In the literature reviewed, it was noted that after using phrases like these, the authors did not continue to define what they meant by “sound assessment”, “appropriate”, “effective” or even “good”. It became clear that a universally agreed-upon definition, principles and quality indicators of what makes an item of summative assessment effective had not yet been identified. It was, therefore, difficult for a teacher to reference guidance as to what constitutes an effective summative assessment in order to reflect on and determine with any certainty whether their assessment practice was, in fact, effective.

To address this gap in the literature, this SLR sought to understand “what principles guide the design of effective teacher-created summative assessment?” The review's focus was limited to primary and secondary school settings, given that different settings (such as higher education) may vary due to the overarching governing and accrediting body requirements. The outcomes of this process resulted in two important outcomes:

  1. the identification and assimilation of a set of agreed-upon principles guiding summative assessment design; and

  2. the use of these principles to develop a set of quality indicators for teachers to use when creating summative assessment items, which is aimed to engender confidence in the effectiveness of the item.

This section will describe the process followed for an SLR as a research method.

In conducting this SLR, the guidelines set out by the Preferred Reporting Items for Systematic Reviews and Meta-Analyses workflow (PRISMA), which instructs the process of systematic reviews (Moher et al., 2009), were followed. A systematic search for relevant literature in online electronic databases was conducted initially in April 2020 and then again in November 2022 to check for any new research since the initial search and analysis. A stringent search string was applied, assessing full-text documents in the following databases, selected for their reputable association with education research in Australia: Scopus, Academic Search Ultimate, ERIC, Education Research Complete, E-Journals, Emerald Insight, Australasian Education Directory, A+ Education, AUSTGUIDE, DELTAA, Humanities and Social Sciences Collection, THESUS, Taylor and Francis Online and SAGE Journals. The search included peer-reviewed journal articles in English but did not include a start date limitation to ensure the inclusion of seminal works written prior to an arbitrary start date being imposed.

The inclusion of “effective” as a key search term was determined based on the critical assumption that reference to effectiveness is a key indicator that guides or encourages adherence to principles in design. Although most of the literature reviewed mentioned effective summative assessment, a seminal publication specifically defining principles characterising effective teacher-created summative assessment design appeared unidentified. It was noted that an author’s defining characteristics of what effective summative assessment meant were generally only mentioned in passing as the authors continued onto another focus of study (Vogt, 2022; Wyatt-Smith et al., 2017). It was, therefore, essential to review as many articles as possible to capture these mentions, either in the abstract, introduction or background of a study. The study’s authors’ previous experience teaching assessment at a tertiary level showed that resources aimed at initial teacher education discussed the assessment design in more depth. For this reason, it was also decided to search for textbooks covering initial teacher education, assessment and curriculum design.

A crucial step in conducting this SLR was identifying key search terms that would identify articles that covered what was meant by teacher-created summative assessment. The terms “summative assessment” and “assessment of learning” were applied first, as these terms are synonymous (Cumming et al., 2019). This included only research, including summative assessment rather than formative or diagnostic. There has been a plethora of current research defining the principles of effective formative assessment; however, this review was focused explicitly on assessment at the conclusion of a teaching sequence. The research question was particularly focused on how summative assessment design makes it effective. Therefore, the concept of principles was captured by the terms “principles”, “characteristics”, “qualities” and “criteria”. Multiple search terms were required for quality indicators, as there does not seem to be a universal term used consistently in assessment research. The literature suggested that “effective” (Remesal, 2011), “assessment literate” (Edwards, 2013), “quality” (Broadfoot and Black, 2004), “improvement” (Wyatt-Smith et al., 2017) or “successful” (Simpson, 2004) all served to identify good practice. These terms were grouped and used based on the understanding that when research focused on the need to improve assessment items, it would typically mention an effective assessment item. This also implied assessment that could be changed or altered by the teacher and eliminated some articles focused only on externally set, standardised testing as summative assessment. It was determined that teacher-created summative assessment principles would emerge from these concept groupings cumulatively rather than definitively when associated with one term only.

The following search string was thus finalised and applied in April 2020 and again in November 2022 to check for recent additions to the body of literature: (“Summative Assess*” OR “Assessment of learning”) AND (“assessment literac*” OR effective OR improv* OR “assessment knowledge”) AND (principle? OR characteristic? OR qualit* OR criteri*). The stages of the literature search and the screening process through which 156 studies were finally identified are provided as supplementary material.

Although few, additional print textbooks owned by the authors were included due to the specific identification of teacher-created summative assessment design principles. Based on their title or abstract, records were excluded at the screening stage if they did not meet the inclusion criteria (see below). Articles were again excluded at the eligibility stage of the PRISMA workflow when the inclusion criteria were not met on the first reading of the full text. Finally, 58 articles were excluded from the SLR because they did not explicitly specify the principles of summative assessment. In these cases, the terms “quality assessment”, “effective assessment” or even “good assessment” were used but neither clarified nor defined by the authors.

Delineations of the research question were ensured by limiting the scope of the literature reviewed. Articles were included or excluded based on the following criteria. The review included studies that focused only on summative assessment. Assessment for learning and assessment as learning were excluded, as there has already been significant research in recent years to determine principles indicating effective assessment in these genres. Studies that looked at teacher-created summative assessment were included, while externally designed and set summative assessment (such as standardised testing) was excluded. Only articles published in English were considered, as they were accessible for the authors to access the research. During the initial scan, notable differences were identified between the purposes and uses of summative assessment in primary and secondary schools compared with higher education. Therefore, only primary or secondary school classroom studies were included.

The review considered studies that focused on qualitative and quantitative data, as it was determined that what constitutes effective assessment was not dependent on a particular research approach. It considered interpretive studies that drew on teachers’ experiences with the creation of summative assessment. Finally, published textbooks and reference material on summative assessment design were also included, given that these texts were more explicit in defining effective summative assessment.

Consequently, studies were included if the article’s focus included reference to summative assessment, teacher-created assessment design and primary or secondary educational contexts. Further, the study had to be published in a ranked journal, government report or initial teacher education textbook where principles of assessment design were explicitly stated. Studies were excluded if they referred to formative peer or self-assessment, externally set assessment, tertiary education contexts, or if the article was published in an unranked journal or editorial.

Given the identified gap in the literature identifying the principles of effective assessment, the scope for inclusion was quite broad. This was to allow for any article that mentioned “quality”, “successful”, “good” or “effective” assessment and then went on to clarify these terms to be included in this article. It is acknowledged that terms such as “validity” and “reliability” are also used in research when looking at assessment to describe the research process; therefore, the full papers were read carefully to determine the context of the terms used.

Following the search, all identified citations (n = 2,353) were collated and uploaded into EndNote X9, where duplicates were removed. The remaining 1976 titles and abstracts were screened for assessment against the inclusion criteria. Relevant studies (n = 312) were fully retrieved and assessed in detail against the inclusion criteria. Reasons for the exclusion of full-text studies were recorded. In total, 154 studies were identified as meeting the criteria.

Validity and reliability were commonly used in relation to defending the quality of the research process of the study. These terms were also used to describe the qualities of effective summative assessment. Consequently, all studies were read and coded by hand to ensure the accuracy of theme identification, allowing the authors’ immersion in the data (Castleberry and Nolen, 2018). A spreadsheet to record like terms and keep detailed notes of specific terms was used for future theme identification (see Supplementary File).

An inductive approach to thematic analysis was used to identify emerging themes. This was deemed suitable for the SLR in exploring to what extent consensus exists when defining effective summative assessment (Yin, 2016). Definitions or descriptions of what constitutes good practice in assessment creation were noted. It was here that a broader approach to the theme generation needed to be adopted to capture the nuanced differences in perspectives. Groups were created where the same terms were used or if terms were synonymous, based on the definition or from interpreting the explanation provided. Finally, overarching themes were synthesised, resulting in principles defining best practice, which may be used as a basis for evidence-based practice.

In total, 95 studies which looked at effective teacher-created summative assessment in the primary or secondary classroom were eligible for this review and were analysed accordingly to explore the principles collectively agreed upon, which identified an effective summative assessment item.

When discussing principles of teacher-created summative assessment, different authors refer to other principles and terms used individually or in combination. Based on the patterns and trends emerging from the research, the following six categories were identified regarding identified principles of summative assessment:

  • validity;

  • reliability;

  • fairness;

  • authenticity;

  • flexibility; and

  • other.

Each of these principles will now be discussed.

The most commonly occurring principle when defining effective summative assessment was the concept of validity (91.58% or n = 87), which a synthesis of the literature demonstrated as a visible relationship between the created assessment item and the content which has been previously taught. In defining validity, this term, or synonymous terms, were identified, including:

The definitions provided across the literature highlighted that, for an assessment item to have high levels of validity, it must clearly and explicitly align with the content taught and, therefore, any externally set curriculum requirements. At this point, it is important to note that validity in this context refers to the creation of the assessment item rather than the validity of the process of implementing and administering the assessment. Within the articles mentioning validity as a characteristic of effective assessment, some mentioned specific types of validity. “Content validity” and “face validity” were the most represented.

After validity, the next most common theme was reliability, which was mentioned in 80% of studies (n = 76). The literature highlighted that the reliability of an assessment item considers the ability of the marker to make a consistent and dependable decision when looking at a student’s response. As with validity, some articles used the term reliability (n = 65), while 11 articles used synonymous terms, including:

The literature defined a highly reliable assessment item as demonstrating explicit links between the task sheet and rubric and written using language which is easily understood by both the marker and the student. High levels of reliability are present when factors are in place to ensure consistency of marking and dependable overall judgement of student results (Woolfolk and Margetts, 2007).

Fairness was the next most common theme in 44.21% of studies (n = 42). Synthesising definitions of this concept revealed a fair assessment as one that provides all students with an equal opportunity to demonstrate achievement and produces scores that are comparably valid from one person or group to another. In total, 34 articles used the term fairness, whereas eight used terms that came under the umbrella of fairness, including:

The literature further highlighted that if, through equitable treatment and opportunities, each student has an equal chance to demonstrate their knowledge or understanding, it can be possible to compare achievement and draw accurate conclusions (Darling-Hammond et al., 2013). For an assessment item to be fair, there must be:

  • equitable access to the assessment; and

  • an absence of bias both on the part of the teacher and inherent in the assessment item (McMillan, 2014).

Modifications to ensure fair access to the assessment item and equity in the ability to respond are also key considerations when creating a fair assessment.

Authenticity was considered a key component of summative assessment being effective in 30.53% (n = 29) of the articles reviewed. The literature suggests that authenticity in summative assessment considers the extent to which the assessment instrument has relevance to a real-life context and demonstrates meaning in the students’ lives (Gulikers et al., 2004). Once again, most articles used either authentic or “real-world” specifically (n = 24), and a further five used terms deemed to be synonymous when considering the definition of the characteristic of authenticity, such as:

As revealed in the literature, some features which may indicate that an assessment item is authentic are:

  • an item based on a real-life situation;

  • an open-ended task that does not only have one correct solution;

  • a task requiring collaboration; or

  • is performance-based (Gulikers et al., 2004).

It is important to note that the perception of authenticity may differ between teacher and student, depending on personal background. An item deemed highly authentic by a teacher for one cohort of students may be less authentic for the next.

The concepts of student voice and student choice were grouped to form the principle of flexibility, which was mentioned in 18.95% of studies (n = 18). In total, 14 studies referred to these ideas as flexibility, and the remaining four referred to flexibility in other terms:

The literature highlighted the concept of flexibility in assessment as the “wiggle room” in a course syllabus, where a student may negotiate certain parts of the assessment item. Flexibility allows for input from the student, leading to motivation and engagement with the topic, which is more likely to result in deep and long-term content retention (Prianto et al., 2022).

Other principles were mentioned as indicators of effective summative assessment once (16%) but were not repeated by a second study. Some of these included:

In addition to involving teacher judgement and manageability, feedback to the student emerged as an important facet of summative assessment. Feedback is acknowledged as imperative to a quality assessment cycle but has been determined to be outside the scope of this review. Only the creation of the summative assessment item has been considered in this article. Quality teaching, ample time to practice new skills, feedback and feedforward are all acknowledged as essential to a successful teaching, learning and assessment cycle.

Although these principles have been defined separately, it is recognised that they are not independent of each other in practice (Harlen, 2005). It was thus also prudent to look at the combinations of principles, as most studies intimated that it was the combination of principles that led to the effectiveness of the assessment item. Where this was not explicitly stated, it was inferred due to the number of studies listing a combination of principles (n = 82) rather than just identifying a single principle (n = 13). Figure 1 shows the combinations authors recommended in identifying an effective summative assessment item.

The most commonly occurring combination of principles mentioned in the studies was validity and reliability (28.42%, n = 27). This was followed by validity, reliability and fairness (16.84%, n = 16) and validity, reliability, authenticity and fairness (10.53%, n = 10). Significantly, the combination of principles was influenced by the focus of the original article.

If the articles reviewed focused on the assessment item (n = 48), the principles of validity and reliability were most mentioned (n = 25). Fairness was the next most cited in this group (n = 18), combined with validity or reliability. When fairness was mentioned in this context (particularly when paired with validity or reliability), it typically referred to the absence of bias (for example, DeLuca et al., 2016). The absence of bias aspect of fairness would also increase the reliability of an item, as it is commonly through a well-constructed rubric that more objective (and therefore less biased) judgements can be made.

In the same way, when the articles under review focused on learning (n = 13), the effect of the assessment on the learner (n = 10), or the impact of assessment on teaching (n = 8), the principles of authenticity, flexibility and fairness were more commonly identified as principles of effective summative assessment. These terms often still included validity (n = 27); however, this was not always the first principle mentioned (for example, Baird et al., 2017). Interestingly, when fairness was identified within these articles, the focus tended towards avoiding unfair advantages or disadvantages for particular students or groups of students (Klenowski, 2014), as well as equitable instruction leading to fair assessment being conducted (DeLuca et al., 2016).

Only two articles out of the 95 reviewed mentioned all five principles of teacher-created summative assessment (Brookhart, 1997; Klenowski and Wyatt-Smith, 2014) as effectiveness indicators. These articles focused on “a theoretical framework for the role of classroom assessment in motivating student effort and achievement” (p. 161) and “assessment for education”. Rather than considering the principles of effective summative assessment, they focused on the positive effects of carefully designed assessment by drawing together the classroom assessment environment as well as classroom assessment events. All other articles identified one or a combination of principles to define effective assessment.

A rigorous review of the extant literature reveals that for an item of teacher-created summative assessment to be seen as effective, multiple principles must be considered and balanced: validity, reliability, fairness, authenticity and flexibility. The difficulty of ensuring consistent and high levels of all principles in each assessment item is acknowledged. Considering this complexity, a set of quality indicators is therefore now proposed to assist teachers with the decision-making process before creating a summative assessment item, based upon the synthesis of core principles of effective summative assessment and the application of these as revealed across research.

Table 1 depicts the principles as a set of quality indicators, which, when aligned, can ensure the quality of a piece of teacher-created summative assessment. The results of this SLR overwhelmingly identified that the principle of validity is the foundational component that must be ensured when creating summative assessment. The other principles are immaterial without alignment with what has been taught, what is required by curriculum bodies and the breadth, depth and levels of thinking required. Therefore, validity is the foundational indicator upon which all other principles depend.

Validity is the visible alignment of the assessment task with the curriculum requirements and what has been taught. Content validity refers to the alignment of the assessment with externally set curriculum requirements, such as a syllabus or alignment to the content taught (Vogt, 2022; Darling-Hammond et al., 2013). Not only this, but the breadth, depth and thinking levels must explicitly match the teaching phase and assessment requirements (Bloom, 1956). Face validity, however, is seen from the perspective of the student rather than the teacher. An item is said to have a high level of face validity when the student can see the item’s relevance to the previous unit of work taught in class.

Reliability measures the consistency, dependability and accuracy of marking and decision-making in assessment. Two predominant types of reliability are mentioned in the literature: intra-rater reliability, which considers the consistency of one marker marking across the entire cohort, and inter-rater reliability, which examines how to ensure consistency and dependability of marking across multiple markers (Vogt, 2022). As this research studied the creation of a summative assessment item rather than marking it, it is assumed that thorough moderation processes would be undertaken at a school or district level to ensure the reliability and dependability of results across broader cohorts of students.

A highly reliable rubric (or marking criteria) may be characterised by such things as quantifiable terms in the descriptors, a clear delineation between standards and the indication of the weighting of each section of the item. These allow a marker to make more objective, rather than subjective, decisions and minimise rater bias. An item with high levels of reliability ensures students receive a result most clearly aligned with their performance and demonstration of understanding, regardless of when it was marked, how many were marked before it, or by whom it was marked.

Fairness ensures all students and groups have an equitable opportunity to demonstrate their knowledge, skills and understanding of the topic taught in the assessment item. Quality indicators of fairness include equitable access to the assessment item and the absence of bias. These occur when one individual or group of students is not unfairly advantaged or disadvantaged by the teaching and learning process and the conditions of the assessment item (this could be based on gender, race or socioeconomic status, to name a few). The absence of bias considers if a potential advantage or disadvantage exists for a particular group of students when completing the assessment (Vogt, 2022). If an item's context, requirements or wording does not advantage or disadvantage students based on gender, ethnicity, religion, socio-economic status, involvement in extra-curricular activities or other potential sources of bias, the item may be deemed fair.

Avoiding the halo effect (Gipps and Stobart, 2009) is also fundamental to fairness. The halo effect can be explained as either a positive or negative stereotype a teacher may have in their mind, often subconsciously, regarding a particular student or group of students. This stereotype can bias the marking of a student’s work unless a clearly constructed, specific and reliable rubric is present with the assessment item.

Modifications to assessment conditions or requirements to increase the equity of the item are also included within the construct of fairness. Length of required response, increased scaffolding, modification to the genre of the item or increased assistance from staff or technology may all assist a student to equitably demonstrate their understanding of the topic without being unfairly disadvantaged by disability or extenuating circumstances such as prolonged illness.

Authentic items are often problem-based and may have multiple solutions, depending on the approach taken by the student. It may also include life-long skills such as conflict resolution, communication, creativity and higher-order thinking skills to complete the assessment successfully. The communication and justification of the process are often just as important as the final solution. Through authentic assessment items, students can see the skills’ relevance, allowing content to be retained in longer-term memory (Murphy et al., 2017).

The most common approaches allowing for flexibility are usually either the subject or mode of assessment (Australian Curriculum, Assessment and Reporting Authority, 2020). For example, if an item must assess elements and demonstrate persuasive speech (mode), students may choose the topic (subject) they want to persuade the audience. Or, if the correct calculations of simple and compound interest must be demonstrated (subject), a student may choose to present this in a television sales advertisement or create a loan payment schedule to buy the latest television (mode).

Flexibility could also include co-constructing an assessment item with the student cohort (Panadero et al., 2022). The submission date, response length and group or individual items may all be negotiated. Some teachers also allow student input into the rubric and the assessable elements (while ensuring the mandatory components are included). One of the most significant benefits of including flexibility in assessment is the interest and effort a student may subsequently invest in the item. Suppose a student is motivated and engaged in an assessment item. In that case, they may invest more time and energy into producing a higher quality submission, which will more accurately demonstrate the student’s skills and understanding of the topic.

By adhering to the proposed set of quality indicators when creating a piece of summative assessment, teachers can be confident in the assessment item’s effectiveness and the resultant data from students completing the assessment. Teachers can justify the assessment item according to these quality indicators, assuring all stakeholders (including students, parents and themselves) of the validity, reliability, fairness, authenticity and flexibility of the summative assessment item.

The authors acknowledge the limitations of this study. The first limitation is the scope of the review. Although every effort was made to access every article mentioning principles of teacher-created summative assessment within the primary and secondary school context, some could be missed due to the journal rank or language in which they were published. New literature may have also been published since the last search was conducted.

The second limitation is the interpretation of principles. Due to the interconnected nature of the principles discussed earlier in this paper, a reader may determine that one characteristic belongs to a different “principle” group. For some of the findings cited in this paper, there was not a clear definition provided for the terms used, and the definition needed to be inferred from the context. Therefore, some terms were deemed aligned with and placed in the “principle” groups as presented in this paper. Definitions of concepts are rarely universally agreed upon, and it is because of further discussion that new research emerges and grows. Given this, it is hoped that this paper will provide guidance and assist classroom teachers in creating future summative assessment items.

This SLR has successfully synthesised what has been meant by the term “effective summative assessment” in research, teaching and training. After a rigorous analysis, a set of quality indicators has been proposed to guide the creation of effective summative assessment in primary and secondary classrooms. The resultant data from these items can be trusted to be accurate and truly representative of student knowledge, understanding and skills. They can be used confidently when making decisions for the future education of the next generation.

As a result of undertaking this SLR, it is proposed that teacher-created summative assessment items should be guided by the principles emerging from the literature as quality indicators. The literature suggests that an item of effective summative assessment should be valid, reliable, fair, authentic and flexible. These are described as:

  • assessing what has been taught, which can be easily identified against the curriculum and by the student (valid);

  • allowing for a consistent judgement to be made, limiting subjectivity (reliable);

  • not advantaging or disadvantaging a student or a group of students (fair);

  • being seen as relevant to the student (authentic); and

  • allowing for student choice (flexible).

After reaching the above definition of effective teacher-created summative assessment, a set of quality indicators has been proposed (see Table 1). This table is offered as a guide to assist teachers in developing assessment design, prompting the consideration of the most salient principles for the resultant item and their specific context. By being aware of these principles and aiming to create assessment with high levels of each, a teacher may have more confidence in the effectiveness of the assessment item. Not only should the item be an accurate culminating assessment of the student’s learning, but it may also allow the student to be confident that their result is accurate. Further, students may retain the content of the unit of work through relevant research, feel as though they have been equitably treated and experience some sense of agency within aspects of the assessment.

By adhering to a set of quality indicators in the development of summative assessment items, students will be assured that their assessment is an accurate reflection of what has been learnt and to what extent, and teachers, parents and schools can also be confident in the results. At a regional, state, national and even international level, confidence can be gained regarding the validity and reliability of the resultant data. Knowing the data used for high-stakes decision-making is based on effective summative assessment.

Abell
,
S.
and
Siegel
,
M.
(
2011
), “
Assessment literacy: what science teachers need to know and be able to do
”, in
Corrigan
,
D.
,
Dillon
,
J.
and
Gunstone
,
R.
(Eds),
The Professional Knowledge Base of Science Teaching
,
Springer
, pp.
205
-
221
, doi: .
Australian Curriculum, Assessment and Reporting Authority
(
2020
), “
The shape of the Australian curriculum
”,
version 4.0, Australian Government
,
Baird
,
J.
,
Andrich
,
D.
,
Hopfenbeck
,
T.
and
Stobart
,
G.
(
2017
), “
Assessment and learning: fields apart?
”,
Assessment in Education: Principles, Policy and Practice
, Vol.
24
No.
3
, pp.
317
-
350
, doi: .
Bloom
,
B.S.
(
1956
), “
Taxonomy of educational objectives
”,
Handbook I: The Cognitive Domain
,
David McKay Co Inc
,
New York, NY
.
Bolden
,
B.
and
DeLuca
,
C.
(
2022
), “
Nurturing student creativity through assessment for learning in music classrooms
”,
Research Studies in Music Education
, Vol.
44
No.
1
, pp.
273
-
289
, doi: .
Brady
,
L.
and
Kennedy
,
K.
(
2019
),
Assessment and Reporting: celebrating Student Achievement
, (5th ed.)
Pearson Australia
,
Melbourne
.
Broadfoot
,
P.
and
Black
,
P.
(
2004
), “
Redefining assessment? The first ten years of assessment in education
”,
Assessment in Education: Principles, Policy & Practice
, Vol.
11
No.
1
, pp.
7
-
26
, doi: .
Brookhart
,
S.
(
1997
), “
A theoretical framework for the role of classroom assessment in motivating student effort and achievement
”,
Applied Measurement in Education
, Vol.
10
No.
2
, pp.
161
-
180
, doi: .
Castleberry
,
A.
and
Nolen
,
A.
(
2018
), “
Thematic analysis of qualitative data
”,
Currents in Pharmacy Teaching and Learning
, Vol.
10
No.
6
, pp.
807
-
815
, doi: .
Chappuis
,
S.
,
Chappuis
,
J.
and
Stiggins
,
R.
(
2009
), “
The quest for quality
”,
Educational Leadership
, Vol.
67
No.
3
, pp.
14
-
19
,
Christoforidou
,
M.
,
Kyriakides
,
L.
,
Antoniuo
,
P.
and
Creemers
,
B.P.M.
(
2014
), “
Searching for stages of teacher’s skills in assessment
”,
Studies in Educational Evaluation
, Vol.
40
, pp.
1
-
11
, doi: .
Coombs
,
A.
,
DeLuca
,
C.
and
MacGregor
,
S.
(
2020
), “
A person-centered analysis of teacher candidates’ approaches to assessment
”,
Teaching and Teacher Education
, Vol.
87
, p.
102952
, doi: .
Craven
,
G.
,
Beswick
,
K.
,
Fleming
,
J.
,
Fletcher
,
T.
,
Green
,
M.
,
Jensen
,
B.
,
Leinonen
,
E.
and
Rickards
,
F.
(
2014
), “
Action now: classroom ready teachers
”,
Teacher Education Ministerial Advisory Group
,
Cumming
,
J.J.
,
van der Kleigh
,
F.M.
and
Adie
,
L.
(
2019
), “
Contesting educational assessment policies in Australia
”,
Journal of Education Policy
, Vol.
34
No.
6
, pp.
836
-
857
, doi: .
Darling-Hammond
,
L.
,
Herman
,
J.
,
Pellegrino
,
J.
,
Abedi
,
J.
,
Lawrence Aber
,
J.
,
Baker
,
E.
, …
Steele
,
C.
(
2013
), “
Criteria for high-quality assessment
”,
Stanford Center for Opportunity Policy in Education
,
DeLuca
,
C.
and
Bellara
,
A.
(
2013
), “
The current state of assessment education: aligning policy, standards, and teacher education curriculum
”,
Journal of Teacher Education
, Vol.
64
No.
4
, pp.
356
-
372
, doi: .
DeLuca
,
C.
,
LaPointe-McEwan
,
D.
and
Luhanga
,
U.
(
2016
), “
Approaches to classroom assessment inventory: a new instrument to support teacher assessment literacy
”,
Educational Assessment
, Vol.
21
No.
4
, pp.
248
-
266
, doi: .
Donnelly
,
K.
and
Wiltshire
,
K.
(
2014
), “
Review of the Australian curriculum
”,
Australian Government, Canberra, ACT
,
Edwards
,
F.
(
2013
), “
Quality assessment by science teachers: five focus areas
”,
Science Education International
, Vol.
24
No.
2
, pp.
212
-
226
,
Ewing
,
R.
(
2013
),
Curriculum and Assessment
, (2nd ed.)
Oxford University Press
,
Melbourne, Vic
.
Gebril
,
A.
(
2017
), “
Language teachers’ conceptions of assessment: an Egyptian perspective
”,
Teacher Development
, Vol.
21
No.
1
, pp.
81
-
100
, doi: .
Gipps
,
C.
and
Stobart
,
G.
(
2009
), “
Fairness in assessment
”, in
Wyatt-Smith
,
C.
and
Cumming
,
J.J.
(Eds’),
Educational Assessment in the 21st Century
, pp.
105
-
118
, doi: .
Gonski
,
D.
,
Arcu
,
T.
,
Boston
,
K.
,
Gould
,
V.
,
Johnson
,
W.
,
O’Brien
,
L.
,
Perry
,
L.-A.
and
Roberts
,
M.
(
2018
),
Through Growth to Achievement: report of the Review to Achieve Educational Excellence in Australian Schools
,
Commonwealth of Australia
,
Canberra
.
Gulikers
,
J.
,
Bastiaens
,
T.
and
Kirschner
,
P.
(
2004
), “
A five-dimensional framework for authentic assessment
”,
Educational Technology Research and Development
, Vol.
52
No.
3
, pp.
67
-
85
, doi: .
Harlen
,
W.
(
2005
), “
Teachers’ summative practices and assessment for learning – tensions and synergies
”,
The Curriculum Journal
, Vol.
16
No.
2
, pp.
207
-
223
, doi: .
Klenowski
,
V.
(
2014
), “
Towards fairer assessment
”,
The Australian Educational Researcher
, Vol.
41
No.
4
, pp.
445
-
470
, doi: .
Klenowski
,
V.
and
Wyatt-Smith
,
C.
(
2014
),
Assessment for Education: standards, Judgement and Moderation
,
SAGE
,
Los Angeles
.
Leung
,
C.
and
Rea-Dickens
,
P.
(
2007
), “
Teacher assessment as policy instrument: contradictions and capacities
”,
Language Assessment Quarterly
, Vol.
4
No.
1
, pp.
6
-
36
, doi: .
Levy-Vered
,
A.
and
Nasser-Abu Alhija
,
F.
(
2015
), “
Modelling beginning teachers’ assessment literacy: the contribution of training, self-efficacy, and conceptions of assessment
”,
Educational Research and Evaluation
, Vol.
21
Nos
5/6
, pp.
378
-
406
, doi: .
McMillan
,
J.
(
2014
), “
High-quality classroom assessment
”,
Classroom Assessment: Principles and Practice for Effective Standards-Based Instruction
, (6th ed.) ,
Pearson Education
,
Melbourne, Vic
, pp.
57
-
91
.
Moher
,
D.
,
Liberati
,
A.
,
Tetzlaff
,
J.
and
Altman
,
D.G.
(
2009
), “
Preferred reporting items for systematic reviews and meta-analysis: the PRISMA statement
”,
Annals of Internal Medicine
, Vol.
151
No.
4
, pp.
264
-
269
, doi: .
Murphy
,
V.
,
Fox
,
J.
,
Freeman
,
S.
and
Hughes
,
N.
(
2017
), “
Keeping it real’: a review of the benefits, challenges and steps towards implementing authentic assessment
”,
All Ireland Journal of Higher Education
, Vol.
9
No.
3
, pp.
3231
-
3243
,
Panadero
,
E.
,
Fraile
,
J.
,
Pinedo
,
L.
,
Rodriguez-Hernandez
,
C.
and
Diez
,
F.
(
2022
), “
Changes in classroom assessment practices during emergency remote teaching due to COVID-19
”,
Assessment in Education: Principles, Policy and Practice
, Vol.
29
No.
3
, pp.
361
-
382
, doi: .
Prianto
,
A.
,
Qomariyah
,
U.N.
and
Firman
,
F.
(
2022
), “
Does student involvement in practical learning strengthen deeper learning competencies?
”,
International Journal of Learning, Teaching and Educational Research
, Vol.
21
No.
2
, pp.
211
-
231
, doi: .
Rao
,
N.J.
and
Banerjee
,
S.
(
2023
), “
Classroom assessment in higher education
”,
Higher Education for the Future
, Vol.
10
No.
1
, pp.
1
-
20
, doi: .
Remesal
,
A.
(
2011
), “
Primary and secondary teachers’ conceptions of assessment: a qualitative study
”,
Teaching and Teacher Education
, Vol.
27
No.
2
, pp.
472
-
482
, doi: .
Simpson
,
G.
(
2004
), “
Assessing learning in a student-centred classroom environment
”,
Science Education Review
, Vol.
3
No.
3
, pp.
85
-
88
,
Solomonidou
,
G.
and
Michaelides
,
M.
(
2017
), “
Students’ conceptions of assessment purposes in a low stakes secondary school: a mixed methodology approach
”,
Studies in Educational Evaluation
, Vol.
52
, pp.
35
-
41
, doi: .
Tveit
,
S.
(
2014
), “
Educational assessment in Norway
”,
Assessment in Education: Principles, Policy and Practice
, Vol.
21
No.
2
, pp.
221
-
237
, doi: .
Vogt
,
B.
(
2022
), “
Supportive assessment strategies as curriculum events in a performance-oriented classroom context
”,
European Educational Research Journal
, Vol.
21
No.
6
, pp.
1023
-
1040
, doi: .
Wiliam
,
D.
(
2008
), “
Quality in assessment
”, in
Swaffield
,
S.
(Ed.),
Unlocking Assessment: understanding for Reflection and Application
,
David Fulton Publishers
,
Abingdon, Oxon
, pp.
123
-
137
.
Woolfolk
,
A.
and
Margetts
,
K.
(
2007
),
Educational Psychology
,
Pearson Australia
,
French Forest
.
Wyatt-Smith
,
C.
,
Alexander
,
C.
,
Fishburn
,
D.
and
McMahon
,
P.
(
2017
), “
Standards of practice to standards of evidence: developing assessment capable teachers
”,
Assessment in Education: Principles, Policy and Practice
, Vol.
24
No.
2
, pp.
250
-
270
, doi: .
Yin
,
R.
(
2016
),
Qualitative Research from Start to Finish
, (2nd ed.) ,
The Guilford Press
,
New York
.
Department of Education and Training
(
2018
),
Through Growth to Achievement: Report of the Review to Achieve Educational Excellent in Australian Schools
,
Australian Government
,
Canberra, ACT
.
Teacher Education Ministerial Advisory Group
(
2014
), “
Action now: classroom ready teachers
”,

Supplementary material for this article can be found online.

Licensed re-use rights only

Supplementary data

Data & Figures

Figure 1.

Occurrence of summative assessment principles in combination

Figure 1.

Occurrence of summative assessment principles in combination

Close modal
Table 1.

Set of quality indicators for the creation of effective teacher-created summative assessment

QualityIndicators of quality in summative assessment creation
Valid
  • Clear and explicit alignment to the content taught and externally set curriculum requirements

  • Covers all content required to be assessed

  • Weighting of the assessment aligns to the weighting of treatment throughout teaching phase

  • Depth of thinking expected in assessment aligns to the depth of thinking practiced in the teaching phase

  • Relevance of assessment to the teaching phase is clear to student

Reliable
  • Clear and explicit alignment of the task sheet to the rubric

  • Written using language easily understood by both marker and student

  • Rubric created for objective marking, including:Quantifiable terms in descriptorsClear delineation between standardsIndication of weighting in each section

Fair
  • Context, requirements and wording of task considers accessibility of all students by eliminating bias due to gender, ethnicity, disability, culture, religion, socio-economic status, involvement in extra-curricular activities or any other bias specific to the subject

  • Clearly constructed and specific rubric to eliminate possible halo effect

Authentic
  • Task based on real-life situation

  • Open-ended which may have more than one solution

  • Requires collaboration or other interpersonal skills

  • Performance-based

Flexible
  • Student choice or voice in some aspect of the assessment

Source: Authors’ own work

Supplements

Supplementary data

References

Abell
,
S.
and
Siegel
,
M.
(
2011
), “
Assessment literacy: what science teachers need to know and be able to do
”, in
Corrigan
,
D.
,
Dillon
,
J.
and
Gunstone
,
R.
(Eds),
The Professional Knowledge Base of Science Teaching
,
Springer
, pp.
205
-
221
, doi: .
Australian Curriculum, Assessment and Reporting Authority
(
2020
), “
The shape of the Australian curriculum
”,
version 4.0, Australian Government
,
Baird
,
J.
,
Andrich
,
D.
,
Hopfenbeck
,
T.
and
Stobart
,
G.
(
2017
), “
Assessment and learning: fields apart?
”,
Assessment in Education: Principles, Policy and Practice
, Vol.
24
No.
3
, pp.
317
-
350
, doi: .
Bloom
,
B.S.
(
1956
), “
Taxonomy of educational objectives
”,
Handbook I: The Cognitive Domain
,
David McKay Co Inc
,
New York, NY
.
Bolden
,
B.
and
DeLuca
,
C.
(
2022
), “
Nurturing student creativity through assessment for learning in music classrooms
”,
Research Studies in Music Education
, Vol.
44
No.
1
, pp.
273
-
289
, doi: .
Brady
,
L.
and
Kennedy
,
K.
(
2019
),
Assessment and Reporting: celebrating Student Achievement
, (5th ed.)
Pearson Australia
,
Melbourne
.
Broadfoot
,
P.
and
Black
,
P.
(
2004
), “
Redefining assessment? The first ten years of assessment in education
”,
Assessment in Education: Principles, Policy & Practice
, Vol.
11
No.
1
, pp.
7
-
26
, doi: .
Brookhart
,
S.
(
1997
), “
A theoretical framework for the role of classroom assessment in motivating student effort and achievement
”,
Applied Measurement in Education
, Vol.
10
No.
2
, pp.
161
-
180
, doi: .
Castleberry
,
A.
and
Nolen
,
A.
(
2018
), “
Thematic analysis of qualitative data
”,
Currents in Pharmacy Teaching and Learning
, Vol.
10
No.
6
, pp.
807
-
815
, doi: .
Chappuis
,
S.
,
Chappuis
,
J.
and
Stiggins
,
R.
(
2009
), “
The quest for quality
”,
Educational Leadership
, Vol.
67
No.
3
, pp.
14
-
19
,
Christoforidou
,
M.
,
Kyriakides
,
L.
,
Antoniuo
,
P.
and
Creemers
,
B.P.M.
(
2014
), “
Searching for stages of teacher’s skills in assessment
”,
Studies in Educational Evaluation
, Vol.
40
, pp.
1
-
11
, doi: .
Coombs
,
A.
,
DeLuca
,
C.
and
MacGregor
,
S.
(
2020
), “
A person-centered analysis of teacher candidates’ approaches to assessment
”,
Teaching and Teacher Education
, Vol.
87
, p.
102952
, doi: .
Craven
,
G.
,
Beswick
,
K.
,
Fleming
,
J.
,
Fletcher
,
T.
,
Green
,
M.
,
Jensen
,
B.
,
Leinonen
,
E.
and
Rickards
,
F.
(
2014
), “
Action now: classroom ready teachers
”,
Teacher Education Ministerial Advisory Group
,
Cumming
,
J.J.
,
van der Kleigh
,
F.M.
and
Adie
,
L.
(
2019
), “
Contesting educational assessment policies in Australia
”,
Journal of Education Policy
, Vol.
34
No.
6
, pp.
836
-
857
, doi: .
Darling-Hammond
,
L.
,
Herman
,
J.
,
Pellegrino
,
J.
,
Abedi
,
J.
,
Lawrence Aber
,
J.
,
Baker
,
E.
, …
Steele
,
C.
(
2013
), “
Criteria for high-quality assessment
”,
Stanford Center for Opportunity Policy in Education
,
DeLuca
,
C.
and
Bellara
,
A.
(
2013
), “
The current state of assessment education: aligning policy, standards, and teacher education curriculum
”,
Journal of Teacher Education
, Vol.
64
No.
4
, pp.
356
-
372
, doi: .
DeLuca
,
C.
,
LaPointe-McEwan
,
D.
and
Luhanga
,
U.
(
2016
), “
Approaches to classroom assessment inventory: a new instrument to support teacher assessment literacy
”,
Educational Assessment
, Vol.
21
No.
4
, pp.
248
-
266
, doi: .
Donnelly
,
K.
and
Wiltshire
,
K.
(
2014
), “
Review of the Australian curriculum
”,
Australian Government, Canberra, ACT
,
Edwards
,
F.
(
2013
), “
Quality assessment by science teachers: five focus areas
”,
Science Education International
, Vol.
24
No.
2
, pp.
212
-
226
,
Ewing
,
R.
(
2013
),
Curriculum and Assessment
, (2nd ed.)
Oxford University Press
,
Melbourne, Vic
.
Gebril
,
A.
(
2017
), “
Language teachers’ conceptions of assessment: an Egyptian perspective
”,
Teacher Development
, Vol.
21
No.
1
, pp.
81
-
100
, doi: .
Gipps
,
C.
and
Stobart
,
G.
(
2009
), “
Fairness in assessment
”, in
Wyatt-Smith
,
C.
and
Cumming
,
J.J.
(Eds’),
Educational Assessment in the 21st Century
, pp.
105
-
118
, doi: .
Gonski
,
D.
,
Arcu
,
T.
,
Boston
,
K.
,
Gould
,
V.
,
Johnson
,
W.
,
O’Brien
,
L.
,
Perry
,
L.-A.
and
Roberts
,
M.
(
2018
),
Through Growth to Achievement: report of the Review to Achieve Educational Excellence in Australian Schools
,
Commonwealth of Australia
,
Canberra
.
Gulikers
,
J.
,
Bastiaens
,
T.
and
Kirschner
,
P.
(
2004
), “
A five-dimensional framework for authentic assessment
”,
Educational Technology Research and Development
, Vol.
52
No.
3
, pp.
67
-
85
, doi: .
Harlen
,
W.
(
2005
), “
Teachers’ summative practices and assessment for learning – tensions and synergies
”,
The Curriculum Journal
, Vol.
16
No.
2
, pp.
207
-
223
, doi: .
Klenowski
,
V.
(
2014
), “
Towards fairer assessment
”,
The Australian Educational Researcher
, Vol.
41
No.
4
, pp.
445
-
470
, doi: .
Klenowski
,
V.
and
Wyatt-Smith
,
C.
(
2014
),
Assessment for Education: standards, Judgement and Moderation
,
SAGE
,
Los Angeles
.
Leung
,
C.
and
Rea-Dickens
,
P.
(
2007
), “
Teacher assessment as policy instrument: contradictions and capacities
”,
Language Assessment Quarterly
, Vol.
4
No.
1
, pp.
6
-
36
, doi: .
Levy-Vered
,
A.
and
Nasser-Abu Alhija
,
F.
(
2015
), “
Modelling beginning teachers’ assessment literacy: the contribution of training, self-efficacy, and conceptions of assessment
”,
Educational Research and Evaluation
, Vol.
21
Nos
5/6
, pp.
378
-
406
, doi: .
McMillan
,
J.
(
2014
), “
High-quality classroom assessment
”,
Classroom Assessment: Principles and Practice for Effective Standards-Based Instruction
, (6th ed.) ,
Pearson Education
,
Melbourne, Vic
, pp.
57
-
91
.
Moher
,
D.
,
Liberati
,
A.
,
Tetzlaff
,
J.
and
Altman
,
D.G.
(
2009
), “
Preferred reporting items for systematic reviews and meta-analysis: the PRISMA statement
”,
Annals of Internal Medicine
, Vol.
151
No.
4
, pp.
264
-
269
, doi: .
Murphy
,
V.
,
Fox
,
J.
,
Freeman
,
S.
and
Hughes
,
N.
(
2017
), “
Keeping it real’: a review of the benefits, challenges and steps towards implementing authentic assessment
”,
All Ireland Journal of Higher Education
, Vol.
9
No.
3
, pp.
3231
-
3243
,
Panadero
,
E.
,
Fraile
,
J.
,
Pinedo
,
L.
,
Rodriguez-Hernandez
,
C.
and
Diez
,
F.
(
2022
), “
Changes in classroom assessment practices during emergency remote teaching due to COVID-19
”,
Assessment in Education: Principles, Policy and Practice
, Vol.
29
No.
3
, pp.
361
-
382
, doi: .
Prianto
,
A.
,
Qomariyah
,
U.N.
and
Firman
,
F.
(
2022
), “
Does student involvement in practical learning strengthen deeper learning competencies?
”,
International Journal of Learning, Teaching and Educational Research
, Vol.
21
No.
2
, pp.
211
-
231
, doi: .
Rao
,
N.J.
and
Banerjee
,
S.
(
2023
), “
Classroom assessment in higher education
”,
Higher Education for the Future
, Vol.
10
No.
1
, pp.
1
-
20
, doi: .
Remesal
,
A.
(
2011
), “
Primary and secondary teachers’ conceptions of assessment: a qualitative study
”,
Teaching and Teacher Education
, Vol.
27
No.
2
, pp.
472
-
482
, doi: .
Simpson
,
G.
(
2004
), “
Assessing learning in a student-centred classroom environment
”,
Science Education Review
, Vol.
3
No.
3
, pp.
85
-
88
,
Solomonidou
,
G.
and
Michaelides
,
M.
(
2017
), “
Students’ conceptions of assessment purposes in a low stakes secondary school: a mixed methodology approach
”,
Studies in Educational Evaluation
, Vol.
52
, pp.
35
-
41
, doi: .
Tveit
,
S.
(
2014
), “
Educational assessment in Norway
”,
Assessment in Education: Principles, Policy and Practice
, Vol.
21
No.
2
, pp.
221
-
237
, doi: .
Vogt
,
B.
(
2022
), “
Supportive assessment strategies as curriculum events in a performance-oriented classroom context
”,
European Educational Research Journal
, Vol.
21
No.
6
, pp.
1023
-
1040
, doi: .
Wiliam
,
D.
(
2008
), “
Quality in assessment
”, in
Swaffield
,
S.
(Ed.),
Unlocking Assessment: understanding for Reflection and Application
,
David Fulton Publishers
,
Abingdon, Oxon
, pp.
123
-
137
.
Woolfolk
,
A.
and
Margetts
,
K.
(
2007
),
Educational Psychology
,
Pearson Australia
,
French Forest
.
Wyatt-Smith
,
C.
,
Alexander
,
C.
,
Fishburn
,
D.
and
McMahon
,
P.
(
2017
), “
Standards of practice to standards of evidence: developing assessment capable teachers
”,
Assessment in Education: Principles, Policy and Practice
, Vol.
24
No.
2
, pp.
250
-
270
, doi: .
Yin
,
R.
(
2016
),
Qualitative Research from Start to Finish
, (2nd ed.) ,
The Guilford Press
,
New York
.
Department of Education and Training
(
2018
),
Through Growth to Achievement: Report of the Review to Achieve Educational Excellent in Australian Schools
,
Australian Government
,
Canberra, ACT
.
Teacher Education Ministerial Advisory Group
(
2014
), “
Action now: classroom ready teachers
”,

Languages

or Create an Account

Close Modal
Close Modal