Purpose

The study reports a learning study on subtraction with negative numbers based on Variation Theory (VT), with special focus on exploring the conditions for using diagnostic classification models (DCMs) for assessment. DCMs offer a multidimensional statistical modeling approach where the dimensions may mimic CAs (CAs) in VT. The study aims to within a learning study develop a knowledge test that meets DCM standards and to assess its usefulness in that context.

Design/methodology/approach

Data comes from a learning study based on VT, in grade 6 in three Swedish schools. The developed teaching relied on four identified CAs that were also used to develop a knowledge test meeting the standards for DCM. The methodology involved iterative cycles of test development and statistical model evaluation, alongside qualitative methods such as think-aloud interviews.

Findings

The developed test meets statistical standards implied by DCMs. Challenges in operationalizing CAs in test items are identified. Detailed information about how test-takers discerned CAs becomes available, which also may be used for the development of teaching and the articulation of CAs. In particular, the DCM analysis provided individual profiles regarding CAs, supporting a nuanced understanding of learning outcomes at the group level.

Originality/value

The bridging VT and DCMs offer a comprehensive assessment of learning outcomes, emphasizing the importance of aligning tests with the object of learning. The findings suggest that DCMs can complement qualitative methods, providing a systematic and detailed analysis of students' knowledge.

Learning studies are classroom studies focusing on cyclical collegial development of teaching that have gained increased attention in the field of educational research (Ding et al., 2024; Ming Cheung and Yee Wong, 2014). Learning studies are theoretically informed by variation theory, giving grounds for designing teaching of a specific content or capability, an object of learning, as well as for evaluative analysis of teaching as enacted in the classroom (Kullberg et al., 2024; Marton, 2015). Assessment in this context should provide feedback about the effectiveness of enacted teaching in relation to desirable learning outcomes (Lo, 2012).

The focus is on developing teaching that bridges the gap between learning and instruction (Marton, 2015). A close analysis of aspects discerned during the lessons may lead to an improved teaching structure (Björklund et al., 2021). The aim of teaching, according to variation theory, is facilitating learners to discern the CAs that constitute key different meaningful parts or dimensions of the object of learning. Knowledge tests are in previous research used to assess students' knowledge. However, such tests are mainly designed to give a single value to each student's performance without providing further analysis of learning outcomes, i.e. an analysis of different dimensions or aspects of the object of learning. The possibility of knowledge tests being used as a tool to systematically monitor how CAs are discerned after intervention has not been studied earlier.

Reflecting student learning is central in a learning study (Ming Cheung and Yee Wong, 2014). The path of learning, according to variation theory, includes patterns of variation, facilitating first the separate discernment of each CA and then the fusion of simultaneous experience across several CAs, leading to a more comprehensive and powerful understanding of the object of learning (Marton, 2015). This path of discerning different aspects is often analyzed through qualitative methods such as observation, interviews or based on what has been expressed in the lessons (Kullberg et al., 2024). This work-intensive and time-demanding process aims to cover what knowledge tests usually do not evaluate: the qualitative characteristics of knowledge gained by learners. Thus, a structured statistical analysis of test results with the potential to capture different dimensions or aspects of the object of learning could be of high interest to researchers in this field.

Different statistical models have previously been used to investigate the multi-dimensional character of test results in relation to various mathematical content (Choi et al., 2015; Ren et al., 2021; Nurnberger-Haag et al., 2022), but this is new in the learning study domain. Such models make it possible to evaluate the effects of teaching design through statistical analysis of the test results. While learning study research aims to bridge learning and instruction, a multi-dimensional analysis of the test results in a learning study has the potential to innovatively bridge learning, instruction, and assessment (Björklund et al., 2021; Lee and Sawaki, 2009; Ren et al., 2021).

Diagnostic classification models (DCMs) offer multi-dimensional analysis through quantitative modeling of test results. Due to their unique capability of providing statistical analysis of different dimensions of knowledge instead of a straightforward one-dimensional assessment, DCMs have gained increased attention in the last decades (Javidanmehr and Sarab, 2017). To analyze students' learning with DCMs, the knowledge domain must be divided into attributes that are required for mastery of the domain and test items should directly relate to these attributes. The model then suggests multi-dimensional patterns of individuals' strengths and weaknesses (Paulsen and Valdivia, 2022). DCM also provides a correlation analysis between attributes, e.g. the possibility of mastery of one attribute related to other attributes. This provides a great opportunity for analyzing the test result when detailed information on different dimensions of learning is of interest.

In this paper, we explore the possibility of applying DCM as a structured statistical modeling to the test result from a learning study project. Starting from the object of learning as “subtraction with negative numbers” and its four CAs, we address the following questions:

  1. What are the challenges of structured statistical modeling when operationalizing CAs into items?

  2. In what ways can structured statistical analysis provide detailed information about how test-takers discerned CAs of the object of learning?

  3. To what extent can structured statistical analysis of test result provide feedback to researchers concerning the CAs of the object of learning?

A learning study (LS) based on variation theory (VT) is a classroom study with a design focus on the object of learning defined by its CAs (Kullberg et al., 2024). Those aspects that are required to master the object of learning from the desired perspective but are yet not discerned by learners are labeled as critical (Kullberg, 2010; Pang and Ki, 2016). Even though the CAs in principle may be understood as different for different individuals, in practice CAs in a learning study are analyzed at the classroom or group level, which form the ground for teaching design and theoretically is in line with the phenomenographic notion of analyzing qualitative differences of understanding at the collective level (Marton and Booth, 1997).

Teaching is designed to offer learners opportunities to discern those intended aspects of the object of learning (Marton, 2015). Kullberg et al. (2024) define the iterative learning study process in which assessing students' knowledge is integrated into different stages. The assessment aims to evaluate students' pre-knowledge and find out what is critical for students (step 2), how student knowledge is developed (step 5), and finally to assess how different aspects of the object of learning are discerned by learners (step 6). In most published articles about learning study, the focus is on teaching design and learning outcomes in general. This study focuses mainly on step 6: “Analyze student learning in relation to how the object of learning was enacted in the lesson and identify CAs” (Kullberg et al., 2024, p. 34). Analyzing the CAs after each intervention or at the end of the research process requires a qualitative approach, either through discussing observations or interviews or qualitative analysis of test results.

This study is part of a larger project involving an LS based on VT in three Swedish schools. The object of learning for this LS was “addition and subtraction with negative numbers”. The full LS project is reported elsewhere (Authors, 2025 under review) and summarized below.

According to Bofferding and Wessman-Enzinger (2017), much of the existing research on negative numbers has focused on investigating how various models or metaphors may support students' learning. Further, they argue that the results of these studies have been ambiguous and often focused on student difficulties. Another strand of previous research has examined students' reasoning when solving problems involving negative numbers (e.g. Bishop et al., 2014; Bofferding and Wessman-Enzinger, 2017; Lamb et al., 2018). These studies contribute valuable knowledge for teaching and often include some general recommendations or specific instructional sequences. However, they do not amount to a comprehensive description of teaching. Addressing this gap, the present study aimed at approaching the relationship between teaching and learning from this perspective. Through an iterative process of five cycles, a complete and detailed teaching design on operations with negative numbers was developed, implemented, analyzed and revised. The study was carried out in collaboration between six teachers and four researchers. In total, 128 students aged 12–13 years, enrolled in Grade 6 in Sweden, participated in the study, taking part both in the teaching and in the test analyzed in this article. The teaching design was organized into two 60-min lessons. The first lesson consisted of an introduction and two sections on addition. The second lesson focused on subtraction and was also divided into three sections. This article focuses on subtraction and therefore involves the introductory section of Lesson 1 and the three sections of Lesson 2.

In addition to VT, the teaching design was based on findings from previous research on students' understanding of negative numbers (e.g. Bishop et al., 2014; Bofferding and Wessman-Enzinger, 2017; Kilhamn, 2011; Kullberg, 2010; Lamb et al., 2018; Vlassis, 2004). Specifically, the study drew on Kullberg's (2010) work in the sense that the critical aspects reported in her study were used as a starting point. As described by Lamb et al. (2018), the goal of instruction on integers is to foster students' development of multiple ways of reasoning. Since the present study comprised two 60-min lessons, the critical aspects (Table 1) primarily encompass an introduction to the domain and represent fundamental components of students' understanding of operations with negative numbers that have repeatedly been identified as pivotal in previous research.

Table 1

The critical aspects for subtraction with negative number

Aspect numberDefinition
CA1The number system of integers
CA2The twofold meaning of the minus sign
CA3Subtraction can be interpreted as a difference
CA4The commutative law does not apply in subtraction
Source(s): Authors’ own creation

Aspect 1 (CA1), the number system, refers to the transition from natural numbers to integers, which is something that several studies have reported as problematic for students (e.g. Bishop et al., 2014; Kilhamn, 2011). Aspect 2 (CA2), the dual meaning of the minus sign, is also discussed in previous research (e.g. Vlassis, 2004). The fact that the minus sign can denote both an operation and a symbol for a negative number constitutes a challenge for students. Vlassis (2004) shows that some students interpret the minus sign solely as an operation. The interpretation of subtraction as “take away” or as a movement backward on the number line tends to dominate in the early school years (Kilhamn, 2011; Lamb et al., 2018). Aspect 3 (CA3), subtraction as difference, represents reasoning rather than taking away in solving subtraction, which needs to be complemented by Aspect 4 (CA4), which states that the commutative law does not apply in subtraction (e.g. in both 3–5 and 5–3, the absolute difference is 2). As Kilhamn (2011) argues, this aspect is critical for understanding subtraction with negative numbers. Furthermore, Bishop et al. (2014) describe this as a conceptual obstacle, since many students believe that subtraction cannot make larger. In our LS, the first lesson focused on CA1 and CA2 and the second lesson was aimed at making CA3 and CA4 visible.

The research project followed ethical guidelines described in the Swedish Research Council (Swedish Research Council, 2017). Participating students and teachers gave informed consent, and participation was voluntary. Students in the project were underage, and consent was also given by their legal guardians. The reported study used test data from students and with some student interviews regarding the test. Collected data were made anonymous before being subject to analysis.

Some scholars in recent decades have discussed the essence of knowledge tests being capable to evaluate different aspects of students' knowledge in relation to a specification of the underlying construct (e.g. Chudowsky and Pellegrino, 2003; Messick, 1984).

Several studies in the field of multidimensional analysis of mathematics test results apply multidimensional models to large-scale data, such as TIMSS (e.g. Choi et al., 2015) or on simulated data (e.g. Zhang et al., 2021). In the last decade, some studies have applied multi-dimensional models on relatively small real data (N < 300). Schoppek and Landgraf (2011) applied multidimensional Rasch model (MIRT) to evaluate dimensions of learning arithmetic in elementary school. Nurnberger-Haag et al. (2022) applied the Rasch model to a sample of 150 students to assess each of the three different dimensions: addition, subtraction and multiplication.

While there are different statistical models that potentially can assess different dimensions of skills in test results, diagnostic classification assessment is a multidimensional method that is particularly developed to assess different dimensions of a knowledge domain (Javidanmehr and Sarab, 2017; Shi et al., 2021).

DCMs [1] are psychometric models designed to classify test-takers according to their proficiency or nonproficiency (mostly dichotomously) of specified latent characteristics. Implementing DCMs requires a knowledge test and a matrix that connects the different attributes (latent characteristics) to specific test items. Some models can potentially be adopted with small-scale data (Maas et al., 2022; Paulsen and Valdivia, 2022).

When choosing a DCM in a particular study, it must be considered the prerequisites of the model, the purposes of the study, and the limitations that the data implies, must be considered. Different packages in R, which is an open-source programming environment for statistics (R Core Team, 2024), have been developed for DCMs analysis. The R-package G-DINA (de la Torre, 2011; Ma and de la Torre, 2020) is used in this study as the main package primarily due to functionality and convenience, and the R-package CDM (George et al., 2016) is applied to compare and complete the analysis. Other methods, such as research team discussions based on think-aloud interviews, are primarily used to assess the validity of the analysis. We briefly outline below the fundamentals of DCM, preparation for DCM, model evaluation, and the outputs it can generate. We also compare and discuss key differences between VT and DCMs, highlighting necessary compromises to understand their relationship in the study. Finally, we explain the process and how we evaluate and develop the model in different cycles.

The main output of a DCM analysis is the classification of the mastery level of attributes by test-takers at individual- and group-level. Other outputs, such as attribute correlation and item-level analysis, provide information about the independency of attributes from each other and the accuracy of the model, such as how well the allocation of attributes to test items fits with data from test-takers.

CAs that are then discerned by learners during the lesson(s) define the lived object of learning (Marton, 2015; Runesson, 2005). Assessment in this context should provide feedback about the effectiveness of enacted teaching in relation to desirable learning outcomes (Lo, 2012). Although VT provides some grounds for this, the precise assessment of how the object of learning has been experienced has (mainly) been subject to qualitative evaluations. How CAs are discerned in an LS, as described earlier, is subject to qualitative analysis of observations and interviews. DCM may complement this with a structured statistical analysis of test results. In DCMs, the knowledge domain is divided into attributes defined in line with the purpose of the study, and the model assesses the mastery level of these attributes. An appropriate statistical analysis at the attribute level is dependent on how well the test item reflects attributes as well as identifying appropriate model indices that fit the empirical data. Taking the CAs as attributes of knowledge, DCMs have the potential to suggest statistical analysis of the discernment of CAs. Table 2 describes the different terminologies and how corresponding terms relate when applying DCM and VT.

Table 2

Terminology in DCM and VT

DCMVT
Knowledge domainObject of learning
Domain-specific knowledge and skillsIs to be experienced/discerned from different necessary aspects
AttributeCritical Aspect
Pre-specified sub-skills of the target domainThe necessary aspects that are yet not discerned
Mastery of skillsDiscernment of aspects
Competency level on a number of cognitive attributesBeing able to see/experience new necessary aspects of the objects
Javidanmehr and Sarab (2017) Lo (2012) 
Source(s): Authors’ own creation

Preparing for DCM consists of defining the attributes, constructing a test in relation to those attributes, allocating attributes to test items in a matrix, and finally choosing a model among different DCMs and evaluating the model's performance. Assuming CAs of the object of learning as attributes, our first step was to construct a test that reflected one or several CAs in each item. We considered that each attribute should be addressed by a minimum of three items and that at least one item per attribute solely relates to one single attribute (Gu and Xu, 2019; Paulsen and Valdivia, 2022).

Validity of interpretation and use of test

The test designed and used in this study was specifically developed for the purpose of this learning study, aiming to evaluate the effects of teaching design and assess how the CAs of the object of learning were discerned by learners. Accordingly, the use of the test is limited to the purpose of learning study and the test results were interpreted to evaluate the outcomes of the two lessons designed during the LS process.

During the process of test design, we continuously discussed construct underrepresentation and construct-irrelevant variance, which are the two main threats to construct validity (American Educational Research Association et al., 2014). The construct underrepresentation, the risk of missing a part of the construct in the test, was carefully examined by making sure that each aspect was required in at least a few items. The construct-irrelevant variance, the risk that the student's answer is influenced by nonrelevant attributes, was managed by limiting each item to a single mathematical task with one correct answer. In general, we believe that constructing a test that was meant to precisely reflect CAs of learning decreased the risk of these two threats.

Item difficulty for some of our items is presented in Table 3. We compared the structure and difficulty levels of some items from our test with similar items from an established test on operation with negative numbers (Nurnberger-Haag et al., 2022). The item difficulties are similar for the corresponding tasks in both studies, which may be considered as a partial external validation of the properties of our test. However, there is an important difference between the purposes for which the results are used. In our case, it is about dichotomously modelling proficiency of four latent characteristics of students' ability to subtract with negative numbers, while the purpose of the study by Nurnberger-Haag et al. (2022) is to place student performance on a continuous unidimensional Rasch scale for subtraction as a single factor. DCMs sacrifice the continuous latent trait to afford multidimensionality.

Table 3

Integer structure and item mean for the 12 subtraction items (out of 28 in total)

Integer structure*No of itemsMeanItem ID DCM test
p-N10.3021
N-p20.4317, 23
n-N20.3320,27
N-n20.6219,24
p-P30.712,8,13
P-n20.3418, 22

Note(s): *Capital N or p indicates the greater absolute value

Source(s): Authors’ own creation

The final version of the test included 28 items, with a minimum of 5 items reflecting each CA. Theoretically, the combination of the number of test-takers (N = 128), the number of test items (j = 28), and the number of aspects (k = 4) was adequate for appropriate analysis with G-DINA model (Maas et al., 2022). Test items are presented in Table 4. Four cells on the right of each item with a value of 0 or 1 indicate whether the corresponding aspect is required to answer the item correctly or not. This table that relates items(j) and aspects(k) is called Q-matrix (Qjk) and scaffolds the DCM analysis. The programming environment R is then provided with two set of data, with students' test results and the Q-matrix. The codes are written so that all cells indicate 0 or 1 for each item and connect them to students' responses to items. For example, Q51 is 1, which means that answering item 5 correctly requires CA 1, while Q52, Q53 and Q54 are 0, which means that aspects 2, 3, and 4 are not required for this item. If a test-taker answers this item incorrectly, there is a possibility that she has not discerned aspect 1. However, in taking tests, there are always risks of guessing or slipping the correct answer. Thus, more than one item that assess aspect 1 in combination with other aspects is required to increase the creditability of the evaluation.

Table 4

The final Q-matrix. Test items (condensed version) and identified CAs allocated to each item

ItemCA1CA2CA3CA4
1Sort the numbers from smallest to largest1000
24–6 = _____0001
3Which of the expressions below contain negative numbers?
6 +(– 6), 6–5, 5–6, (−5) + (−6)
0100
4Which of the above expressions contain subtraction sign?0100
5Which number is larger? 8 or (−9)1000
6Which number is larger? (−6) or 01000
7Which number is larger? (−14) or (−12)1000
80–3 = _______1001
9Which number is in the middle of (−20) and 100011
10How many negative numbers are in 3 – (−5) = 80100
11How many negative numbers are in 5–3 = 20100
12How many negative numbers are in (−7) – 2 = (−9)0100
135–25 = __0011
14Use the number line to solve the subtraction 5–30010
15Use the number line to solve the subtraction 2 – (−3)0010
16Use the number line to solve the subtraction (−1) – (−4)0010
17(-5) – 2 = __0011
187 – (−4) = __0011
19(-4) – (−1) = __0111
20(-6) – (−8) = __0011
210 – (−5) = __0011
22Insert missing number: 5 – ____ = 80011
23Insert missing number:__ – 1 = (−5)0011
24Insert missing number: __ – (−2) = (−4)0011
25Insert one positive and one negative number: __ – __ = 50011
26Insert two negative numbers: __ – __ = (−9)0011
27Insert two negative numbers: __ – __ = 40011
28Insert two negative numbers: __ – __ =(-12)0011
Source(s): Authors’ own creation

The Q-matrix was developed based on the research team's judgment, think-aloud interviews with some test-takers, and the analysis that the model provides. Based on these sources, we refined the table a few times, which is described later as the model evaluation. Table 4 is the last version of Q-matrix that was used for the final analysis presented as the result.

Q-matrix misspecification is one of the main threats to the model's performance, and creating and refining the Q-matrix is essential in this modelling (de la Torre et al., 2016; Lei and Li, 2016). G-DINA packages used for our analysis offer an evaluation of the Q-matrix and even suggest a refined Q-matrix based on data. Besides evaluating the Q-matrix, the package generates absolute model-fit statistics at item-level and test-level, which contribute to evaluating model performance (Chen et al., 2013; Han and Johnson, 2019). Absolute model-fit statistics indicate the misfit degree between the data and the model, meaning that values closer to zero indicate a better fit. The model is evaluated at both the item level and the test level. There is no theoretically approved range of model-fit; however, there are historical agreements on what is an appropriate model-fit. At the test level, the two measures MADcor and SRMSR are both expected to be < 0.05 (George et al., 2016; Helm et al., 2022; Liu et al., 2018). At the item level, RMSEA is the indicator that is used to assess how well a proposed model for each item (i.e. allocation of CAs) fits the observed data. RMSEA is also a miss-fit indicator, which ideally is close to zero, and in practice is expected to be < 0.1 (Helm et al., 2022; Wang et al., 2015).

Other criteria for evaluating test items are attribute discrimination, guessing, slipping and classification accuracy. Attribute discrimination is the ability of an item to assess that attribute. The higher value of attribute discrimination means that the item has a higher capacity to assess that specific attribute. Guessing is the probability of answering the item correctly without having all the required attributes, and slipping is the probability of not answering the item correctly, although having all the required attributes; for these two indices, a value closer to zero is desired (Rupp et al., 2010). Classification accuracy is a measure that indicates the accuracy level of how well the model classifies test-takers into latent classes (Shi et al., 2021).

Applying G-DINA model on the first version of Q-matrix is referred to as cycle 0 in Figure 1. We applied the G-DINA R-package on the four CAs and a pre-defined Q-matrix that allocated those aspects to different items. This version of the Q-matrix did not meet the desired value of model-fit, which led to an iterative evaluation of different components, including the Q-matrix, CAs, test items, and the data. The research team discussions during the cycles were driven by their professional judgment on subject-matter and CAs, think-aloud interviews, and statistical analysis provided by the G-DINA package. Table 5 is an example of one of the items that was modified several times in the process of model evaluation for improving model-fit indices, but also to achieve acceptable item values such as guessing and slipping, and a reasonable attribute pattern. In explaining the cycles, we use this item, item 20, as an example to demonstrate changes during different cycles. A similar procedure has been done for several items.

Figure 1
A model shows sequential cycles linking critical aspects, test items, matrices, evaluation, and final statistical analysis.The model includes six rectangular boxes arranged horizontally from left to right. The first box is labeled “Cycle 0.” A rightward arrow connects it to the second box labeled “Critical Aspects,” which is followed by another rightward arrow leading to the third box labeled “Test items.” A rightward arrow connects to the fourth box labeled “Item or Aspect Matrix,” then to the fifth box labeled “Evaluating the model,” and finally to the sixth box labeled “Final statistical analysis.” A downward arrow emerges from the “Evaluating the model” box. The downward arrow continues and bends to the left. The leftward arrow then continues until it reaches the area below “Critical Aspects,” where it extends upward as an arrow labeled “Cycle 4” pointing to the “Critical Aspects” box. Another upward arrow labeled “Cycle 3, 4” emerges from the leftward arrow and points to the “Test items” box. An upward arrow labeled “Cycle 1, 2, 3, 5” emerges from the leftward arrow and points to the “Item or Aspect Matrix” box.

The process of implementing DCM in different cycles. Source: Authors’ own creation

Figure 1
A model shows sequential cycles linking critical aspects, test items, matrices, evaluation, and final statistical analysis.The model includes six rectangular boxes arranged horizontally from left to right. The first box is labeled “Cycle 0.” A rightward arrow connects it to the second box labeled “Critical Aspects,” which is followed by another rightward arrow leading to the third box labeled “Test items.” A rightward arrow connects to the fourth box labeled “Item or Aspect Matrix,” then to the fifth box labeled “Evaluating the model,” and finally to the sixth box labeled “Final statistical analysis.” A downward arrow emerges from the “Evaluating the model” box. The downward arrow continues and bends to the left. The leftward arrow then continues until it reaches the area below “Critical Aspects,” where it extends upward as an arrow labeled “Cycle 4” pointing to the “Critical Aspects” box. Another upward arrow labeled “Cycle 3, 4” emerges from the leftward arrow and points to the “Test items” box. An upward arrow labeled “Cycle 1, 2, 3, 5” emerges from the leftward arrow and points to the “Item or Aspect Matrix” box.

The process of implementing DCM in different cycles. Source: Authors’ own creation

Close modal
Table 5

Model fit and Item values for item 20: (−6) – (−8), mean: 0.36

Q-matrix attribute patternAttribute Discrimination
NItemsCA1CA2CA3CA4CA1CA2CA3CA4Item fit RMSEAGuessingSlippingModel fit SRMSRModel fit MADcor
 Primary Q-matrix1282801010.5730.9540.3230.04550.5730.1350.098
Cycle 1G-DINA Q-matrix1282810010.8660.8420.1930.00000.00000.0810.060
Cycle 2Revised Q-matrix1282811111.0000.8420.6461.0000.0040.00000.19590.1320.097
Cycle 3Exkluding outliers and easy items1222300110.6050.3800.0030.0440.2440.0880.066
Cycle 4Q-matrix with 3 aspects128241110.7060.4300.5000.0000.0000.0560.0950.066
Cycle 5Final Q-matrix1282800110.7310.4630.1390.0620.0060.0910.068
Cycle 6Final Q-matrix N = 1,0001,0002800110.6890.6440.0470.0890.0010.0310.024
Source(s): Authors’ own creation

In cycle 1, we used the Q-matrix suggested by G-DINA during cycle 0 to evaluate whether that would adjust better with the data, even though we didn't find the suggested matrix completely reasonable from a variation theory standpoint. With this Q-matrix, the RMSEA for some items was lower, but MADcor and SRMSR were still above 0.05.

In cycle 2, we chose all four aspects as a prerequisite to solving item 20, with the motivation that difficult items might require all CAs. In this revision, not only did the model fit not improve, but the guessing value for easier items increased. Thus, we removed CA1 and CA2 from this item with the motivation that these aspects do not directly contribute to solve difficult items such as item 20. A review of two test-takers video-recorded interviews confirmed our analysis that test-takers who answered item 20 correctly do not directly refer to CA1 and CA2 to solve the task.

Thus, in cycle 3, a new Q-matrix based on our analysis was constructed. In this cycle, we also focused on other components, such as the quality of test items and test-takers’ total scores. We excluded easy items with more than 90% correct answers. We also excluded test-takers with zero and full scores, aiming to achieve a better model fit, but this revision did not result in significant improvement.

In cycle 4, we decided to decrease the number of attributes from four to three. The argument was to balance the number of items, aspects, and participants, aiming to achieve a better model-fit. With the argument that the CA2, the twofold meaning of the negative sign, might be the prerequisite of discerning all other aspects, we decided to eliminate this aspect and those items that solely indicated this aspect. The item-fit values were slightly better, but the model-fit did not improve, so we decided to revert to four aspects.

In cycle 5, we did the last revision of the Q-matrix and implemented the model for all 4 aspects, 28 items, and 128 test-takers. The final revision of the Q-matrix in cycle 5 resulted in good item fit, RMSEA<0.1, also reasonable guessing and slipping values for all items. The attribute discrimination range for most of the items was also acceptable. Still, absolute model-fit indices (MADcor and SRSMS) were slightly higher than 0.05 (Table 5). However, as Ma (2020) stressed, these cut-off values are not empirically tested and approved, and thus there is no exact value that fits all models. This convinced us that the model fit we had achieved might be adequate and the sample-size might be the dominant factor that influences the model fit negatively (Lei and Li, 2016). One way to increase the sample-size is simulating data based on existing data. The simulated data in R is randomly generated item results based on the results from the current sample. This is a possibility that the R-package offers by following the mean and the variance of the already existing item data. To evaluate whether increasing sample-size, while other factors are constant, will improve our model-fit indices, in cycle 6, we simulated data to increase sample-size to 1,000. The model-fit indicators were strongly improved in this cycle; MADcor and SRMSR decreased to 0.024, respectively 0.031, which was the best model fit we achieved. When the accuracy of the analysis was approved through the procedure of data simulation, we used the output from cycle 5 with real data as the result of the modeling.

For a DCM to generate a reliable statistical analysis, it is important that the model-fit indices are in an acceptable range, the Q-matrix misspecification detected by the model is minimal, and the attributes are not highly correlated. Empirically fulfilling all these requirements at the end of the model improvement process of this study, we argue that the CAs and the test items provide a fair ground for applying DCM and that DCM could be used as an assessment framework to analyze the test result in this LS.

The result and analysis presented here follow the research questions. We first present the challenges we identified and faced when operationalizing CAs into test items. When we achieved appropriate model-fit and item-fit indices, we had a test with 28 dichotomic items, which were related to the 4 CAs (Table 4), and data from 128 test takers. The result breakdown from this set of data is then presented as critical-aspect discernment at group level, and at individual level. Finally, a correlation analysis between CAs and other some item-level analysis is presented to evaluate the accuracy of CAs.

Relating the CAs that are central in VT to test items raises challenges, and making use of DCM further enhances such challenges. DCM requires minimum one item that relates to only one aspect and items that require a combination of different aspects. CAs are the experiential constituents of the target capability and may be less well suited to be assessed individually. Thus, constructing items that reflect CAs individually was the most challenging part of the test construction process. CA3 is an example of an aspect that is difficult to assess individually mainly because it requires other aspects to be discerned simultaneously.

We chose two main strategies to construct a test that meets DCM requirements. First, we constructed some items based on each CA; second, we chose some items on “subtraction with negative numbers” and analyzed the item to find out what aspects were possibly required to answer these items correctly. To confirm our analysis, data from some think-aloud interviews with test-takers, compared with the allocated aspects in the Q-matrix informed us whether the test-taker employed the allocated aspects to solve the item. When modifying the Q-matrix, we also evaluated item-fit indices such as RMSEA, guessing and slipping parameters. Thus, an iterative process of item development, including researchers' analysis, item-fit indicators and think-aloud interviews, led to a test that met the DCM requirements.

The analysis of CAs on the group level (Figure 2) shows that about 90% of test-takers discerned aspects one and two (CA1, CA2), and about 62% discerned aspect four (CA4). Aspect three (CA3) was discerned by 29% of test-takers. Almost 16% of test-takers discerned all four aspects. The pattern 1,101, meaning that the test-taker has discerned CA1, CA2 and CA4, but not CA3, is the most repeated attribute pattern (45%).

Figure 2
A vertical bar chart shows percentages for four critical aspects labeled C A 1 to C A 4.The vertical bar chart shows a vertical axis labeled “Percentage underscore N 128” and ranges from 0.00 to 0.75 in increments of 0.25 units. The horizontal axis shows four categories labeled from left to right as “C A 1,” “C A 2,” “C A 3,” and “C A 4.” Each category contains a single vertical bar. The data for the bars on the graph are as follows: C A 1: 0.886. C A 2: 0.899. C A 3: 0.307. C A 4: 0.647. Note: All numerical data values are approximated.

The analysis of four CAs at the group level. Source: Authors own creation

Figure 2
A vertical bar chart shows percentages for four critical aspects labeled C A 1 to C A 4.The vertical bar chart shows a vertical axis labeled “Percentage underscore N 128” and ranges from 0.00 to 0.75 in increments of 0.25 units. The horizontal axis shows four categories labeled from left to right as “C A 1,” “C A 2,” “C A 3,” and “C A 4.” Each category contains a single vertical bar. The data for the bars on the graph are as follows: C A 1: 0.886. C A 2: 0.899. C A 3: 0.307. C A 4: 0.647. Note: All numerical data values are approximated.

The analysis of four CAs at the group level. Source: Authors own creation

Close modal

Table 6 demonstrates examples of aspect patterns of four test-takers and the probability of those aspects actually being discerned, two test-takers with pattern 1,111 and two with pattern 1,101. As a part of the evaluation of the quality of statistical analysis for individuals, we analyzed four think-aloud protocols for these four test-takers. The two test-takers with pattern 1,111 used different aspects in a consistent way during the interview and we didn't find any discrepancies between the interview and their test performances. Evaluating the accuracy of those with 1,101 pattern required more investigation. For one of the test-takers (Id74, Table 6) with the 1,101 pattern, it was challenging to determine whether the CA3 was completely critical. The test taker answered some items that required CA3 correctly. During the interview, the test taker could also confidently perform CA3 to solve item 18 but not item 20 (see Table 4). This made us investigate the attribute classification accuracy, which indicates the probability of each aspect being discerned by each test-taker. For test-taker Id74, the probability of CA1 and CA2 being discerned was 100%, and CA4 70%. The probability of discerning CA3 is 30% which can be interpreted as this aspect is partially discerned, which also resonates with the participant's performance during the interview. The analysis for test-taker Id127, with 1,101 pattern, indicates a 4% probability that CA3 is discerned. The test-taker's performance during the interview was aligned with this finding, which made us confirm that this test-taker had not yet discerned CA3.

Table 6

Aspect pattern and probability of each aspect for four test takers

Test-takerPatternCA1CA2CA3CA4
Id741,1010.991.000.290.70
Id1271,1010.991.000.040.97
Id931,1111.001.000.990.99
Id901,1110.991.000.990.99
Source(s): Authors’ own creation

With the support from the DCM analysis, we observed that the CAs might not be either critical or discerned, but, for some students, partially discerned. A learner's knowledge may thus be placed on a probability spectrum from how critical the aspect is to how well it is discerned. A qualitative interpretation of this may be that the learner's knowledge is dependent on the context or the situation when a specific item is answered.

Besides the qualitative evaluation of classification accuracy, we generated the class pattern accuracy. The accuracy for all class patterns was above 90%, which indicates that for test-takers classified in a specific pattern, there is a high probability that the pattern is accurate.

The outcome from the modeling is, as mentioned above, a profile consisting of a combination of zeros and ones in four dimensions. Tetrachoric correlations, suitable for dichotomous variables, of the four aspects are presented in Table 7.

Table 7

Latent trait correlations (correlation between CAs)

CA1CA2CA3CA4
CA11
CA20.411
CA30.550.381
CA40.570.69−0.131
Source(s): Authors’ own creation

In Table 7, three of the values are between 0.4 and 0.7, which may be considered as a low/moderate correlation (Ravand, 2016). The one exception, the correlation between aspects 3 and 4, is close to zero, which is interpreted as the two aspects being independent of each other.

DCM results support the relatively independent character of the CAs, i.e. that the qualitatively derived CAs are different dimensions of subtraction with negative numbers, i.e. the analysis statistically approved that the four aspects are four different necessary aspects.

Although the aim of knowledge assessment in learning studies is to evaluate to what extent the teaching design facilitate learners' experiencing and discerning CAs of the object of learning, the results of knowledge tests are still evaluated in the traditional total-score scale. This study is an attempt to fill this gap by using DCM to provide fine-grained information about how the CAs are discerned after the enacted teaching. In that regard, the test and the result presented in this study are related to the research design and goal.

Operationalization of CAs into test items was the main challenge we faced when developing the model. We understand this challenge to have two main parts. One part is theoretical, concerning commensurability between the manifestation of the object of learning as enacted in teaching in the classroom and in meeting test items, which are given individual written answers. Further, CAs typically tend to reflect more of a conceptual understanding, while test items and answering them may reflect more of procedural understanding. The second part is practical and methodical, concerning the construction of test items within the practical considerations and limitations of the study as well as methodical and statistical prerequisites given by DCM, such as the number of items, attributes and responses.

Input from the DCM analysis, such as the correlation matrix (Table 7), can be used to give additional input regarding whether the formulated CAs behave in an independent way when manifesting in how students answer the items, i.e. as a statistical verification of the qualitative hypothesis that those aspects are independent. In our study, such independence was observed because the correlations were low to moderate. LS is an iterative research process in which analyzing CAs is of high interest to evaluate teaching design and what is afforded to be learned (Lo, 2012). A cyclical development of CAs in conjunction with test items with support from this correlation analysis may enhance how attuned the articulation of these aspects is with the constitution of the lived object of learning. When defining and refining CAs in this iterative process, it might be of interest to consider the importance of CAs being feasible and practical to be assessed during the whole process of LS. Making CA feasible to be assessed by test-items or tasks contributes to all steps of LS and plays an essential role in analyzing what is intended to be learned, how it is enacted, and to what extent the object of learning is lived during the lesson(s).

To “analyze student learning in relation to the object of learning” and “present and discuss the findings with other teachers” are the two last steps of LS (Kullberg et al., 2024, p. 34). These two final steps are traditionally carried out qualitatively through discussing observations during the intervention and sometimes interviews with participants in a researcher-teachers team. Structured fine-grained analysis of CAs based on DCM offers another measure of the extent to which CAs are discerned. Applying such structural modeling may complement the assessment of aspect discernment on statistical grounds in addition to qualitative analysis. However, observations and interviews remain important methods to measure the outcome of an LS, while being time-consuming and limited in the number of participants typically involved.

To reach better alignment between the object of learning, enacted teaching and assessment of learning outcome, our conclusion is that changes should be considered in all parts of LS that may increase the quality in both the parts as such due to information from the other parts, as well as in the quality of the alignment. Possible such changes would be to reformulate CAs as well as items, and/or to coordinate this with changes in the teaching that is enacted.

DCM may constitute a possible path to evaluate test results at the aspect-level and may offer another systematic way of handling data with a larger group of students, for which, e.g. interviews quickly become cumbersome. That also contributes to the development of the field to increase the precision of how the learning outcome is assessed, which is the objective in focus of the research approach. In light of that, a discussion regarding the conditions for the quality and validity of an LS is supported.

1.

The model is also known as cognitive diagnostic model (CDM). Here we use the term diagnostic classification models based on The SAGE Encyclopedia of Educational Research, Measurement, and Evaluation, Bradshaw, 2018, “Diagnostic classification models,” The SAGE Encyclopedia of Educational Research, Measurement, and Evaluation SAGE Publications, Inc. pp. 508–512.

American Educational Research Association, American Psychological Association and National Council On Measurement In Education
(
2014
),
Standards for Educational and Psychological Testing
,
American Educational Research Association
,
Washington, DC
.
Bishop
,
J.P.
,
Lamb
,
L.L.
,
Philipp
,
R.A.
,
Whitacre
,
I.
and
Schappelle
,
B.P.
(
2014
), “
Using order to reason about negative numbers: the case of Violet
”,
Educational Studies in Mathematics
, Vol. 
86
No. 
1
, pp. 
39
-
59
, doi: .
Björklund
,
C.
,
Ekdahl
,
A.-L.
and
Runesson Kempe
,
U.
(
2021
), “
Implementing a structural approach in preschool number activities. Principles of an intervention program reflected in learning
”,
Mathematical Thinking and Learning
, Vol. 
23
No. 
1
, pp. 
72
-
94
, doi: .
Bofferding
,
L.
and
Wessman-Enzinger
,
N.
(
2017
), “
Subtraction involving negative numbers: connecting to whole number reasoning
”,
The Mathematics Enthusiast
, Vol. 
14
Nos
1-3
, pp. 
241
-
262
, doi: .
Bradshaw
,
L.
(
2018
), “Diagnostic classification models”, in
Keeves
,
J.P.
and
Watanabe
,
R.
(Eds),
The SAGE Encyclopedia of Educational Research, Measurement, and Evaluation
,
SAGE
,
Thousand Oaks, CA
, pp.
508
-
512
, doi: .
Chen
,
J.
,
de la Torre
,
J.
and
Zhang
,
Z.
(
2013
), “
Relative and absolute fit evaluation in cognitive diagnosis modeling: relative and absolute fit evaluation in CDM
”,
Journal of Educational Measurement
, Vol. 
50
No. 
2
, pp. 
123
-
140
, doi: .
Choi
,
K.M.
,
Lee
,
Y.-S.
and
Park
,
Y.S.
(
2015
), “
What CDM can tell about what students have learned: an analysis of TIMSS eighth grade mathematics
”,
Eurasia Journal of Mathematics, Science and Technology Education
, Vol. 
11
No. 
6
,
1563
, doi: .
Chudowsky
,
N.
and
Pellegrino
,
J.W.
(
2003
), “
Large-scale assessments that support learning: what will it take?
”,
Theory Into Practice
, Vol. 
42
No. 
1
, pp. 
75
-
83
, doi: .
De la Torre
,
J.
(
2011
), “
The generalized DINA model framework
”,
Psychometrika
, Vol. 
76
No. 
2
, pp. 
179
-
199
, doi: .
De la Torre
,
J.
,
Carmona
,
G.
,
Kieftenbeld
,
V.
,
Tjoe
,
H.
and
Lima
,
C.
(
2016
), “
Chapter 3: diagnostic classification models and mathematics education research: opportunities and challenges
”,
Journal for Research in Mathematics Education - Monograph
, Vol. 
15
, pp. 
53
-
72
.
Ding
,
M.
,
Huang
,
R.
,
Pressimone Beckowski
,
C.
,
Li
,
X.
and
Li
,
Y.
(
2024
), “
A review of lesson study in mathematics education from 2015 to 2022: implementation and impact
”,
ZDM
, Vol. 
56
No. 
1
, pp. 
87
-
99
, doi: .
George
,
A.C.
,
Robitzsch
,
A.
,
Kiefer
,
T.
,
Groß
,
J.
and
Ünlü
,
A.
(
2016
), “
The R package CDM for cognitive diagnosis models
”,
Journal of Statistical Software
, Vol. 
74
No. 
2
, pp. 
1
-
24
, doi: .
Gu
,
Y.
and
Xu
,
G.
(
2019
), “
The sufficient and necessary condition for the identifiability and estimability of the DINA model
”,
Psychometrika
, Vol. 
84
No. 
2
, pp. 
468
-
483
, doi: .
Han
,
Z.
and
Johnson
,
M.S.
(
2019
),
Global- and Item-Level Model Fit Indices
,
Springer International Publishing
,
Cham
, pp. 
265
-
285
, doi: .
Helm
,
C.
,
Warwas
,
J.
and
Schirmer
,
H.
(
2022
), “
Cognitive diagnosis models of students' skill profiles as a basis for adaptive teaching: an example from introductory accounting classes
”,
Empirical Research in Vocational Education and Training
, Vol. 
14
No. 
1
, pp. 
1
-
30
, doi: .
Javidanmehr
,
Z.
and
Sarab
,
M.A.
(
2017
), “
Cognitive diagnostic assessment: issues and considerations
”,
International Journal of Language Testing
, Vol. 
7
No. 
2
, pp. 
73
-
98
.
Kilhamn
,
C.
(
2011
),
Making Sense of Negative Numbers
,
Acta Universitatis Gothoburgensis
,
Göteborg
.
Kullberg
,
A.
(
2010
),
What Is Taught and what Is Learned: Professional Insights Gained and Shared by Teachers of Mathematics
,
Acta Universitatis Gothoburgensis
,
Göteborg
.
Kullberg
,
A.
,
Ingerman
,
Å.
and
Marton
,
F.
(
2024
),
Planning and Analyzing Teaching: Using the Variation Theory of Learning
, (1st ed.) ,
Routledge
,
Oxford
, doi: .
Lamb
,
L.
,
Bishop
,
J.P.
,
Philipp
,
R.A.
,
Whitacre
,
I.
and
Schappelle
,
B.P.
(
2018
), “
A cross-sectional investigation of students' reasoning about integer addition and subtraction: ways of reasoning, problem types, and flexibility
”,
Journal for Research in Mathematics Education
, Vol. 
49
No. 
5
, pp. 
575
-
613
, doi: .
Lee
,
Y.-W.
and
Sawaki
,
Y.
(
2009
), “
Cognitive diagnosis approaches to language assessment: an overview
”,
Language Assessment Quarterly
, Vol. 
6
No. 
3
, pp. 
172
-
189
, doi: .
Lei
,
P.-W.
and
Li
,
H.
(
2016
), “
Performance of fit indices in choosing correct cognitive diagnostic models and Q-matrices
”,
Applied Psychological Measurement
, Vol. 
40
No. 
6
, pp. 
405
-
417
, doi: .
Liu
,
R.
,
Huggins-Manley
,
A.C.
and
Bulut
,
O.
(
2018
), “
Retrofitting diagnostic classification models to responses from IRT-based assessment forms
”,
Educational and Psychological Measurement
, Vol. 
78
No. 
3
, pp. 
357
-
383
, doi: .
Lo
,
M.L.
(
2012
),
Variation Theory and the Improvement of Teaching and Learning
,
Acta Universitatis Gothoburgensis
,
Göteborg
.
Ma
,
W.
(
2020
), “
Evaluating the fit of sequential G-DINA model using limited-information measures
”,
Applied Psychological Measurement
, Vol. 
44
No. 
3
, pp. 
167
-
181
, doi: .
Ma
,
W.
and
de la Torre
,
J.
(
2020
), “
GDINA: an R package for cognitive diagnosis modeling
”,
Journal of Statistical Software
, Vol. 
93
No. 
14
, pp. 
1
-
26
, doi: .
Maas
,
L.
,
Brinkhuis
,
M.J.S.
,
Kester
,
L.
and
Wijngaards-de Meij
,
L.
(
2022
), “
Diagnostic classification models for actionable feedback in education: effects of sample size and assessment length
”,
Frontiers in Education
, Vol. 
7
, 802828, doi: .
Marton
,
F.
(
2015
),
Necessary Conditions of Learning
,
Routledge
,
London
.
Marton
,
F.
and
Booth
,
S.
(
1997
),
Learning and Awareness
,
Lawrence Erlbaum Associates
,
Mahwah, NJ
.
Messick
,
S.
(
1984
), “
The psychology of educational measurement
”,
Journal of Educational Measurement
, Vol. 
21
No. 
3
, pp. 
215
-
237
, doi: .
Ming Cheung
,
W.
and
Yee Wong
,
W.
(
2014
), “
Does lesson study work?: a systematic review on the effects of lesson study and learning study on teachers and students
”,
International Journal of Language and Literary Studies
, Vol. 
3
No. 
2
, pp. 
137
-
149
, doi: .
Nurnberger-Haag
,
J.
,
Kratky
,
J.
and
Karpinski
,
A.C.
(
2022
), “
The integer test of primary operations: a practical and validated assessment of middle school students' calculations with negative numbers
”,
International Electronic Journal of Mathematics Education
, Vol. 
17
No. 
1
, em0667, doi: .
Pang
,
M.F.
and
Ki
,
W.W.
(
2016
), “
Revisiting the idea of ‘critical aspects’
”,
Scandinavian Journal of Educational Research
, Vol. 
60
No. 
3
, pp. 
323
-
336
, doi: .
Paulsen
,
J.
and
Valdivia
,
D.S.
(
2022
), “
Examining cognitive diagnostic modeling in classroom assessment conditions
”,
The Journal of Experimental Education
, Vol. 
90
No. 
4
, pp. 
916
-
933
, doi: .
R Core Team
(
2024
), “
R: a language and environment for statistical computing
”.
Ravand
,
H.
(
2016
), “
Application of a cognitive diagnostic model to a high-stakes reading comprehension test
”,
Journal of Psychoeducational Assessment
, Vol. 
34
No. 
8
, pp. 
782
-
799
, doi: .
Ren
,
H.
,
Xu
,
N.
,
Lin
,
Y.
,
Zhang
,
S.
and
Yang
,
T.
(
2021
), “
Remedial teaching and learning from a cognitive diagnostic model perspective: taking the data distribution characteristics as an example
”,
Frontiers in Psychology
, Vol. 
12
, 628607, doi: .
Runesson
,
U.
(
2005
), “
Beyond discourse and interaction. Variation: a CA for teaching and learning mathematics
”,
Cambridge Journal of Education
, Vol. 
35
No. 
1
, pp. 
69
-
87
, doi: .
Rupp
,
A.A.
,
Templin
,
J.
and
Henson
,
R.A.
(
2010
),
Diagnostic Assessment: Theory, Methods, and Applications
,
The Guilford Press
,
New York
.
Schoppek
,
W.
and
Landgraf
,
A.
(
2011
), “
Can a multidimensional hierarchy of skills generate data conforming to the Rasch model? A comparison of methods
”,
Psychology Science
, Vol. 
53
No. 
1
, p.
3
.
Shi
,
Q.
,
Ma
,
W.
,
Robitzsch
,
A.
,
Sorrel
,
M.A.
and
Man
,
K.
(
2021
), “
Cognitively diagnostic analysis using the G-DINA model in R
”,
Psych
, Vol. 
3
No. 
4
, pp. 
812
-
835
, doi: .
Swedish Research Council
(
2017
),
Good Research Practice
,
Swedish Research Council
,
Stockholm
.
Vlassis
,
J.
(
2004
), “
Making sense of the minus sign or becoming flexible I negativity
”,
Learning and Instruction
, Vol. 
14
No. 
7
, pp. 
469
-
484
, doi: .
Wang
,
C.
,
Shu
,
Z.
,
Shang
,
Z.
and
Xu
,
G.
(
2015
), “
Assessing item-level fit for the DINA model
”,
Applied Psychological Measurement
, Vol. 
39
No. 
7
, pp. 
525
-
538
, doi: .
Zhang
,
J.
,
Lu
,
J.
,
Yang
,
J.
,
Zhang
,
Z.
and
Sun
,
S.
(
2021
), “
Exploring multiple strategic problem solving behaviors in educational psychology research by using mixture cognitive diagnosis model
”,
Frontiers in Psychology
, Vol. 
12
, 568348, doi: .
Published by Emerald Publishing Limited. This article is published under the Creative Commons Attribution (CC BY 4.0) licence. Anyone may reproduce, distribute, translate and create derivative works of this article (for both commercial and non-commercial purposes), subject to full attribution to the original publication and authors. The full terms of this licence may be seen at Link to the terms of the CC BY 4.0 licence.

or Create an Account

Close Modal
Close Modal