The 'fuzzy computer-test' problem
The "fuzzy computer-test"problem
Ketevan M. PanchvidzeInstitute of Control Systems, Georgian Academy of Sciences, Tbilisi, Republic of Georgia (FSU)
Keywords Artificial intelligence, Computer testing, Cybernetics, Fuzzy approach
Abstract The problem of fuzzy computer-testing, as an alternative to ordinary testing, is discussed. Demonstrative examples, considered in the paper, emphasize the main aspects of the problem. Only one example for the topic of arithmetic is formalized in detail but, taking the problem from the wide angle, it is shown that the investigation of the various fuzzy directions in computer-testing is very important.
1. Introduction
Nowadays, a test method is used for examination of pupils and students. This approach, being naturally limited by the possibilities of computers, sometimes is effective enough and includes a training factor, especially for language and grammar teaching.
Positive aspects of the test problem are well-known and further developed. The consideration of computer teaching and test problems from the point of the qualitative extension would be more interesting.
Computer teaching programs undoubtedly help in mastering knowledge but, at the same time, they should assist in the treatment and comprehension of studied facts, and in finding out the regularities which studied topics are based on. Hence, the development of this kind of teaching program is of special interest.
The problem is similar to the one of the human everlasting search for the best pedagogical methods, and as the pedagogics develops (undoubtedly on the basis of traditional experience) the computer teaching science should develop as well. Moreover, the latter could cause an appearance of new teaching methods based on peculiarities of education by computer programs.
In the present paper some problems of computer testing are discussed. A significant shortcoming of an ordinary computer-examination is the "certainty" of a "pupil-computer" dialog, i.e. a question is usually followed by possible answers, and a pupil should select the one (only and true) of them. In this way of expressing the problem statement free thinking is limited, on the one hand, and does not permit us to estimate the common knowledge and erudition of a pupil, on the other hand.
How does a teacher usually realize an examination? If a pupil is not able to give any particular answer a teacher makes "fuzzification" of a question with the purpose of learning what the pupil knows about the discussed topic; he or she puts leading questions, draws parallels, etc. A good teacher is never limited by asking facts only but tries to reveal how much sensible knowledge a pupil has and what he or she knows in general.
The problem is close to the one of acquisition of expert knowledge during the construction of a knowledge base for decision support and expert systems. In this case, a programmer, who is making a computer application of a system, is trying to get the information (as much as possible) from experts. Then he/she makes its formalization for the further use in a decision-making method. The process is very difficult since knowledge in the expert memory is kept not as separate facts but is of a complicated structure including intuition and experience. Moreover, each specialist comprehends a problem from his own professional view, while an expert formalized opinion is needed to be used in decision support systems.
Thus, the investigation of the methods, especially "fuzzy" methods (i.e. those based on the theory of fuzzy sets (Zadeh, 1965)), for the questioning of experts, the knowledge treatment and formalization, is an important problem of artificial intelligence (AI).
Considering these methods in respect of a pupil test problem, the various possibilities and directions for their application and development can appear.
One kind of fuzzy testing could be realized in the following way (Panchvidze, 1998). At first, "experts", i.e. specialists, are to be asked with the purpose of constructing a "knowledge base"about the topic under consideration. Then each pupil has to answer the same questions. The results should be estimated by comparing them with the "expert knowledge base".
To base a fuzzy approach to a test problem more completely, a demonstrative example should be given. Moreover, the purpose of the examples, presented in this paper, is to reveal the common aspects that would further help in formalization of the fuzzy test problems concerning the various topics.
2. Common aspects
Let us consider an example from arithmetic which is quite simple for clarity and apprehension of the fuzzy approach to the computer test problem.
Let a pupil be examined in the multiplication of two numbers that are large. Certainly he or she could give an exact answer. That would point to the pupil's special ability for arithmetic. But an ordinary pupil is expected to give an answer by range, segment, in which an exact result could be arranged. A degree of pupil training could be estimated by a measure of closeness of a decision and exact result. If a pupil is giving answers, close to the true values and quickly enough, this would show his knowledge and experience in arithmetic. The problem of controlling the interval of time, put for decision making (DMT), could be considered separately.
In the presented example fuzzy methods of AI seem to be necessary for the result-data treatment. Let us pick out the main aspects.
At first, a fuzzy-answer should be estimated by a criterion which, by its nature, is close to the notion of a"membership function".
Second, the result information,derived from the questioning of pupils, is of a fuzzy nature and requires fuzzy decision-making methods for the knowledge estimation.
Third, a fuzzy aspect arises in a DMT control problem, e.g. how could it be estimated that a pupil replies"quickly enough", "not very quickly", "very late", etc?. These notions could be considered as linguistic variables.
Let {Q} be a set of questions with cardinality CQ.
Therefore, the problem is to define a criterion for a fuzzy answer estimation. With this purpose all questions should be divided into typical "classes", "types", etc. For each question a membership function (MF) of its answer accuracy (from the interval [0,1]) should be constructed. An effort could be made to search for a common formulation for estimation of the same class, type, etc. answers. Then, during testing, the values of MF for each fuzzy-answer should be calculated, and they could be considered as accuracy estimations of the latter.
The possible estimation scale could be the following:
(2) (non-satisfactory)(3) (satisfactory)(4) (good)(5) (excellent)
But let us consider, for convenience, the continuous range [0,1] for a fuzzy estimation of separate answers (denoted by EAq, q = 1, ..., CQ),so that instead of the discrete set
(1)
the unit interval is considered that would put fewer limits on the estimations EAq. Therefore, we would have:
"2" corresponds to EAq= 0;
"5" corresponds to EAq= 1; and
"3", "4" corresponds to EAq× [0,1].
Besides EAq,the DMT estimation should be taken into account. The problem is quite important. An exact a priori definition of a value of DMT for each question would not be appropriate, as a pupil could give the more exact answer for theperiod of time,large enough, and vice versa the less exact answer for less period of time.
One possible solution of the DMT problem could be the investigation of "experts", i.e. specialists' answers in the aspect of the spent DMT. The results should be treated with the purpose of revealing regularities and construct an MF of the DMT for the different classes and types of questions.
Since a DMT for each question is derived (it could be some kind of function of a result fuzzy-answer), a common estimation Eq of a fuzzy-answer could be calculated, e.g. by the formula:
(2)
where EAqis an estimation for the fuzzy-answer accuracy, and ETqan estimation for its DMT.
Below, in Section 3, the superscript "A" would be omitted, assuming that Eq is an estimation of accuracy.
Since all Eq are derived, an entire test-estimation (denoted by MD) should be calculated by either taking an average of the values of Eq, q =1,..., CQ, or by the use of the statistical cumulative law(SCL) (Panchvidze and Gachechiladze, 1996). It should be noted that for the different topics, the value of M could be derived by the different methods, e.g. if the test questions are not equivalent by significance, then an MF of "significance" for each Eq should be constructed and taken into account during the calculation.
In the end, the value of Mshould be reflected into the set (1), e.g. by the following rules:
(3)
It should be noted that the fuzzy aspects considered above could be spread on the different topics (probably by some modification). The emphasizing of the common directions for carrying out computer programs for fuzzy testing, generally, is of great importance. Since a many-sided analysis is made, a possibility of using the existing, extended or modified fuzzy approaches and methods of AI for the special problems could be considered.
3. Estimation of fuzzy-answer
Since the example presented below is a particular case, the introduced notions would be corresponding to the topic under consideration. Other topics could require other notions and designations but the description of the methods should always follow the common aspects,considered in Section 2.
Let us consider four classes of questions for arithmetic (the answers would be classified respectively):
(1) multiplication;(2) addition;(3) division;(4) subtraction.
Each class consists of several types, e.g. for the class of multiplication ("M") we have:
multiplication of two 2-digit numbers;
multiplication of two 3-digit numbers;
multiplication of two 2- and 3-digit numbers;
multiplication of three 2-, 3- and 4-digit numbers, etc.
Generally, test questions could be of the different types and they would require different methods for treatment. Naturally, the greater an order of multiplied numbers the greater deviation from the exact answer is allowed. This factor causes a problem of checking answers of the same type by a common estimation. Though, as it was noted above, it could be made an effort to search for a single formulation for estimation of the same type and even same class answers.
Let us introduce some notions and designations:
Multiplying numbers ni, i = 1, 2, 3, ...
For two number (for convenience) aNb.
"Order" of the number ni ri, where
(4)
Exact product P.
"Order-segment" (OS) of the product [Olow, Ohigh], where
(5)
"Dominant-segment" (DS) of the product [Dlow, Dhigh]:
(6)
where
(7)
In (4) and (7), "[ ]" truncates a fractional part of an argument.
"Answer-segment" (AS) [Ylow,Yhigh].
Keeping as general as possible, let us consider one type of the class "M", namely, multiplication of two 2-digit numbers, e.g. a = 98 and b = 59.
The "easy" numbers like 10, 11, 20,25, etc., should be considered separately.
The exact product is:
(8)
Now consider an AS, [Ylow,Yhigh].
The most rough estimation could be made by the OS:
(9)
If AS "covers" the OS, i.e.
(10)
then the estimation Eqcould be considered as "non-satisfactory".
At the next step AS should be compared with the DS:
(11)
If AS is arranged "between" the OS and DS, i.e.
(12)
then the estimation Eqcould be considered as "satisfactory".
If the following condition is fulfilled:
(13)
i.e.
(14)
the segment bounds for the estimations "good" and "excellent" should be found out.
Thus, the purpose at this step is to define a degree of accuracy of the AS. In the considered case, a problem of an "interval argument" of an MF could arise so that the direct use of the notion would not be appropriate.
Let us consider the function µ(x), x×R+, from Figure 1. The function consists of three parts:
(1)increasing;
(2)constant; and
(3)decreasing ones.
[P ×low, P ×high] interval could be interpreted as the highest estimation for a true "exact answer". ×low, ×highdepend on the order and value of P. The greater an order of a product,the larger constant interval should be taken. Proceeding from this, the following formulas for ×low and ×high could be used:

Figure 1.Membership function µ(x)
(15)
(16)
The domain of the first part of the MF is the interval [0, P ×low], including Dlow,and the domain of the third part is the interval [P + ×high,+ƒ],including Dhigh. Proceeding from (13-14) and the above discussion of an AS for the estimation "satisfactory", for the arguments Dlowand Dhigh, the MF should satisfy the conditions:
(17)
Several ways for an AS estimation could be considered:
(a)One possible formula for Eq calculation could be simply the following:
(18)
Note that it is not possible to estimate the "compactness" of an AS by (18). Therefore, it seems to be more appropriate to treat the values of Ylow and Yhighin more detail.
(b)With this purpose, let us consider the following auxiliary estimations:
(19)
where
(20)
is a middle point of an AS,and
(21)
Therefore, for the values e1, e2×[0,1], the following conditions are fulfilled:
(22)
Let us consider an S-form decreasing MF µsmall: [0,1]× [0,1] from Figure 2. The value µsmall (e1)could be interpreted as a measure of "accuracy" with regard to an exact product,and the value µsmall (e2) as a measure of "compactness" for an AS.
Taking into account the conditions (13-14), a common estimation Eq could be calculated by the formula:
(23)
Note that (23) is used only if the conditions (13-14) are not satisfied. The coefficients "0.3" and "0.35"are picked out with the purpose to provide the proper values for the minimum and maximum of Eq. Even if µ(e1) and µ(e2) take the zero-values, Eq estimation should not be lower than the threshold "0.3"(because of (13-14)), and if µ(e1) = µ(e2) = 1, the highest estimation (Eq = 1) is derived.
(c)But even now, there could arise difficulties in the cases when (13-14) conditions are satisfied, an AS is quite "compact", but P×[Ylow,Yhigh].
Taking into account this aspect, Ylowand Yhigh seem to be treated as separately as possible. With this purpose let us consider two membership functions:
(1)µlow for Ylow(Figure 3), and
(2)µhigh for Yhighassessment (Figure 4).
µlow and µhigh, for the values of Dlowand Dhigh, should satisfy the conditions similar to (17):
(24)

Figure 2.Membership function µsmall (x)

Figure 3."MF" for lower bound of "AS"

Figure 4."MF" for upper bound of "AS"
Thus, a common estimation Eqcould be calculated by the formula:
(25)
In the end, taking into account a factor of DMT, the value (ETq) of which could be depended on the order of P, the final common estimation of the fuzzy answer would be derived by the formula (2), where EAqcould be calculated by (18), (23) or (25).
During testing, for the classes considered and the types of arithmetical questions, the various kinds of interface could be used to be displayed. Certainly, the direct questions like"what is the product of a by b?" could be asked but, also, the structured questions could be used asking about, e.g.:
- a.
the result of multiplication, addition, subtraction and division of fractions;
- b.
the areas or perimeters of triangles, rectangles and other geometrical figures, where the use of the arithmetical operations is necessary; etc.
It would be preferable to pick up the problem questions in a game form so that a final fuzzy solution of task chains would lead to the prize with some accuracy. In this way it would make a pupil pursue a "target" with full concentration.
The problem of DMT could be solved in a game form as well, e.g. if a "target" is achieved in a short enough time,the pupil could be rewarded by some additional advantage.
4. Demonstrative examples from different topics
a) Example from chemistry
Let a pupil be given a name of some substance, e.g. some kind of alcohol. Using appropriate hardware the question could be displayed by illustrative examples from nature or human activity. That could extend pupil knowledge about the use of the particular substance.
A pupil could answer by giving the formula and structure of the substance and its chemical properties, displaying the corresponding chemical reactions. In difficult cases a friendly computer interface could provide the prompts, e.g. by the visualization of the natural examples.
If a pupil is not able to answer the question, he/she could be asked (as a teacher does) about the elements and its groups which the particular substance could consist of, about a class of the compound (if it is organic), at least, about its type (organic or inorganic),etc.
Naturally, this kind of answer could not be considered as an exact one but it would reveal all the knowledge that a pupil has. Estimating a training level by fuzzy measures (similar to those introduced above), a pupil does not get an excellent estimation but he could get "3" or "4", depending on the accuracy of his answers, i.e. deviation from the exact ones. This kind of estimation, by its nature, differs from the one derived by the ordinary testing, which is able only to count the wrong and right answers.
Fuzzy problem testing is also appropriate for the topics of literature and grammar.
b) Example from literature
Let a pupil be given a quotation, a passage from a novel or poem. He/she is to make a decision about its author and title. If a pupil is not able to answer exactly, he/she could be asked about:
a literary school of the given quotation;
a list of possible authors;
a possible period of activity of the author, etc.
This would reveal pupil erudition,especially if the given quotation does not belong to a well-known author. The latter case should be taken into account in estimation.
c) Example from grammar
Let a pupil be given a word-form or word group (e.g. from a sentence). The request is to find out its grammar features, its characteristics: which part of speech and/or structure of sentence it is. Then, e.g.:
- a.
For the verbs of which number, person (especially for the languages for which different persons have different forms), tense, mood, voice, aspect,versia, causative, skriva ("mckrivi", for the Georgian language) it is;
- b.
For the nouns of which number, gender, case; of which kind: common,proper, abstract, collective, it is; etc.
The more information is derived from a pupil the more knowledge he has. Each right answer to each question item would bring its share in an entire estimation, depending on its significance.
5. Conclusion
Each particular topic, and sometimes even each test-question, would require its own way of formalization and application. The questions for each topic should form the "classes","types", etc., typical for the one considered. Thus, each applied problem should be described and carried out separately.
Undoubtedly, the problem of visualization of a "pupil-computer" dialog is of great importance but, at the same time, the investigation of a computer test problem from the point of knowledge estimation (i.e. treatment of pupils' fuzzy-answers), should extend to an applied area of the problem in general.
Editor's note: Communications and forum contributions are not sent to referees and consequently will be available to readers much more quickly. Comments and alternative viewpoints on all matters pertaining to cybernetics and systems are sought, particularly on the many issues that are raised in this section.
