CHAPTER 4
Reliability of Assessment Results
ITEMS
1. Reliability refers to the
(a) consistency of the test scores.
(b) relevancy of the test scores.
(c) usefulness of the test scores.
2. Test scores that have a high degree of reliability (e.g., .90) are also valid.
(a) True
(b) False
3. Reliability is a(an) ____________ condition for validity of the results of a test.
(a) irrelevant
(b) necessary
(c) sufficient
4. What is meant by the statement “reliability is a necessary condition for validity“?
(a) Reliability and validity are both necessary concepts.
(b) Test scores can be valid without being reliable.
(c) Test scores that are valid have a significant degree of reliability.
5. A student’s true score on a test is not related to the specific questions appearing on the test.
(a) True
(b) False
6. A student’s true score in a subject does not change with the test questions the teacher asks.
(a) True
(b) False
7. Suppose a student is given the same math test in two consecutive days. Which score will remain
the same?
(a) The error score will be the same.
(b) The obtained score will be the same.
(c) The true score will be the same.
8. A test-retest reliability coefficient is a measure of the change in students’ relative standing in a
group from one administration of a test to another.
(a) True
(b) False
9. The standard error of measurement refers to the standard deviation of
(a) persons’ true score about their observed score.
(b) persons’ observed score about their true score.
(c) persons’ error about their observed score.
10. Which of the following procedures of estimating test reliability would yield the highest reliability
coefficient on a test comprised of 30 heterogeneous items?
(a) Coefficient alpha
(b) KR-20
(c) Split-halves
11. A teacher wants to know if students’ scores on the same assessment tasks remain stable over a
two-week period of time. Which of the following procedures of estimating reliability would be
appropriate for this purpose?
(a) Alternate forms reliability (different occasions)
(b) Alternate forms reliability (same occasion)
(c) Test-retest reliability (different occasions)
(d) Internal consistency reliability
The following instructions apply to Items 12 – 14.
Each numbered statement below describes a type of assessment data collection. Read each
statement and decide which type of reliability estimate it supports. On your answer sheet mark the
letter:
A — if the data are used for internal consistency reliability estimation
B — if the data are used for test-retest reliability estimation
C — if the data are used for alternate form (same occasion) reliability estimation
D — if the data are used for alternate form (different occasions) reliability estimation
12. A group of students was assessed on two occasions with the same assessment tasks with a week’s
interval between assessments. The assessor wanted to find out whether there was a marked
change in students’ ranking in the group.
13. A group of students was administered a different but similar set of assessment tasks, one set in the
morning and one later in the day. The assessor wanted to find out whether the relative ordering
of the students was similar on both sets of assessment tasks.
14. A test was administered to a group of students and the score of each student on each test task were
recorded. The assessor wanted to know whether students responded in a similar way from one
task to another.
15. Three multiple-choice tests were constructed to cover the same curriculum area. Test A had 20
items, Test B had 40 items, and Test C had 60 items. What is the most accurate statement that can
be made about the reliability of these tests?
(a) All the tests will have the same reliability.
(b) Test C will be three times as reliable as Test B and Test B will be twice as reliable as Test A.
(c) Test C will have the highest reliability and Test A will have the lowest reliability.
16. Suppose that a teacher wanted to determine how objective she had been in scoring her students’
essays. Which of the following procedures would be appropriate for her purpose?
(a) Retest her students with a similar test and correlate the two set of scores.
(b) Request a colleague to score the essays and correlate her scores with her colleagues’ scores.
(c) Separate by score the even and odd-numbered essays, creating two scores for each student.
Then correlate the two sets of scores.
17. A principal wanted to award prizes to students who demonstrated the best sense of citizenship.
The principal asked the students’ teachers and their immediate past teachers to rate the students.
With which type of reliability should the principal be most concerned?
(a) Internal consistency reliability
(b) Alternate forms reliability
(c) Scorer reliability
18. The standard error of measurement of a test is 3.5. A student obtained a score of 70 on the test.
How should the student’s score be interpreted?
(a) The student’s test scores will always be between 66.5 and 73.5.
(b) The student’s true score probably lies between 66.5 and 73.5.
(c) The student’s obtained scores should be raised to 73.5 or lowered to 66.5 depending on
information from other assessments.
(d) None of these is an appropriate interpretation.
19. Test A and Test B each have the same value for their standard deviations, 8.0. The reliability
coefficients of the two tests are .75 and .90, respectively. Which, if any, of the two tests has the
smaller standard error of measurement (SEM)?
(a) Test A
(b) Test B
(c) Both would have the same SEM.
20. Jane and John were assessed in reading. The standard error of measurement (SEM) for the scores
is 3.5. Jane obtained a score of 80 and John obtained 78 on the reading assessment. Is it
reasonable to conclude that Jane and John have the same reading ability?
(a) Yes
(b) No
21. Which of the following methods of estimating reliability would be most affected by the
speededness of the test?
(a) Alternate forms
(b) Test-retest
(c) Split-half
22. A high school mathematics teacher gave a test and used it to classify her students into two groups:
Those who have mastered the solutions to quadratic equations and those who have not. All those
who were classified as masters will be provided the same new instruction. Which of the following
reliability concerns should be of most concern to the teacher?
(a) Consistency of the students’ scores
(b) Consistency of her scoring
(c) Consistency of her classifications
23. A teacher of geography was interested in categorizing his students as masters and non-masters of
what they were taught. He constructed five essay items to match what he taught the class and
administered the items. What factor(s) could account for errors in his classification of students as
masters and non-masters?
(a) Setting a high pass score.
(b) Using only five essays in the assessment.
(c) Poor reliability in scoring the essays.
(d) All of the above.
(e) Only (b) and (c).
24. Which of the following decisions requires the highest reliability coefficient for the assessment
results?
(a) Placing students in instructional groups in the classroom.
(b) Granting admission to a college.
(c) Deciding whether subject-matter content should be retaught.
25. Which of the following decision-making situations is the LEAST high-stakes?
(a) Interviewing for job selection
(b) Testing for certification of competency at the end of a course.
(c) Testing for placement in a sequence of mathematics courses.
26. A 50-item multiple-choice test on the addition of three digit numbers was administered to a group
of students. The KR20 reliability estimate will yield a higher coefficient than the split-halves
procedure.
(a) True
(b) False
27. A perfectly reliable test is
(a) perfectly valid.
(b) perfectly consistent.
(c) free from bias.
(d) All of the above.
(e) Only (a) and (c).
28. Sue obtained a raw score of 75 in an algebra test. The standard error of measurement for the test is
5.5. What is the 68% uncertainty interval for Sue’s score on the test?
(a) 69.5 through 80.5
(b) 67.0 through 84.0
(c) 58.5 through 91.5
29. When a teacher crafts, administers, and scores an essay test, which of the following is the LEAST
likely source of measurement error?
(a) day-to-day fluctuations in student performance
(b) lack of strict parallel forms
(c) low scorer reliability
30. If you want to generalize a student’s performance on a group of tasks today to how the student is
likely performing on similar tasks, you need assessment results that are
(a) consistent over multiple occasions of testing.
(b) consistent across different items sampled from the same domain.
(c) consistent over multiple occasions and across different items from the same domain.
31. Which of the following will most likely improve the reliability of your assessments?
(a) applying strict time limits on your assessments
(b) including more questions on your assessments
(c) using performance assessments scored by a single well-trained rater
CHAPTER 4
Answers and Classification of Items