CHAPTER 7
TESTS OF STATISTICAL SIGNIFICANCE
TEACHING ACTIVITIES FOR CHAPTER 7
Activity 1. Statistical significance of the difference between two means.
Even if students have taken a statistics course, we find it helpful to introduce them to important statistical
concepts in an intuitive, hands-on manner. Students tell us that they appreciate the opportunity to get multiple
perspectives on these concepts, especially if the statistics course emphasizes the mathematical logic and
computational procedures for statistics.
Students need to understand that samples lead to inconclusive results. Researchers might find a difference
ID
Control Group
Posttest Scores
Experimental
Group 1 Posttest
Scores
Experimental
Group 2
Posttest
Scores
1.
2
2
4
2.
5
5
7
3.
8
8
10
4.
3
3
5
5.
1
1
3
6.
7
7
9
7.
8
8
10
8.
8
8
10
9.
9
9
11
10.
3
3
5
11.
5
5
7
12.
5
5
7
13.
6
6
8
14.
2
2
4
15.
7
7
9
16.
9
9
11
17.
4
4
6
18.
6
6
8
19.
8
8
10
20.
3
3
5
Tell students that the control group and Experimental Group 1 participated in an experiment to improve their
learning. The total number of students was 40 (20 in each group). The posttest scores for the two groups revealed
that they performed identically, as shown in the table. The mean for both groups is 5.45, and the standard deviation
is 2.52.
Have students select another five random numbers and use them to select five students in Experimental Group
1. Ask them to record the posttest scores for these five students and then compute the mean. Finally, ask them to
subtract the Control Group mean from the Experimental Group 1 mean. Emphasize this step, lest they mistakenly
subtract the Experimental Group I mean from the Control Group mean.
Ask each student to say aloud the mean difference score that they obtained. Record their responses on a
blackboard, preferably using a line that has marks on it: -4, -3, -2, -1, 0, +1, +2, +3, +4. You can place an X or some
other symbol above this line corresponding to their response. You can round fractional numbers to whole numbers
for this purpose.
Students should find that the largest number of marks are around the 0 point, with an increasingly smaller
number of marks as the line goes closer to +4 and –4.
Now that students have collected and arrayed the data, you can ask them to think about what the data mean.
Ask the students to suppose that the 20 experimental group students and the 20 control group students constitute the
total population of students, and they have studied a random sample of 10 students from each group. Had they
studied all 40 students, half assigned to the control group and half assigned to Experimental Group 1, they would
have found no difference between them. In other words, the intervention had no effect. However, they know only
Activity 2. Confidence intervals.
A confidence interval is an important concept for education students. It teaches them that statistics, calculated
from samples, are only an estimate of what the population parameters are. You might remind students that this is
why pollsters usually report margins of error, which are the same thing as confidence intervals, to remind the public
that their poll results are only estimates based on their sampling of the population, and therefore are subject to error.
If you have done Activity 1 above, you can demonstrate confidence intervals intuitively by asking your students
to look at all the means that they computed for samples of N = 5 for the control group. You can ask each student to
Activity 3. Inferential statistics in the context of a research study.
You can help students learn how and why inferential statistics are used in research by forming them into small
groups in class. Ask each group to think of an educational intervention that they think is likely to be more effective
than conventional practice. Then ask them to think of an appropriate control group.
MULTIPLE-CHOICE TEST ITEMS FOR CHAPTER 7
1. Tests of statistical significance are used to
a. determine how important a statistical result is for educational practice.
b. make inferences about a population’s scores on a measure.
c. decide whether a statistical result is valid.
d. decide whether a statistical result is reliable.
2. Parameters provide information about score distributions for a
a. population.
b. sample.
c. subgroup within a sample.
d. sample with missing data.
3. If the 95 percent confidence interval around a sample mean of 8.24 is 7.27 and 9.68, we can conclude
that
a. the population mean is likely to be either 7.27 or 9.68.
b. the population mean is likely to be the midpoint between 7.27 and 9.68.
c. the population parameter for the sample mean is likely to be between 7.27 and 9.68.
d. the population parameter for the sample mean is likely to be lower than 7.27 or higher than 9.68.
4. Confidence intervals
a. become larger as the sample size increases.
b. become smaller as the sample size increases.
c. remain the same no matter what the sample size is.
d. do not contain confidence limits.
5. A margin of error has the same meaning as a(n)
a. effect size.
b. p value.
c. t value.
d. confidence interval.
6. Inferential statistics are used to make inferences about
a. populations from data collected on samples.
b. samples from data collected on populations.
c. margins of error.
d. whether to use a parametric or nonparametric test of significance.
7. If a t test is done to determine whether a sample of boys and a sample of girls perform differently on a
creativity test, the null hypothesis would be that the two samples
a. come from different populations.
b. come from the same population.
c. have the same confidence interval.
d. have different confidence intervals.
8. A p value is calculated to determine whether to
a. use a parametric or nonparametric test of statistical significance.
b. use a t test for independent means or for related means.
c. reject the null hypothesis.
d. reject a population parameter.
9. A p value of ___ is usually considered statistically significant.
a. .50
b. .10
c. .05
d. .55
10. A Type I error involves
a. the acceptance of a confidence interval when it is false.
b. the rejection of a confidence interval when it is true.
c. the acceptance of a null hypothesis when it is false.
d. the rejection of a null hypothesis when it is true.
11. A directional hypothesis can involve a prediction that
a. the population represented by Sample A will have a higher score on a test than the population
represented by Sample B.
b. the population represented by Sample A will have a lower score on a test than the population
represented by Sample B.
c. the populations represented by Samples A and B will have a higher score on a test than the
population represented by Sample C.
d. All of the above.
12. The t test for independent means is used to determine whether
a. two populations, each represented by a sample, have identical means on a variable.
b. two populations, each represented by a sample, have the same effect size.
c. a confidence interval for a population represented by a sample is statistically significant.
d. All of the above.
13. Analysis of variance is the appropriate statistical test to determine whether
a. the difference between two sample means is statistically significant.
b. the difference between three sample means is statistically significant.
c. two sample means have the same score distribution.
d. a Type I error occurred.
14. In an experiment, an interaction effect means that
a. an intervention increases in effectiveness the longer it is in place.
b. the effectiveness of an intervention depends upon characteristics of the individuals who receive it.
b. the test of statistical significance is dependent upon sample size.
d. the test of statistical significance is dependent upon whether the posttest is a ratio scale or
categorical scale.
15. Analysis of covariance is useful if an experimental and control group
a. are being compared on two posttest measures.
b. have different standard deviations on two posttest measures.
c. have different means on a pretest measure.
d. have missing data on the pretest measure for some members of the sample.
16. A chi-square test is appropriate if
a. there is an unequal number of individuals in each of the groups being compared.
b. the standard deviations of the posttest scores differ markedly for the groups being compared.
c. the sample size of each group being compared is below 10.
d. the posttest measures yield frequency counts.
17. The most common null hypothesis for a correlation coefficient obtained from a sample is that
the correlation coefficient for the population represented by the sample is
a. zero.
b. greater than zero.
c. less than zero.
d. – 1.00 or +1.00.
18. Parametric tests of statistical significance make the assumption that
a. there are equal intervals between the scores on the measures.
b. the scores are normally distributed about the mean.
c. the scores of the different comparison groups have equal variances.
d. All of the above.
19. Researchers generally recommend that a statistically significant finding should be viewed as
a. conclusive proof or refutation of a null hypothesis.
b. tentative evidence of the validity or invalidity of a null hypothesis.
c. more persuasive than evidence from replication studies.
d. evidence that the statistical result has practical significance.
SHORT-ANSWER TEST ITEMS FOR CHAPTER 7
1. Briefly explain the meaning of this statement: “Based on the experimental results, the null hypothesis of
no difference was rejected at a probability level of .05.”
2. Identify the appropriate test of statistical significance for each of the following examples.
a. The researcher compared the mean scores of Caucasian, Hispanic, and Asian-American students on
a measure of self-esteem.
b. The researcher compared the mean scores of Caucasian and Hispanic students on a measure of self
esteem.
c. The researcher compared the post-treatment mean scores of Caucasian and Hispanic students on a
measure of self-esteem after adjusting for differences in students’ self-esteem scores on a measure
administered before the treatment.
d. The researcher compared the number of Caucasian, Hispanic, and Asian-American students
who reported living in single-parent, two-parent, and extended-family households.
3. Identify the appropriate statistic for each of the following situations.
a. The typical performance of the students in a teacher’s class on a spelling test.
b. The amount of variability in the spelling test scores of a teacher’s students.
c. Whether there is a significant difference among the mean scores of music educators, language arts
educators, and science educators on a measure of satisfaction with an inservice program.
d. Whether a correlation coefficient is statistically significant.
e. Whether the proportion of teachers from each of five ethnic groups who are enrolled in an early
childhood education program and the proportion of teachers from each of the same five ethnic
groups who are enrolled in a secondary teacher education program is significantly different.
f. Whether a computer software program is more effective for students with good keyboarding skills
than for students with poor keyboarding skills.
CHAPTER 7
TESTS OF STATISTICAL SIGNIFICANCE
Multiple-Choice Test Items
Short-Answer Test Items