(1) A researcher gathers data over n=1500 months on X: Average Price of New Cars sold
in the state of California, and Y: Number of Highway Accidents in the state of California.
He discovers a negative association/correlation between X and Y.
(a) These data are [Pick one and explain your answer]: (1) cross-sectional data, (2) panel
data, (3) time-series data, (4) cross-sectional time-series data.
Time-series data. As with this example, in time series data the cases are temporal units
(here months), for some entity (here, California).
(b) Assume that there is a causal relationship between X and —that is, that X has a causal
effect on Y. Explain what it means to say that the average price of s new car (X) has a
causal effect on the number of highway accidents (Y).
Counterfactually, were X to have been larger or smaller for any given case (here, any given
month), then Y would have been different by some amount. That difference in Y—which
can never be observed—is the causal effect of the change in X.
(c) Even if we are confident that X has a (negative) causal effect on Y, we still may have
questions about the causal mechanism: Why do higher new car prices produce lower
numbers of highway traffic fatalities? Generate a hypothesis regarding the causal
mechanism by identifying an intervening variable (W). Explain your reasoning and
illustrate it with a causal diagram.
Here are two possible intervening variables: W1: number of miles driven (by Californians,
on average), and W2: typical driving speed (among Californians, on average). It could be
that as new car prices go up people respond by driving less (because they are driving their
old junkers or cannot afford to buy a car and must take public transportation) which results
in the lower level of traffic accidents. It could also be that as new car prices go up people
respond by driving more slowly (because their old junkers are not as powerful or as safe at
high speed) which results in the lower level of traffic accidents. Other possibilities might
also be imagined, of course.
(d) Now consider the possibility that the observed correlation is spurious—not, in fact, due
to the effect of X on Y, but a correlation that arises due to the operation of a confounding
variable, Z. Generate a hypothesis about how the X-Y relationship might be spurious by
identifying a confounding variable, Z. Explain your reasoning, indicate whether your
argument fits the common-cause or correlated-cause scenario, and illustrate it with a causal
diagram. Remember that the confounding variable you identify cannot be plausibly
considered an outcome of X or Y (i.e., it cannot be affected by X or Y).
Again, lots of possibilities might be imagined here. One possibility is that average new car
price happens to be negatively correlated with average rainfall (Z)—i.e., prices are lower
in months when the weather tends to be rainy and higher when the weather tends to be
sunny—and rainy weather causes higher traffic accidents. This could have produced the
observed correlation between gas prices and highway accidents. This would be a correlated
cause spuriousness scenario:
The curved, double-sided arrow above signifies a correlation; the straight-line arrow
designates a causal effect.
The most common mistake students make in developing a spuriousness scenario is that the
supposed “Z” they concoct is actually a “W”—a variable that is correlated with X because
X affects it! For a variable to be a confound, Z, (1) it cannot be affected by either X or Y,
(2) it must have a causal effect on Y, and it must either affect X (common cause scenario)
or be correlated with X for some unspecified reason (correlated cause scenario).
(2) Are the following statements true or false? Explain your answers—i.e. if it a statement
is true, explain why it is true, and if it is false, explain why it is false and what, instead, is
true.
A. The difference between an experiment and an observational study is that in an
experiment the researcher assigns subjects to the treatment group at random and in an
observational study the researcher has no control over which subjects get the treatment.
FALSE. What distinguishes an experiment from an observational study is that in an
experiment the researcher (1) has control over the assignment of subjects to experimental
groups and (2) has control over what X consists of and the setting of X’s administration.
The statement above is wrong because it gets #1 wrong and ignores #2. It gets #1 wrong
because the researcher’s assignment of subjects may or may not involve random
assignment (instead, the researcher could using matching).
B. To test the hypothesis that some X is sufficient to some Y, one need only gather data on
cases where Y is present.
FALSE. If X is sufficient to Y, then if X occurs Y must occur. To test it, then, one takes
cases where X occurs. One expects each of these cases to show Y occurring. If one finds a
case with X but without Y, the hypothesis is invalidated.
C. Suppose that in one dataset you find a correlation between X and Y. To replicate that
finding, one could either (a) see if the XY correlation is evident when analyzing random
subsets of cases in that dataset, or (b) look to see if the XY correlation emerges in entirely
new dataset.
TRUE. The essence of replication is showing whether or not a finding holds up with new
data. A finding demonstrated with one dataset is considered replicated if it shows up again