Chapter Fourteen: Basic Data Analysis
Chapter 14
Basic Data Analysis
AT-A-GLANCE
I. Introduction
II. Coding Qualitative Responses
A. Structured qualitative responses and dummy variables
III. The Nature of Descriptive Analysis
IV. Creating and Interpreting Tabulation
A. Cross-tabulation
2. Percentage cross-tabulations
4. How many cross-tabulations?
V. Data Transformation
A. Simple transformations
B. Problems with data transformations
C. Index numbers
D. Tabular and graphic methods of displaying data
VI. Hypothesis Testing Using Basic Statistics
A. Hypothesis testing procedure
VII. Significance Levels and p-Values
A. Type I and type II errors
1. Type I error
2. Type II error
VIII. Univariate Tests of Means
LEARNING OUTCOMES
2. Know what descriptive statistics are and why they are used.
4. Perform basic data transformations.
6. Be able to use a p-value to make statistical inferences.
7. Conduct a univariate t-test.
Chapter Fourteen: Basic Data Analysis
CHAPTER VIGNETTE: Don’t Tell Me!
Most people think that the last thing businesses like to hear is a consumer complaint. However, a
bigger problem may occur when consumers don’t complain. Complaints can provide key data that
allows businesses to improve the way they deal with customers. By improving service, the
businesses keep more customers and become more profitable. Thus, when consumers truly have a
problem, management should welcome complaints. From another perspective, a relatively large
SURVEY THIS!
Students are asked to compute the appropriate descriptive statistic for a question shown using the
data from the results for their class or school (across sections of marketing research classes taking
the survey) and to draw conclusions about dummy variables and they are asked to compute basic
descriptive statistics.
RESEARCH SNAPSHOTS
Wine Index Can Help Retailers
Indexes can be very useful and researchers are sometimes asked to create index values from
secondary data. For example, if a U.S. grocer is considering wine merchandising in another
country, he or she may be interested to know wine indexes in other countries. Using 1968
The Law and Type I and Type II Errors
While most attorneys and judges do not use statistical terminology such as Type I and Type II
errors, they do follow this logic. For example, the null hypothesis is that an innocent
TIPS OF THE TRADE
A frequency table can be a very useful way to depict basic tabulations.
Cross-tabulations and contingency tables are a simple and effective way to examine
relationships among less than interval variables.
When a distinction can be made between independent and dependent variables (that
are nominal or ordinal), the convention is that rows are independent variables and
columns are dependent variables.
Chapter Fourteen: Basic Data Analysis
A continuous variable that displays a bimodal distribution is appropriate for a median split.
Median splits should be performed on variables that display a normal distribution
OUTLINE
I. INTRODUCTION
A. As researchers, we infer whether or not some condition exists in a population based
on what we observe in a sample.
B. Alternatively, the research could be more exploratory and the researcher could be
using statistics simply to search for some pattern within the data.
II. CODING QUALITATIVE RESPONSES
A. Coding represents the way a specific meaning is assigned to a response within
previously edited data.
1. Codes represent the meaning in data by assigning some measurement symbol to
2. The proper form of coding relates back to the level of scale measurement.
a. Researchers code nominal data by using a word, letter, or any identifying
3. Any mistakes in coding can dramatically change the conclusions.
B. Structured Qualitative Responses And Dummy Variables
1. Qualitative responses to structured questions such as “yes” or “no” can be stored
in a data file with literally or with letters such as “Y” or “N.”
3. Since this represents a nominal numbering system, the actual numbers used are
arbitrary.
4. For statistical purposes the research may consider adopting dummy coding for
dichotomous responses like yes or no.
a. Dummy coding assigns a 0 to one category and a 1 to the other.
5. An alternative to dummy coding is effects coding.
© 2016 Cengage Learning. All Rights Reserved. May not be scanned, copied or duplicated, in whole or in
part, except for use as permitted in a license distributed with a certain product or service or otherwise on a
password-protected website for classroom use.
4. Today’s researcher has many convenient tools to quickly produce charts, graphs,
5. Bar charts (histograms), pie charts, curve/line diagrams and scatter plots are
among the most widely used tools.
VI. HYPOTHESIS TESTING USING BASIC STATISTICS
A. Descriptive research and causal research designs often climax with hypotheses tests.
1. Empirical testing typically involves inferential statistics.
3. Statistical analysis can be divided into several groups based on how many
variables are involved:
B. Hypothesis Testing Procedure
1. Steps:
a. First, the hypothesis is derived from the research objectives. The hypothesis
should be stated as specifically as possible and should be theoretically sound.
2. The exact point where the hypothesis changes from not being supported to being
3. Univariate hypotheses are typified by tests comparing some observed sample
mean against a benchmark value.
a. The test addresses the question, is the sample mean truly different from the
VII. SIGNIFICANCE LEVELS AND p-VALUES
A. A significance level is a critical probability associated with a statistical hypothesis
B. The term p-value stands for probability-value and is essentially another name for an
observed or computed significance level.
1. The probability in a p-value is that the statistical expectation (null) for a given
test is true.
2. So, low p-values mean there is little likelihood that the statistical expectation is
true.
Chapter Fourteen: Basic Data Analysis
a. This means the researcher’s hypothesis positing (suggesting) a difference
between an observed mean and a population mean, or between an observed
frequency and a population frequency, or for a relationship between two
C. Type I and Type II Errors
1. Because we cannot make any statement about a sample with complete certainty,
there is always the chance that an error will be made.
2. Type I Error
a. A Type I error occurs when a condition that is true in the population is
rejected based on statistical observations.
3. Type II Error
a. A Type II error is the probability of failing to reject a false hypothesis.
b. This incorrect decision is called beta (β).
4. Unfortunately, without increasing sample size the researcher cannot
simultaneously reduce Type I and Type II errors.
a. They are inversely related.
b. Thus, reducing the probability of a Type II error increases the probability of
Chapter Fourteen: Basic Data Analysis
VIII. UNIVARIATE TESTS OF MEANS
A. A univariate t-test is appropriate for testing hypotheses involving some observed
mean against some specified value such as a sales target.
2. When sample size (n) is larger than 30, the t-distribution and Z-distribution are
almost identical.
a. Therefore, while the t-test is strictly appropriate for tests involving small
sample sizes with unknown standard deviations, researchers commonly apply
3. To calculate a t statistic, the following formula is used:
a. 𝑡 = 𝑋−𝜇
𝑆𝑋
b. The researcher calculates the sample mean and standard deviation.
c. The standard error is computed.
B. The Z-distribution and the t-distribution are very similar, and thus the Z-test and t-
test will provide much the same result in most situations.
1. However, when the population standard deviation (σ) is known, the Z-test is most
appropriate.
Chapter Fourteen: Basic Data Analysis
QUESTIONS FOR REVIEW AND CRITICAL THINKING/ANSWERS
1. How does coding allow qualitative data to become useful to the researcher?
2. What are five descriptive statistics used to describe the basic properties of variables?
3. What is a histogram? What is the advantage of overlaying a normal distribution over a
histogram?
4. A survey asks respondents to respond to the statement “My work is interesting.” Interpret the
frequency distribution shown (taken from an SPSS output):
Relative Adjusted CUM
Absolute Freq Freq Freq
Not Very True 3 61 2.2 5.9 97.3
Not At All True 4 28 1.0 2.7 100.0
This table shows that 650 of the respondents in a survey of 2,715 responded that it was very true
that their work is interesting. The dot and the corresponding relative frequency column indicate
61.6 percent of the respondents did not answer this question, most likely because they were not
© 2016 Cengage Learning. All Rights Reserved. May not be scanned, copied or duplicated, in whole or in
part, except for use as permitted in a license distributed with a certain product or service or otherwise on a
password-protected website for classroom use.
91.5 percent.
5. Use the data in the following frequency table to:
a. Prepare a frequency distribution of the respondents’ ages
b. Cross-tabulate the respondents’ genders with cola preference
c. Identify any outliers
Weekly
Cola Unit
Individual Sex Age Preference Purchases
Mary F 20 Coke 2
Jim M 18 Coke 4
Sassi F 22 Pepsi 6
Amie F 20 Pepsi 2
a. Prepare a frequency distribution of the respondents’ ages.
Age Number Percent
22 1 10
19 2 20
18 1 10
Total 10 100
b. Cross-tabulate the respondents’ sex with cola preference.
Coke Pepsi Total
c. Identify any outliers.
Sassi (22 years old) could be considered an outlier because there are two years between Sassi’s
age and the next youngest age, but there is only one year between age categories under 20 years