McDaniel & Gates Marketing Research, 9th Edition Instructor’s Manual
Instructor’s Manual
CHAPTER 15
Data Processing and Fundamental Data Analysis
LEARNING OBJECTIVES
1. To get an overview of the data analysis procedures.
3. To understand the data entry process and data entry alternatives.
5. To learn how to set up and interpret cross tabulations.
6. To comprehend basic techniques of statistical analysis.
KEY TERMS
Validation
Editing
Skip pattern
Marginal report
One-way frequency table
Cross tabulation
CHAPTER SCAN
This chapter examines the data processing procedure and fundamental data analysis. Data
processing begins with the validation of the interviewing process. The researcher must ensure
that the data was collected and recorded accurately and correctly. Once that has been
Instructor’s Manual
research because of technological advances. Once data has been entered, it should be submitted
to machine cleaning. Once the data has been correctly entered, it can be tabulated. This can be
done in several ways. One is a sample one-way frequency table that shows the number of
CHAPTER OUTLINE
1. Overview of the Data Analysis Procedure
2. Step One: Validation and Editing
I. Validation
3. Step Two: Coding
I. Coding
4. Step Three: Data Entry
I. Data Entry
5. Step Four: Logical Cleaning of Data
6. Step Five: Tabulation and Statistical Analysis
7. Graphic Representations of Data
I. Graphic Representations of Data Defined
8. Descriptive Statistics
I Descriptive Statistics
9. Summary
CHAPTER SUMMARY
1. OVERVIEW OF THE DATA ANALYSIS PROCEDURE
Once data collection is completed and questionnaires returned, here’s a 5-step procedure for data
analysis:
2. STEP ONE: VALIDATION AND EDITING
I. Validation
A. Validation Definedthe process of ascertaining that interviews were conducted as specified
goal of validation is solely to detect interviewer fraud or failure to follow key instructions
1. Telephone Validationcovers four areas
a. Was the person actually interviewed?
b. Did the person who was interviewed actually qualify to be interviewed?
c. Was the interview conducted in the required manner?
d. Did the interviewer cover the entire survey?
2. Checking for Other Kinds of Problems
a. Was the interviewer courteous?
b. Did the interviewer speculate about the client’s identity or the purpose of the survey?
c. Was the interviewer neat in appearance?
3. Purpose of the Validationprocess is to ensure that interviews were administered properly
and completely
II. Editing
A. Editing Definedinvolves checking for interviewer and respondent mistakes. The process of
ascertaining that questionnaires were filled out properly and completely.
1. Editing Processinvolves manual checking for a number of problems including the following:
a. Whether the interviewer failed to ask certain questions or record answers for certain questions.
See Exhibit 15.1 Sample Questionnaire (p 438 – 440)
b. Questionnaires are checked to make sure that skip patterns were followed.
See Exhibit 15.1 Sample Questionnaire (p 438 – 440)
c. Whether the interviewer paraphrased respondent’s answers to open-ended questions are
checked.
See Exhibit 15.2 Recording of Open-Ended Questions (p 441)
3. STEP TWO: CODING
I. Coding
A. Coding Definedrefers to the process of grouping and assigning numeric codes to the
responses to a question
B. Coding Process
1. Four steps in the coding of responses
See Exhibit 15.3 Sample of Responses to Open-Ended Question (p 442)
c. Set codes.
See Exhibit 15.4 Consolidated Response Categories and Codes of Open-Ended Responses
from Beer Study (p 443)
d. Enter codes
1) Read responses to individual open-ended questions on questionnaires
3) Write numeric code on the questionnaire for the response to the particular question
See Exhibit 15.5 Example Questionnaire Setup For Open-Ended Questions (p 443)
C. Automated Coding Systems
1. CATI and Internet Surveysdata entry and coding are completely eliminated for closed-
ended questions
a. Open-ended Questionstext is captured electronically but the coding process is still required
2. TextSmart Module of SPSSusing algorithms based on semiotics will do automated coding
4. STEP THREE: DATA ENTRY
I. Data Entry
Instructor’s Manual
b. System can be programmed to avoid certain types of errors at the point of data entry
C. The Data Entry Process
1. Process of going directly from the questionnaire to the data entry device and the associated
storage medium has proven to be more accurate and efficient than transferring data from
questionnaires to coding sheets.
D. Scanning
1. Technology of this Typehas been around for years. It has not been economically feasible to
use in marketing research until recently. Now a computer can generate the forms needed and
2. Electronically Captured Data is Increasing
a. Computer-assisted telephone interviewing
5. STEP FOUR: LOGICAL CLEANING OF DATA
I. Logical (Machine) Cleaning of Data Defined
A. Final Computerized Error Check of the Data
1. Error Checking Routinescomputer programs that accept instructions from the user to check
for logical errors in the data
2. Marginal Reporta computer-generated table of the frequencies of the responses to each
question, used to monitor entry of valid codes and correct use of skip patterns
3. Final Error Check in the Processwhen this step is completed, the computer data file should
be ready for tabulation and statistical analysis
6. STEP FIVE: TABULATION AND STATISTICAL ANALYSIS
I. Tabulation and Statistical Analysis Defined
1. One-way frequency table is the first summary of survey results. These tables typically indicate
the percentage of those responding who gave each possible response to a question
2. An Issue that must be dealt with when one-frequency tables are generated is what base to use
for the percentages for each table. There are three options for a base:
a. Total respondents
b. Number of people asked the particular question
c. Number answering the question
See Exhibit 15.7 One-Way Frequency Table Using Three Different Bases for Calculating
Percentages (p 447)
3. The question is whether percentages in frequency tables showing the results for these question
be based on the number of respondents or the number of responses
See Exhibit 15.8 Percentages for a Multiple-Response Question Calculated on the Bases of
Total Respondents and Total Responses (p 448)
B. Cross Tabulationsexamination of the responses to one question relative to the responses to
one or more other questions
See Exhibit 15.9 Sample Cross Tabulation (p 449). This cross tabulation includes frequencies
and percentage, with the percentages based on column totalsit shows an interesting relationship
between age and likelihood of choosing Minneapolis or St. Paul for hospitalization.
1. Following are a number of considerations regarding the setup of cross tabulation tables and the
2) row percentagebases on the row total
3) total percentagesbased on the table total
See Exhibit 15.10 Cross-Tabulation Table with Column, Row, and Total Percentages (p
449)
2. A common way of setting up cross-tabulation tables is to use columns to represent factors
3. A complex cross-tabulation, generated using the UNCLE software package is referred to as a
stub and banner table
See Exhibit 15.13 A Stub and Banner Table (p 500)
See Practicing Marketing Research: Six Practical Tips for Easier Cross Tabulations (p 451)
Here are six practical tips to improve your cross tabulation gleaning from Custom Insight, a
provider of Web-based survey software located in Carson City, Nevada
(www.custominsight.com).
2. Look for What is Not There
4. Keep Your Mind Open
5. Trust the Data
6 Watch the “n”–small totals should raise suspicions
See From the Front Line: Solid Set of Cross- Tabulations Critical (p 452-453)
1. Were all the responses to rating scale questions entered to ascending order (lowest rating is
lowest number, highest rating is highest number) with the exception of two or three questions?
2. Did an unusual number of respondents fail to answer only one particular open-end question?
7. GRAPHIC REPRESENTATIONS OF DATA
I. Graphic Representations of Data Defined
A. Line Charts
1. Simplest form of graphs
2. Useful for presenting a given measurement taken at several points over time
1. Appropriate for displaying percentages or proportions of a whole
See Exhibit 15.13 Three-Dimensional Pie Chart for Types of Music Listened to Most Often
(p 454)
C. Bar Charts
1. Plain Bar Chartsimplest form of bar chart
See Exhibit 15.14 Simple Two-Dimensional Bar Chart for Types of Music Listened to Most
Often (p 455)
See Exhibit 15.15 Simple Three-Dimensional Bar Chart for Types of Music Listened to
Most Often (p 456)
2. Clustered Bar Chartone of three types of bar charts useful for showing the results of cross
(p 456)
3. Stacked Bar Charthelpful in graphically representing cross tabulation results
See Exhibit 15.17 Stacked Bar Chart for Types of Music Listened to Most Often by Age (p
457)
4. Multiple-Row, Three-Dimensional Bar Chartmost visually appealing way of presenting
cross-tabulation information
See Exhibit 15.18 Multiple-Row, Three-Dimensional Bar Chart for Types of Music
Listened to Most Often by Age (p 457)
See Practicing Marketing Research: Expert Tips on Making Bad Graphics Every Time (p
455)
1. Occultation: hide everything important and avoid showing the relevant data.
Reduce to a minimum the data density (put miniscule data on a large chart).
Minimize the data/ink ratio (use a little ink for the data but lots for the axis, reference grid, and
Instructor’s Manual
Ignore the codification (make the lengths or data areas non-proportional to its values).
Discussion Questions
1. Come up with two more ways to hide data and two more ways to be inconsistent in presenting
data graphics.
8. DESCRIPTIVE STATISTICS
I. Descriptive Statistics
A. Descriptive Statistics Definedmost efficient means of summarizing characteristics of large
sets of data
B. Measures of Central Tendency
1. Four Basic Types of Measurement Scales
a. Nominal and ordinalreferred to as nonmetric scales
b. Interval and ratio scalesmetric scales
2. Three measures of central tendency
a. Meanproperly computed only from interval or ratio (metric) datasum of the values for all
observations of a variable divided by the number of observations (see p 506 for formula for
1. Measures of Dispersion Defined – measure how “spread out” the data are. They include the
standard deviation, variance, and range.
2. Measures of Central Tendency Defined indicate typical values for particular variable,
measures of dispersion indicate how spread out the data is.
a. The dangers associated with relying only on measures of central tendency are suggested by the
example in Exhibit 15.20 (p 459)
See Exhibit 15.20 Measures of Dispersion and Measures of Central Tendency (p 459)
D. Percentages and Statistical Tests
1. The research analyst makes the decision of whether to use measures of central tendency
(mean, median, mode) or percentages (one-way frequency tables, cross-tabulations)
2. Responses to questions either are categorical or take the form of continuous variables
a. Categorical variables–categories such as “Occupation” (coded 1 for professional/managerial,
2) If categoriesone-way frequency tables and cross tabulations are used for analysis
9. SUMMARY
QUESTIONS FOR REVIEW AND CRITICAL THINKING
1. What is the difference between measurement validity and interview validation?
2. Assume that Sally Smith, an interviewer, completed 50 questionnaires. Ten of the
questionnaires were validated by calling the respondents and asking them one opinion
question and two demographic questions over again. One respondent claimed that his age
category was 30-40, when the age category marked on the questionnaire was 20-30. On
another questionnaire, in response to the question, “What is the most important problem
facing our city government?” the interviewer had written, “The city council is too eager to
raise taxes.” When the interview was validated, the respondent said, “The city tax rate was
too high.” As a valuator would you assume that these were honest mistakes and accept the
entire lot of 50 interviews as valid? If not, what would you do?
The error on the first questionnaire may have been a judgment call on the part of the interviewer
or it may have simply been a missed stroke of the pencil or computer. The second mistake,
3. What is meant by the editing process? Should editors be allowed to fill in what they
think a respondent meant in open-ended questions if the information seems incomplete?
Why or why not?
4. Give an example of a skip pattern on a questionnaire. Why is it important to always
follow the skip patterns correctly?
1. Have you eaten at a fast food restaurant in the last two weeks? Yes No
5. It has been said that, to some degree, coding of open-ended questions is an art. Would
you agree or disagree? Why? Suppose that after coding a large number of questionnaires,
the researcher notices that many responses end up in the “Other” category, what might
this imply? What could be done to correct it?
6. Describe an intelligent data entry system. Why are data typically entered directly from
the questionnaire into the data entry device?
Intelligent data entry refers to the use of a computer and a software package to assist a keyboard
7. What is the purpose of machine cleaning the data? Give some examples of how data can
be machine cleaned. Do you think that machine cleaning is an expensive and unnecessary
step in the data tabulation process? Why or why not?
The purpose of machine cleaning is to ensure that all keypunches are in the proper columns and
8. It has been said that a cross-tabulation of two variables offers the researcher more
insightful information than does a one-way frequency table. Why might this true? Give an
example.
A cross-tabulation between two variables provides the researcher with data that looks at the