1
Chapter 07 Measurement and Scaling
Chapter 7
Measurement and Scaling
Learning Objectives (PPT slide 7-2)
1. Understand the role of measurement in marketing research.
2. Explain the four basic levels of scales.
Key Terms and Concepts
1. Behavioral intention scale
2. Comparative rating scale
3. Constant-sum scales
4. Construct
5. Construct development
Chapter Summary by Learning Objectives
Understand the role of measurement in marketing research.
2
Chapter 07 Measurement and Scaling
Measurement is the process of developing methods to systematically characterize or quantify
information about persons, events, ideas, or objects of interest. As part of the measurement process,
researchers assign either numbers or labels to phenomena they measure. The measurement process
consists of two tasks: construct selection/development and scale measurement. A construct is an
unobservable concept that is measured indirectly by a group of related variables. Thus, constructs
are made up of a combination of several related indicator variables that together define the concept
being measured. Construct development is the process in which researchers identify
characteristics that define the concept being studied by the researcher.
Explain the four basic levels of scales.
The four basic levels of scales are nominal, ordinal, interval, and ratio. Nominal scales are the most
basic and provide the least amount of data. They assign labels to objects and respondents but do
not show relative magnitudes between them. Nominal scales ask respondents about their religious
affiliation, gender, type of dwelling, occupation, last brand of cereal purchased, and so on. To
analyze nominal data researchers use modes and frequency distributions. Ordinal scales require
respondents to express relative magnitude about a topic. Ordinal scales enable researchers to
create a hierarchical pattern among the responses (or scale points) that indicate “greater than/less
than” relationships. Data derived from ordinal scale measurements include medians and ranges as
Describe scale development and its importance in gathering primary data.
There are three important components to scale measurement: (1) question/setup; (2) dimensions of
the object, construct, or behavior; and (3) the scale point descriptors. Some of the criteria for scale
3
Chapter 07 Measurement and Scaling
development are the intelligibility of the questions, the appropriateness of the primary descriptors,
and the discriminatory power of the scale descriptors. Likert scales use agree/disagree scale
descriptors to obtain a person’s attitude toward a given object or behavior. Semantic differential
Discuss comparative and noncomparative scales.
Comparative scales require the respondent to make a direct comparison between two products or
services, whereas noncomparative scales rate products or services independently. Data from
comparative scales is interpreted in relative terms. Both types of scales are generally considered
interval or ratio and more advanced statistical procedures can be used with them. One benefit of
comparative scales is they enable researchers to identify small differences between attributes,
constructs, or objects. In addition, comparative scales require fewer theoretical assumptions and
are easier for respondents to understand and respond to than are many noncomparative scales.
Chapter Outline
Opening Vignette
Santa Fe Grill Mexican Restaurant: Predicting Customer Loyalty
The opening vignette in this chapter describes the problems facing Santa Fe Grill Mexican
Restaurant as the owners seek to better understand the factors leading to customer loyalty. That is,
what would motivate customers to return to their restaurant more often?
Several insights about the importance of construct and measurement developments can be gained
from the Santa Fe Grill experience. First, not knowing the critical elements that influence
customers’ restaurant loyalty can lead to intuitive guesswork and unreliable sales predictions.
Second, developing loyal customers requires identifying and precisely defining constructs that
predict loyalty (i.e., customer attitudes, emotions, behavioral factors).
I. Value of Measurement in Information Research (PPT slide 7-3)
4
Chapter 07 Measurement and Scaling
Measurement is an integral part of the modern world, yet the beginnings of measurement lie in the
distant past. Because accurate measurement is essential to effective decision making, this chapter
provides a basic understanding of the importance of measuring customers’ attitudes and behaviors
and other marketplace phenomena.
II. Overview of the Measurement Process (PPT slide 7-3)
Measurement is the integrative process of determining the intensity (or amount) of information
about constructs, concepts, or objects. As part of the measurement process, researchers assign
either numbers or labels to phenomena they measure.
The measurement process consists of two tasks.
Construct selection/development: The goal is to precisely identify and define what is to be
measured.
III. What is a Construct? (PPT slides 7-4 to 7-6)
A construct is an abstract idea or concept formed in a person’s mind. This idea is a combination of
a number of similar characteristics of the construct. The characteristics are the variables that
collectively define the concept and make measurement of the concept possible.
A. Construct Development (PPT slide 7-4)
Marketing constructs must be clearly defined. A construct is an unobservable concept that is
measured indirectly by a group of related variables. Thus, constructs are made up of a
Construct development is an integrative process in which researchers determine what specific
data should be collected for solving the defined research problem. The process begins with an
5
Chapter 07 Measurement and Scaling
IV. Scale Measurement (PPT slides 7-7 to 7-12)
The quality of responses associated with any question or observation technique depends directly
on the scale measurements used by the researcher. Scale measurement is the process of assigning
descriptors to represent the range of possible responses to a question about a particular object or
construct. The scale descriptors are a combination of labels, such as “Strongly Agree” or
“Strongly Disagree” and numbers, such as 1 to 7, which are assigned using a set of rules.
Scale points reflect designated degrees of intensity assigned to the responses in a given
questioning or observation method.
All scale measurements can be classified as one of four basic scale levels as given below (PPT
slide 7-10).
Nominal
A. Nominal Scales (PPT slides 7-8 to 7-9)
A nominal scale is a type of scale in which the questions require respondents to provide only
some type of descriptor as the raw response. It is the most basic and least powerful scale design.
Responses do not contain a level of intensity. Thus, a ranking of the set of responses is not
B. Ordinal Scales (PPT slides 7-8 to 7-10)
6
Chapter 07 Measurement and Scaling
An ordinal scale is a scale that allows a respondent to express relative magnitude between the
answers to a question. Ordinal scales are more powerful than nominal scales. This type of scale
enables respondents to express relative magnitude between the answers to a question and
responses obtained can be rank-ordered in a hierarchical pattern. Thus, relationships between
responses can be determined such as “greater than/less than,” “higher than/lower than,” “more
often/less often,” “more important/less important,” or “more favorable/less favorable.”
Following are the mathematical calculations that can be applied with ordinal scales.
Mode
Ordinal scales cannot be used to determine the absolute difference between rankings. Exhibit
7.3 provides several examples of ordinal scales (PPT slide 7-10).
C. Interval Scales (PPT slides 7-8 to 7-11)
An interval scale is a scale that demonstrates absolute differences between each scale point.
The intervals between the scale numbers tell the researchers how far apart the measured objects
are on a particular attribute.
D. Ratio Scales (PPT slides 7-8 to 7-12)
A ratio scale is a scale that allows the researcher not only to identify the absolute differences
between each scale point but also to make comparisons between the responses. Ratio scales are
V. Evaluating Measurement Scales (PPT slides 7-13 and 7-14)
All measurement scales should be evaluated for reliability and validity.
7
Chapter 07 Measurement and Scaling
A. Scale Reliability (PPT slide 7-13)
Scale reliability refers to the extent to which a scale can reproduce the same or similar
measurement results in repeated trials. Thus, reliability is a measure of consistency in
measurement. Random error produces inconsistency in scale measurements that leads to lower
scale reliability. But researchers can improve reliability by carefully designing scaled questions.
Following are the two techniques that help researchers assess the reliability of scales.
Test-retest
Equivalent form
Some of the potential problems associated with the test-retest approach are listed below.
Respondents who completed the scale the first time might be absent for the second
administration of the scale.
Some researchers believe the problems associated with test-retest reliability technique can be
avoided by using the equivalent form technique. In this technique, researchers create two
similar yet different (e.g., equivalent) scale measurements for the given construct (e.g., teaching
effectiveness) and administer both forms to either the same sample of respondents or to two
samples of respondents from the same defined target population.
There are two potential drawbacks with the equivalent form reliability technique that are given
below.
Even if equivalent versions of the scale can be developed, it might not be worth the time,
effort, and expense of determining that two similar yet different scales can be used to
measure the same construct.
8
Chapter 07 Measurement and Scaling
teaching effectiveness.
The previous approaches to examining reliability are often difficult to complete in a timely and
accurate manner. As a result, marketing researchers most often use internal consistency
reliability. Internal consistency is the degree to which the individual questions of a construct are
correlated. Following are two popular techniques that are used to assess internal consistency.
Split-half test: The scale questions are divided into two halves (odd versus even, or
randomly) and the resulting halves’ scores are correlated against one another. High
correlations between the halves indicate good (or acceptable) internal consistency.
B. Validity (PPT slide 7-14)
Since reliable scales are not necessarily valid, researchers also need to be concerned about
validity. Scale validity assesses whether a scale measures what it is supposed to measure. Thus,
validity is a measure of accuracy in measurement. A construct with perfect validity contains no
measurement error. An easy measure of validity would be to compare observed measurements
with the true measurement. The problem is that we very seldom know the true measure.
Validation, in general, involves determining the suitability of the questions (statements) chosen
to represent the construct. Two approaches to assess scale validity are listed below.
Face validity: It is based on the researcher’s intuitive evaluation of whether the
statements look like they measure what they are supposed to measure. Establishing the
face validity of a scale involves a systematic but subjective assessment of a scale’s ability
to measure what it is supposed to measure. Thus, researchers use their expert judgment to
determine face validity.
Convergent validity: It is evaluated with multi-item scales and represents a situation in
9
Chapter 07 Measurement and Scaling
which the multiple items measuring the same construct share a high proportion of
variance, typically more than 50 percent.
Discriminant validity: It is the extent to which a single construct differs from other
constructs and represents a unique construct.
Following are two approaches that are typically used to obtain data to assess validity.
If sufficient resources are available, a pilot study is conducted with 100 to 200
respondents believed to be representative of the defined target population.
VI. Developing Scale Measurements (PPT slides 7-15 and 7-16)
Designing measurement scales requires the following:
understanding the research problem
A. Criteria for Scale Development (PPT slide 7-15)
Questions must be phrased carefully to produce accurate data. To do so, the researcher must
develop the following appropriate scale descriptors to be used as the scale points.
Understanding of the Questions: The researcher must consider the intellectual capacity and
language ability of individuals who will be asked to respond to the scales. Researchers should
Discriminatory Power of Scale Descriptors: The discriminatory power of scale descriptors
is the scale’s ability to differentiate between the scale responses. Researchers must decide how
Balanced versus Unbalanced Scales: Researchers must consider whether to use a balanced or
unbalanced scale. A balanced scale has an equal number of positive (favorable) and negative
(unfavorable) response alternatives.
Forced or Nonforced Choice Scales: A scale that does not have a neutral descriptor to divide
the positive and negative answers is referred to as a forced-choice scale. It is forced because the
respondent can only select either a positive or a negative answer, and not a neutral one. In
contrast, a scale that includes a center neutral response is referred to as a nonforced or
free-choice scale. Exhibit 7.6 presents several different examples of both “even-point,
forced-choice” and “oddpoint, nonforced” scales.
Negatively Worded Statements: Scale development guidelines traditionally suggested that
negatively worded statements should be included to verify that respondents are reading the
questions. In more than 40 years of developing scaled questions, the authors have found that
negatively worded statements almost always create problems for respondents in data collection.
As a result, inclusion of negatively worded statements should be minimized and even then
approached with caution.
Desired Measures of Central Tendency and Dispersion (PPT slide 7-16): The type of
11
Chapter 07 Measurement and Scaling
Mean: It is the arithmetic average of all the raw data responses.
Measures of dispersion describe how the data are dispersed around a central value. These
statistics enable the researcher to report the variability of responses on a particular scale.
Following are the measures of dispersion.
Frequency distribution: It is a summary of how many times each possible response to a
scale question/setup was recorded by the total group of respondents. This distribution can
be easily converted into percentages or histograms.
Given the important role these statistics play in data analysis, an understanding of how different
levels of scales influence the use of a particular statistic is critical in scale design. Exhibit 7.7
displays these relationships (PPT slide 7-16).
Nominal scales can only be analyzed using frequency distributions and the mode.
B. Adapting Established Scales
There are literally hundreds of previously published scales in marketing. Following are the
most relevant sources of these scales.
William Bearden, Richard Netemeyer, and Kelly Haws, Handbook of Marketing Scales,
3rd ed., Sage Publications, 2011
VII. Scales to Measure Attitudes and Behaviors (PPT slide 7-17 to 7-18)
12
Chapter 07 Measurement and Scaling
Scales are the “rulers” that measure customer attitudes, behaviors, and intentions. Well-designed
scales result in better measurement of marketplace phenomena, and thus provide more accurate
information to marketing decision makers. Several types of scales have proven useful in many
different situations. This section discusses three scale formats.
Likert scales
Exhibit 7.8 shows the general steps in the construct development/scale measurement process (PPT
slide 7-19).
A. Likert Scale (PPT slide 7-17)
A Likert scale is an ordinal scale format that asks respondents to indicate the extent to which
they agree or disagree with a series of mental belief or behavioral belief statements about a
given object. Named after its original developer, Rensis Likert, this scale initially had five scale
descriptors as given below.
Strongly agree
Agree
The Likert scale is often expanded beyond the original 5-point format to a 7-point scale, and
most researchers treat the scale format as an interval scale. Likert scales are best for research
designs that use self- administered surveys, personal interviews, or online surveys. Exhibit 7.9
provides an example of a 6-point Likert scale in a self-administered survey.
B. Semantic Differential Scale (PPT slide 7-17)
A semantic differential scale is a unique bipolar ordinal scale format that captures a person’s
attitudes or feelings about a given object (PPT slide 7-17). Only the endpoints of the scale are
labeled. In most cases, semantic differential scales use either 5 or 7 scale points.
Chapter 07 Measurement and Scaling
A problem encountered in designing semantic differential scales is the inappropriate narrative
expressions of the scale descriptors. In a well-designed semantic differential scale, the
individual scales should be truly bipolar. Sometimes researchers use a negative pole descriptor
that is not truly an opposite of the positive descriptor. This creates a scale that is difficult for the
respondent to interpret correctly.
C. Behavioral Intention Scale (PPT slide 7-18)
A behavioral intention scale is a special type of rating scale designed to capture the likelihood
that people will demonstrate some type of predictable behavior intent toward purchasing an
object or service in a future time frame.
Behavioral intentions are often a key variable of interest in marketing research studies. To make
scale points more specific, researchers can use descriptors that indicate the percentage chance
they will buy a product, or engage in a behavior of interest. The following set of scale points
could be used.
Definitely will (90 to 100 percent chance)
Probably will (50 to 89 percent chance)
Exhibit 7.12 shows what a shopping intention scale might look like.
No matter what kind of scale is used to capture people’s attitudes and behaviors, there often is
no one best or guaranteed approach. While there are established scale measures for obtaining
the components that make up respondents’ attitudes and behavioral intentions, the data
provided from these scale measurements should not be interpreted as being completely
predictive of behavior. Unfortunately, knowledge of an individual’s attitudes may not predict