⌘ KStart free
0%
Skip to lesson

Biostatistics

Measurement Scales and Ordinal Data

Distinguish labels, ranks, equal intervals, and true zeros, then choose summaries and analyses that preserve what clinical measurements actually mean.

A heart failure registry records functional class as I, II, III, or IV. Class IV indicates greater limitation than class II. Does it mean twice the limitation? No. The ordering belongs to the clinical categories; the arithmetic belongs to the coding scheme. Before calculating anything, ask what a difference of one unit actually represents.

Four scales, four different permissions

Nominal data identify categories. ABO blood group, pathogen species, genotype category, and assigned treatment arm distinguish one group from another without ranking them. Coding O as 1 and AB as 4 does not make AB greater than O. Counts, proportions, and the most frequent category are meaningful; the average blood-group code is not. A different coding scheme could change that average without changing a single specimen.

Ordinal data add order. None, mild, moderate, and severe describe increasing symptom burden. NYHA functional classes and ordered tumor stages also contain rank information. However, the change from one category to the next need not represent an equal clinical change. Rank permits statements about higher and lower values, not an automatic claim about how much higher. An ordinal scale can be useful precisely because it captures meaningful clinical distinctions without pretending to measure a physical distance.

Interval data add equal differences. A change of 5 degrees Celsius has the same temperature unit throughout that scale. Its zero is a reference point, so a ratio of Celsius readings does not express a ratio of thermodynamic temperatures. Ratio data add a meaningful zero. Mass, elapsed time, and concentration have meaningful units and zeros for the quantities being measured. Two kilograms is twice one kilogram, although twice the mass does not imply twice the clinical risk. These distinctions concern mathematical interpretation, not a ranking of clinical importance. [1]

Read each row as a permission granted by the recorded variable

Nominal

Same or different category
Example. Blood group A versus O

Ordinal

Same, different, higher, or lower
Example. Mild versus severe symptoms

Interval

Order plus meaningful differences
Example. 36 to 38 degrees Celsius

Ratio

Order, differences, and meaningful ratios
Example. 2 versus 4 hours elapsed

The layout compares mathematical operations. It does not depict equal distances between symptom categories.

A binary variable coded 0 and 1 is still categorical. There is an important exception to the warning about averages of codes. If 1 means the event occurred and 0 means it did not, the arithmetic mean is exactly the event proportion. That mathematical identity does not turn diagnosis into a continuous biological quantity.

Classify the recording, not the disease

The same patient can supply variables on several scales. Disease present or absent is binary nominal. Ordered severity is ordinal. Elapsed days since diagnosis is ratio. Calendar year has meaningful differences but an arbitrary origin. The disease name alone cannot select a scale, and an electronic form's numeric field cannot establish equal spacing.

Age measured in years retains distance. Grouping age as child, adult, and older adult retains order but loses exact differences. A 17-year-old and a 2-year-old may now share one category. Assigning those categories codes 1, 2, and 3 cannot recover the lost ages. Retain the original numeric data when possible, then derive clinically justified groups for a particular analysis.

Similarly, pain recorded as none, mild, moderate, or severe is ordinal. A position measured in millimeters along a visual analog line is a quantitative recording. Its physical distance can be measured precisely, but claiming that twice the marked distance means twice the subjective pain requires evidence about the instrument. The ruler measures the mark directly; it does not by itself validate the latent construct.

Units and zeros need separate checks. Serum sodium concentration in mmol/L is a ratio quantity even though zero is incompatible with ordinary physiology. A true zero need not be a healthy or observed value. Conversely, a questionnaire minimum of zero may only indicate the lowest reportable score. It does not necessarily represent the absence of every aspect of disability, anxiety, or satisfaction. Instrument interpretation must be supported in its intended population and setting. [1] [2]

Preserve the distribution before compressing it

For a short ordinal outcome, show the number and proportion in each category. This retains the observed information and makes floor effects, ceiling effects, and bimodal responses visible. A median and quartile categories can add a compact summary, but they do not replace the distribution. When the middle observations lie in different categories, report the median convention or middle categories instead of inventing a clinical category halfway between them.

Two symptom distributions with the same median category

Clinic A, 10 patients

None 0
Mild 6
Moderate 4
Severe 0

Clinic B, 10 patients

None 4
Mild 2
Moderate 0
Severe 4

Both middle observations are mild in each clinic. Clinic B nevertheless includes four patients with severe symptoms. The median alone conceals that difference.

A mean of arbitrary ordinal codes can change if the categories are recoded with unequal gaps, even when the ordering and every response stay identical. Rank summaries retain that ordering. A rank test, however, is not automatically a test of medians. Differences in distributional shape can matter. To interpret a rank-sum result as a location or median shift, additional assumptions about the distributions are needed. [3]

Dichotomizing the outcome creates another compression. Combining none, mild, and moderate as one category discards the distinctions within them. A prespecified clinical threshold can answer a useful question, such as the proportion reaching an established response definition. A threshold chosen after inspecting which split gives the smallest p value creates a different problem, selective analysis. Report the rationale for the threshold and preserve category-level results. [4]

The design chooses the comparison

First identify the scientific quantity of interest, sometimes called the estimand. Are you comparing event proportions, an ordered disability distribution, a mean laboratory value, or time until an event? Then establish whether observations come from different people, the same people twice, repeated visits, or clustered sites. Scale alone cannot answer these questions.

For independent groups with a short ordinal outcome, a rank-sum comparison or an ordinal regression model may preserve ordering. Mann-Whitney uses ranks from the combined samples; it does not require invented equal category gaps. For three or more independent groups, Kruskal-Wallis provides an omnibus rank comparison. A significant omnibus result does not identify every differing pair. Planned comparisons and multiplicity remain relevant. [3] [5]

Repeated ordinal assessments need a method that respects the pairing. If only the direction of change is defensible, a sign-based analysis can compare improvement with worsening among non-ties. [7] Wilcoxon signed-rank additionally uses the ranks of difference magnitudes and requires defensible differences and appropriate symmetry assumptions. Merely assigning numbers to symptom categories does not establish those conditions. For continuous paired data, a paired t test analyzes within-person differences, with assumptions applying to those differences. [6]

A proportional-odds model compares cumulative odds across the ordered cut points and assumes a common effect across those splits. Always state the category direction. If lower values are better, an odds ratio for being at or below a threshold has a different direction from one for exceeding it. Inspect whether the common-odds assumption is credible; if it fails, consider a more flexible model and display category probabilities. The model does not assume equal category spacing. [4]

Censoring is another design issue. A patient last seen alive at six months has at least six months of observed survival, not a known death at six months. The ratio scale of elapsed time does not make an ordinary mean of recorded follow-up times an adequate survival analysis. Account for censoring and its assumptions. [4]

One item and a composite are different measured objects

A single Likert item from strongly disagree to strongly agree is ordered categorical. A multi-item sum can take many values and may support an approximately continuous analysis. That is a modeling decision supported by the instrument's construction, validity, reliability, distribution, and intended use. It is neither forbidden merely because the items are ordinal nor guaranteed merely because many items were added. [2] [4]

Reliability asks about consistency. Validity concerns whether the evidence supports the intended interpretation. A highly consistent score can still measure the wrong construct. Check whether validation applies to this population, language, mode of administration, and clinical setting. Translation, shortening a questionnaire, or changing its response categories can alter what the score means. Follow the scoring algorithm for missing items rather than silently replacing unanswered questions with zero. [2]

If linear and ordinal analyses lead to materially different conclusions, disclose that sensitivity and investigate why. A large sample can improve precision without repairing an invalid interpretation of the units. Report category distributions or appropriate score summaries, an effect estimate with uncertainty, missingness, and the assumptions that matter. A p value alone cannot show the clinical size of a change.

Read the variable definition, identify its valid operations, preserve its information, account for dependence, and then choose the analysis. Numbers are useful only when their interpretation survives the calculation.

Practice interpreting the recorded variable

Case 1

A heart failure registry records NYHA class I through IV at a clinic visit. An investigator proposes that class IV represents twice the limitation of class II. Which interpretation is justified?

Show answer and explanations for case 1
  1. A. The classes rank limitation without establishing equal gaps (Best answer)

    The categories are ordered, but neither equal spacing nor meaningful ratios follows from their Roman numerals.

  2. B. Class IV is twice class II (Why this does not fit)

    Dividing category codes does not quantify functional limitation.

  3. C. The classes are nominal because they use letters (Why this does not fit)

    Roman numerals label a clinically ordered classification.

  4. D. Changing the codes to 10 through 40 creates an interval scale (Why this does not fit)

    Recoding preserves order but adds no evidence about distances.

Takeaway: Order is meaningful; a ratio of class codes is not.

Case sources: [1]

Case 2

A transfusion service codes O, A, B, and AB as 1, 2, 3, and 4. The mean code differs between two hospitals. What should the analyst report to describe their blood-group mix?

Show answer and explanations for case 2
  1. A. Mean code with a narrower confidence interval (Why this does not fit)

    More precision does not give blood-group codes quantitative meaning.

  2. B. Median code as the typical blood group (Why this does not fit)

    The chosen numeric ordering has no intrinsic blood-group rank.

  3. C. A ratio of mean codes (Why this does not fit)

    The ratio would change under a harmless relabeling of the same groups.

  4. D. Counts and proportions for each blood group (Best answer)

    These show the actual category distribution without imposing an arbitrary numerical order.

Takeaway: Use category distributions for nominal variables.

Case sources: [1]

Case 3

A laboratory compares specimen storage at 4 degrees Celsius and 8 degrees Celsius. A report calls the latter twice the thermodynamic temperature. Which correction is best?

Show answer and explanations for case 3
  1. A. The temperature ratio is valid because both Celsius readings are positive (Why this does not fit)

    Positive numbers do not make an arbitrary zero absolute.

  2. B. Celsius becomes ordinal when specimens are assigned to two storage groups (Why this does not fit)

    Recording two values does not change the underlying temperature scale.

  3. C. The difference is 4 degrees, but the Celsius ratio is not a thermodynamic ratio (Best answer)

    Celsius has equal increments and an arbitrary zero.

  4. D. Celsius readings provide temperature labels without a meaningful ordering (Why this does not fit)

    Eight is warmer than four on this ordered quantitative scale.

Takeaway: Equal increments do not guarantee meaningful ratios.

Case sources: [1]

Case 4

A rehabilitation clinic records pain as none, mild, moderate, or severe. Most reports are mild, but a smaller cluster is severe. Which display best preserves that pattern?

Show answer and explanations for case 4
  1. A. A coefficient of variation calculated from the assigned pain-category codes (Why this does not fit)

    That calculation needs interpretable quantitative ratios, which these labels do not supply.

  2. B. Category counts and proportions, with median and quartile categories if helpful (Best answer)

    The full distribution retains the severe subgroup that a median alone may conceal.

  3. C. The mean of codes 0 through 3, without reporting individual pain categories (Why this does not fit)

    This hides the category pattern and assumes the code gaps are informative.

  4. D. The percentage reporting any pain, without separating the severity categories (Why this does not fit)

    This combines mild, moderate, and severe, losing the distinction at issue.

Takeaway: Show the distribution before compressing an ordinal outcome.

Case sources: [1] [4]

Case 5

A fatigue trial uses a validated 24-item questionnaire. Its prespecified total score has many values and a roughly symmetric distribution. Which analysis statement is defensible?

Show answer and explanations for case 5
  1. A. A continuous model may be used with justified score interpretation and checked assumptions (Best answer)

    The composite can support an approximation even though the individual responses are ordinal.

  2. B. Each individual response acquires equal intervals when included in the total score (Why this does not fit)

    Properties of the total score do not automatically establish item-level spacing.

  3. C. Parametric analysis is prohibited for a composite formed from ordinal item responses (Why this does not fit)

    That absolute ban ignores the evidence supporting the composite and model.

  4. D. A minimum total score of zero establishes a ratio scale (Why this does not fit)

    A scoring minimum need not mean absence of all fatigue.

Takeaway: Evaluate the total score as its own measured object.

Case sources: [2] [4]

Case 6

A disability trial records five ordered categories. After seeing the data, an analyst chooses the split producing the smallest p value and reports only that binary result. What is the main concern?

Show answer and explanations for case 6
  1. A. An ordinal outcome cannot be analyzed in a randomized trial (Why this does not fit)

    Randomization and ordinal outcome analysis are compatible.

  2. B. A significant split establishes clinical validity of the cutoff (Why this does not fit)

    Statistical separation does not validate a clinical threshold.

  3. C. The split creates continuous data (Why this does not fit)

    Combining categories produces a binary outcome, not a continuous measurement.

  4. D. Information loss combined with selective choice of the analysis (Best answer)

    The split discards distinctions and was selected for its result rather than a prior clinical rationale.

Takeaway: Prespecify defensible cutoffs and retain the ordered results.

Case sources: [4]

Case 7

A clinic codes vaccination completed as 1 and not completed as 0 for 200 eligible patients. The mean of the indicator is 0.72. Which interpretation is correct?

Show answer and explanations for case 7
  1. A. Vaccination status is continuous because its mean is decimal (Why this does not fit)

    The individuals still have binary categorical outcomes.

  2. B. The average patient completed 72% of a vaccine dose (Why this does not fit)

    The indicator concerns completion status, not dose volume.

  3. C. Seventy-two percent completed vaccination (Best answer)

    A 0/1 indicator mean equals the proportion coded 1.

  4. D. The mean is meaningless for every nominal variable (Why this does not fit)

    The indicator encoding has a specific arithmetic interpretation as a proportion.

Takeaway: The mean of a defined binary indicator is a proportion.

Case sources: [1]

Case 8

A clinic retains ages only as child, adult, or older adult. It later needs the mean age in years. Which statement is most accurate?

Show answer and explanations for case 8
  1. A. The groups are nominal because boundaries were chosen by investigators (Why this does not fit)

    The categories still have a meaningful younger-to-older ordering.

  2. B. The exact mean cannot be recovered from those groups alone (Best answer)

    Many different ages produce the same grouping, so exact distances were lost.

  3. C. The mean of codes 1, 2, and 3 equals mean age (Why this does not fit)

    Group codes are not years.

  4. D. Recoding the groups as 10, 20, and 30 restores age (Why this does not fit)

    New labels cannot recover discarded individual ages.

Takeaway: Coarsening can preserve order while discarding distance.

Case sources: [1]

Case 9

A laboratory records sodium concentration in mmol/L. A trainee argues that it cannot be ratio data because no living participant has zero sodium. Which response is best?

Show answer and explanations for case 9
  1. A. Sodium has a meaningful zero even if physiologically unobserved (Best answer)

    Ratio classification concerns the measured quantity, not whether every possible value is compatible with life.

  2. B. Laboratory concentrations provide ordered ranks rather than quantitative measurements (Why this does not fit)

    Concentration has quantitative units, not just ordered labels.

  3. C. A measurement qualifies as ratio only when the sample includes an observed zero (Why this does not fit)

    A sample need not include zero for the scale to have a true origin.

  4. D. Twice the concentration necessarily means twice the clinical danger (Why this does not fit)

    Meaningful measurement ratios do not establish linear biological risk.

Takeaway: A meaningful zero need not occur in the study sample.

Case sources: [1]

Case 10

Two clinics each evaluate 10 patients. Clinic A records 0 none, 6 mild, 4 moderate, and 0 severe. Clinic B records 4 none, 2 mild, 0 moderate, and 4 severe. Both medians are mild. What follows?

Show answer and explanations for case 10
  1. A. The matching medians establish that both clinics have identical symptom distributions (Why this does not fit)

    The category counts explicitly differ.

  2. B. Every rank-based comparison must have p equal to 1 (Why this does not fit)

    Equal medians alone do not determine a rank-test result.

  3. C. Clinic B necessarily has a larger mean pain intensity (Why this does not fit)

    The category data alone do not establish quantitative intensity gaps.

  4. D. Equal medians can accompany different symptom distributions (Best answer)

    Clinic B includes four severe reports while Clinic A includes none.

Takeaway: An equal median does not establish equal distributions.

Case sources: [3] [4]

Case 11

A clinic records four ordered mobility categories before and after rehabilitation in the same patients. Equal gaps are not established, and the question is whether improvement is more common than worsening among patients who changed. Which analysis matches that question?

Show answer and explanations for case 11
  1. A. A paired t test on arbitrary category-code differences (Why this does not fit)

    The clinical meaning of those numerical differences is not established.

  2. B. An unpaired chi-square test treating visits as different patients (Why this does not fit)

    The observations are dependent because the same patients contributed both visits.

  3. C. A paired sign-based comparison of improvement and worsening (Best answer)

    It uses within-person direction and can exclude ties without inventing difference magnitudes.

  4. D. An independent-samples rank-sum test ignoring patient identity (Why this does not fit)

    This discards the within-person pairing.

Takeaway: Pairing and the intended comparison matter beyond scale labels.

Case sources: [1] [7]

Case 12

A trial grades disability from 1, best, to 4, worst. An ordinal model reports a common odds ratio for being at or below each cut point. Which assumption should be checked?

Show answer and explanations for case 12
  1. A. The odds ratio is also the risk ratio at every split (Why this does not fit)

    Odds and probabilities are distinct, especially for common outcomes.

  2. B. Compatible treatment odds ratios across cumulative splits (Best answer)

    That common-effect restriction defines the proportional-odds model.

  3. C. Adjacent disability categories represent equally large clinical differences (Why this does not fit)

    Ordinal logistic modeling does not require equal category spacing.

  4. D. Every category has the same number of patients (Why this does not fit)

    Equal frequencies are not the proportional-odds assumption.

Takeaway: Proportional odds concerns cumulative effects, not equal score gaps.

Case sources: [4]

Case 13

A survival study records a participant's last contact at six months, when she is alive. The analyst enters six months as her time of death. What is the error?

Show answer and explanations for case 13
  1. A. Censored follow-up was classified as an observed death time (Best answer)

    The record establishes survival through six months, not death at that time.

  2. B. Elapsed survival time lacks a meaningful zero at its defined origin (Why this does not fit)

    Time measured from a defined origin has meaningful elapsed units and zero.

  3. C. Time-to-event records are nominal categories rather than quantitative elapsed times (Why this does not fit)

    Elapsed time is quantitative even when some event times are censored.

  4. D. The error is the choice of months rather than days as the time unit (Why this does not fit)

    Changing units leaves the mistaken event classification unchanged.

Takeaway: A quantitative time scale does not eliminate censoring.

Case sources: [4]

Case 14

A trial translates a symptom questionnaire into another language and changes several response labels. The original version had high reliability. What is needed before assuming the same score interpretation?

Show answer and explanations for case 14
  1. A. A larger sample of patients using the adapted version, with no further assessment (Why this does not fit)

    Precision alone does not establish equivalent interpretation.

  2. B. No new assessment, since reliability of the original transfers to the adapted version (Why this does not fit)

    Reliability and validity evidence are context-dependent.

  3. C. Replace unanswered items with zero to maintain comparability (Why this does not fit)

    Zero may encode no symptom, so this would invent responses rather than follow a supported scoring rule.

  4. D. Evidence of suitability for the intended population and use (Best answer)

    Language and response changes can affect understanding and measurement properties.

Takeaway: Adaptation requires evidence, not inherited reassurance.

Case sources: [2]

Case 15

A prespecified linear analysis of a short satisfaction scale favors a new clinic process, while an ordinal sensitivity analysis gives materially different estimates. What should the report do?

Show answer and explanations for case 15
  1. A. Declare both analyses wrong solely because they differ (Why this does not fit)

    Different assumptions can produce different answers without a coding error.

  2. B. Convert category codes to larger numbers to resolve the disagreement (Why this does not fit)

    Rescaling labels cannot supply missing measurement justification.

  3. C. Present both and investigate the assumptions driving the discrepancy (Best answer)

    Model sensitivity is relevant evidence about how strongly the conclusion is supported.

  4. D. Publish only the smaller p value (Why this does not fit)

    That selects the result after analysis and conceals uncertainty.

Takeaway: Report sensitivity to modeling assumptions.

Case sources: [2] [4]

Case 16

A trial prespecifies an established threshold defining a clinically important response on an ordinal scale. It reports both the response proportion and the full category distribution. Which appraisal is best?

Show answer and explanations for case 16
  1. A. Suppress the category distribution because reporting it alongside the endpoint adds multiple testing (Why this does not fit)

    Descriptive reporting does not require hiding the outcome distribution.

  2. B. The binary endpoint can answer a useful question while category reporting preserves additional information (Best answer)

    A justified threshold and transparent full reporting serve complementary purposes.

  3. C. Dichotomizing the outcome makes the trial uninterpretable despite a prespecified clinical threshold (Why this does not fit)

    A prespecified, defensible clinical threshold can be informative.

  4. D. A prespecified binary endpoint establishes equal distances between categories of the original scale (Why this does not fit)

    Thresholding does not establish category spacing.

Takeaway: A justified threshold is useful when its limits remain visible.

Case sources: [4]

Search Bone Wizardry

Quick links