Distinguish labels, ranks, equal intervals, and true zeros, then choose summaries and analyses that preserve what clinical measurements actually mean.
A heart failure registry records functional class as I, II, III, or IV. Class IV indicates greater limitation than class II. Does it mean twice the limitation? No. The ordering belongs to the clinical categories; the arithmetic belongs to the coding scheme. Before calculating anything, ask what a difference of one unit actually represents.
Four scales, four different permissions
Nominal data identify categories. ABO blood group, pathogen species, genotype category, and assigned treatment arm distinguish one group from another without ranking them. Coding O as 1 and AB as 4 does not make AB greater than O. Counts, proportions, and the most frequent category are meaningful; the average blood-group code is not. A different coding scheme could change that average without changing a single specimen.
Ordinal data add order. None, mild, moderate, and severe describe increasing symptom burden. NYHA functional classes and ordered tumor stages also contain rank information. However, the change from one category to the next need not represent an equal clinical change. Rank permits statements about higher and lower values, not an automatic claim about how much higher. An ordinal scale can be useful precisely because it captures meaningful clinical distinctions without pretending to measure a physical distance.
Interval data add equal differences. A change of 5 degrees Celsius has the same temperature unit throughout that scale. Its zero is a reference point, so a ratio of Celsius readings does not express a ratio of thermodynamic temperatures. Ratio data add a meaningful zero. Mass, elapsed time, and concentration have meaningful units and zeros for the quantities being measured. Two kilograms is twice one kilogram, although twice the mass does not imply twice the clinical risk. These distinctions concern mathematical interpretation, not a ranking of clinical importance. [1]
Read each row as a permission granted by the recorded variable
Nominal
Same or different category Example. Blood group A versus O
Ordinal
Same, different, higher, or lower Example. Mild versus severe symptoms
Interval
Order plus meaningful differences Example. 36 to 38 degrees Celsius
Ratio
Order, differences, and meaningful ratios Example. 2 versus 4 hours elapsed
The layout compares mathematical operations. It does not depict equal distances between symptom categories.
A binary variable coded 0 and 1 is still categorical. There is an important exception to the warning about averages of codes. If 1 means the event occurred and 0 means it did not, the arithmetic mean is exactly the event proportion. That mathematical identity does not turn diagnosis into a continuous biological quantity.
Classify the recording, not the disease
The same patient can supply variables on several scales. Disease present or absent is binary nominal. Ordered severity is ordinal. Elapsed days since diagnosis is ratio. Calendar year has meaningful differences but an arbitrary origin. The disease name alone cannot select a scale, and an electronic form's numeric field cannot establish equal spacing.
Age measured in years retains distance. Grouping age as child, adult, and older adult retains order but loses exact differences. A 17-year-old and a 2-year-old may now share one category. Assigning those categories codes 1, 2, and 3 cannot recover the lost ages. Retain the original numeric data when possible, then derive clinically justified groups for a particular analysis.
Similarly, pain recorded as none, mild, moderate, or severe is ordinal. A position measured in millimeters along a visual analog line is a quantitative recording. Its physical distance can be measured precisely, but claiming that twice the marked distance means twice the subjective pain requires evidence about the instrument. The ruler measures the mark directly; it does not by itself validate the latent construct.
Units and zeros need separate checks. Serum sodium concentration in mmol/L is a ratio quantity even though zero is incompatible with ordinary physiology. A true zero need not be a healthy or observed value. Conversely, a questionnaire minimum of zero may only indicate the lowest reportable score. It does not necessarily represent the absence of every aspect of disability, anxiety, or satisfaction. Instrument interpretation must be supported in its intended population and setting. [1][2]
Preserve the distribution before compressing it
For a short ordinal outcome, show the number and proportion in each category. This retains the observed information and makes floor effects, ceiling effects, and bimodal responses visible. A median and quartile categories can add a compact summary, but they do not replace the distribution. When the middle observations lie in different categories, report the median convention or middle categories instead of inventing a clinical category halfway between them.
Two symptom distributions with the same median category
Clinic A, 10 patients
None 0 Mild 6 Moderate 4 Severe 0
Clinic B, 10 patients
None 4 Mild 2 Moderate 0 Severe 4
Both middle observations are mild in each clinic. Clinic B nevertheless includes four patients with severe symptoms. The median alone conceals that difference.
A mean of arbitrary ordinal codes can change if the categories are recoded with unequal gaps, even when the ordering and every response stay identical. Rank summaries retain that ordering. A rank test, however, is not automatically a test of medians. Differences in distributional shape can matter. To interpret a rank-sum result as a location or median shift, additional assumptions about the distributions are needed. [3]
Dichotomizing the outcome creates another compression. Combining none, mild, and moderate as one category discards the distinctions within them. A prespecified clinical threshold can answer a useful question, such as the proportion reaching an established response definition. A threshold chosen after inspecting which split gives the smallest p value creates a different problem, selective analysis. Report the rationale for the threshold and preserve category-level results. [4]
The design chooses the comparison
First identify the scientific quantity of interest, sometimes called the estimand. Are you comparing event proportions, an ordered disability distribution, a mean laboratory value, or time until an event? Then establish whether observations come from different people, the same people twice, repeated visits, or clustered sites. Scale alone cannot answer these questions.
For independent groups with a short ordinal outcome, a rank-sum comparison or an ordinal regression model may preserve ordering. Mann-Whitney uses ranks from the combined samples; it does not require invented equal category gaps. For three or more independent groups, Kruskal-Wallis provides an omnibus rank comparison. A significant omnibus result does not identify every differing pair. Planned comparisons and multiplicity remain relevant. [3][5]
Repeated ordinal assessments need a method that respects the pairing. If only the direction of change is defensible, a sign-based analysis can compare improvement with worsening among non-ties. [7] Wilcoxon signed-rank additionally uses the ranks of difference magnitudes and requires defensible differences and appropriate symmetry assumptions. Merely assigning numbers to symptom categories does not establish those conditions. For continuous paired data, a paired t test analyzes within-person differences, with assumptions applying to those differences. [6]
A proportional-odds model compares cumulative odds across the ordered cut points and assumes a common effect across those splits. Always state the category direction. If lower values are better, an odds ratio for being at or below a threshold has a different direction from one for exceeding it. Inspect whether the common-odds assumption is credible; if it fails, consider a more flexible model and display category probabilities. The model does not assume equal category spacing. [4]
Censoring is another design issue. A patient last seen alive at six months has at least six months of observed survival, not a known death at six months. The ratio scale of elapsed time does not make an ordinary mean of recorded follow-up times an adequate survival analysis. Account for censoring and its assumptions. [4]
One item and a composite are different measured objects
A single Likert item from strongly disagree to strongly agree is ordered categorical. A multi-item sum can take many values and may support an approximately continuous analysis. That is a modeling decision supported by the instrument's construction, validity, reliability, distribution, and intended use. It is neither forbidden merely because the items are ordinal nor guaranteed merely because many items were added. [2][4]
Reliability asks about consistency. Validity concerns whether the evidence supports the intended interpretation. A highly consistent score can still measure the wrong construct. Check whether validation applies to this population, language, mode of administration, and clinical setting. Translation, shortening a questionnaire, or changing its response categories can alter what the score means. Follow the scoring algorithm for missing items rather than silently replacing unanswered questions with zero. [2]
If linear and ordinal analyses lead to materially different conclusions, disclose that sensitivity and investigate why. A large sample can improve precision without repairing an invalid interpretation of the units. Report category distributions or appropriate score summaries, an effect estimate with uncertainty, missingness, and the assumptions that matter. A p value alone cannot show the clinical size of a change.
Read the variable definition, identify its valid operations, preserve its information, account for dependence, and then choose the analysis. Numbers are useful only when their interpretation survives the calculation.
Practice interpreting the recorded variable
Case 1
Show answer and explanations for case 1
A. The classes rank limitation without establishing equal gaps (Best answer)
The categories are ordered, but neither equal spacing nor meaningful ratios follows from their Roman numerals.
B. Class IV is twice class II (Why this does not fit)
Dividing category codes does not quantify functional limitation.
C. The classes are nominal because they use letters (Why this does not fit)
Roman numerals label a clinically ordered classification.
D. Changing the codes to 10 through 40 creates an interval scale (Why this does not fit)
Recoding preserves order but adds no evidence about distances.
Takeaway: Order is meaningful; a ratio of class codes is not.