⌘ KStart free
0%
Skip to lesson

Biostatistics

Screening Tests, Predictive Values, and Study Results

Build screening metrics from patient counts, update disease odds, assess screening benefit, and interpret risk reductions, study designs, and uncertainty.

A test can detect 95% of people with a disease while most positive results in a screening clinic are false positives. Those statements answer different questions. Start by naming the group in the denominator, then ask how much the result changes this patient's probability and whether acting on it improves outcomes.

First decide whether you know disease status from a reference standard, the test result, or both. A positive screen is not itself a diagnosis.

Four cells tell you which question was answered

Imagine 1,000 people tested with a defined reference standard. One hundred have the target condition and 900 do not. The new test detects 90 of the 100 affected people and labels 80 unaffected people positive. These are illustrative counts, not performance claims about a named clinical assay.

One dark tile among 100 represents 1,000 affected people. Ninety-nine light tiles represent 99,000 unaffected people. Arrows select 900 true positives and 4,950 false positives. The final strip shows their exact shares among positive results.Open whole image
Count from the population into the positive-result group. Teaching assumptions are prevalence 1%, sensitivity 90%, specificity 95%. Tile colors identify reference status; strip colors identify true versus false positive results.Image: Bone Wizardry original mathematical illustration, assistant-authored. Original artwork; no third-party license asserted. Original source.
Whole image
One dark tile among 100 represents 1,000 affected people. Ninety-nine light tiles represent 99,000 unaffected people. Arrows select 900 true positives and 4,950 false positives. The final strip shows their exact shares among positive results.

Count from the population into the positive-result group. Teaching assumptions are prevalence 1%, sensitivity 90%, specificity 95%. Tile colors identify reference status; strip colors identify true versus false positive results.

Image: Bone Wizardry original mathematical illustration, assistant-authored. Original artwork; no third-party license asserted. Original source.

Open the image directly

Reference status determines the columns; the new test determines the rows
New testCondition presentCondition absent
Positive90 true positives80 false positives
Negative10 false negatives820 true negatives
Read down a condition column for sensitivity or specificity. Read across a result row for a predictive value.

Sensitivity is TP/(TP + FN). Here it is 90/100, or 90%. Its complement is the false-negative rate among people with the condition. Specificity is TN/(TN + FP). Here it is 820/900, about 91.1%. Its complement is the false-positive rate among people without the condition. A 5% false-negative rate is not a 5% probability of disease after a negative result. Those denominators differ. [1]

Positive predictive value is TP/(TP + FP). Here 90/170, or 52.9%, of positive results are true positives. Negative predictive value is TN/(TN + FN). Here 820/830, or 98.8%, of negative results are true negatives. Prevalence is 100/1,000, or 10%. Accuracy is (TP + TN)/total, here 91%. Accuracy alone can mislead when one condition status dominates the sample. A test that always says negative can look highly accurate in a population where disease is rare.

These calculations assume reference status is established appropriately. Comparing two imperfect tests establishes agreement, not automatically sensitivity and specificity. Verifying only positive screens can distort accuracy estimates because disease among negative screens is not assessed comparably. When complete verification is impractical, use a justified probability sample with known sampling probabilities, the same appropriate reference standard, and estimation adjusted for the sampling design. Prespecification alone does not make a sample representative. Include relevant patient subgroups, account for indeterminate results, and report uncertainty around performance estimates. [1]

Try it here · Checkpoint 1 of 3

Make your prediction before reading the choices. A first attempt is just a starting point.

Case 2

Among 500 negative results from an imperfect non-reference comparator, a new test is negative for 460 and positive for 40. Among 100 comparator-positive results, the new test is positive for 90 and negative for 10. No participant has a true disease reference result. What does 460/500 mean, and what follow-up can estimate the new test's accuracy?

Show answer and explanations for case 2
  1. A. 92% specificity; reference-verify probability samples from all cells and apply sampling weights (Why this does not fit)

    Read the complete explanation

    The verification design is useful, but comparator-negative people are not known to be truly disease-free. Thus 460/500 is not specificity.

  2. B. 92% negative percent agreement; reference-verify only discordant cells and estimate accuracy from those results (Why this does not fit)

    Read the complete explanation

    The agreement label is right, but disease status among 550 concordant results remains unknown. Discordance-only adjudication cannot identify accuracy.

  3. C. 92% specificity; reference-verify only discordant cells and estimate accuracy from those results (Why this does not fit)

    Read the complete explanation

    Neither the imperfect comparator identifies true negatives nor discordance-only verification covers the agreement cells.

  4. D. 92% negative percent agreement; reference-verify probability samples from all cells and apply sampling weights (Best answer)

    Read the complete explanation

    The denominator is comparator negatives, so 460/500 is negative percent agreement. Reference verification across all cells with known selection probabilities permits design-adjusted accuracy estimates.

Takeaway: Comparator agreement is not accuracy; reference verification needs coverage of concordant and discordant cells.

Case sources: [1]

The same result means different things in different populations

Holding sensitivity and specificity fixed, increasing prevalence raises PPV and lowers NPV. The fixed-performance assumption is a calculation condition, not a promise that test performance never changes between settings. Disease severity, competing diagnoses, specimen collection, timing, and the reference standard can change observed sensitivity and specificity. A validation sample containing only severe disease and very healthy controls can overstate performance in ordinary patients. This is spectrum bias. [1]

Before testing, disease to no-disease odds are 1 to 99. A positive likelihood ratio of 18 changes odds to 18 to 99, equivalent to 900 to 4,950. A partitioned bar adds both groups to make the probability denominator, 18 plus 99.Open whole image
The positive likelihood ratio multiplies odds, not probability. Strip widths show disease and no-disease proportions. Count units change from 1/99 to 18/99; the matching cohort ratio is 900/4,950. These are illustrative calculations, not real-assay performance.Image: Bone Wizardry original mathematical illustration, assistant-authored. Original artwork; no third-party license asserted. Original source.
Whole image
Before testing, disease to no-disease odds are 1 to 99. A positive likelihood ratio of 18 changes odds to 18 to 99, equivalent to 900 to 4,950. A partitioned bar adds both groups to make the probability denominator, 18 plus 99.

The positive likelihood ratio multiplies odds, not probability. Strip widths show disease and no-disease proportions. Count units change from 1/99 to 18/99; the matching cohort ratio is 900/4,950. These are illustrative calculations, not real-assay performance.

Image: Bone Wizardry original mathematical illustration, assistant-authored. Original artwork; no third-party license asserted. Original source.

Open the image directly

Natural frequencies for a 95% sensitive, 95% specific test in 10,000 people

Prevalence 1%

100 have disease
95 true positives
495 false positives
PPV 95/590 = 16.1%

Prevalence 10%

1,000 have disease
950 true positives
450 false positives
PPV 950/1,400 = 67.9%

A larger affected group supplies more true positives. The large unaffected group at low prevalence supplies many false positives even with good specificity.

Worked cohort, without a calculator: In 100,000 people at 1% prevalence, an illustrative test with 90% sensitivity and 95% specificity yields 900 true positives, 4,950 false positives, 100 false negatives, and 94,050 true negatives. Among 5,850 positive results, 900 represent disease. At 20% prevalence with the same operating point, the four counts are 18,000, 4,000, 2,000, and 76,000, so 18,000 of 22,000 positives represent disease. At 1% prevalence, a different illustrative lower cutoff with 98% sensitivity and 80% specificity yields 980, 19,800, 20, and 79,200. These are teaching assumptions, not measurements of a real assay. [1]

For an individual, use the relevant pretest probability, informed by the population and clinical findings. Likelihood ratios provide an odds update. LR+ = sensitivity/(1 − specificity). LR− = (1 − sensitivity)/specificity. Convert probability p to odds p/(1 − p), multiply by the LR for the observed result, then convert odds o back to probability o/(1 + o).

At 10% pretest probability, odds are 1/9. A positive result with LR+ 12 gives odds 12/9 and probability 57.1%. At 80% pretest probability, odds are 4. A negative result with LR− 0.05 gives odds 0.20 and probability 16.7%. Even a strong negative likelihood ratio leaves substantial residual probability when the starting probability is high. Likelihood ratios multiply odds, not probabilities. [1] [24]

SnNOut and SpPIn are reminders about favorable test characteristics, not stand-alone clinical verdicts. Whether a result crosses a testing or treatment threshold depends on the full performance, pretest probability, and consequences of error. For example, current CDC guidance calls for backup throat culture after a negative rapid antigen detection test in symptomatic children aged three years or older. A reassuring mnemonic does not replace that specific testing pathway. [11]

Try it here · Checkpoint 2 of 3

Make your prediction before reading the choices. A first attempt is just a starting point.

Case 8

A 9-year-old has fever, exudates, tender anterior nodes, no cough or viral features, and a negative rapid antigen test for group A strep. A follow-up system can contact the family. What plan addresses both test limitations and age-specific risk?

Show answer and explanations for case 8
  1. A. Arrange follow-up only if symptoms persist, without culture now (Why this does not fit)

    Read the complete explanation

    Symptom follow-up is useful, but it does not replace the recommended backup culture after negative RADT in this age group.

  2. B. Repeat RADT tomorrow and treat if that repeat is positive (Why this does not fit)

    Read the complete explanation

    A repeat antigen test is not the CDC-specified backup throat culture for this symptomatic child.

  3. C. Start antibiotics now and culture only if symptoms persist (Why this does not fit)

    Read the complete explanation

    Compatible symptoms are insufficient to confirm GAS; backup culture should be obtained rather than empirical treatment after the negative RADT.

  4. D. Obtain backup throat culture and follow its result (Best answer)

    Read the complete explanation

    A symptomatic child older than three with negative RADT needs backup culture; the contact system enables follow-up if positive.

Takeaway: Negative RADT in symptomatic children aged three or older needs backup culture.

Case sources: [11]

A cutoff changes classification, not the underlying measurements

For a test where higher values indicate disease, lowering the positivity threshold labels more results positive. Sensitivity cannot decrease and specificity cannot increase in the same dataset. Some cutoff changes produce no change if no observations lie between them. Raising the threshold has the opposite direction. If lower values indicate disease, reverse the numerical direction of the cutoff rule.

Three pairs of bars compare illustrative operating points. Sensitivity decreases from 98% to 90% to 70% as the cutoff rises; false-positive rate decreases from 20% to 5% to 1%. Each bar uses the same zero to 100 percent scale.Open whole image
Dark bars use affected people as their denominator. Purple bars use unaffected people. All bars use a zero-to-100% scale. Raising the illustrative cutoff reduces detection and false positives; it does not change prevalence.Image: Bone Wizardry original mathematical illustration, assistant-authored. Original artwork; no third-party license asserted. Original source.
Whole image
Three pairs of bars compare illustrative operating points. Sensitivity decreases from 98% to 90% to 70% as the cutoff rises; false-positive rate decreases from 20% to 5% to 1%. Each bar uses the same zero to 100 percent scale.

Dark bars use affected people as their denominator. Purple bars use unaffected people. All bars use a zero-to-100% scale. Raising the illustrative cutoff reduces detection and false positives; it does not change prevalence.

Image: Bone Wizardry original mathematical illustration, assistant-authored. Original artwork; no third-party license asserted. Original source.

Open the image directly

The ROC curve plots sensitivity against the false-positive rate, 1 − specificity, across cutoffs. The upper-left region combines high sensitivity with few false positives. ROC area under the curve summarizes discrimination across thresholds. An AUC of 0.5 indicates no rank discrimination and 1 indicates perfect rank discrimination in the evaluated population; values below 0.5 are possible, for example when score direction is reversed. AUC is not PPV, accuracy at one cutoff, calibration, or proof of clinical benefit. [1]

Precision-recall displays answer a related question using PPV, also called precision, against sensitivity, also called recall. Because PPV changes with prevalence, interpret these plots in the evaluated population. Neither an appealing curve nor a statistically selected balance of sensitivity and specificity determines the best clinical threshold. Missed disease, false-positive procedures, treatment effects, and resource use must be considered.

Screening often places substantial weight on sensitivity, but false positives are not harmless. A diagnostic pathway may use subsequent testing to improve certainty, yet repeated tests are not automatically independent. Do not multiply likelihood ratios for correlated tests without a justified model. Check what outcome defines disease, what threshold defines a positive screen, and what subsequent action is supported.

Finding disease earlier is not the same as improving health

Lead-time bias changes the starting date. In a hypothetical example, symptom diagnosis occurs at age 66 and death at 70. Screening diagnoses the same disease at 63 without changing death at 70. Survival from diagnosis grows from four to seven years while lifespan is unchanged. This arithmetic example does not assert that a particular screening program lacks benefit.

Length-time bias changes which disease is detected. A slowly progressing tumor remains detectable before symptoms for longer, so periodic screening is more likely to find it. Aggressive tumors may become symptomatic between screens. The screen-detected group can therefore have more favorable biology independent of screening benefit. Overdiagnosis identifies a real condition that would never have caused symptoms or death during the person's lifetime. It is distinct from a false-positive result. [2]

A randomized screening comparison with appropriate mortality follow-up in the entire assigned population addresses these interpretive traps more directly than comparing survival only among diagnosed cases. Disease-specific mortality is informative, with careful cause-of-death assessment; all-cause mortality and harms provide complementary information. Randomization does not make screen-detected tumors representative of all tumors, and it does not erase missed follow-up, contamination, or other study limitations.

Evaluate the complete program. Benefits depend on whether earlier detection permits effective action; harms include false-positive workup and unnecessary treatment of overdiagnosed disease. A high detection count, small average tumor size, or improved survival from diagnosis is insufficient by itself to establish net benefit. [2]

Read the absolute difference and its uncertainty

For an adverse event, absolute risk reduction is control risk minus treatment risk. Relative risk is treatment risk divided by control risk, and relative risk reduction is 1 − RR. If five-year event risks are 8% and 4%, ARR is four percentage points, RR is 0.5, RRR is 50%, and NNT is 1/0.04 = 25 over five years. For a positive noninteger NNT, round upward to a whole patient for conventional reporting. The time horizon, outcome, comparator, and population belong with the number. [4] [23]

For harm, absolute risk increase is treatment risk minus comparison risk and NNH is its reciprocal. An NNH of 50 over one year corresponds to two additional events per 100 treated over that year, not a total adverse-event rate of 2% or 50%. NNT and NNH are average comparative summaries. They neither identify which individual benefits nor show that every other patient receives no benefit of any kind.

Comparing their numerical sizes alone cannot weigh a mild benefit against a severe harm. Use comparable follow-up periods and an explicit valuation or resource rule when a decision requires trading benefit against harm. [23]

A p value describes how incompatible the observed result is with a specified null model using the chosen test statistic. It is not the probability that the null is true, that the result is a mistake, or that treatment works. Statistical significance does not establish clinical importance or freedom from bias. A small blood-pressure difference may matter differently by risk, cost, harms, and duration; there is no universal minimum useful change implied by p. [5]

A confidence interval shows estimation uncertainty under the method's assumptions. For matched two-sided methods, a 95% interval excluding the null corresponds to rejection at alpha 0.05. The null is 1 for ratios and 0 for differences. A frequentist 95% confidence procedure covers the fixed true parameter in 95% of repeated applications under its assumptions. An interval including the null does not prove equivalence. Equivalence requires excluding effects beyond both prespecified margins; noninferiority excludes unacceptable inferiority on one specified side. [3] [10]

Type I error rejects a true null. Type II error fails to reject a false null. Power is 1 − beta for a specified alternative and design. More participants or less measurement noise generally improves power, holding other conditions fixed. A nonsignificant small study may be imprecise; its result alone cannot establish that a Type II error actually occurred.

Try it here · Checkpoint 3 of 3

Make your prediction before reading the choices. A first attempt is just a starting point.

Case 17

A randomized prevention trial reports five-year primary-event risks of 4% in controls and 3% with treatment. An advertisement calls this a 25% reduction. In a prespecified lower-risk subgroup, the treatment risk ratio is 0.90 (95% CI 0.55 to 1.47); the treatment-by-subgroup interaction test has p = 0.60. What does the advertisement quantify, and does the subgroup result establish a smaller relative benefit there?

Show answer and explanations for case 17
  1. A. The 25% is the overall relative risk reduction, not the one-percentage-point absolute reduction; the subgroup results do not establish a difference in relative effects (Best answer)

    Read the complete explanation

    Overall risk ratio is 3/4 = 0.75, so relative reduction is 25% and absolute reduction is 1 percentage point. The subgroup CI includes 1 and interaction p = 0.60 provides no evidence of different effects, not proof of equality.

  2. B. The 25% is the overall relative risk reduction; the nonsignificant subgroup result establishes a smaller relative benefit there (Why this does not fit)

    Read the complete explanation

    The overall relative reduction is 25%. A nonsignificant result within one subgroup cannot establish that its effect is smaller than another subgroup effect; the interaction p = 0.60 does not show a difference.

  3. C. The 25% is the overall absolute risk reduction; the subgroup results do not establish a difference in relative effects (Why this does not fit)

    Read the complete explanation

    The absolute reduction is 4% - 3% = 1 percentage point, not 25%; 25% divides that difference by the 4% control risk.

  4. D. The 25% is the overall relative risk reduction; the nonsignificant interaction establishes equal relative effects (Why this does not fit)

    Read the complete explanation

    The overall relative reduction is 25%. Interaction p = 0.60 does not establish equality of subgroup effects.

Takeaway: An overall relative reduction differs from the absolute reduction; within-subgroup nonsignificance does not demonstrate effect modification.

Case sources: [4] [7]

Try these without looking

With 80% pretest probability and a negative LR of 0.05, is residual probability 4%?

No. Convert to odds first: 4 times 0.05 is 0.20. Convert back: 0.20 divided by 1.20 is about 16.7%.

Revisit this explanation [1]

Five-year risks fall from 4% to 3%. Is the absolute decrease 25 percentage points?

No. The absolute decrease is one percentage point. Dividing it by the original 4% risk gives a 25% relative reduction.

Revisit this explanation [4]

For those 4% and 3% five-year risks, what is the number needed to treat?

The absolute risk reduction is 0.01. Its reciprocal is 100 over five years for this outcome and comparison.

Revisit this explanation [4]

Match the analysis to how the data were obtained

An RCT assigns interventions randomly. A cohort follows an exposure-defined population through outcome time, prospectively or using records, and can estimate risks. A conventional case-control study samples by outcome and compares prior exposures, usually using an odds ratio without directly estimating population incidence. Cross-sectional studies measure a population at a point or period and often estimate prevalence. Case series lack a comparison group but can identify signals worth investigating. Matching controls is optional, not definitional. [18] [3]

For independent categorical counts, chi-square uses expected cell counts; sparse 2 by 2 tables often favor Fisher exact. Expected counts are row total times column total divided by the grand total; they are not the observed counts. [22] For two independent continuous means, use an appropriate t procedure, often Welch when variances need not be equal.

Paired t tests use within-person differences. One-way ANOVA compares independent group means under its assumptions; significant omnibus evidence does not identify which pairs differ. Specific contrasts need an appropriate prespecified analysis with multiplicity control when several claims form a family. A significant omnibus test is not a universal prerequisite for a prespecified contrast. Bonferroni intervals allocate alpha across a finite planned family to control simultaneous coverage. [19] [20] [12] [13] [14]

Mann-Whitney compares ranks in two independent samples. When the distributions differ only by a location shift, that difference can be described as a difference in medians; different shapes make a median-only reading inadequate. [15] Kruskal-Wallis compares rank sums across independent groups under a common-distribution null. Rejection does not by itself establish a median difference or identify the differing pairs. [17] [21] Pearson correlation describes linear association of measured values; Spearman describes association of their ranks and can capture monotonic relationships. Neither establishes causation or measurement agreement. Paired, repeated, clustered, or censored data need methods respecting those structures. [15] [16]

Bias may enter through recruitment, recall, outcome assessment, treatment selection, or missing outcomes. Hospital admission can distort associations when it depends on both exposure and disease, the Berkson mechanism. Awareness of observation can alter behavior, the Hawthorne effect, without necessarily inflating a between-group effect. Intention-to-treat preserves original assignment but cannot recover unobserved outcomes without assumptions. There is no universally safe dropout percentage. [3] [6]

In evidence synthesis, forest plots display estimates and intervals; a pooled result inherits the limitations of its studies. I-squared is not a rule that automatically prohibits pooling above 50%. Funnel asymmetry can arise from missing evidence, heterogeneity, methodological differences, or chance. Prespecified interim monitoring can support ethical early stopping, but repeated examination and effect estimation require appropriate statistical control. [7] [8] [9]

The bias and study-design lesson develops recruitment, confounding, trial follow-up and evidence synthesis in more detail.

Name the denominator, update the odds, identify the design, inspect absolute effects and uncertainty, then decide what clinical action the evidence supports.

When assignment changes treatment uptake, the effect of assignment and the effect of receiving treatment are different questions. With random assignment, an exclusion restriction, no defiers and a nonzero uptake difference, divide the assignment outcome contrast by the uptake contrast. This Wald ratio estimates the effect among people whose uptake changes with assignment, not automatically the whole population. [27]

Practice interpreting tests and study results

Case 1

A screening validation verifies every one of 100 positive screens: 90 have disease and 10 do not. It randomly verifies half of 100 negative screens: 5 have disease and 45 do not. The same valid reference standard is used, and verification is representative within each result stratum. All diseased validation patients have advanced disease; intended screening mostly encounters early disease. What sensitivity estimate follows after accounting for verification sampling, and does it establish early-disease sensitivity?

Show answer and explanations for case 1
  1. A. 90% in this advanced-disease validation setting; obtain evidence in the intended early-disease spectrum (Best answer)

    Read the complete explanation

    The five verified false negatives represent ten after weighting, so 90/(90+10) = 90%. Advanced-only disease validation cannot establish early-disease sensitivity.

  2. B. 94.7% in this validation setting; obtain evidence in the intended early-disease spectrum (Why this does not fit)

    Read the complete explanation

    Five is only the sampled false-negative count. The random half-sample represents ten false negatives before estimating sensitivity.

  3. C. 90% here; apply it unchanged to early-disease screening (Why this does not fit)

    Read the complete explanation

    The weighting gives 90%, but the validation contains no diseased patients with early disease. Its sensitivity may not transport.

  4. D. 94.7% here; apply it unchanged to early-disease screening (Why this does not fit)

    Read the complete explanation

    This uses the unweighted five false negatives and also transports an advanced-only estimate to an unstudied spectrum.

Takeaway: Verification weighting corrects the sampled result stratum; disease spectrum limits transport.

Case sources: [1]

Case 3

An illustrative assay has 85% sensitivity and 90% specificity. In a screening population of 10,000 with 1% prevalence, every positive screen prompts an invasive follow-up. A service will adopt this pathway only if it needs no more than 12 follow-ups per true disease found. Using unrounded expected counts, which appraisal follows?

Show answer and explanations for case 3
  1. A. About 12.65 follow-ups per detected case; do not adopt (Best answer)

    Read the complete explanation

    The 85 true positives and 990 false positives require 1,075 follow-ups; 1,075/85 = 12.647 exceeds the maximum of 12.

  2. B. About 12.65 follow-ups per detected case; adopt (Why this does not fit)

    Read the complete explanation

    The ratio is correct, but 12.65 exceeds the service maximum of 12; rounding or treating the limit as 13 changes the supplied rule.

  3. C. About 10.75 follow-ups per detected case; adopt (Why this does not fit)

    Read the complete explanation

    1,075/100 uses all affected people, including 15 missed cases; detected cases number 85.

  4. D. About 11.65 follow-ups per detected case; adopt (Why this does not fit)

    Read the complete explanation

    The 990 false-positive follow-ups divided by 85 true positives gives 11.65, but omits follow-ups for the 85 true positives.

Takeaway: Rare-condition prevalence changes positive-result workload.

Case sources: [1]

Case 4

A case-control validation deliberately enrolls 100 diseased and 100 nondiseased people. It finds 80 true positives, 20 false negatives, 90 true negatives and 10 false positives. Every diseased validation patient has advanced disease. The intended clinic sees mostly mild disease at roughly 2% prevalence. Which reported NPV and additional evidence justify a clinic rule-out estimate?

Show answer and explanations for case 4
  1. A. 81.82% in the selected sample; use it in the clinic after checking that the same assay threshold is used (Why this does not fit)

    Read the complete explanation

    The threshold alone does not transport predictive value: deliberate 50% prevalence differs from the 2% clinic, and disease severity also differs.

  2. B. 90% in the selected sample; measure clinic prevalence and performance in its disease spectrum (Why this does not fit)

    Read the complete explanation

    The additional evidence is appropriate, but 90% is specificity, 90/(90+10), not NPV among negative tests.

  3. C. 81.82% in the selected sample; obtain accuracy in the intended severity spectrum and representative clinic prevalence before a clinic rule-out estimate (Best answer)

    Read the complete explanation

    The selected-sample negative results give 90/(90+20). Neither its predictive value nor advanced-disease sensitivity transports to mostly mild disease; clinic prevalence is also needed.

  4. D. 81.82% in the selected sample; combine its 80% sensitivity and 90% specificity with 2% prevalence without further accuracy evidence (Why this does not fit)

    Read the complete explanation

    Clinic prevalence alone is insufficient: advanced-only validation does not establish mild-disease sensitivity at this threshold.

Takeaway: Selected case-control predictive values and advanced-disease accuracy do not establish mild-disease clinic NPV.

Case sources: [1] [24]

Case 5

At one visit, two schools each screen 1,000 students with complete reference verification. School A has 90 true positives, 45 false positives, 10 false negatives and 855 true negatives. School B has 360 true positives, 30 false positives, 40 false negatives and 570 true negatives. No prior infection histories or incident follow-up are available. How do the reliability of positive results and the evidence about infection incidence compare?

Show answer and explanations for case 5
  1. A. Positive results are more reliable in B (92.31% versus 66.67%) because current prevalence is higher; B also has higher infection incidence (Why this does not fit)

    Read the complete explanation

    PPV and current prevalence are correctly inferred, but a one-visit table has no new-case onset rate or person-time.

  2. B. Both PPVs are 90% because sensitivity is the same; infection incidence cannot be compared (Why this does not fit)

    Read the complete explanation

    The incidence caution is sound, but 90% conditions on disease, not positive results; false-positive counts differ.

  3. C. Positive results are more reliable in B (92.31% versus 66.67%) with greater current prevalence despite matched accuracy; incidence is not identified (Best answer)

    Read the complete explanation

    PPVs are 360/390 and 90/135. Both sensitivities are 90% and specificities 95%; these one-visit data do not give incident onsets.

  4. D. Positive results are more reliable in B (92.31% versus 66.67%) because its sensitivity is higher; incidence is not identified (Why this does not fit)

    Read the complete explanation

    The PPV and incidence conclusion fit, but B sensitivity is 360/400 = 90%, equal to A 90/100. Current prevalence differs.

Takeaway: Current prevalence can change PPV without revealing incidence.

Case sources: [1] [26]

Case 6

Clinical assessment gives a patient 10% pretest probability. A positive illustrative test has LR+ 12, and a proposed action threshold is 50% post-test probability. Which choice follows?

Show answer and explanations for case 6
  1. A. Cross threshold at about 57% (Best answer)

    Read the complete explanation

    Pretest odds 1/9 become 12/9, yielding probability 12/(9+12) = 57.1%, above 50%.

  2. B. Remain below threshold at about 12% (Why this does not fit)

    Read the complete explanation

    Treating the LR+ value 12 as a post-test percentage discards the 10% pretest probability and the odds conversion.

  3. C. Cross threshold at about 92% (Why this does not fit)

    Read the complete explanation

    Starting from pretest odds 1 rather than 1/9 gives posterior odds 12 and probability 12/13 = 92.3%.

  4. D. Cross threshold at about 55% (Why this does not fit)

    Read the complete explanation

    Converting 10% x 12 = 1.2 as if it were posterior odds yields 1.2/2.2 = 54.5%; likelihood ratios multiply pretest odds 1/9, not pretest probability.

Takeaway: Likelihood ratios update odds, then the posterior guides the stated action.

Case sources: [1] [24]

Case 7

Clinical findings yield 80% pretest probability. An illustrative negative test has LR- 0.05; a plan would withhold further evaluation only below 10% residual risk. What decision follows?

Show answer and explanations for case 7
  1. A. Continue: residual risk is about 20% (Why this does not fit)

    Read the complete explanation

    Prior odds 4 become 0.20, but 0.20 is odds, not 20% probability; converting gives 16.7%.

  2. B. Stop: residual risk is about 4% (Why this does not fit)

    Read the complete explanation

    Multiplying 80% probability directly by LR- 0.05 produces 4%; the likelihood ratio multiplies odds, not probability.

  3. C. Stop: residual risk is about 16.7% (Why this does not fit)

    Read the complete explanation

    The probability is calculated correctly, but 16.7% is above rather than below the required 10% threshold.

  4. D. Continue: residual risk is about 16.7% (Best answer)

    Read the complete explanation

    Pretest odds .8/.2 = 4 become .2, corresponding to .2/1.2 = 16.7%, above the 10% stopping threshold.

Takeaway: Negative likelihood ratio does not by itself establish rule-out.

Case sources: [1] [24]

Case 9

A score labels values 12 or higher positive. In a verified sample, 4 affected and 6 unaffected patients score 8 to 11. The cutoff is lowered to 8. What cell changes follow?

Show answer and explanations for case 9
  1. A. Six additional true positives and four additional false positives (Why this does not fit)

    Read the complete explanation

    This swaps affected and unaffected counts in the reclassified band.

  2. B. Four additional true positives and six additional false positives (Best answer)

    Read the complete explanation

    The 8-to-11 group becomes positive: affected people move from FN to TP and unaffected from TN to FP.

  3. C. Four additional true positives and six additional true negatives (Why this does not fit)

    Read the complete explanation

    The six unaffected people move from true negative to false positive at the lower cutoff, not to additional true negatives.

  4. D. Four additional false negatives and six additional true negatives (Why this does not fit)

    Read the complete explanation

    These were the old cells for 8-to-11 scores; lowering the threshold moves patients out of them, not into them.

Takeaway: Lowering a cutoff reclassifies the intervening score band.

Case sources: [1]

Case 10

In a 1:1 diseased/nondiseased validation sample, piecewise-linear ROC curves have vertices (false-positive rate, sensitivity): A (0,0), (.02,.50), (.10,.90), (1,1); B (0,0), (.02,.65), (.10,.70), (1,1). Their trapezoidal AUCs are .916 and .8255. The intended clinic tolerates a 2% false-positive rate and has 2% disease prevalence. At that point, B is advertised with 97.01% PPV from validation. Assuming only for this projection that the operating-point sensitivity and false-positive rate transport, which score fits the constrained region and what clinic PPV follows?

Show answer and explanations for case 10
  1. A. A despite its lower 50% sensitivity at 2% false positives; about 40% clinic PPV (Why this does not fit)

    Read the complete explanation

    A has larger global AUC, but B detects 65% rather than 50% at the clinic false-positive constraint. The roughly 40% projection pertains to B, not A.

  2. B. A despite its lower local sensitivity; about 97% PPV from the 1:1 sample (Why this does not fit)

    Read the complete explanation

    A has lower sensitivity at 2% false positives, and the 1:1 predictive value cannot be carried to 2% prevalence.

  3. C. B for 65% versus 50% sensitivity at 2% false positives; about 97% clinic PPV (Why this does not fit)

    Read the complete explanation

    B is the right local choice, but 97.01% conditions on the validation sample with 50% disease prevalence.

  4. D. B for 65% versus 50% sensitivity at 2% false positives; about 39.88% projected clinic PPV (Best answer)

    Read the complete explanation

    B wins at the constrained operating point despite lower AUC. With assumed transport, (.65*.02)/(.65*.02+.02*.98) = 39.88%, not the 1:1 sample PPV.

Takeaway: Monotone piecewise-linear vertices make both same-cohort ROC summaries feasible. Local performance and projected PPV answer different questions; descriptive AUCs alone do not establish population superiority.

Case sources: [1] [24]

Case 11

A test appears 95% sensitive in a study restricted to severe cases and 95% specific in exceptionally healthy controls. In the intended clinic, mild cases and mimicking conditions are common. Which validation change best addresses the likely overestimate before adoption?

Show answer and explanations for case 11
  1. A. Recruit consecutive specialty referrals and verify everyone (Why this does not fit)

    Read the complete explanation

    Full verification helps, but specialty referrals need not represent the intended clinic with mild cases and mimics.

  2. B. Recruit intended-use patients and verify only positive screens (Why this does not fit)

    Read the complete explanation

    The intended spectrum is represented, but selective verification leaves false negatives among negative screens unmeasured.

  3. C. Recruit consecutive intended-use patients and verify disease status independently (Best answer)

    Read the complete explanation

    This includes mild disease and mimics, while independent reference assessment permits accuracy estimation across that spectrum.

  4. D. Recruit more severe cases and healthy controls, with independent verification (Why this does not fit)

    Read the complete explanation

    Independent verification preserves internal classification, but more extreme patients only tighten an estimate that may not transport.

Takeaway: Validation should represent intended clinical spectrum and ascertain reference status.

Case sources: [1]

Case 12

A new assay is positive in 100 of 1,000 patients. All positives undergo an appropriate reference procedure, but the 900 negatives are assumed unaffected. Investigators report sensitivity from only the verified patients. Which revision permits an adjusted estimate without verifying every negative?

Show answer and explanations for case 12
  1. A. Reference-test a random subset of negatives but calculate sensitivity only among all verified patients without adjustment (Why this does not fit)

    Read the complete explanation

    The negative sample can reveal false negatives, but the raw verified subset overrepresents positives because all positives versus only some negatives were verified.

  2. B. Reference-test a random subset of negatives with known sampling probabilities, then weight by those probabilities (Best answer)

    Read the complete explanation

    Random negative verification using the same reference standard reveals otherwise unobserved false negatives; design-weighted estimation accounts for unequal verification.

  3. C. Repeat the index assay in a random subset of negatives and use inverse-probability weighting (Why this does not fit)

    Read the complete explanation

    Random sampling and weights do not solve the absence of disease status if the same index assay is repeated instead of an appropriate reference standard.

  4. D. Reference-test the first 100 negatives and weight each as nine negatives (Why this does not fit)

    Read the complete explanation

    A fixed first 100 are not a known-probability random sample of the 900; weighting cannot remove order-dependent selection bias.

Takeaway: Partial verification can support adjusted accuracy estimates when sampling is known.

Case sources: [1]

Case 13

Without screening, a patient would be diagnosed at age 66 and die of the cancer at 70. In a hypothetical screening pathway, diagnosis moves to age 63 but cancer death remains at 70. A registry calls survival from diagnosis beyond five years a success. Which paired interpretation follows from this timeline?

Show answer and explanations for case 13
  1. A. Survival rises from 4 to 7 years but neither pathway crosses the five-year mark (Why this does not fit)

    Read the complete explanation

    The no-screen four-year interval does not cross five years; the screen-diagnosed seven-year interval does.

  2. B. Survival remains 4 years in both pathways because age at death is unchanged (Why this does not fit)

    Read the complete explanation

    The unchanged death age does not fix diagnosis-to-death time when diagnosis shifts three years earlier.

  3. C. Survival rises from 4 to 7 years and crosses the five-year mark, implying three additional life-years (Why this does not fit)

    Read the complete explanation

    The survival durations and threshold crossing are right, but the death age is 70 in both pathways, so the three added measured years are lead time.

  4. D. Survival rises from 4 to 7 years and crosses the five-year mark, while age at death is unchanged (Best answer)

    Read the complete explanation

    Diagnosis-to-death duration is 70 - 66 = 4 versus 70 - 63 = 7 years; death remains at 70, so apparent five-year survival improves without life extension.

Takeaway: Lead time can change diagnosis-anchored survival without changing death time.

Case sources: [2]

Case 14

Suppose slow and aggressive tumors arise equally often. Before symptoms, slow tumors remain screen-detectable for four years and aggressive tumors for three months. Screening occurs once a year; onset is uniformly timed relative to visits, and each tumor is detectable throughout its stated window. Ignoring other detection differences, what mix is expected among screen-detected tumors, and what does longer survival among that group establish?

Show answer and explanations for case 14
  1. A. About 80% slow; longer diagnosed-case survival establishes a screening treatment effect (Why this does not fit)

    Read the complete explanation

    The 4:1 detection ratio is right, but enrichment of slower disease can explain favorable case survival without an intervention effect.

  2. B. About 50% slow; longer diagnosed-case survival could reflect duration selection (Why this does not fit)

    Read the complete explanation

    Equal incidence is not equal representation after annual sampling: slow tumors are seen with probability 1, aggressive with .25.

  3. C. About 94% slow; longer diagnosed-case survival could reflect duration selection (Why this does not fit)

    Read the complete explanation

    Dividing raw windows, four years/.25 years = 16:1 (94.1%), ignores saturation: the slow detection probability is capped at 1 over annual visits.

  4. D. About 80% slow; longer diagnosed-case survival could reflect duration selection (Best answer)

    Read the complete explanation

    A four-year window guarantees at least one annual screen, whereas three months gives 3/12 = 25% chance. Equal onset gives slow:aggressive detection 1:.25 = 4:1, or 80% slow; selection can lengthen observed survival.

Takeaway: Annual screens preferentially sample prolonged detectable disease.

Case sources: [2]

Case 15

In comparable randomized groups, cumulative pathologically confirmed diagnoses per 10,000 are 180 with screening versus 100 without at year 5 and 200 versus 150 at year 15. Target-disease mortality at year 15 is similar, while treatment-related complications are more frequent in the screening group. Follow-up is adequate for many, but not necessarily all, delayed presentations. Which interpretation best uses the pattern?

Show answer and explanations for case 15
  1. A. The remaining 50 diagnoses could include overdiagnosis or later presentations; treatment adds measured harm (Best answer)

    Read the complete explanation

    At year 15, 200 - 150 = 50 extra real diagnoses remain after partial catch-up. This is compatible with overdiagnosis as a contributor, but delayed cases and uncertainties prevent labeling any one lesion; treatment complications matter.

  2. B. The 80-diagnosis early excess is fully explained by advancing diagnosis, since the control group partially catches up (Why this does not fit)

    Read the complete explanation

    At year 5 the gap is 180 - 100 = 80; at year 15 it narrows but remains 200 - 150 = 50, so pure completed catch-up does not account for all diagnoses under this follow-up.

  3. C. The remaining 50-diagnosis excess estimates 50 people harmed by treatment during follow-up (Why this does not fit)

    Read the complete explanation

    The incidence difference is 50, but it is not a count of treatment complications or a certain classification of individual overdiagnoses; harm requires its own outcome data.

  4. D. The excess is best counted as 50 false-positive pathology reports because mortality is unchanged (Why this does not fit)

    Read the complete explanation

    The excess counts pathologically confirmed diagnoses, not false-positive test results; unchanged mortality alone does not classify the pathology.

Takeaway: Persistent excess of real diagnoses with incomplete catch-up can support overdiagnosis without proving each case.

Case sources: [2]

Case 16

A trial assigns 10,000 adults to screening invitation and 10,000 to usual care, with equal follow-up and near-complete outcome ascertainment. Target-disease deaths are 40 versus 60, respectively; serious diagnostic-workup complications are 90 versus 30. Five-year survival among diagnosed cancers also rises with screening. Which assigned-group interpretation should inform the program decision?

Show answer and explanations for case 16
  1. A. Screening prevents about 20 disease deaths but causes about 60 extra complications, a net loss of 40 lives per 10,000 (Why this does not fit)

    Read the complete explanation

    Both absolute differences are correct, but a serious workup complication is not automatically equivalent to a death; subtracting counts as lives is invalid.

  2. B. Per 10,000 invited: 20 fewer disease deaths and 60 extra serious complications; value the outcomes separately (Best answer)

    Read the complete explanation

    The assigned-group differences are 60 - 40 = 20 fewer disease deaths and 90 - 30 = 60 extra serious complications per 10,000. Their counts cannot be simply netted as equal outcomes.

  3. C. Screening prevents about 60 disease deaths and causes about 60 extra complications per 10,000 invited (Why this does not fit)

    Read the complete explanation

    Sixty is the usual-care disease-death count, not the reduction; the reduction is 20.

  4. D. Screening prevents about 20 disease deaths and causes about 90 extra serious complications per 10,000 invited (Why this does not fit)

    Read the complete explanation

    The mortality difference is correct; 90 is total complications in screening, not the excess over 30 in usual care, which is 60.

Takeaway: Compare mortality benefit and diagnostic harms on the same randomized denominator without treating unlike outcomes as interchangeable.

Case sources: [2] [6]

Case 18

A randomized trial followed 1,000 people per arm for two years. Preventable events occurred in 100 controls and 80 treated patients; major bleeding occurred in 40 controls and 50 treated patients. People taking an anticoagulant were excluded. A current patient takes an anticoagulant. What trial NNT and NNH follow, and can its NNH be used as that patient's personal bleeding estimate?

Show answer and explanations for case 18
  1. A. NNT 50 and NNH 100 over two years in eligible trial patients; bleeding risk for the excluded anticoagulant regimen needs separate evidence (Best answer)

    Read the complete explanation

    Event risk falls from 100/1000 to 80/1000, an absolute 2% reduction (NNT 50). Bleeding rises from 40/1000 to 50/1000, an absolute 1% increase (NNH 100). Exclusion leaves the anticoagulant interaction unknown.

  2. B. NNT 50 and NNH 100 over two years; use NNH 100 as the anticoagulated patient's personal bleeding estimate (Why this does not fit)

    Read the complete explanation

    The trial differences yield NNT 50 and NNH 100, but anticoagulant use was excluded; an interaction could change bleeding risk for this patient.

  3. C. NNT 10 and NNH 20 over two years; bleeding risk for the excluded anticoagulant regimen needs separate evidence (Why this does not fit)

    Read the complete explanation

    NNT 10 reciprocates the 10% control event risk, and NNH 20 reciprocates the 5% treatment bleeding risk. Neither reciprocates the incremental difference; exclusion caution is right.

  4. D. NNT 100 and NNH 50 over two years; bleeding risk for the excluded anticoagulant regimen needs separate evidence (Why this does not fit)

    Read the complete explanation

    The two absolute differences are 2% event reduction and 1% bleeding increase, so their reciprocals are 50 and 100, not 100 and 50. Exclusion caution is right.

Takeaway: NNT and NNH invert comparable absolute risk differences, not arm risks; trial exclusions limit personal transport.

Case sources: [4] [23]

Case 19

With the same one-year follow-up and comparator, a trial estimates major bleeding in 4% of controls and 6% of treated patients. The estimated treatment-minus-control bleeding risk difference is 2 percentage points (95% CI -1 to 5 percentage points). What is the point NNH, and what can be concluded about increased bleeding?

Show answer and explanations for case 19
  1. A. Point NNH 17; the 95% risk-difference interval does not establish increased bleeding (Why this does not fit)

    Read the complete explanation

    About 17 reciprocates the 6% treated-arm rate, not the 2% increase versus control. The uncertainty statement is right.

  2. B. Point NNH 50; the 95% risk-difference interval establishes increased bleeding (Why this does not fit)

    Read the complete explanation

    NNH 50 is the correct point summary, but the risk-difference interval includes zero and negative values, so an increase is not established.

  3. C. Point NNH 50; the 95% risk-difference interval does not establish increased bleeding (Best answer)

    Read the complete explanation

    The absolute point increase is 6% - 4% = 2%, whose reciprocal is 50. Its interval crosses zero, including possible benefit, so it cannot be expressed as one finite harm-only NNH interval.

  4. D. Point NNH 100; the 95% risk-difference interval does not establish increased bleeding (Why this does not fit)

    Read the complete explanation

    The -1-point lower interval endpoint is opposite-direction benefit, not a 1-point harm increase; the point NNH comes from 2 points, not that limit. The uncertainty statement is right.

Takeaway: Point NNH uses the comparative absolute harm increase, while an interval crossing zero does not establish harm.

Case sources: [4] [23]

Case 20

For an adverse event, a trial reports experimental versus reference RR 0.82 with a matched two-sided 95% CI of 0.65 to 1.04. Before the trial, noninferiority was defined as ruling out RR 1.10 or greater using this interval; superiority uses the matched two-sided alpha 0.05 test against RR 1. Which conclusion follows under these prespecified rules?

Show answer and explanations for case 20
  1. A. Noninferiority is shown, but superiority is not (Best answer)

    Read the complete explanation

    The upper CI limit 1.04 is below the harmful margin 1.10, excluding unacceptable inferiority. Because the CI includes 1, two-sided superiority is not established.

  2. B. Neither noninferiority nor superiority is shown because the CI includes 1 (Why this does not fit)

    Read the complete explanation

    Including 1 prevents superiority, but it does not prevent noninferiority: the entire interval is below the 1.10 margin.

  3. C. Superiority is shown but noninferiority is not because the CI extends below 0.90 (Why this does not fit)

    Read the complete explanation

    For an adverse event, lower RR favors experimental treatment; the harmful-side noninferiority boundary is 1.10, not 0.90.

  4. D. Both noninferiority and superiority are shown because RR 0.82 favors experimental treatment (Why this does not fit)

    Read the complete explanation

    The point estimate favors treatment and the margin is excluded, but CI 0.65 to 1.04 includes the no-effect RR 1 and cannot establish matched two-sided superiority.

Takeaway: A ratio interval can support noninferiority without demonstrating superiority.

Case sources: [3] [5] [10]

Case 21

A trial's prespecified hospitalization endpoint has p = 0.03. The treatment prevents 1.5 hospitalizations per 100 patients (95% CI 0.1 to 2.9) but causes 2.5 additional dizziness episodes per 100 (95% CI 1.4 to 3.6). The committee requires evidence of at least 1 hospitalization prevented per 100 and no more than 3 additional dizziness episodes per 100 across the reported intervals before routine adoption. Which interpretation meets its rule?

Show answer and explanations for case 21
  1. A. The null has a 3% probability, so defer because both intervals miss the required bounds (Why this does not fit)

    Read the complete explanation

    The interval decision is right, but p = 0.03 is not the probability that the null is true.

  2. B. The result provides evidence against the hospitalization null, but neither interval establishes the required net profile (Best answer)

    Read the complete explanation

    Under the null, p = 0.03 describes results this extreme or more; the benefit lower bound is 0.1, below 1, and the harm upper bound is 3.6, above 3.

  3. C. The result provides evidence against the hospitalization null, so adopt because both point estimates satisfy the rule (Why this does not fit)

    Read the complete explanation

    The point estimates 1.5 and 2.5 meet the thresholds, but the rule explicitly requires the interval bounds.

  4. D. The result provides evidence against the hospitalization null, so defer only because the benefit interval misses the rule (Why this does not fit)

    Read the complete explanation

    The harm interval also extends to 3.6, exceeding the specified limit of 3.

Takeaway: Interpret p values under the null separately from a prespecified benefit-harm decision rule.

Case sources: [5] [23]

Case 22

A planned trial had 80% power at two-sided alpha 0.05 to detect a 5-point reduction in disease risk. Its observed treatment-minus-control difference is -2 points, with a 95% CI from -8 to +4 and a nonsignificant superiority test. Which interpretation respects both the design and observed result?

Show answer and explanations for case 22
  1. A. The 20% design miss rate identifies this nonsignificant result as a Type II error (Why this does not fit)

    Read the complete explanation

    The power interpretation is right, but the true effect is unknown; the interval cannot diagnose an actual Type II error.

  2. B. The interval includes -5 and zero, so benefit remains possible; 80% power is the posterior chance that -5 is true (Why this does not fit)

    Read the complete explanation

    The interval reading is right, but design power is not a posterior probability for this trial.

  3. C. At a true -5 effect, 20% miss significance; this CI supports neither superiority nor exclusion of -5 (Best answer)

    Read the complete explanation

    Power is conditional on the specified alternative; the interval includes zero and -5, so the observed data leave both possibilities compatible.

  4. D. Because the observed effect is -2 rather than -5, planned 80% power establishes that a 5-point reduction is ruled out (Why this does not fit)

    Read the complete explanation

    The CI extends to -8 and includes -5; planned power does not negate an interval-compatible effect.

Takeaway: Planned power is conditional on an alternative; the observed interval governs what this trial remains compatible with.

Case sources: [3] [9]

Case 23

For a new minus standard cure-rate difference, investigators prespecified equivalence margins of -5 to +5 percentage points as clinically acceptable. They prespecified the same 90% interval for both equivalence and a one-sided 0.05 noninferiority test against a 5-point loss. Their 90% CI is -2 to +8 points, and a conventional superiority test is nonsignificant. What conclusion follows under these prespecified margins?

Show answer and explanations for case 23
  1. A. The new treatment is superior because the CI allows an 8-point gain (Why this does not fit)

    Read the complete explanation

    Possible gain is not demonstrated superiority when the interval also includes zero and loss.

  2. B. The new treatment meets noninferiority against a 5-point loss, but not equivalence (Best answer)

    Read the complete explanation

    The lower CI bound -2 exceeds -5, but the upper bound +8 exceeds +5, so the interval is not wholly within both margins.

  3. C. The new treatment fails noninferiority because the CI includes a 2-point loss (Why this does not fit)

    Read the complete explanation

    The prespecified allowable loss is 5 points; -2 excludes losses worse than -5.

  4. D. The treatments are equivalent because the CI crosses zero and superiority is nonsignificant (Why this does not fit)

    Read the complete explanation

    Crossing zero says neither superiority direction was established; the upper bound exceeds the equivalence margin.

Takeaway: Equivalence requires the interval within both prespecified margins; one-sided noninferiority can hold without it.

Case sources: [10]

Case 24

Three independent treatment arms have continuous outcomes with reasonable ANOVA residual assumptions. The omnibus equal-means test has p = 0.01. Prespecified pairwise comparisons with multiplicity control report adjusted p values A versus B = 0.04, A versus C = 0.12, and B versus C = 0.32. At a familywise 0.05 threshold, which report is supported?

Show answer and explanations for case 24
  1. A. The means are not all equal; each pair differs because the omnibus result is significant (Why this does not fit)

    Read the complete explanation

    The omnibus result does not identify pairs; adjusted p values 0.12 and 0.32 do not meet 0.05.

  2. B. The omnibus result is not significant after controlling three pairwise comparisons (Why this does not fit)

    Read the complete explanation

    The separately reported omnibus p = 0.01 rejects joint equality; the pairwise adjusted p values apply to their individual contrasts.

  3. C. Only A versus B differs, and its inference requires a significant omnibus gate (Why this does not fit)

    Read the complete explanation

    The A-B result is supported, but prespecified contrasts need not universally wait for an omnibus rejection.

  4. D. The means are not all equal; only A versus B is established by the adjusted comparisons (Best answer)

    Read the complete explanation

    The omnibus rejects joint equality and only the adjusted A-B comparison is below 0.05; the other comparisons remain uncertain.

Takeaway: An omnibus result does not locate contrasts, and planned contrasts need no universal omnibus gate.

Case sources: [14] [19] [20]

Case 25

Two independent prospective hospital cohorts have 40 patients each, with one complication in hospital A and eight in hospital B over 30 days. Each hospital uses its own protocol, assigned by hospital rather than randomized; no patients are matched. Investigators want a two-sided comparison of complication counts. Which table analysis and interpretation fit?

Show answer and explanations for case 25
  1. A. Fisher exact test; the observed comparison does not identify a causal protocol benefit (Best answer)

    Read the complete explanation

    Nine pooled complications imply 4.5 expected events per group, favoring Fisher exact inference for sparse independent counts. Protocol is inseparable from hospital, so this contrast cannot isolate its causal effect.

  2. B. Fisher exact test; the observed comparison identifies a causal protocol benefit (Why this does not fit)

    Read the complete explanation

    Fisher fits the sparse independent table, but hospitals differ along with protocols; the observed association cannot isolate protocol benefit.

  3. C. Exact McNemar test; the observed comparison does not identify a causal protocol benefit (Why this does not fit)

    Read the complete explanation

    There are distinct patients with no matching, so paired McNemar inference does not fit. Hospital confounding caution is right.

  4. D. Pearson chi-square test; the observed comparison does not identify a causal protocol benefit (Why this does not fit)

    Read the complete explanation

    Pearson uses a large-sample reference with only 4.5 expected events in each group; Fisher exact is better suited. Hospital confounding caution is right.

Takeaway: Sparse independent binary tables call for a suitable exact comparison; hospital-level protocol allocation remains confounded.

Case sources: [3] [22]

Case 26

Thirty patients have blood pressure measured before and after treatment. Patient-level before-minus-after differences are approximately normal; their mean is a 4-mm Hg reduction with a 95% CI from 1 to 7 mm Hg. The clinical target is to establish a mean reduction greater than 3 mm Hg. Which analysis and conclusion fit?

Show answer and explanations for case 26
  1. A. Use a paired t analysis; the mean reduction exceeds 3 mm Hg because its point estimate is 4 (Why this does not fit)

    Read the complete explanation

    The method is appropriate, but the CI lower bound 1 does not exclude clinically insufficient reductions.

  2. B. Use a paired t analysis; a positive mean reduction is supported, but exceeding 3 mm Hg is not established (Best answer)

    Read the complete explanation

    The paired differences target mean within-person change. The CI excludes zero but includes reductions below 3.

  3. C. Use a paired signed-rank analysis; it establishes mean reduction above 3 because the CI upper bound is 7 (Why this does not fit)

    Read the complete explanation

    A signed-rank test does not directly target the stated mean and an upper bound above 3 cannot establish the lower threshold.

  4. D. Use an independent-samples t analysis; a positive mean reduction is supported but exceeding 3 is not established (Why this does not fit)

    Read the complete explanation

    The interval interpretation is right, but treating each before-after pair as independent discards the linked observations.

Takeaway: Analyze patient-level differences and distinguish evidence of change from evidence of a clinically specified minimum.

Case sources: [13]

Case 27

Two independent recovery-time groups have a significant Mann-Whitney rank result (p = 0.02). Sample medians are nearly equal, but the reported interquartile ranges differ substantially. No distribution plot or individual observations are supplied. Which interpretation is warranted?

Show answer and explanations for case 27
  1. A. Report evidence against the rank-based null and attribute it specifically to wider spread (Why this does not fit)

    Read the complete explanation

    The IQRs differ, but these summaries alone cannot assign the rank result entirely to spread.

  2. B. Report a difference in population medians because the rank p value is below 0.05 (Why this does not fit)

    Read the complete explanation

    A rank result with differing shapes does not by itself establish a population-median difference.

  3. C. Report the rank result with the median and IQR summaries, without isolating a median or spread effect (Best answer)

    Read the complete explanation

    The rank result indicates a distributional contrast under the test assumptions, while the unlike spreads make a simple common-shape location interpretation unsafe; the summaries do not isolate the cause.

  4. D. Report that most individuals in one arm recover more slowly because the rank p value is below 0.05 (Why this does not fit)

    Read the complete explanation

    A significant group-level rank test does not determine ordering for most individual pairs from the supplied summaries.

Takeaway: Rank evidence is not automatically a median difference when distribution shapes differ.

Case sources: [15]

Case 28

Patients with a rare cancer and source-population controls are interviewed about medication used years earlier. Cases search old records more diligently than controls, and the same archived pharmacy database covers both groups for the years of interest. Which change most directly reduces the threatened differential exposure measurement?

Show answer and explanations for case 28
  1. A. Ask every participant identical interview questions but allow cases and controls to bring whatever records they locate (Why this does not fit)

    Read the complete explanation

    Question wording is standardized, but the stated disparity in record-search effort remains.

  2. B. Match cases and controls on age and interview year, retaining the original interviews (Why this does not fit)

    Read the complete explanation

    Matching can improve covariate comparability but still leaves cases exerting more effort to recover exposure.

  3. C. Apply the same blinded pharmacy-record abstraction to cases and controls (Best answer)

    Read the complete explanation

    The common archival source and blinded standardized abstraction reduce outcome-dependent effort in recalling and documenting exposure.

  4. D. Use separate disease-specific interviewers trained to pursue past medications thoroughly in each group (Why this does not fit)

    Read the complete explanation

    Thorough interviewing can improve recall, but separate procedures do not remove differential case-control ascertainment as directly as a common blinded archive.

Takeaway: Comparable blinded records reduce outcome-dependent exposure ascertainment.

Case sources: [3]

Case 29

A registry samples 200 patients with a rare neurologic disease and 200 source-population controls. Prior occupational exposure is documented in 60 cases and 40 controls. Controls represent the population that produced the cases, but the fraction of all diseased people sampled and exposed-worker denominators are unavailable. What can these sampled data directly support?

Show answer and explanations for case 29
  1. A. Exposure odds ratio about 1.71 and disease risk ratio exactly 1.71 (Why this does not fit)

    Read the complete explanation

    The sampled exposure odds comparison is valid, but a disease risk ratio is not directly identified by these outcome-sampled counts.

  2. B. Exposure odds ratio about 1.71, but not the exposed workers’ disease risk (Best answer)

    Read the complete explanation

    Exposure odds are 60/140 in cases and 40/160 in controls; their ratio is about 1.71. Outcome-selected sampling lacks exposure-group risk denominators.

  3. C. Disease risk ratio 1.5 using 60 exposed cases and 40 exposed controls (Why this does not fit)

    Read the complete explanation

    60/40 compares exposed counts in outcome-selected samples, not disease risks among all exposed and unexposed workers.

  4. D. Exposure odds ratio about 0.58 because controls are the denominator (Why this does not fit)

    Read the complete explanation

    That reverses the requested case-to-control exposure odds ratio; 0.58 is the reciprocal of about 1.71.

Takeaway: Outcome-based sampling supports an exposure odds ratio, not direct exposure-group risk.

Case sources: [3]

Case 30

A trial randomizes 100 patients to exercise and 100 to usual care. Thirty assigned exercise stop attending; nevertheless, one-year hospitalization is known for everyone: 10 assigned exercise and 20 assigned usual care are hospitalized. The policy question is whether offering exercise lowers hospitalization, not the effect of actually completing sessions. Which estimate and analysis answer it?

Show answer and explanations for case 30
  1. A. Exclude the 30 nonattenders and compare 10 events with the usual-care group (Why this does not fit)

    Read the complete explanation

    Attendance is postrandomization and the 10 hospitalizations are not specified by attendance; exclusion would no longer estimate the offer effect.

  2. B. Compare original assignment groups: 10% versus 20%, a 10-point lower risk for the exercise offer (Best answer)

    Read the complete explanation

    With all outcomes available, original-group risks are 10/100 and 20/100; the -10-point difference estimates assignment to the offer despite nonattendance.

  3. C. Compare original assignment groups: 10% versus 20%, a 10-point effect of completing exercise (Why this does not fit)

    Read the complete explanation

    The arithmetic is right, but random assignment identifies the offer effect, not directly the effect of adherence.

  4. D. Reassign the 30 nonattenders to usual care, then estimate a 10-point offer effect (Why this does not fit)

    Read the complete explanation

    Reassignment changes original allocation and its denominators, losing the randomized offer comparison.

Takeaway: With complete outcomes, original assignment estimates the offer effect rather than the effect of adherence.

Case sources: [6]

Case 31

A randomized trial targets the effect of assignment on a final symptom score regardless of whether medication is stopped. Follow-up assessments stopped when medication stopped: 20% assigned drug and 5% assigned control discontinued. Treated patients' last observed scores tended to worsen before toxicity-related discontinuation. Baseline severity, earlier scores and toxicity were recorded. Which improved primary collection and analysis, plus sensitivity strategy, best addresses the target?

Show answer and explanations for case 31
  1. A. Collect scores after discontinuation and retain assigned groups, but rely only on imputation conditional on observed history without a departure analysis (Why this does not fit)

    Read the complete explanation

    The collection and assigned-group strategy targets assignment, but observed-history MAR alone does not probe plausible worse scores after toxicity-related withdrawal.

  2. B. Analyze only patients who kept taking medication, and use pattern-mixture delta shifts for their unavailable final scores (Why this does not fit)

    Read the complete explanation

    Receipt-group analysis changes the assignment estimand and selects on postrandomization adherence. Delta shifts do not repair that estimand change.

  3. C. Continue collecting scores after discontinuation and retain assigned groups; for residual missing scores, assess plausible worse outcomes with conditional pattern-mixture delta shifts (Best answer)

    Read the complete explanation

    Stopping medication need not end outcome follow-up for the assignment effect, so collect post-stop scores and analyze assigned groups. Worsening before selective withdrawal makes observed-history MAR extrapolation uncertain; conditional delta shifts test plausible worse missing outcomes.

  4. D. Continue collecting scores after discontinuation but analyze only complete scores in assigned groups; reserve delta shifts for withdrawn participants (Why this does not fit)

    Read the complete explanation

    Collecting after stop is right, but restricting primary analysis to complete scores can select on toxicity-related missingness; use available assigned outcomes with explicit missing-data assumptions.

Takeaway: For an assignment estimand, continue outcomes after discontinuation and test departures from observed-history missingness assumptions.

Case sources: [6] [28]

Case 32

A complete town register has 250 people in each exposure (E) by disease (D) cell: E+D+, E+D-, E-D+, E-D-. Hospital admissions from those cells number 100, 100, 100 and 25, respectively. Independent complete administrative records establish these exact cell-specific admission probabilities; all are positive. The hospital analysis uses only admitted people and reports an odds ratio of .25. What explains the contrast with the town, and how can admitted data be reanalyzed under these known selection probabilities?

Show answer and explanations for case 32
  1. A. The .25 association indicates a protective exposure effect; adjust for hospital-stay duration (Why this does not fit)

    Read the complete explanation

    Admission selection depends jointly on E and D, so the admitted OR need not be a causal town effect. Stay duration does not invert entry probabilities.

  2. B. Admission is a common-effect selection filter; multiply admitted cell counts by their admission probabilities .4, .4, .4 and .1 (Why this does not fit)

    Read the complete explanation

    The mechanism is recognized, but multiplying by selection probability selects again. Inverse probabilities, not probabilities, recover the town table.

  3. C. Admission is a common-effect selection filter; weight admitted records 2.5, 2.5, 2.5 and 10 by cell to recover a town OR of 1 (Best answer)

    Read the complete explanation

    Admission depends on both E and D. Dividing each admitted count by its known positive admission probability recovers 250 in every town cell and OR 1.

  4. D. Admission is a common-effect selection filter; apply the same 2.5 admission weight to all four cells (Why this does not fit)

    Read the complete explanation

    A common weight leaves the hospital OR .25 intact. The E-D- cell has only 25/250 = .1 inclusion and needs weight 10.

Takeaway: Known positive selection probabilities permit inverse-probability recovery of the source table, not causal proof.

Case sources: [25]

Case 33

An observational drug study begins at first prescription. More severe patients preferentially receive the drug, and pretreatment severity predicts subsequent admission. The drug can also change a severity score measured two weeks later, before many admissions. Which comparison best addresses confounding by indication when estimating the drug’s total effect on admission?

Show answer and explanations for case 33
  1. A. Compare treated and untreated patients after matching on the two-week severity score (Why this does not fit)

    Read the complete explanation

    The post-treatment score can be changed by the drug and may lie on the causal pathway; matching there does not repair baseline indication and can distort total effects.

  2. B. Compare only patients with the same baseline severity without checking whether both treatments occur at that level (Why this does not fit)

    Read the complete explanation

    Baseline comparability is useful, but without treatment overlap at those severity levels no supported comparison is available.

  3. C. Compare all treated and untreated patients adjusting for the two-week score but not baseline severity (Why this does not fit)

    Read the complete explanation

    Post-treatment adjustment cannot replace baseline confounding control and may remove part of the drug effect.

  4. D. Compare treated and untreated patients with overlapping pretreatment severity, adjusting for baseline severity (Best answer)

    Read the complete explanation

    Pretreatment severity influences both choice and prognosis; restricting to comparable baseline patients addresses measured indication without conditioning on a possible treatment mediator.

Takeaway: Compare on pretreatment indication, respecting treatment overlap and avoiding adjustment for downstream variables.

Case sources: [3] [6]

Case 34

A review finds that small trials report larger benefits than large trials. Small trials also enrolled more severe patients, and no comparison with trial registries has yet been made. Which investigation best addresses both live explanations for the funnel asymmetry?

Show answer and explanations for case 34
  1. A. Compare registered protocols and unpublished results, then pool all severities under a single fixed effect (Why this does not fit)

    Read the complete explanation

    The registry check addresses missing evidence but ignores the documented severity difference that could modify effects.

  2. B. Compare registered protocols and unpublished results, then assess whether severity modifies effects across studies (Best answer)

    Read the complete explanation

    Registry checks address possible missing results; severity-stratified analysis probes clinical heterogeneity without declaring either explanation proven.

  3. C. Apply a funnel-asymmetry correction and retain one pooled estimate without investigating severity (Why this does not fit)

    Read the complete explanation

    A statistical correction cannot distinguish selection from a real difference in populations when severity differs.

  4. D. Stratify by severity and exclude unregistered trials without seeking their results (Why this does not fit)

    Read the complete explanation

    Severity is relevant, but exclusion by registration alone does not investigate whether completed studies or outcomes are missing.

Takeaway: Funnel asymmetry has multiple explanations; check missing studies and clinical effect modification.

Case sources: [8]

Case 35

A completed survival trial examined five accumulating datasets without prespecified monitoring boundaries, stopping when an unadjusted p value first fell below 0.05. The team now wants to report the finding responsibly and plan the next trial. Which combined response addresses both repeated-testing error and a favorable-stopping effect estimate?

Show answer and explanations for case 35
  1. A. Treat the five past looks as if a conventional boundary had been prespecified, then report only the fifth-look effect and interval (Why this does not fit)

    Read the complete explanation

    A later boundary cannot retroactively make unplanned monitoring prespecified, nor does selective fixed-sample reporting address estimation bias.

  2. B. Disclose all looks; assess sequential-testing sensitivity and stopping-aware estimates now; prespecify next-trial boundaries (Best answer)

    Read the complete explanation

    The past looks cannot be erased or retroactively prespecified. Disclosure and sensitivity to repeated looks plus stopping-aware estimation address the current report; the future design controls its error.

  3. C. Disclose every look and prespecify controlled boundaries for the next trial, but report the stopped point estimate as an ordinary fixed-sample estimate (Why this does not fit)

    Read the complete explanation

    The next design and disclosure help, but the selected favorable stopping estimate remains vulnerable to exaggeration.

  4. D. Disclose every look and adjust only the stopped p value for five looks, while presenting the ordinary effect interval as if no interim selection occurred (Why this does not fit)

    Read the complete explanation

    Repeated-testing adjustment addresses one concern, but the ordinary interval and estimate do not address favorable stopping.

Takeaway: Completed unplanned looks require disclosure and qualified inference; future error control cannot rewrite the past.

Case sources: [9]

Case 36

A trial in high-risk patients reports p < 0.001 for a mean 1-mm Hg pressure reduction and estimates a 20% relative reduction in cardiovascular events. A patient considering the drug has untreated one-year event risk 1%, versus 10% over the same one-year horizon in the trial, and the drug adds 1.5 dizziness episodes per 100 patients over that same year. Assume the relative effect transports. This patient values preventing one event as four times avoiding one dizziness episode and wants the larger expected benefit. Which decision follows from these stated quantities?

Show answer and explanations for case 36
  1. A. Decline: benefit is 0.2 events prevented per 100, weighted to 0.8 dizziness episodes, below 1.5 added (Best answer)

    Read the complete explanation

    At 1% baseline risk, a 20% relative reduction is 0.2 percentage points; weighted benefit 4 times 0.2 = 0.8, less than 1.5 dizziness episodes per 100.

  2. B. Accept: benefit is 2 events prevented per 100, weighted to 8 dizziness episodes, above 1.5 added (Why this does not fit)

    Read the complete explanation

    Two events prevented per 100 uses the trial baseline risk of 10%, not this patient's 1%. Applying the stipulated relative reduction to 1% gives 0.2 events prevented per 100.

  3. C. Accept: benefit is 20 events prevented per 100, weighted to 80 dizziness episodes, above 1.5 added (Why this does not fit)

    Read the complete explanation

    A 20% relative reduction is not 20 prevented events per 100 people. Multiply the patient's 1% baseline risk by 20% to obtain a 0.2-per-100 absolute benefit.

  4. D. Decline: benefit is 0.02 events prevented per 100, weighted to 0.08 dizziness episodes, below 1.5 added (Why this does not fit)

    Read the complete explanation

    The risk reduction is 0.01 times 0.20 = 0.002 per person, or 0.2 per 100. Reporting 0.02 per 100 introduces a factor-of-ten conversion error, although declining is the correct decision.

Takeaway: Patient baseline risk converts a potentially transportable relative effect into a different absolute benefit.

Case sources: [5] [23]

Case 37

A trial targets a fixed 4-unit treatment effect on its validated clinical score, with two-sided alpha and equal allocation held fixed. Feasible clinical-score plans are 200 patients per arm with SD 24 or 120 per arm with SD 18. A laboratory biomarker can be measured in 120 per arm at SD 12, but its treatment effect and link to the clinical score are unknown. Another proposal merely declares a smaller true clinical effect without changing measurements. Which plan gives more information for the original clinical question?

Show answer and explanations for case 37
  1. A. Use the biomarker in 120 per arm at SD 12 because its numerical signal proxy is greatest (Why this does not fit)

    Read the complete explanation

    For the fixed clinical effect, sqrt(120)/18 = 0.6086 exceeds sqrt(200)/24 = 0.5893. The biomarker has a different, unknown effect and cannot inherit the clinical 4-unit effect.

  2. B. Use the clinical score in 200 per arm at SD 24 because the larger sample outweighs its extra noise (Why this does not fit)

    Read the complete explanation

    The larger clinical sample is noisier: sqrt(200)/24 = 0.5893 is below sqrt(120)/18 = 0.6086.

  3. C. Retain the same physical data but declare a smaller true clinical effect to increase power (Why this does not fit)

    Read the complete explanation

    A smaller actual effect with unchanged n and SD reduces the fixed-alpha signal; declaring it does not increase information.

  4. D. Use the clinical score in 120 per arm at SD 18 because its signal proxy exceeds the larger noisier clinical plan (Best answer)

    Read the complete explanation

    Both clinical plans preserve the endpoint; sqrt(120)/18 = 0.6086 exceeds sqrt(200)/24 = 0.5893. The smaller biomarker SD does not answer the same clinical endpoint.

Takeaway: At a fixed clinical endpoint and effect, compare sample size against variance rather than inheriting a surrogate effect.

Case sources: [3] [9]

Case 38

Adults were randomly assigned to a screening invitation or no invitation. Screening uptake was 50% with invitation and 10% without; two-year mortality was 2.0% and 2.8%, respectively, by assignment with complete ascertainment. Assume random assignment, no effect of invitation on mortality except through screening, and no one who would screen without invitation but refuse when invited. Under these assumptions, what does the Wald risk-difference ratio identify?

Show answer and explanations for case 38
  1. A. A 0.8-percentage-point mortality reduction among invitation compliers, because the assignment difference already describes them (Why this does not fit)

    Read the complete explanation

    The -0.8-point assignment difference is the invitation ITT, not the complier effect; divide by the 0.40 uptake difference under the stated assumptions.

  2. B. A 2-percentage-point mortality reduction in the entire population, because dividing by uptake recovers a population effect (Why this does not fit)

    Read the complete explanation

    The Wald ratio is -0.02, but it is a local complier effect. The randomized population invitation effect is -0.8 point; a population effect of screening receipt is not identified here.

  3. C. A 2-percentage-point mortality reduction among invitation compliers, not the overall population effect (Best answer)

    Read the complete explanation

    The assignment risk difference is 2.0% - 2.8% = -0.8 percentage point and the uptake difference is 50% - 10% = 40 points. Wald ratio -0.008/0.40 = -0.02 identifies a local effect among invitation compliers under the stated assumptions, not the population effect.

  4. D. A 0.8-percentage-point per-protocol reduction among recipients, because assignment comparisons measure actual receipt (Why this does not fit)

    Read the complete explanation

    The -0.8-point difference compares invitation assignments rather than screened recipients; per-protocol grouping would not preserve randomization.

Takeaway: Under explicit instrumental-variable assumptions, the uptake-adjusted Wald contrast is local to invitation compliers.

Case sources: [27]

Case 39

A complete 2015 factory roster includes 1,000 solvent-exposed and 2,000 unexposed workers, all kidney-disease-free then. Over a defined two-year follow-up, kidney disease is completely ascertained for everyone, with no deaths or losses before an outcome: 50 exposed and 40 unexposed develop disease. Person-time is not tabulated. Which measure is directly estimable?

Show answer and explanations for case 39
  1. A. Incidence-rate ratio 2.5 (Why this does not fit)

    Read the complete explanation

    The numerical ratio of cumulative risks is 2.5, but no person-time was supplied to estimate rates.

  2. B. Two-year cumulative-risk ratio 1.25 (Why this does not fit)

    Read the complete explanation

    50/40 compares case counts, overlooking the exposed denominator 1,000 and unexposed denominator 2,000.

  3. C. Two-year cumulative-risk ratio 2.5 (Best answer)

    Read the complete explanation

    Cumulative risks are 50/1000 = 5% and 40/2000 = 2%; their ratio is 2.5. The disease-free baseline and complete follow-up support risks, not an incidence-rate ratio without person-time.

  4. D. Two-year cumulative-risk ratio 0.4 (Why this does not fit)

    Read the complete explanation

    0.4 reverses the exposed-to-unexposed comparison; 2%/5% is the unexposed-to-exposed ratio.

Takeaway: Complete outcome follow-up in baseline disease-free exposure groups permits cumulative risks; rates require person-time.

Case sources: [3]

Case 40

A community study follows initially pain-free joggers and nonjoggers for first episodes of persistent knee pain. The incidence rate is 20 per 1,000 person-years in each group. Separate complete follow-up of episodes finds a mean duration of six months in joggers and three months in nonjoggers. Assume a rare condition, approximate steady state, comparable ascertainment and no differential detection. What prevalences does the incidence-duration relationship predict, and does this observational comparison identify a causal jogging effect on onset?

Show answer and explanations for case 40
  1. A. Predicted prevalence is 20 versus 20 per 1,000; the observational comparison does not identify a causal onset effect (Why this does not fit)

    Read the complete explanation

    This substitutes an annual onset rate for the proportion currently affected. Multiply 20 per 1,000 per year by 0.5 and 0.25 years to obtain 10 and 5 per 1,000. The causal caution is appropriate.

  2. B. Predicted prevalence is 10 versus 5 per 1,000; equal observed onset identifies no causal jogging effect on onset (Why this does not fit)

    Read the complete explanation

    The predicted prevalences are correct, but equal rates in observational groups do not identify a zero causal effect. Differences in other causes or selection can mask a causal effect.

  3. C. Predicted prevalence is 20 versus 10 per 1,000; the observational comparison does not identify a causal onset effect (Why this does not fit)

    Read the complete explanation

    These predictions double the stated episode durations, using one year and half a year instead of six and three months. The causal caution is appropriate.

  4. D. Predicted prevalence is 10 versus 5 per 1,000; the observational comparison does not identify a causal onset effect (Best answer)

    Read the complete explanation

    Under the supplied rare-condition steady-state approximation, 20 per 1,000 per year times 0.5 or 0.25 years predicts 10 or 5 per 1,000. Observational equality in onset does not identify a causal effect.

Takeaway: Prevalence reflects onset and duration; observational equal onset rates do not identify causal effects.

Case sources: [26]

Case 41

Seven injury reports after a drug first raised a possible hepatotoxicity signal. A complete two-year new-user cohort then follows 2,000 exposed and 2,000 comparable unexposed people, all injury-free at baseline. With preexisting stable liver disease: exposed 50 new injuries among 1,000; unexposed 20 among 400. Without it: exposed 10 among 1,000; unexposed 16 among 1,600. Standardize both groups to a population with 20% preexisting liver disease. What changes when denominators and this common case mix are used?

Show answer and explanations for case 41
  1. A. Standardized risks are 1.8% versus 1.8%; measured case mix accounts for the crude contrast without establishing a causal effect (Best answer)

    Read the complete explanation

    The complete cohort supplies denominators: 60/2000 versus 36/2000. Standardization gives .2*.05+.8*.01 = 1.8% for each; residual confounding or other uncertainty remains possible.

  2. B. Standardized risks are 3.0% versus 1.8%; the 1.2-point contrast identifies a causal drug-related excess (Why this does not fit)

    Read the complete explanation

    The crude difference is 1.2 points, but exposed and unexposed groups have different liver-disease composition; stratum risks match.

  3. C. Standardized risks are 1.8% versus 1.8%; the zero contrast identifies absence of a causal drug-related excess (Why this does not fit)

    Read the complete explanation

    The measured-stratum standardized contrast is zero, but observational equality does not prove causal absence or rule out residual confounding.

  4. D. Standardized risks are 3.0% versus 1.8%; the 1.2-point contrast persists after common-mix adjustment (Why this does not fit)

    Read the complete explanation

    Using 50% for exposed and 20% for unexposed retains different mixtures and reproduces crude risks. An adjustment comparison needs the same target distribution.

Takeaway: Cohort denominators and common-mix standardization separate a crude safety signal from measured case-mix imbalance.

Case sources: [3]

Search Bone Wizardry

Quick links