The V Lab
AudienceLearners in medicine, psychology, pharmacy, and public health
Study timeApproximately 180–240 minutes
PrerequisitesConfidence intervals, regression models, and basic R

Teaching Case and Scope of Responsibility The 360 participants, study sites, efficacy outcomes, and adverse events on this page are simulated with a fixed random seed and do not correspond to any real patient or product. The code explains design and analysis principles; it cannot replace a trial protocol, statistical analysis plan, ethics review, data monitoring committee, or validated regulatory analysis workflow.

How to Use This Tutorial

A clinical trial is not a matter of “collecting all the data and then choosing a test.” We recommend studying the following causal and operational sequence:

Clinical question → estimand → design safeguards → sample size → data quality → prespecified model → sensitivity analysis → transparent reporting

The running example is a simulated multicenter, double-blind, two-arm, parallel-group superiority trial. The primary outcome is the symptom score at week 12, and lower scores are better. The tutorial also illustrates response, counts of acute episodes, adverse events, and recurrence within one year. All effects are expressed as Treatment − Control: negative values favor Treatment for continuous outcomes; response RDs and RRs above their reference values favor Treatment; an adverse-event RR above 1 indicates increased risk; and a recurrence HR below 1 favors Treatment.

Learning Objectives

After completing this tutorial, you should be able to:

  • distinguish trial phases, study objectives, and specific designs;
  • express a PICO question as a five-attribute estimand under ICH E9(R1);
  • distinguish the randomization sequence, allocation concealment, and blinding, and assess baseline balance without “testing randomization”;
  • work backward from the primary estimand to calculate sample size for continuous, binary, and time-to-event outcomes;
  • select models for continuous, repeated-measures, binary, count, and survival outcomes and report effect sizes with 95% CIs;
  • distinguish ITT/FAS, per-protocol, safety, and as-treated analysis sets;
  • address randomization stratification factors, covariate adjustment, missing data, and intercurrent events;
  • correctly interpret multiplicity, interim analyses, subgroup interactions, and noninferiority margins;
  • identify the key unit of analysis for parallel, cluster, crossover, factorial, and adaptive trials;
  • report participant flow, analysis sets, effects, uncertainty, and harms in accordance with CONSORT 2025.

1 What Questions Do Clinical Trials Answer?

1.1 Interventional and Observational Questions

In a clinical trial, investigators assign interventions, usually with the aim of estimating the causal effect of a clearly defined treatment strategy; observational studies record exposures that have already occurred and require additional methods to address confounding. The central value of randomization is that it makes known and unknown baseline prognostic factors comparable in probability, not that it guarantees numerical equality for every variable in every small sample.

“Explanatory” trials focus more on whether an intervention can work under ideal conditions (efficacy), whereas “pragmatic” trials more closely examine how it works in routine care pathways (effectiveness). These form a continuum: eligibility criteria, intervention flexibility, the control condition, follow-up intensity, and the analysis strategy should all serve the intended use setting.

1.2 Phase Does Not Determine the Statistical Model

Phase Primary objective Common focus Statistical reminder
Phase 0 / early exploration Microdosing, target engagement, or feasibility PK/PD, mechanism Generally not used to confirm efficacy
Phase I Safety, tolerability, and dose DLT, PK/PD, dose escalation A small sample does not mean “no design is needed”
Phase II Dose selection and preliminary activity Dose–response, feasible endpoints Selection rules affect downstream uncertainty
Phase III Confirm benefit–risk Prespecified primary outcome, comparator, multicenter conduct Control type I error and protect the estimator
Phase IV Postmarketing effectiveness and safety Rare harms, real-world implementation Long-term follow-up, adherence, and selection become prominent concerns

Medical-device, behavioral-intervention, psychotherapy, and rare-disease studies do not necessarily follow the same drug-development phase labels. The phase describes the development objective; the design and estimand determine the analysis.

1.3 Protocol, Registration, and Statistical Analysis Plan

Before viewing results by randomized group, at minimum lock down the eligibility criteria, time zero, intervention version, comparator, primary and key secondary outcomes, measurement times, estimand, randomization, sample size, analysis sets, covariates, missing-data methods, multiplicity, interim rules, and sensitivity analyses.

  • Protocol: explains why and how the trial will be conducted; it can be organized according to SPIRIT 2025.
  • Trial registration: makes the research question, primary outcome, and timeline traceable; registration is not a complete protocol.
  • SAP: specifies variable derivations, models, contrasts, exceptional circumstances, and outputs before unblinding or accessing unblinded results by randomized group.
  • Amendments: justified changes are permitted, but they must include the date, rationale, and responsible person, and must distinguish changes made before versus after unblinding.

2 From PICO to an Estimand

2.1 Five Essential Attributes

ICH E9(R1) requires the treatment question to be expressed as an estimand. One operational version for the running example is:

Attribute Primary estimand in this example
Population Eligible, randomized adults in the multicenter target population
Treatment conditions Assignment to Treatment versus assignment to Control, including protocol-specified concomitant care
Variable Symptom score from baseline through week 12, with the week 12 score as the primary analysis variable
Intercurrent-event strategy A treatment-policy strategy for treatment discontinuation or switching to rescue therapy; a separate strategy is needed for events such as death that make the scale undefined
Population-level summary Mean difference at week 12 adjusted for baseline and randomization stratification factors (Treatment − Control)

Common intercurrent-event strategies are not synonyms for methods of “filling in” missing values:

  • Treatment policy: compare the outcomes of the originally assigned strategies regardless of events such as treatment discontinuation or switching;
  • Hypothetical: estimate what would have happened if a specified event had not occurred;
  • Composite: incorporate the event itself into the outcome, for example, classify treatment discontinuation as nonresponse;
  • While on treatment: consider only outcomes before the event occurs;
  • Principal stratum: target the population who would have the specified event status under both potential treatments; this usually requires strong assumptions.

An ICE Is Not the Same as Missingness Treatment discontinuation, rescue medication, death, or pregnancy are post-randomization events; an unobserved scale measurement is missing data. An intercurrent event can lead to missingness, but the clinical question must be defined first, followed by data-handling and estimation methods aligned with that question.

2.2 Primary, Sensitivity, and Supplementary Analyses

An estimator is the rule used to obtain an estimate of the estimand from the data. The primary analysis depends on a clearly specified set of assumptions; a sensitivity analysis changes key, unverifiable assumptions (for example, systematically shifting missing values in the Treatment arm toward worse outcomes); and a supplementary analysis answers a related but distinct question (for example, the per-protocol effect). “Running several models to see whether any are significant” is not a prespecified sensitivity framework.

3 Design Safeguards: Randomization, Concealment, and Blinding

3.1 Three Components Address Three Types of Bias

Component Question Practical requirement
Random-sequence generation Which group should the next participant enter? A reproducible algorithm, appropriate stratification/blocking, and a controlled randomization system
Allocation concealment Can the assignment be predicted or changed before enrollment? Recruiters cannot see upcoming assignments; use varying block sizes and restrict access
Blinding Are post-assignment behavior or assessments influenced by knowledge of treatment assignment? Report separately for participants, care providers, outcome assessors, and analysts

Blinding cannot be described adequately by the single phrase “double-blind.” In some surgical or psychological interventions, care providers cannot be blinded, but independent blinded outcome assessment, standardized co-interventions, blinded endpoint adjudication, and prespecified analysis can still be used. The triggers, permissions, and records for emergency unblinding are also part of the design.

3.2 Simulating Stratified Permuted-Block Randomization

The code below generates 360 simulated participants. Randomization is stratified by study site × baseline severity, with random block sizes of 4 or 6 within each stratum; in a real trial, block sizes should not be disclosed to recruiters.

set.seed(20260811)

make_block_sequence <- function(n, block_sizes = c(4L, 6L)) {
  allocation <- character(0)
  while (length(allocation) < n) {
    block_size <- sample(block_sizes, 1L)
    block <- rep(c("Control", "Treatment"), each = block_size / 2L)
    allocation <- c(allocation, sample(block))
  }
  allocation[seq_len(n)]
}

n <- 360L
trial <- data.frame(
  id = sprintf("P%03d", seq_len(n)),
  site = factor(sample(paste0("Site ", 1:4), n, replace = TRUE,
                       prob = c(0.28, 0.26, 0.24, 0.22))),
  sex = factor(sample(c("Female", "Male"), n, replace = TRUE,
                      prob = c(0.58, 0.42))),
  age = round(pmin(pmax(rnorm(n, 45, 13), 18), 75), 1)
)

site_baseline <- c("Site 1" = 0, "Site 2" = 0.8,
                   "Site 3" = -0.6, "Site 4" = 0.4)
trial$baseline <- round(
  pmin(pmax(rnorm(n, 24, 5.5) + site_baseline[trial$site], 10), 40), 1
)
trial$severity <- factor(ifelse(trial$baseline >= 24, "High", "Low"),
                         levels = c("Low", "High"))
trial$stratum <- interaction(trial$site, trial$severity, drop = TRUE)

trial$arm <- NA_character_
for (stratum_i in levels(trial$stratum)) {
  index <- which(trial$stratum == stratum_i)
  trial$arm[index] <- make_block_sequence(length(index))
}
trial$arm <- factor(trial$arm, levels = c("Control", "Treatment"))
treatment <- as.integer(trial$arm == "Treatment")

  # Week 12 continuous outcome: lower scores are better.
site_week12 <- c("Site 1" = 0, "Site 2" = 0.5,
                 "Site 3" = -0.5, "Site 4" = 0.2)
trial$week12_full <- 5 + 0.58 * trial$baseline - 2.5 * treatment +
  site_week12[trial$site] + 0.25 * (trial$sex == "Male") + rnorm(n, 0, 5.2)
trial$week12_full <- round(pmin(pmax(trial$week12_full, 0), 45), 1)

  # Poorer outcomes and the Treatment arm have higher missingness probabilities;
  # complete values are used only as simulated truth for teaching purposes.
p_missing <- plogis(-2.25 + 0.35 * treatment +
                     0.085 * (trial$week12_full - 18) +
                     0.20 * (trial$severity == "High"))
trial$missing12 <- rbinom(n, 1, p_missing)
trial$week12 <- ifelse(trial$missing12 == 1, NA_real_, trial$week12_full)
trial$change <- trial$week12 - trial$baseline

  # Response: at least a 30% reduction from baseline; treating missing observations
  # as nonresponse is only an example of a composite strategy.
trial$responder_full <- as.integer(
  (trial$baseline - trial$week12_full) / trial$baseline >= 0.30
)
trial$responder_nri <- ifelse(is.na(trial$week12), 0L, trial$responder_full)

  # Safety outcome.
p_ae <- plogis(qlogis(0.13) + log(1.75) * treatment +
                0.012 * (trial$age - 45))
trial$adverse_event <- rbinom(n, 1, p_ae)

  # Recurrence/progression within one year, with independent censoring for loss to follow-up.
event_time <- rexp(n, rate = 0.0022 *
                    exp(log(0.68) * treatment + 0.020 * (trial$baseline - 24)))
dropout_time <- rexp(n, rate = 0.00075)
trial$time_days <- pmin(event_time, dropout_time, 365)
trial$status <- as.integer(event_time <= pmin(dropout_time, 365))

  # Repeated scores and an overdispersed 12-week episode count; simulate these after
  # the core outcomes to preserve the fixed primary results.
trial$week4 <- pmax(0, 3.5 + 0.78 * trial$baseline - 0.9 * treatment +
                    rnorm(n, 0, 4.7))
trial$week8 <- pmax(0, 4.0 + 0.66 * trial$baseline - 1.8 * treatment +
                    rnorm(n, 0, 5.0))
trial$week4[rbinom(n, 1, plogis(-3 + 0.2 * treatment)) == 1] <- NA
trial$week8[rbinom(n, 1, plogis(-2.7 + 0.25 * treatment)) == 1] <- NA
trial$exposure_weeks <- round(runif(n, 8, 12), 1)
individual_frailty <- rgamma(n, shape = 2, rate = 2)
episode_mean <- exp(log(0.16) + log(0.75) * treatment +
                    0.018 * (trial$baseline - 24)) *
  trial$exposure_weeks * individual_frailty
trial$episode_count <- rpois(n, episode_mean)

stopifnot(nrow(trial) == n, !anyNA(trial$arm),
          all(trial$status %in% 0:1), all(trial$time_days > 0))

allocation_table <- addmargins(table(trial$stratum, trial$arm))
allocation_table
##              
##               Control Treatment Sum
##   Site 1.Low       22        24  46
##   Site 2.Low       18        19  37
##   Site 3.Low       28        28  56
##   Site 4.Low       16        16  32
##   Site 1.High      27        28  55
##   Site 2.High      28        27  55
##   Site 3.High      17        18  35
##   Site 4.High      22        22  44
##   Sum             178       182 360

3.3 Describe Baseline Characteristics with SMDs, Without Significance Hunting

After randomization, baseline differences are random variables. Item-by-item p-values cannot prove that randomization succeeded and should not determine which covariates are adjusted for; adjustment variables should be prespecified in the SAP based on prognostic value and the stratified design. Standardized mean differences (SMDs) help describe magnitude, but they are not “pass/fail” tests.

smd_continuous <- function(x, arm) {
  x0 <- x[arm == "Control"]
  x1 <- x[arm == "Treatment"]
  (mean(x1) - mean(x0)) /
    sqrt((stats::var(x1) + stats::var(x0)) / 2)
}

smd_binary <- function(x, arm) {
  p0 <- mean(x[arm == "Control"])
  p1 <- mean(x[arm == "Treatment"])
  (p1 - p0) / sqrt((p1 * (1 - p1) + p0 * (1 - p0)) / 2)
}

baseline_table <- data.frame(
  variable = c("Sample size", "Baseline score, mean (SD)", "Age, mean (SD)",
               "Male, n (%)", "Missing at week 12, n (%)"),
  Control = c(
    sum(trial$arm == "Control"),
    sprintf("%.1f (%.1f)", mean(trial$baseline[trial$arm == "Control"]),
            sd(trial$baseline[trial$arm == "Control"])),
    sprintf("%.1f (%.1f)", mean(trial$age[trial$arm == "Control"]),
            sd(trial$age[trial$arm == "Control"])),
    sprintf("%d (%.1f%%)", sum(trial$sex == "Male" & trial$arm == "Control"),
            100 * mean(trial$sex[trial$arm == "Control"] == "Male")),
    sprintf("%d (%.1f%%)", sum(trial$missing12 == 1 & trial$arm == "Control"),
            100 * mean(trial$missing12[trial$arm == "Control"]))
  ),
  Treatment = c(
    sum(trial$arm == "Treatment"),
    sprintf("%.1f (%.1f)", mean(trial$baseline[trial$arm == "Treatment"]),
            sd(trial$baseline[trial$arm == "Treatment"])),
    sprintf("%.1f (%.1f)", mean(trial$age[trial$arm == "Treatment"]),
            sd(trial$age[trial$arm == "Treatment"])),
    sprintf("%d (%.1f%%)", sum(trial$sex == "Male" & trial$arm == "Treatment"),
            100 * mean(trial$sex[trial$arm == "Treatment"] == "Male")),
    sprintf("%d (%.1f%%)", sum(trial$missing12 == 1 & trial$arm == "Treatment"),
            100 * mean(trial$missing12[trial$arm == "Treatment"]))
  ), check.names = FALSE
)

smd_table <- data.frame(
  variable = c("Baseline score", "Age", "Male"),
  SMD = c(smd_continuous(trial$baseline, trial$arm),
          smd_continuous(trial$age, trial$arm),
          smd_binary(trial$sex == "Male", trial$arm))
)

knitr::kable(baseline_table, caption = "Baseline characteristics and follow-up completeness in the simulated trial")
Baseline characteristics and follow-up completeness in the simulated trial
variable Control Treatment
Sample size 178 182
Baseline score, mean (SD) 24.2 (5.4) 24.4 (6.0)
Age, mean (SD) 45.1 (13.1) 44.9 (13.2)
Male, n (%) 78 (43.8%) 73 (40.1%)
Missing at week 12, n (%) 20 (11.2%) 25 (13.7%)
knitr::kable(smd_table, digits = 3,
             caption = "Treatment − Control standardized differences; absolute values describe magnitude only")
Treatment − Control standardized differences; absolute values describe magnitude only
variable SMD
Baseline score 0.030
Age -0.016
Male -0.075

Permuted-block randomization produces an overall allocation close to 1:1, and the absolute values of all three SMDs are small. This observation is not the primary evidence of randomization quality; what matters more is whether sequence generation, concealment, implementation logs, and post-assignment exclusions are auditable.

4 Superiority, Noninferiority, and Equivalence

4.1 Establish the Scientific Framework Before Examining Confidence Intervals

Suppose lower values of a continuous outcome are better, with effect θ=μT−μC\theta=\mu_T-\mu_C:

Objective Core null hypothesis 95% CI position supporting the conclusion Invalid shortcut
Superiority Treatment is not superior to Control The two-sided CI excludes 0 in the favorable direction Reporting only that “the Treatment arm changed significantly”
Noninferiority Treatment is worse than Control by at least Δ\Delta The one-sided 95% upper bound is below +Δ+\Delta “The difference is not significant, so Treatment is noninferior”
Equivalence The difference lies outside the acceptable interval The two-sided 90% CI lies entirely within [−Δ,+Δ][-\Delta,+\Delta] Treating a high p-value as evidence of equivalence

Before results are examined, the noninferiority margin must be justified jointly by the clinically acceptable loss and reliable historical evidence. The discussion must also address the active control’s assay sensitivity in the current trial, the constancy assumption supporting transportability of the historical effect, and biocreep caused by successively “relaxing the margin.” Noninferiority trials usually present FAS/ITT and PP results side by side; consistency strengthens interpretation but cannot automatically eliminate biases shared by both analyses.

  # Teaching demonstration: reuse the simulated superiority data to illustrate CI
  # decision logic; this does not constitute a real noninferiority design.
fit_ancova_preview <- lm(
  week12 ~ arm + baseline + site + severity,
  data = trial
)
ni_margin <- 2
ni_estimate <- coef(fit_ancova_preview)["armTreatment"]
ni_se <- sqrt(vcov(fit_ancova_preview)["armTreatment", "armTreatment"])
ni_upper <- ni_estimate + qt(0.95, df.residual(fit_ancova_preview)) * ni_se

ni_table <- data.frame(
  comparison = "Treatment − Control (lower scores are better)",
  estimate = ni_estimate,
  one_sided_95_upper = ni_upper,
  margin = ni_margin,
  noninferior = ni_upper < ni_margin
)
knitr::kable(ni_table, digits = 3,
             caption = "Demonstration of noninferiority CI logic; the +2 margin is for teaching purposes only")
Demonstration of noninferiority CI logic; the +2 margin is for teaching purposes only
comparison estimate one_sided_95_upper margin noninferior
armTreatment Treatment − Control (lower scores are better) -1.76 -0.775 2 TRUE

5 Common Designs and Their Units of Analysis

Design Appropriate setting Key analysis Common pitfalls
Individually randomized parallel-group Most acute and chronic conditions Compare groups as randomized and adjust for prespecified stratification factors Post-allocation exclusions, selective outcome switching
Cluster-randomized Substantial contamination or interventions delivered at the institutional level Use a cluster-correlated structure; account for the ICC and number of clusters in the sample size Treating individuals as independent, too few clusters
Crossover Stable condition, reversible effects, and a feasible washout Include period, sequence, and within-participant correlation Carryover effects, irreversible outcomes
Factorial Simultaneous evaluation of two interventions Prespecify main effects and interactions according to the estimand Interpreting only main effects despite an interaction
Adaptive/platform Prespecified adaptations to dose, sample size, or trial arms Simulate operating characteristics and control the type I error rate Ad hoc rules, operational bias, time trends

The simple design effect for a cluster-randomized design is DE=1+(m‾−1)ICCDE=1+(\bar m-1)ICC, but unequal cluster sizes, stratified matching, and a small number of clusters can make this formula inadequate. Further derivations for crossover, factorial, and adaptive designs are provided in Experimental Design; trials must also safeguard allocation concealment, participant safety, an interpretable estimand, and a regulatory audit trail.

Adaptive does not mean changing the trial as you go Adaptation rules, information timing, decision-makers, simulations, error rates, and data access must be prespecified. Current adaptive-design principles and the draft ICH E20 can be consulted; do not present a draft that has not yet been finally adopted as a standard already in effect.

6 Sample Size: Work Backward from the Primary Estimand

6.1 Sample size is not a software button

Before calculation, specify the scale of the primary outcome, target difference, variance or control-group risk, two-sided or one-sided α\alpha, power, allocation ratio, loss to follow-up, covariate efficiency, clustering or repeated-measures correlation, interim design, and multiplicity. The target difference should be clinically meaningful rather than simply the most optimistic point estimate from a pilot study.

For a continuous outcome with equal allocation to two groups, an approximation is:

nper arm≈2σ2(z1−α/2+z1−β)2δ2. n_{\text{per arm}}\approx \frac{2\sigma^2(z_{1-\alpha/2}+z_{1-\beta})^2}{\delta^2}.

A binary outcome requires risks for both groups, not just a relative effect; a time-to-event trial is driven primarily by the number of events; and a cluster-randomized trial also depends on the number of clusters and the ICC. Inflation for loss to follow-up is usually calculated as n/(1−d)n/(1-d), not n(1+d)n(1+d).

continuous_ss <- power.t.test(
  delta = 2.5, sd = 6, sig.level = 0.05, power = 0.80,
  type = "two.sample", alternative = "two.sided"
)
binary_ss <- power.prop.test(
  p1 = 0.35, p2 = 0.50, sig.level = 0.05, power = 0.80,
  alternative = "two.sided"
)

sample_size_table <- data.frame(
  endpoint = c("Continuous score: difference 2.5, SD 6", "Response: 35% vs 50%"),
  per_arm = c(ceiling(continuous_ss$n), ceiling(binary_ss$n)),
  per_arm_with_15pct_attrition = c(
    ceiling(continuous_ss$n / 0.85), ceiling(binary_ss$n / 0.85)
  )
)
knitr::kable(sample_size_table,
             caption = "Illustrative calculations with 80% power, two-sided alpha=0.05, and 1:1 allocation")
Illustrative calculations with 80% power, two-sided alpha=0.05, and 1:1 allocation
endpoint per_arm per_arm_with_15pct_attrition
Continuous score: difference 2.5, SD 6 92 108
Response: 35% vs 50% 170 200
scenario_grid <- expand.grid(
  target_difference = c(2.0, 2.5, 3.0),
  sd = c(5, 6),
  attrition = c(0.10, 0.20)
)
scenario_grid$per_arm <- mapply(function(delta, sd, loss) {
  base_n <- power.t.test(delta = delta, sd = sd, sig.level = 0.05,
                         power = 0.80, type = "two.sample")$n
  ceiling(base_n / (1 - loss))
}, scenario_grid$target_difference, scenario_grid$sd, scenario_grid$attrition)
knitr::kable(scenario_grid, caption = "Continuous-outcome sample-size scenarios as key assumptions vary")
Continuous-outcome sample-size scenarios as key assumptions vary
target_difference sd attrition per_arm
2.0 5 0.1 111
2.5 5 0.1 71
3.0 5 0.1 50
2.0 6 0.1 159
2.5 6 0.1 102
3.0 6 0.1 71
2.0 5 0.2 124
2.5 5 0.2 80
3.0 5 0.2 56
2.0 6 0.2 178
2.5 6 0.2 115
3.0 6 0.2 80

Do not report post hoc power After data have been observed, “observed power” is essentially a recoding of the p-value and adds nothing to the effect estimate and confidence interval. After a trial, report the actual information, effect, CI, number of events, and precision.

7 Analysis Sets, Participant Flow, and Data Auditing

7.1 Analysis sets must serve the estimand

Analysis set Typical definition Primary use and risks
ITT / Full analysis set Include all participants according to randomized assignment insofar as possible Preserves randomization; missing outcomes still require an appropriate method
Modified ITT Adds a post-randomization condition, such as receiving at least one dose May introduce selection bias if affected by prognosis or treatment
Per protocol Meets prespecified adherence criteria and has no major protocol deviations Not a randomized comparison; often used as complementary evidence in noninferiority trials
As treated Groups participants according to treatment actually received Vulnerable to confounding from treatment switching and adherence; does not estimate the randomized effect
Safety set Usually includes participants who received at least one study treatment and classifies them by actual exposure Denominators, exposure time, and treatment classification must be explicit

ITT is a randomized-assignment strategy, not “filling in the data with LOCF.” Estimating a strictly per-protocol effect usually requires addressing adherence, deviations, and post-randomization selection, rather than merely deleting “protocol violators.”

7.2 Minimum audit and flow table

Before modeling, verify unique IDs, nonmissing group assignments, temporal order, variable ranges, primary-outcome derivation, duplicate records, denominators, exposure, protocol deviations, treatment discontinuations, and reasons for loss to follow-up. The table below distinguishes “did not complete treatment,” “outcome not measured,” and “not included in the analysis”; an actual report also requires a CONSORT flow diagram.

flow_table <- data.frame(
  stage = c("Randomized", "Week 12 outcome observed", "Week 12 outcome missing",
            "Included in NRI response analysis", "Included in safety analysis", "One-year survival outcome available"),
  Control = c(
    sum(trial$arm == "Control"),
    sum(trial$arm == "Control" & !is.na(trial$week12)),
    sum(trial$arm == "Control" & is.na(trial$week12)),
    sum(trial$arm == "Control" & !is.na(trial$responder_nri)),
    sum(trial$arm == "Control" & !is.na(trial$adverse_event)),
    sum(trial$arm == "Control" & !is.na(trial$time_days))
  ),
  Treatment = c(
    sum(trial$arm == "Treatment"),
    sum(trial$arm == "Treatment" & !is.na(trial$week12)),
    sum(trial$arm == "Treatment" & is.na(trial$week12)),
    sum(trial$arm == "Treatment" & !is.na(trial$responder_nri)),
    sum(trial$arm == "Treatment" & !is.na(trial$adverse_event)),
    sum(trial$arm == "Treatment" & !is.na(trial$time_days))
  )
)
knitr::kable(flow_table, caption = "Analysis denominators and follow-up flow in the simulated trial")
Analysis denominators and follow-up flow in the simulated trial
stage Control Treatment
Randomized 178 182
Week 12 outcome observed 158 157
Week 12 outcome missing 20 25
Included in NRI response analysis 178 182
Included in safety analysis 178 182
One-year survival outcome available 178 182

8 Outcome Type Determines the Effect Measure and Model

Outcome Preferred description Common primary model Report at minimum
Continuous, single time point Group-specific mean/SD and distribution ANCOVA: follow-up value ~ group + baseline + stratification factors Adjusted mean difference, 95% CI, units
Repeated continuous n and mean/SD at each visit MMRM or an appropriate mixed model/GEE Difference at each time point, CI, covariance structure, and missing-data assumptions
Binary Number of events/risk in each group Binomial model, standardized risk; exact/penalized methods when sparse Risk, RD, and RR; label the OR explicitly
Count/rate Number of events and observed person-time Poisson/negative binomial, with log person-time as the offset Group-specific rates, IRR, and an overdispersion check
Ordinal Proportion in each category Proportional-odds model or a prespecified alternative Full distribution and common OR, with an assumption check
Time-to-event KM, survival probabilities, and numbers at risk Log-rank and Cox; alternatives such as RMST under non-PH Number of events, HR/CI, and absolute risk at fixed time points

“Normal/non-normal” is not the only switch for selecting a method. First align the estimand, unit of randomization, outcome scale, follow-up structure, and missingness mechanism; then assess model assumptions and evaluate robustness with prespecified alternative models.

9 Continuous Primary Outcome: ANCOVA

9.1 Why baseline adjustment is usually appropriate

For a continuous outcome with a baseline measurement, analyzing the Week 12 value with adjustment for baseline is usually more efficient than comparing change scores alone and can also address chance baseline imbalances. Prognostic baseline covariates can improve precision; the FDA’s 2023 guidance on covariate adjustment emphasizes prespecifying their use in randomized trials. Randomization stratification factors should generally also be included in the primary model.

fit_ancova <- lm(
  week12 ~ arm + baseline + site + severity,
  data = trial
)
ancova_row <- coef(summary(fit_ancova))["armTreatment", ]
ancova_ci <- confint(fit_ancova, "armTreatment")

change_test <- with(
  subset(trial, !is.na(change)),
  t.test(change[arm == "Treatment"], change[arm == "Control"])
)

continuous_results <- data.frame(
  analysis = c("ANCOVA: Week 12 value + baseline/stratification factors",
               "Unadjusted Welch comparison of change scores"),
  estimate = c(ancova_row["Estimate"], unname(change_test$estimate[1] -
                                                change_test$estimate[2])),
  lower = c(ancova_ci[1], change_test$conf.int[1]),
  upper = c(ancova_ci[2], change_test$conf.int[2]),
  p_value = c(ancova_row["Pr(>|t|)"], change_test$p.value)
)
continuous_results$effect_95_CI <- with(
  continuous_results, fmt_ci(estimate, lower, upper)
)
continuous_results$p_value <- vapply(continuous_results$p_value, format_p,
                                     character(1))
knitr::kable(continuous_results[c("analysis", "effect_95_CI", "p_value")],
             caption = "Week 12 continuous outcome: Treatment - Control; negative values favor Treatment")
Week 12 continuous outcome: Treatment - Control; negative values favor Treatment
analysis effect_95_CI p_value
Estimate ANCOVA: Week 12 value + baseline/stratification factors -1.76 (-2.93, -0.59) 0.00342
Unadjusted Welch comparison of change scores -1.72 (-3.00, -0.44) 0.00863

The available-case ANCOVA estimates an adjusted mean difference of approximately -1.76 points (95% CI -2.93 to -0.59). This is a useful modeling demonstration, but because the Week 12 outcome was not observed for every participant, it cannot be called a complete ITT analysis solely because participants were modeled according to randomized group. The conclusion also depends on the missingness mechanism; MI and delta sensitivity analyses are used below.

old_par <- par(mfrow = c(1, 2), mar = c(4.2, 4.2, 2, 1))
plot(fitted(fit_ancova), residuals(fit_ancova),
     pch = 19, col = grDevices::adjustcolor(trial_palette["treatment"], 0.55),
     xlab = "Fitted values", ylab = "Residuals", main = "Variance and functional form")
abline(h = 0, lty = 2, col = trial_palette["harm"])
qqnorm(residuals(fit_ancova), pch = 19,
       col = grDevices::adjustcolor(trial_palette["navy"], 0.55),
       main = "Residual Q-Q plot")
qqline(residuals(fit_ancova), col = trial_palette["harm"], lwd = 2)
Two diagnostic plots showing ANCOVA residuals versus fitted values and a normal Q-Q plot of the residuals.

ANCOVA residual diagnostics: residuals versus fitted values on the left and a normal Q-Q plot on the right.

par(old_par)

When diagnostics reveal problems, consider robust standard errors, transformation, an appropriate generalized model, or a prespecified robust estimator rather than repeatedly deleting outliers in response to the results. The treatment coefficient is the adjusted between-group mean difference, not the within-participant change for each participant.

10 Repeated Continuous Outcomes: A Compact Application of MMRM

Repeated follow-up can describe the timing of onset and use partially observed records. A common MMRM treats post-baseline visit as a categorical variable, includes treatment × visit, and specifies a flexible covariance structure for residuals from the same participant; it does not require missing values to be filled in with LOCF first. Under the MAR assumption, the likelihood uses the available observations, but MAR still cannot be “tested and shown to be true” from the data.

long_trial <- reshape(
  trial[c("id", "arm", "baseline", "site", "severity",
          "week4", "week8", "week12")],
  varying = c("week4", "week8", "week12"),
  v.names = "score", timevar = "visit_index",
  times = 1:3, direction = "long"
)
long_trial$visit <- factor(
  long_trial$visit_index, levels = 1:3,
  labels = c("Week 4", "Week 8", "Week 12")
)
long_trial <- long_trial[order(long_trial$id, long_trial$visit_index), ]

fit_mmrm <- nlme::gls(
  score ~ baseline + site + severity + visit * arm,
  correlation = nlme::corSymm(form = ~ visit_index | id),
  weights = nlme::varIdent(form = ~ 1 | visit),
  data = long_trial, method = "REML", na.action = na.omit,
  control = nlme::glsControl(opt = "optim")
)

mmrm_beta <- coef(fit_mmrm)
mmrm_vcov <- vcov(fit_mmrm)
contrast_matrix <- matrix(0, 3, length(mmrm_beta),
                          dimnames = list(levels(long_trial$visit),
                                          names(mmrm_beta)))
contrast_matrix[, "armTreatment"] <- 1
contrast_matrix["Week 8", "visitWeek 8:armTreatment"] <- 1
contrast_matrix["Week 12", "visitWeek 12:armTreatment"] <- 1
mmrm_estimate <- drop(contrast_matrix %*% mmrm_beta)
mmrm_se <- sqrt(diag(contrast_matrix %*% mmrm_vcov %*% t(contrast_matrix)))
  # A large-sample normal approximation is used here; a small-sample degrees-of-freedom method should be prespecified for confirmatory analysis.
mmrm_table <- data.frame(
  visit = rownames(contrast_matrix),
  estimate = mmrm_estimate,
  lower = mmrm_estimate - qnorm(0.975) * mmrm_se,
  upper = mmrm_estimate + qnorm(0.975) * mmrm_se
)
mmrm_table$effect_95_CI <- with(mmrm_table, fmt_ci(estimate, lower, upper))
knitr::kable(mmrm_table[c("visit", "effect_95_CI")],
             caption = "MMRM-adjusted mean difference at each visit: Treatment - Control; negative values favor Treatment")
MMRM-adjusted mean difference at each visit: Treatment - Control; negative values favor Treatment
visit effect_95_CI
Week 4 Week 4 -1.57 (-2.57, -0.57)
Week 8 Week 8 -2.15 (-3.22, -1.08)
Week 12 Week 12 -1.75 (-2.92, -0.58)

The covariance structure, degrees-of-freedom approximation, and numerical convergence criteria should be specified in the SAP. If the primary estimand is the mean difference at Week 12, the other visits can describe the trajectory, but a smaller p-value at another time point cannot be used to change the primary conclusion. More complete treatments of random effects, GEE, and missing-data structures are provided in Longitudinal Data Analysis.

11 Binary Outcomes: Present Absolute Risk and Relative Effects Together

11.1 RD, RR, and OR answer different questions

The risk difference (RD) directly gives the number of additional or fewer events per 100 participants, the risk ratio (RR) gives the proportional change, and the odds ratio (OR) arises from a logistic model but may differ substantially from the RR when events are common. The OR is also non-collapsible: adjusted and unadjusted ORs can differ even without confounding, so this difference should not automatically be interpreted as bias having been “corrected.”

binary_effects <- function(y, arm, conf.level = 0.95) {
  keep <- !is.na(y) & !is.na(arm)
  y <- as.integer(y[keep])
  arm <- droplevels(arm[keep])
  stopifnot(nlevels(arm) == 2L, all(y %in% 0:1))
  z <- qnorm(1 - (1 - conf.level) / 2)
  n0 <- sum(arm == levels(arm)[1]); n1 <- sum(arm == levels(arm)[2])
  e0 <- sum(y[arm == levels(arm)[1]]); e1 <- sum(y[arm == levels(arm)[2]])
  p0 <- e0 / n0; p1 <- e1 / n1
  rd <- p1 - p0
  se_rd <- sqrt(p1 * (1 - p1) / n1 + p0 * (1 - p0) / n0)
  cells <- c(e1, n1 - e1, e0, n0 - e0)
  if (any(cells == 0)) cells <- cells + 0.5
  p1_cc <- cells[1] / sum(cells[1:2])
  p0_cc <- cells[3] / sum(cells[3:4])
  rr <- p1_cc / p0_cc
  se_log_rr <- sqrt(1 / cells[1] - 1 / sum(cells[1:2]) +
                    1 / cells[3] - 1 / sum(cells[3:4]))
  odds_ratio <- (cells[1] * cells[4]) / (cells[2] * cells[3])
  se_log_or <- sqrt(sum(1 / cells))
  data.frame(
    measure = c("Risk difference", "Risk ratio", "Odds ratio"),
    estimate = c(rd, rr, odds_ratio),
    lower = c(rd - z * se_rd, exp(log(rr) - z * se_log_rr),
              exp(log(odds_ratio) - z * se_log_or)),
    upper = c(rd + z * se_rd, exp(log(rr) + z * se_log_rr),
              exp(log(odds_ratio) + z * se_log_or)),
    control_risk = p0, treatment_risk = p1
  )
}

responder_effects <- binary_effects(trial$responder_nri, trial$arm)
responder_effects$effect_95_CI <- with(
  responder_effects, fmt_ci(estimate, lower, upper, digits = 3)
)
knitr::kable(
  responder_effects[c("measure", "effect_95_CI", "control_risk",
                       "treatment_risk")], digits = 3,
  caption = "Week 12 response effects with missing responses treated as nonresponse"
)
Week 12 response effects with missing responses treated as nonresponse
measure effect_95_CI control_risk treatment_risk
Risk difference 0.108 (0.009, 0.207) 0.315 0.423
Risk ratio 1.345 (1.021, 1.771) 0.315 0.423
Odds ratio 1.598 (1.037, 2.461) 0.315 0.423
fit_responder <- glm(
  responder_nri ~ arm + baseline + site + severity,
  family = binomial(), data = trial
)
adjusted_or <- exp(c(
  estimate = coef(fit_responder)["armTreatment"],
  confint.default(fit_responder, "armTreatment")
))
data.frame(
  measure = "Adjusted OR",
  estimate = adjusted_or[1], lower = adjusted_or[2], upper = adjusted_or[3]
) |>
  knitr::kable(digits = 3, caption = "Logistic model adjusted for prespecified covariates")
Logistic model adjusted for prespecified covariates
measure estimate lower upper
estimate.armTreatment Adjusted OR 1.6 1.04 2.48

Treating missing responses as nonresponse (NRI) combines missingness and failure and is an example of a composite strategy; it is not universally “conservative.” If missingness rates or reasons for missingness differ between groups, the direction can be complex. A formal analysis should also present the original numerator and denominator for each group, absolute effects, and sensitivity analyses. For sparse events or complete separation, consider exact or penalized methods rather than relying on an ordinary Wald CI.

12 Counts and Incidence Rates: Include Observation Time in the Offset

fit_poisson <- glm(
  episode_count ~ arm + baseline + site + severity +
    offset(log(exposure_weeks)),
  family = poisson(), data = trial
)
dispersion <- sum(residuals(fit_poisson, type = "pearson")^2) /
  df.residual(fit_poisson)
fit_quasipoisson <- update(fit_poisson, family = quasipoisson())
count_row <- coef(summary(fit_quasipoisson))["armTreatment", ]
irr <- exp(c(
  estimate = count_row["Estimate"],
  lower = count_row["Estimate"] - qnorm(0.975) * count_row["Std. Error"],
  upper = count_row["Estimate"] + qnorm(0.975) * count_row["Std. Error"]
))
rate_table <- aggregate(
  cbind(events = episode_count, person_weeks = exposure_weeks) ~ arm,
  data = trial, FUN = sum
)
rate_table$rate_per_100_person_weeks <-
  100 * rate_table$events / rate_table$person_weeks
knitr::kable(rate_table, digits = 2,
             caption = "Acute episodes, observed person-weeks, and crude incidence rates by group")
Acute episodes, observed person-weeks, and crude incidence rates by group
arm events person_weeks rate_per_100_person_weeks
Control 271 1784 15.2
Treatment 207 1806 11.5
knitr::kable(data.frame(
  IRR = irr[1], lower = irr[2], upper = irr[3],
  Pearson_dispersion = dispersion
), digits = 3, caption = "Treatment/Control IRR from the quasi-Poisson model")
Treatment/Control IRR from the quasi-Poisson model
IRR lower upper Pearson_dispersion
estimate.Estimate 0.746 0.587 0.948 1.75

The offset coefficient is fixed at 1, converting differing observation times into rates. When Pearson dispersion is substantially greater than 1, ordinary Poisson standard errors are too small; quasi-Poisson adjusts precision without changing the point estimate, whereas a negative binomial model also changes the mean-variance relationship. If recurrent events involve within-participant correlation, terminal events, or event dependence, GEE, frailty models, or dedicated recurrent-event models should also be considered.

13 Time-to-Event Outcomes

13.1 KM, log-rank, and Cox each have a role

Kaplan-Meier estimates survival probabilities over time, the log-rank test assesses overall differences between curves, and the Cox model estimates a conditional hazard ratio. An HR is neither a risk ratio nor evidence that “time to recurrence increased by 34%.” Censoring must be defensibly independent conditional on the covariates and model assumptions.

survival_outcome <- with(trial, survival::Surv(time_days, status))
fit_km <- survival::survfit(survival_outcome ~ arm, data = trial)
fit_logrank <- survival::survdiff(survival_outcome ~ arm, data = trial)
logrank_p <- pchisq(fit_logrank$chisq,
                    df = length(fit_logrank$n) - 1, lower.tail = FALSE)
fit_cox <- survival::coxph(
  survival_outcome ~ arm + baseline + site + severity,
  data = trial
)
cox_row <- summary(fit_cox)$coefficients["armTreatment", ]
cox_ci <- exp(c(coef(fit_cox)["armTreatment"],
                confint(fit_cox, "armTreatment")))

km_summary <- summary(fit_km, times = c(90, 180, 365), extend = TRUE)
km_table <- data.frame(
  arm = sub("^arm=", "", km_summary$strata),
  day = km_summary$time,
  survival = km_summary$surv,
  lower = km_summary$lower,
  upper = km_summary$upper
)
knitr::kable(km_table, digits = 3,
             caption = "KM recurrence-free survival probabilities at Days 90, 180, and 365")
KM recurrence-free survival probabilities at Days 90, 180, and 365
arm day survival lower upper
Control 90 0.857 0.806 0.910
Control 180 0.696 0.630 0.769
Control 365 0.485 0.413 0.568
Treatment 90 0.892 0.848 0.939
Treatment 180 0.782 0.721 0.847
Treatment 365 0.615 0.543 0.696
knitr::kable(data.frame(
  analysis = c("Log-rank", "Adjusted Cox HR"),
  estimate = c(NA, cox_ci[1]),
  lower = c(NA, cox_ci[2]), upper = c(NA, cox_ci[3]),
  p_value = c(logrank_p, cox_row["Pr(>|z|)"])
), digits = 3, caption = "Between-group comparison of recurrence/progression within one year")
Between-group comparison of recurrence/progression within one year
analysis estimate lower upper p_value
Log-rank NA NA NA 0.014
armTreatment Adjusted Cox HR 0.658 0.474 0.913 0.012
plot(fit_km, col = c(trial_palette["control"], trial_palette["treatment"]),
     lwd = 2, mark.time = TRUE, conf.int = FALSE,
     xlab = "Days since randomization", ylab = "Recurrence-free survival probability", ylim = c(0.35, 1))
legend("bottomleft", legend = c("Control", "Treatment"),
       col = c(trial_palette["control"], trial_palette["treatment"]),
       lwd = 2, bty = "n")
Kaplan-Meier recurrence-free survival curves for Control and Treatment, with the Treatment curve generally higher.

One-year Kaplan-Meier recurrence-free survival curves for the simulated trial; short vertical marks indicate censoring.

ph_check <- survival::cox.zph(fit_cox)
knitr::kable(as.data.frame(ph_check$table), digits = 3,
             caption = "Schoenfeld-residual test of proportional hazards; with low power, do not rely on the p-value alone")
Schoenfeld-residual test of proportional hazards; with low power, do not rely on the p-value alone
chisq df p
arm 0.000 1 0.998
baseline 0.252 1 0.616
site 4.198 3 0.241
severity 0.688 1 0.407
GLOBAL 4.885 6 0.559

When proportional hazards is implausible, prespecified survival differences at fixed times, restricted mean survival time (RMST), piecewise effects, or time-varying coefficients are often more direct. When competing events are present, distinguish cause-specific hazards from cumulative incidence risks; treating competing events as ordinary censoring does not necessarily answer the target question. See Survival Analysis for details.

14 Missing Data and Sensitivity Analyses

14.1 Prevent First, Model Second, Then Challenge the Assumptions

The first priority is to minimize missing data: even after treatment discontinuation, make every effort to continue collecting outcomes relevant to the estimand, together with the timing of and reasons for intercurrent events. Complete-case analysis is unbiased only under strong conditions; LOCF generally distorts trajectories and variances; and “MAR” is not an imputation algorithm, but an assumption that, conditional on observed information, missingness no longer depends on unobserved values.

The EMA guideline on missing data emphasizes prevention, assumptions for the primary analysis, and sensitivity analyses. The MI model should include predictors of the outcome, predictors of missingness, treatment group, stratification variables, and important structures from the analysis model, with results combined using Rubin’s rules; single mean imputation spuriously increases precision.

14.2 MAR Multiple Imputation and a Delta-Adjusted Pattern-Mixture Model

The analysis below first fits an imputation model within each group using observed Week 12 outcomes. It then uniformly increases imputed values in the treatment group by δ\delta points, representing the assumption that, given the same observed information, participants with missing data in the treatment group have worse outcomes than predicted under MAR. The same random numbers are used for each value of δ\delta (common random numbers), so that changes primarily reflect the assumption rather than Monte Carlo noise.

mi_ancova <- function(data, delta = 0, m = 100L, seed = 41000L) {
  set.seed(seed)
  estimates <- variances <- numeric(m)
  for (j in seq_len(m)) {
    completed <- data
    for (group in levels(completed$arm)) {
      observed <- completed$arm == group & !is.na(completed$week12)
      missing <- completed$arm == group & is.na(completed$week12)
      if (!any(missing)) next
      imputation_fit <- lm(
        week12 ~ baseline + site + severity + sex,
        data = completed, subset = observed
      )
      beta_hat <- coef(imputation_fit)
      beta_vcov <- vcov(imputation_fit)
      beta_draw <- drop(beta_hat +
                          t(chol(beta_vcov)) %*% rnorm(length(beta_hat)))
      x_missing <- model.matrix(
        ~ baseline + site + severity + sex,
        data = completed[missing, , drop = FALSE]
      )
      imputed <- drop(x_missing %*% beta_draw) +
        rnorm(sum(missing), 0, summary(imputation_fit)$sigma)
      if (group == "Treatment") imputed <- imputed + delta
      completed$week12[missing] <- imputed
    }
    analysis_fit <- lm(
      week12 ~ arm + baseline + site + severity,
      data = completed
    )
    estimates[j] <- coef(analysis_fit)["armTreatment"]
    variances[j] <- vcov(analysis_fit)["armTreatment", "armTreatment"]
  }
  q_bar <- mean(estimates)
  u_bar <- mean(variances)
  between <- var(estimates)
  total <- u_bar + (1 + 1 / m) * between
  df <- if (between < .Machine$double.eps) Inf else
    (m - 1) * (1 + u_bar / ((1 + 1 / m) * between))^2
  critical <- qt(0.975, df)
  data.frame(
    delta = delta, estimate = q_bar, se = sqrt(total), df = df,
    lower = q_bar - critical * sqrt(total),
    upper = q_bar + critical * sqrt(total),
    p_value = 2 * pt(-abs(q_bar / sqrt(total)), df)
  )
}

delta_grid <- seq(0, 8, by = 0.5)
tipping <- do.call(rbind, lapply(delta_grid, function(delta_i) {
  mi_ancova(trial, delta = delta_i, m = 100L, seed = 41000L)
}))

selected_tipping <- tipping[tipping$delta %in% c(0, 2, 4, 5, 5.5, 6, 8), ]
selected_tipping$effect_95_CI <- with(
  selected_tipping, fmt_ci(estimate, lower, upper, digits = 3)
)
selected_tipping$p_value <- vapply(selected_tipping$p_value, format_p,
                                   character(1))
knitr::kable(selected_tipping[c("delta", "effect_95_CI", "p_value")],
             caption = "MI sensitivity analysis after worsening missing values in the treatment group by delta")
MI sensitivity analysis after worsening missing values in the treatment group by delta
delta effect_95_CI p_value
1 0.0 -1.877 (-3.047, -0.708) 0.00166
5 2.0 -1.605 (-2.777, -0.434) 0.00726
9 4.0 -1.333 (-2.516, -0.151) 0.0271
11 5.0 -1.197 (-2.389, -0.006) 0.0489
12 5.5 -1.129 (-2.326, 0.067) 0.0643
13 6.0 -1.061 (-2.263, 0.141) 0.0835
17 8.0 -0.789 (-2.019, 0.440) 0.208
plot(tipping$delta, tipping$estimate, type = "b", pch = 19,
     col = trial_palette["treatment"], lwd = 2,
     ylim = range(tipping$lower, tipping$upper, 0),
     xlab = expression(paste("Worsening adjustment for missing values in the treatment group ", delta, " (points)")),
     ylab = "Treatment − Control adjusted mean difference")
segments(tipping$delta, tipping$lower, tipping$delta, tipping$upper,
         col = grDevices::adjustcolor(trial_palette["treatment"], 0.55))
abline(h = 0, lty = 2, col = trial_palette["harm"])
The treatment effect and its 95% confidence interval move toward zero as missing outcomes in the treatment group are worsened by delta; the confidence interval crosses zero at a delta of approximately 5.5.

Delta-adjusted pattern-mixture sensitivity curve; the horizontal line marks the null value of 0.

Under MAR (δ=0\delta=0), the estimated effect is approximately -1.88 points. The two-sided 95% CI first includes 0 only after each missing outcome in the treatment group is worsened by an average of approximately 5.5 points. The meaning of the tipping point depends on whether 5.5 points is clinically plausible; it is not a basis for simply declaring the result “robust.” A real trial should also align delta values with reasons for treatment discontinuation, visit timing, and expert knowledge, and should assess the compatibility of the imputation model with the primary analysis.

15 Multiplicity and Interim Analyses

15.1 Multiple Endpoints Are Not Three Independent Stories

Multiple primary endpoints, multiple doses, multiple time points, and repeated looks at the data all increase false positives. Fixed-sequence/hierarchical testing, gatekeeping, Holm, or Bonferroni procedures control the familywise error rate (FWER); FDR is more commonly used for exploratory settings with many hypotheses. The multiplicity adjustment plan should define hypothesis families in the SAP rather than mechanically placing every p value in the same list.

p_continuous <- coef(summary(fit_ancova))["armTreatment", "Pr(>|t|)"]
p_responder <- coef(summary(fit_responder))["armTreatment", "Pr(>|z|)"]
p_survival <- summary(fit_cox)$coefficients["armTreatment", "Pr(>|z|)"]
multiplicity <- data.frame(
  endpoint = c("Week 12 continuous score", "Week 12 response", "One-year recurrence/progression"),
  raw_p = c(p_continuous, p_responder, p_survival)
)
multiplicity$holm_p <- p.adjust(multiplicity$raw_p, method = "holm")
multiplicity$reject_at_0.05 <- multiplicity$holm_p < 0.05
knitr::kable(multiplicity, digits = 4,
             caption = "Example of Holm adjustment when three hypotheses are prespecified as one family")
Example of Holm adjustment when three hypotheses are prespecified as one family
endpoint raw_p holm_p reject_at_0.05
Week 12 continuous score 0.0034 0.0103 TRUE
Week 12 response 0.0327 0.0327 TRUE
One-year recurrence/progression 0.0124 0.0248 TRUE

15.2 Interim Looks Must Spend Alpha

An independent DMC reviews efficacy, futility, and safety in accordance with its charter; the sponsor’s operational team generally should not see unblinded trends. The information fraction should, whenever possible, be based on the number of events or estimated information rather than calendar time alone. The example below only converts cumulative alpha to an “equivalent two-sided z threshold” to illustrate why early thresholds are stringent; it is not an exact sequential boundary that accounts for correlations among tests at different looks.

alpha <- 0.05
information_fraction <- c(0.33, 0.67, 1.00)
z_alpha <- qnorm(1 - alpha / 2)
spend_obf <- 2 - 2 * pnorm(z_alpha / sqrt(information_fraction))
spend_pocock <- alpha * log(1 + (exp(1) - 1) * information_fraction)
spending <- rbind(
  data.frame(method = "O'Brien–Fleming-like", look = 1:3,
             information_fraction = information_fraction,
             cumulative_alpha = spend_obf),
  data.frame(method = "Pocock-like", look = 1:3,
             information_fraction = information_fraction,
             cumulative_alpha = spend_pocock)
)
spending$incremental_alpha <- ave(
  spending$cumulative_alpha, spending$method,
  FUN = function(x) c(x[1], diff(x))
)
spending$cumulative_z_equivalent <-
  qnorm(1 - spending$cumulative_alpha / 2)
knitr::kable(spending, digits = 4,
             caption = "Illustrative cumulative alpha spending; not an exact correlated sequential boundary")
Illustrative cumulative alpha spending; not an exact correlated sequential boundary
method look information_fraction cumulative_alpha incremental_alpha cumulative_z_equivalent
O’Brien–Fleming-like 1 0.33 0.0006 0.0006 3.41
O’Brien–Fleming-like 2 0.67 0.0166 0.0160 2.39
O’Brien–Fleming-like 3 1.00 0.0500 0.0334 1.96
Pocock-like 1 0.33 0.0225 0.0225 2.28
Pocock-like 2 0.67 0.0383 0.0158 2.07
Pocock-like 3 1.00 0.0500 0.0117 1.96

Confirmatory designs should use validated software to calculate correlated boundaries, conditional power, and operating characteristics, and should prespecify nonbinding/binding futility boundaries, sample-size re-estimation, and the handling of overrunning data. Action may be taken at any time for safety reasons, but early stopping for efficacy still requires adjusted estimation and reporting of the information time.

16 Subgroup Analyses: Test Interactions Directly

fit_interaction <- lm(
  week12 ~ arm * sex + baseline + site + severity,
  data = trial
)
interaction_beta <- coef(fit_interaction)
interaction_vcov <- vcov(fit_interaction)
interaction_df <- df.residual(fit_interaction)
main_term <- "armTreatment"
interaction_term <- "armTreatment:sexMale"

female_est <- interaction_beta[main_term]
female_se <- sqrt(interaction_vcov[main_term, main_term])
male_est <- interaction_beta[main_term] + interaction_beta[interaction_term]
male_se <- sqrt(interaction_vcov[main_term, main_term] +
                  interaction_vcov[interaction_term, interaction_term] +
                  2 * interaction_vcov[main_term, interaction_term])
critical <- qt(0.975, interaction_df)
subgroup <- data.frame(
  subgroup = c("Female", "Male"),
  estimate = c(female_est, male_est),
  lower = c(female_est - critical * female_se,
            male_est - critical * male_se),
  upper = c(female_est + critical * female_se,
            male_est + critical * male_se)
)
interaction_p <- coef(summary(fit_interaction))[
  interaction_term, "Pr(>|t|)"
]
subgroup$effect_95_CI <- with(subgroup, fmt_ci(estimate, lower, upper))
knitr::kable(subgroup[c("subgroup", "effect_95_CI")],
             caption = paste0("Exploratory sex subgroups; interaction p = ", format_p(interaction_p)))
Exploratory sex subgroups; interaction p = 0.432
subgroup effect_95_CI
Female -1.37 (-2.91, 0.17)
Male -2.33 (-4.16, -0.50)
plot(subgroup$estimate, seq_len(nrow(subgroup)),
     xlim = range(subgroup$lower, subgroup$upper, 0),
     ylim = c(0.5, nrow(subgroup) + 0.5), yaxt = "n",
     pch = 19, col = trial_palette["treatment"],
     xlab = "Adjusted mean difference (Treatment − Control)", ylab = "")
segments(subgroup$lower, seq_len(nrow(subgroup)),
         subgroup$upper, seq_len(nrow(subgroup)),
         col = trial_palette["treatment"], lwd = 2)
axis(2, at = seq_len(nrow(subgroup)), labels = subgroup$subgroup, las = 1)
abline(v = 0, lty = 2, col = trial_palette["harm"])
Treatment minus Control adjusted mean differences and 95% confidence intervals for the Female and Male subgroups; one interval crosses zero and the other does not, but the interaction test is not significant.

Adjusted mean differences for exploratory sex subgroups; the dashed line marks the null value of 0.

The CI includes 0 for Female participants but not for Male participants, whereas the p value for the treatment × sex interaction is approximately 0.43: “significant in one group but not in the other” does not mean that the difference between groups is significant. Subgroups should be few and prespecified, interactions should be tested directly, continuous effect modifiers should preferably remain continuous, and findings should be interpreted in light of direction, credibility, multiplicity, and biological mechanisms. The EMA guideline on subgroups provides further reading.

17 Safety Analyses and Benefit–Risk

ae_effects <- binary_effects(trial$adverse_event, trial$arm)
ae_rr <- ae_effects[ae_effects$measure == "Risk ratio", ]
ae_fisher <- fisher.test(table(trial$arm, trial$adverse_event))
ae_summary <- data.frame(
  arm = levels(trial$arm),
  participants = as.integer(table(trial$arm)),
  with_event = as.integer(tapply(trial$adverse_event, trial$arm, sum)),
  risk = as.numeric(tapply(trial$adverse_event, trial$arm, mean))
)
knitr::kable(ae_summary, digits = 3,
             caption = "Participants with at least one simulated adverse event")
Participants with at least one simulated adverse event
arm participants with_event risk
Control 178 20 0.112
Treatment 182 38 0.209
knitr::kable(data.frame(
  measure = "Treatment/Control RR",
  estimate = ae_rr$estimate, lower = ae_rr$lower, upper = ae_rr$upper,
  Fisher_p = ae_fisher$p.value
), digits = 3, caption = "Relative risk of adverse events and exact test")
Relative risk of adverse events and exact test
measure estimate lower upper Fisher_p
Treatment/Control RR 1.86 1.13 3.06 0.015

Safety analyses should describe the safety set, actual exposure, risk window, event severity, serious adverse events (SAEs), treatment discontinuations, deaths, and causality assessments. Common events may be reported using risk differences or risk ratios, with rates reported when exposure times differ; rare serious harms are generally evaluated using CIs, cumulative case counts, narrative review, and evidence across sources. Performing an unadjusted p-value test for each of dozens of AEs and selecting “significant findings” both lacks power and misleads benefit–risk assessment.

18 From the SAP to CONSORT 2025 Reporting

18.1 A Minimum Checklist for an Actionable SAP

  • Version, date, blinding status, and handling of deviations;
  • The five attributes of each primary/key secondary estimand;
  • Data snapshot, analysis sets, variable derivations, visit windows, and baseline definition;
  • Randomization stratification factors, covariates, models, contrasts, direction of effect, and CIs;
  • Intercurrent events, primary assumptions about missing data, sensitivity analyses, and supplementary analyses;
  • Multiplicity, interim analyses, subgroups, safety, model failure, and software versions;
  • Prespecified tables and figures, with traceable quality control/independent review.

18.2 Report Effects, Not Only p Values

An adequate results sentence could read:

Among simulated participants with a Week 12 measurement, the symptom score was an average of 1.76 points lower with Treatment than with Control after adjustment for baseline score, center, and baseline severity and according to randomized group (95% CI, 0.59 to 2.93 points lower; two-sided p=0.003). This available-case estimate depends on the missingness mechanism; MAR MI and prespecified MNAR delta analyses were used to assess robustness.

The results section should also provide the denominator and descriptive values for each group, actual follow-up, number of events, absolute effects, the model, and the direction of effect. Statistical significance is not equivalent to clinical importance; the CI simultaneously expresses the range of compatible benefits and harms.

CONSORT 2025 has superseded CONSORT 2010. At its core are a 30-item checklist and a participant flow diagram, with greater emphasis on registration; access to the protocol and SAP; data and code sharing; patient and public involvement; harms; actual delivery of interventions; analysis sets; and missing data. Reporting guidelines improve transparency, but they cannot repair a poor design.

19 Common Errors and Corrections

Common error Why it is wrong Better approach
Test every baseline characteristic with a p value and adjust only if significant Random imbalances should not drive the model Prespecify based on prognostic value, stratification, and the SAP
Perform only within-group pre–post tests Within-group significance is not a between-group difference Estimate the between-group effect and CI directly
ITT = LOCF An analysis strategy is conflated with an imputation method Define the estimand first, then select a missing-data method
Declare “equivalence” when p>0.05 Lack of precision does not demonstrate similarity Use a prespecified margin and the appropriate one- or two-sided CI
One subgroup is significant and another is not The difference in effects has not been tested Test the treatment × subgroup interaction directly
Treat the HR as a risk ratio An instantaneous conditional effect differs from cumulative risk Also report KM risk at a fixed time point; assess PH
Examine p values weekly and stop when significant Type I error and estimation bias increase Prespecify boundaries, a DMC, and alpha spending
Report only the OR and p value The clinical meaning on the absolute scale is unclear Also report each group’s risk, RD/RR, and CIs
Exclude participants who discontinue treatment to obtain a “clean ITT” analysis This breaks randomization and selects the population Use FAS for the primary analysis; use PP as a prespecified supplementary analysis or apply causal methods
Search item by item for significant safety findings Power is low and multiplicity is substantial Describe denominators, exposure, severity, CIs, and cases

Knowledge Check and Exercises

  1. If Week 12 scores are still collected after treatment discontinuation, which estimand strategy does this represent?

If the score is used in the analysis regardless of treatment discontinuation, this generally corresponds to a treatment-policy strategy. Treatment discontinuation is an intercurrent event; whether the score is missing is a separate issue.

  1. Why are allocation concealment and blinding not interchangeable?

Allocation concealment prevents selective enrollment before randomization; blinding reduces the influence of knowledge of treatment assignment on behavior, co-interventions, and outcome assessment after randomization. They operate at different times and address different biases.

  1. Can noninferiority be concluded when p>0.05 in a noninferiority trial?

No. The one-sided confidence bound in the direction of harm must be below a prespecified margin with clinical and historical justification, and assay sensitivity, FAS/PP analyses, and deviations must be evaluated.

  1. Why can an ANCOVA result not automatically be described as a full ITT analysis?

The model compares participants according to randomized group but includes only those with a Week 12 value. Outcomes for participants with missing data still require a method aligned with the estimand and sensitivity analyses.

  1. Does OR=1.60 mean that the response probability increased by 60%?

No. The OR compares odds; in this example, the observed risks, RD, and RR should also be examined. When events are common, the OR and RR can differ substantially.

  1. Female is not significant and Male is significant. Does this prove that sex modifies the treatment effect?

No. The direct interaction test in this example gives p≈0.43; the interaction estimate and CI should be interpreted while accounting for the exploratory nature of the analysis and multiplicity.

  1. Is a delta tipping point a “pass/fail” threshold?

No. It shows how large an unobserved departure is needed to change the conclusion; the key question is whether that delta is plausible relative to clinical knowledge, reasons for missingness, and the range of the data.

  1. Why can the three p values from interim analyses not all be compared with 0.05?

Repeated opportunities increase the familywise type I error rate. Correlated sequential boundaries or alpha spending must be prespecified and implemented by an independent DMC in accordance with its charter.

  1. Design your own primary estimand.

In turn, specify the population, treatment conditions, variable, strategy for each intercurrent event, and population-level summary, and then describe the corresponding primary estimator and at least one key sensitivity analysis.

Quick Reference and Final Checklist

Before Analysis

  • The clinical question, direction of effect, primary time point, and MCID are clearly defined;
  • The five attributes of the estimand and the strategy for each intercurrent event are explicitly specified;
  • Randomization, allocation concealment, blinding, emergency unblinding, and audit responsibilities are operationally feasible;
  • The sample size accounts for realistic parameters, loss to follow-up, design effects, multiplicity, and the interim-analysis plan;
  • The protocol, registration, and SAP are finalized and version-controlled before unblinded results become available.

Data and Models

  • IDs, denominators, visits, exposure, events, deviations, and reasons for missingness have been verified;
  • The model matches the outcome scale, unit of randomization, stratification, and repeated-measures structure;
  • Every result includes an effect measure, 95% CI, units, direction, and descriptive values for each group;
  • The primary assumption about missing data is paired with at least one sensitivity analysis targeting an unverifiable assumption;
  • Rules for multiplicity, interim analyses, subgroups, and safety were not changed after seeing the results;
  • Model diagnostics, convergence, alternative models, and software versions have an audit trail.

Reporting and Interpretation

  • Report the participant flow, allocation, actual interventions received, numbers analyzed, and harms in accordance with CONSORT 2025;
  • Report the protocol/SAP/registration number, amendments, code, and paths to shareable data;
  • Distinguish statistical significance, clinical importance, imprecision, and absence of evidence;
  • Do not treat the OR as the RR, the HR as a cumulative risk ratio, or a high p value as equivalence;
  • Limit conclusions to the prespecified population, interventions, comparator, outcome, time point, and implementation setting.
sessionInfo()
## R version 4.6.1 (2026-06-24)
## Platform: aarch64-apple-darwin23
## Running under: macOS Tahoe 26.5.1
## 
## Matrix products: default
## BLAS:   /Library/Frameworks/R.framework/Versions/4.6/Resources/lib/libRblas.0.dylib 
## LAPACK: /Library/Frameworks/R.framework/Versions/4.6/Resources/lib/libRlapack.dylib;  LAPACK version 3.12.1
## 
## locale:
## [1] C.UTF-8/C.UTF-8/C.UTF-8/C/C.UTF-8/C.UTF-8
## 
## time zone: America/Edmonton
## tzcode source: internal
## 
## attached base packages:
## [1] stats     graphics  grDevices utils     datasets  methods   base     
## 
## loaded via a namespace (and not attached):
##  [1] digest_0.6.39   R6_2.6.1        fastmap_1.2.0   Matrix_1.7-5    xfun_0.60      
##  [6] lattice_0.22-9  splines_4.6.1   cachem_1.1.0    knitr_1.51      htmltools_0.5.9
## [11] rmarkdown_2.31  stats4_4.6.1    lifecycle_1.0.5 cli_3.6.6       grid_4.6.1     
## [16] sass_0.4.10     jquerylib_0.1.4 compiler_4.6.1  tools_4.6.1     nlme_3.1-169   
## [21] evaluate_1.0.5  bslib_0.12.0    survival_3.8-6  yaml_2.3.12     rlang_1.3.0    
## [26] jsonlite_2.0.0  MASS_7.3-65