Scope. This is an educational guide. It cannot replace subject-matter expertise, community consultation, a statistician or epidemiologist, an ethics review, or jurisdiction-specific legal and regulatory advice. Sample-size calculations are planning aids, not guarantees of scientific value.
A design is the architecture that connects a question to an answer. The study label alone—“cohort,” “case-control,” or “trial”—does not make an investigation valid. Validity depends on who could enter, how exposure or treatment began, how follow-up was aligned, which outcomes were measured, which contrasts were estimated, and which people or clusters remained observable.
This guide can be read in three passes:
All executable examples use base R, stats, or
datasets. No external data or analysis package is
required.
After completing the guide, you should be able to:
| Purpose | Example target question | Design priority |
|---|---|---|
| Descriptive | What was hypertension prevalence in adults in 2026? | Coverage and measurement |
| Etiologic/causal | Would reducing occupational dust lower 10-year COPD risk? | Counterfactual comparability |
| Intervention | Does an outreach program increase vaccination uptake? | Assignment, adherence, and implementation |
| Diagnostic | How accurately does a rapid test classify current infection? | Representative spectrum and reference standard |
| Prognostic | Among diagnosed patients, who develops complications in one year? | Time origin, validation, calibration |
| Policy evaluation | Did a smoke-free law change admissions beyond the prior trend? | Credible counterfactual trend |
Descriptive and predictive goals are not inferior to causal goals; they answer different questions. A representative prevalence survey can be excellent for description yet unsuitable for estimating a treatment effect. A prediction model can forecast risk accurately without identifying an intervention that changes it.
For an exposure question, PECO(T) is useful:
For a trial, PICOT replaces exposure with intervention. For a diagnostic question, specify the intended use, index test, reference standard, positivity threshold, and clinical setting. For a prognostic question, specify the time origin, prediction horizon, predictors available at that time, and intended decision.
Question template. Among [eligible population and setting] at [time zero], what is the [risk/rate/mean/survival contrast] by [time horizon] under [strategy A] versus [strategy B], for [precisely defined outcome]?
Eligibility, assignment or exposure classification, and start of follow-up should be aligned. If a person must survive long enough to be classified as exposed, but their earlier survival time is credited to the exposed group, immortal-time bias can result. Draw a timeline before writing a model:
eligibility assessed -> strategy/exposure assigned -> follow-up -> outcome horizon
all three left-hand events should share a defensible time zero
Designs can be classified along several independent axes.
| Axis | Possibilities | Why it matters |
|---|---|---|
| Investigator assignment | Assigned intervention / observed exposure | Supports different causal assumptions |
| Sampling anchor | Population / exposure / outcome / cluster | Determines estimable quantities and weights |
| Time | Single occasion / repeated / longitudinal | Separates prevalence from incidence and sequence |
| Unit | Person / event / household / facility / community | Determines dependence and target of inference |
| Comparison | Concurrent / historical / synthetic / self-controlled | Changes vulnerability to secular trends |
| Objective | Description / causation / prediction / diagnosis | Determines performance criteria |
The same data source may support several designs. For example, electronic health records can define a retrospective cohort, a nested case-control study, an interrupted time series, or a diagnostic-accuracy study. The protocol—not the database name— defines the design.
Evidence-synthesis designs answer questions by identifying and appraising existing studies. A systematic review uses a prespecified, reproducible protocol to search for, select, assess, and synthesize evidence for a focused question. A scoping review maps the range and characteristics of evidence when concepts, methods, or gaps are the main interest. A meta-analysis is a statistical synthesis that may be included in either type of review; it is an analysis method, not proof that the underlying review was systematic or that its studies are comparable.
Here the sampled units are reports or studies rather than people. Define databases, dates, search terms, eligible designs and populations, duplicate screening, data extraction, risk-of-bias assessment, and synthesis before seeing the results. Major threats include incomplete searches, selective inclusion, publication and outcome- reporting bias, duplicate populations, incompatible estimands, and unexplained heterogeneity. A precise pooled estimate cannot repair biased or incommensurable source studies.
A case report describes one unusual observation; a case series describes several. They can detect new syndromes, adverse events, and unexpected clinical patterns. Without a comparison group or defined denominator, they generally cannot estimate risk or establish that an exposure caused the outcome. State how cases were found to avoid implying completeness.
Surveillance is the ongoing, systematic collection, analysis, interpretation, and dissemination of health data for action. Systems may be passive, active, sentinel, syndromic, laboratory-based, or event-based. Evaluate timeliness, completeness, sensitivity, positive predictive value, representativeness, stability, simplicity, and acceptability. A trend can reflect changes in reporting or case definition rather than disease occurrence.
Ecological studies compare groups rather than individuals—for example, regional air-pollution averages and regional mortality rates. They are efficient for policies or contextual exposures that truly act at group level. An ecological association does not establish that exposed individuals are those with the outcome (ecological fallacy). Conversely, individual-level results need not describe contextual effects (the atomistic fallacy).
Independent probability samples taken at successive times estimate population trends without following the same people. They avoid individual attrition but cannot measure within-person change. Stable wording, mode, frame, season, and weighting are essential for a defensible trend.
Exposure and outcome are measured for a sampled population at one period. This is well suited to prevalence, burden, service use, and hypothesis generation.
Strengths: relatively fast; multiple outcomes and exposures; probability sampling can support population description.
Limits: temporal order may be unclear; prevalent cases overrepresent longer duration; participation can depend on both exposure and outcome. A prevalence ratio or difference is usually easier to interpret than a prevalence odds ratio when the outcome is common.
A cohort defines eligible people at time zero, classifies exposure or intervention, and observes incident outcomes. A prospective cohort plans measurement before outcomes occur; a retrospective cohort reconstructs follow-up from existing data. Both can be valid if eligibility, exposure, follow-up, and outcomes are well defined.
Strengths: establishes temporal sequence; estimates risks or rates; studies multiple outcomes; useful for rare exposures.
Limits: inefficient for very rare or delayed outcomes; loss to follow-up; time-varying exposure and confounding; changes in measurement or care.
Special forms include:
Investigators sample people with the outcome (cases) and people representing the source population that produced those cases (controls), then compare prior exposure. Controls are not simply “healthy people.” They should be eligible to become cases and should represent the exposure distribution in the source population under study.
| Control-sampling scheme | When controls are sampled | Interpretation of odds ratio |
|---|---|---|
| Incidence-density/risk-set | While each case occurs, from those at risk | Estimates an incidence rate ratio without requiring a rare outcome |
| Cumulative-incidence | From noncases at the end of a fixed period | Approximates a risk ratio mainly when outcome is rare |
| Case-cohort | From the cohort at baseline | Can estimate several outcome associations using design-aware methods |
Matching can improve efficiency or control strong confounders, but the analysis must honor the matching. Overmatching on an exposure consequence or strong exposure proxy can reduce validity or efficiency.
cc <- matrix(c(90, 60,
210, 240),
nrow = 2, byrow = TRUE,
dimnames = list(Outcome = c("Case", "Control"),
Exposure = c("Exposed", "Unexposed")))
cc## Exposure
## Outcome Exposed Unexposed
## Case 90 60
## Control 210 240
odds_ratio <- (cc["Case", "Exposed"] * cc["Control", "Unexposed"]) /
(cc["Case", "Unexposed"] * cc["Control", "Exposed"])
se_log_or <- sqrt(sum(1 / cc))
or_ci <- exp(log(odds_ratio) + c(-1, 1) * qnorm(0.975) * se_log_or)
c(odds_ratio = odds_ratio, lower_95 = or_ci[1], upper_95 = or_ci[2])## odds_ratio lower_95 upper_95
## 1.714 1.178 2.496
The odds ratio is a valid measure for the sampled design, but its causal interpretation still requires control selection, exposure measurement, confounding, and other assumptions to be defensible.
Eligible participants are randomly allocated to intervention strategies and followed concurrently. Randomization makes treatment assignment independent of baseline prognostic factors in expectation; allocation concealment prevents foreknowledge of assignment. Blinding, when feasible, can reduce differential behavior, co-intervention, outcome ascertainment, and analysis decisions.
Distinguish:
Schools, clinics, communities, or other groups are randomized when an intervention is delivered collectively or contamination is likely. Outcomes within a cluster are correlated, so effective information is lower than for the same number of independent participants. The number and diversity of clusters can matter more than the raw number of people. Analysis and sample-size planning must respect clustering.
Clusters cross from control to intervention at randomized times until all receive the intervention. Calendar-time effects, learning, anticipation, and changing cluster composition require careful modeling. A stepped wedge is not automatically ethical or more efficient merely because all clusters eventually receive the intervention.
These designs estimate intervention or policy effects without investigator randomization. Their credibility depends on a transparent assignment mechanism and a testable or at least arguable counterfactual.
| Design | Core comparison | Central assumption/threat |
|---|---|---|
| Controlled before-after | Change in intervention group versus comparison group | Groups would otherwise change similarly |
| Difference-in-differences | Difference between group-specific before-after changes | Parallel counterfactual trends; no differential co-intervention |
| Interrupted time series | Level/slope after intervention versus projected pretrend | No concurrent event or measurement change explains break |
| Regression discontinuity | Units just above versus below an assignment cutoff | No precise manipulation; continuity near cutoff |
| Instrumental variable | Outcome differences induced by instrument-linked exposure | Relevance, independence, exclusion, and estimand assumptions |
| Natural experiment | External process produces exposure variation | Assignment process is plausibly as-if random conditional on design |
Plot data and pre-intervention trends. Use enough time points for interrupted time series. Avoid treating a single pre/post comparison as an interrupted time series. Negative-control outcomes or exposures, falsification dates, alternative bandwidths, and unaffected comparison groups can probe—but not prove—assumptions.
| Situation | Often useful | Usually less suitable | Key check |
|---|---|---|---|
| Estimate current prevalence | Probability cross-sectional survey | Case series | Does the frame cover the target population? |
| Rare outcome | Case-control or very large linked cohort | Small cohort | Do controls represent the case source? |
| Rare exposure | Cohort of exposed and comparable unexposed | General-population case-control | Is outcome ascertainment comparable? |
| Multiple outcomes after one exposure | Cohort | Basic case-control | Are time zero and follow-up aligned? |
| Multiple prior exposures for one outcome | Case-control | Small cohort for rare outcome | Can exposure history be measured comparably? |
| Individual intervention feasible | Randomized trial | Uncontrolled before-after | Are concealment and follow-up adequate? |
| Group-delivered intervention | Cluster trial | Individual randomization with contamination | Are there enough clusters? |
| Population policy at known date | Controlled ITS or difference-in-differences | One-group pre/post | Is the counterfactual trend credible? |
| Diagnostic accuracy | Paired index/reference-standard study | Case-control with extreme cases and healthy controls | Is the clinical spectrum representative? |
No design dominates on every dimension. Feasibility can change the question, but the protocol should state that change rather than quietly overclaim what the data answer.
| Population/set | Definition | Example |
|---|---|---|
| Target population | Group to which the scientific conclusion should apply | Adults with hypertension in a province |
| Source population | Population that could give rise to observed participants | Adults listed in participating primary-care practices |
| Sampling frame | Operational list or mechanism used to select units | Current practice rosters |
| Eligible population | Source members meeting protocol criteria | Rostered adults with two qualifying readings |
| Invited/sample selected | Units chosen according to sampling design | Stratified random selection |
| Enrolled/study population | People who consent or otherwise enter | Responding eligible adults |
| Analytic population | Records included in a specified analysis | Enrolled adults with defined outcome data |
At each arrow, ask who is lost, added, duplicated, or misclassified. Coverage error arises when the frame excludes eligible target members (undercoverage), includes ineligible members (overcoverage), or contains duplicates. Eligibility criteria should protect validity and safety without excluding groups merely for convenience.
A household may be sampled, one adult selected, and individual blood pressure analyzed. A clinic may be randomized, patients sampled, and clinic-level policy effects estimated. Document every stage. Standard errors and weights need the actual selection structure, not only the final spreadsheet rows.
Known, nonzero inclusion probabilities support design-based population inference when the frame, response, and measurement assumptions are adequate.
| Method | How it works | Main use/caution |
|---|---|---|
| Simple random sample | Every sample of fixed size is equally likely | Requires an enumerated frame |
| Systematic sample | Random start, then every kth ordered unit | Hidden periodicity can distort selection |
| Stratified sample | Sample independently within prespecified groups | Guarantees subgroup representation; needs stratum weights |
| Cluster sample | Sample natural groups, then all or some units within groups | Reduces field cost; intracluster correlation reduces precision |
| Multistage sample | Sample areas, households, then people | Practical for large populations; record probabilities at each stage |
| Probability proportional to size | Larger clusters receive higher selection probability | Useful when cluster sizes vary; may be self-weighting with later sampling |
# A finite synthetic population; y is a binary health outcome.
population <- data.frame(
id = 1:5000,
region = rep(c("North", "Central", "South"), c(1000, 2500, 1500))
)
region_risk <- c(North = 0.18, Central = 0.10, South = 0.14)
population$y <- rbinom(nrow(population), 1, region_risk[population$region])
# Simple random sample
srs_id <- sample(population$id, size = 300, replace = FALSE)
srs <- population[population$id %in% srs_id, ]
c(sample_estimate = mean(srs$y), population_truth = mean(population$y))## sample_estimate population_truth
## 0.150 0.129
# Disproportionate stratified sample: equal n per region.
selected_rows <- unlist(lapply(split(seq_len(nrow(population)), population$region),
sample, size = 100, replace = FALSE))
strat <- population[selected_rows, ]
Nh <- table(population$region)
nh <- table(strat$region)
strat$base_weight <- as.numeric(Nh[strat$region] / nh[strat$region])
c(unweighted = mean(strat$y),
design_weighted = weighted.mean(strat$y, strat$base_weight),
population_truth = mean(population$y))## unweighted design_weighted population_truth
## 0.1333 0.1160 0.1290
Equal allocation increases the sample share from the smallest stratum. The unweighted mean therefore targets the artificial equal-stratum mixture. The base-weighted mean reconstructs the population mixture because within stratum. In one finite random sample, the unweighted estimate can happen to lie closer to the population truth. That coincidence does not change its target: correct design weights recover the intended population mixture over repeated samples under the stated sampling assumptions, while any single estimate still has sampling error.
stratum_plan <- data.frame(
stratum = c("Urban", "Rural", "Remote"),
N = c(60000, 25000, 5000),
anticipated_sd = c(12, 18, 25)
)
total_n <- 600
stratum_plan$proportional_n <- round(total_n * stratum_plan$N / sum(stratum_plan$N))
score <- with(stratum_plan, N * anticipated_sd)
stratum_plan$neyman_n <- round(total_n * score / sum(score))
stratum_planRounded allocations may not sum exactly to the requested total; reconcile remaining units explicitly. Anticipated variability should come from defensible pilot or prior data, not from inspecting outcomes after selection.
Nonprobability samples can be appropriate for feasibility work, qualitative inquiry, rare or hidden populations, rapid surveillance, instrument development, and mechanistic studies. They do not provide known inclusion probabilities, so a large sample alone does not guarantee population representativeness.
| Method | Description | Primary limitation |
|---|---|---|
| Convenience | Enroll accessible volunteers | Unknown selection mechanism |
| Consecutive | Enroll every eligible unit encountered in a period | Setting and period may be unrepresentative |
| Quota | Fill cells to resemble selected population margins | Selection within cells remains nonrandom |
| Purposive | Select information-rich or theoretically relevant cases | Inference is analytic, not frequency-based |
| Snowball/chain referral | Participants recruit contacts | Network and seed dependence |
| Respondent-driven sampling | Structured peer recruitment with network information | Relies on strong network, recruitment, and weighting assumptions |
Report the recruitment mechanism, sites, dates, incentives, referral chains, duplicate prevention, inclusion criteria, refusals when observable, and why the sample is fit for the intended inference. Poststratification can align measured margins; it cannot automatically correct selection driven by unmeasured determinants of the outcome.
Common error: describing an online open-call survey as “random” because many people saw the link. Random exposure to a link is not the same as known random selection, and viewing, clicking, eligibility, consent, and completion each add a selection step.
Predefine and count:
frame -> sampled -> contact attempted -> reached -> screened -> eligible
-> invited -> consented -> baseline complete -> followed -> analyzed
Record mutually exclusive reasons for exclusion and nonparticipation where ethical and feasible. Use accessible materials, appropriate languages, multiple contact modes, community partners, flexible scheduling, and reasonable reimbursement. Avoid undue influence, coercion, misleading promises, or recruiting through authority figures without safeguards.
Overall response rate is not enough. Compare response across frame variables related to exposure or outcome, examine contact versus cooperation failures, and assess whether reasons differ by site or mode. A low response rate does not mathematically prove large bias, and a high rate does not prove its absence; bias depends on how respondents and nonrespondents differ for the estimand.
A common sequence is:
# Synthetic invited sample with a response mechanism known for demonstration.
n_invited <- 1200
invited <- data.frame(
age = round(runif(n_invited, 18, 80)),
rural = rbinom(n_invited, 1, 0.30)
)
invited$outcome <- rbinom(n_invited, 1,
plogis(-2.4 + 0.025 * invited$age + 0.45 * invited$rural))
invited$response_probability <- plogis(1.1 - 0.012 * invited$age -
0.55 * invited$rural)
invited$responded <- rbinom(n_invited, 1, invited$response_probability)
respondents <- invited[invited$responded == 1, ]
# In real work, response probabilities are estimated from frame variables; here the
# known simulation probabilities let us isolate the weighting idea.
respondents$nonresponse_weight <- 1 / respondents$response_probability
c(full_invited_truth = mean(invited$outcome),
respondent_unweighted = mean(respondents$outcome),
respondent_weighted = weighted.mean(respondents$outcome,
respondents$nonresponse_weight))## full_invited_truth respondent_unweighted respondent_weighted
## 0.2800 0.2514 0.2639
The correction works only to the extent that the response model captures the relevant selection process and has adequate overlap. Extreme inverse probabilities signal weak overlap and unstable extrapolation.
Collect durable, consented contact options; confirm preferred modes; schedule reminders; reduce burden; keep measures consistent; and track moves, withdrawals, deaths, competing events, and administrative censoring separately. Do not turn a participant’s right to withdraw into pressure to remain. Predefine whether existing data may be retained after withdrawal according to consent and law.
A causal estimand should identify:
An odds ratio from a fitted model is not a complete estimand. State the population and time horizon, then prefer an absolute effect alongside a relative measure when possible.
Suppose baseline severity affects treatment and outcome :
C -> E -> Y
| ^
+---------+
The backdoor path motivates measuring and controlling . A mediator on should not be adjusted for when the target is the total effect. A collider in should generally not be conditioned on, because doing so can open a noncausal path.
A DAG records subject-matter assumptions; data cannot reveal the correct DAG by themselves. Include selection and measurement nodes when those processes threaten the study. Use the diagram to choose variables before inspecting effect estimates, and pair it with a written bias analysis.
For a simple causal contrast, investigators commonly consider:
Randomization addresses baseline exchangeability for assignment, not loss to follow-up, nonadherence, measurement error, interference, or chance imbalance in a small trial.
| Bias/threat | Design-stage question | Possible prevention or probe |
|---|---|---|
| Coverage/selection | Who cannot enter the frame or study? | Multiple frames, probability selection, flow counts |
| Nonresponse | Does participation share causes with the outcome? | Accessible recruitment, frame data, weighting, bounds |
| Loss to follow-up | Does remaining observed depend on prognosis? | Retention, outcome linkage, censoring analysis |
| Confounding | What common causes influence exposure and outcome? | Randomization, restriction, measurement, design-based control |
| Exposure misclassification | Is measurement differential by outcome or time? | Blinding, validation substudy, repeated objective measures |
| Outcome misclassification | Is ascertainment differential by exposure? | Standard definitions, blinded adjudication, equal surveillance |
| Recall/interviewer bias | Does knowledge alter reporting or probing? | Memory aids, records, standardized blinded interviews |
| Immortal-time bias | Must one survive to become exposed? | Align eligibility, assignment, and follow-up |
| Reverse causation | Could early disease alter exposure? | New-user design, lag, repeated measures, biological timing |
| Informative missingness | Why is each value missing? | Prevention, auxiliary variables, sensitivity analysis |
| Selective reporting | Were outcomes/models chosen after results? | Registration, protocol, SAP, version history |
Classify each threat by expected direction if possible, likely magnitude, affected estimand, information available to address it, and residual uncertainty. Avoid the empty phrase “all studies have limitations.” Explain how each limitation could alter the result.
Even a simple threshold analysis is more informative than a ritual limitation. Ask:
Sensitivity analyses should vary uncertain assumptions; they should not be a search for the model that produces the preferred conclusion.
Start from the primary estimand and design. A sample size may be chosen to:
State all inputs, their sources, and a range of plausible values. Statistical power is not the probability that a completed significant result is true. Failure to reject a null hypothesis is not proof of equivalence. Very large samples can estimate a biased quantity with great precision.
For a large simple random sample, anticipated proportion , two-sided confidence level , and margin of error :
If sampling without replacement from a finite population of size , an approximate finite-population correction gives:
planning <- data.frame(
scenario = c("Unknown prevalence, very large population",
"Expected prevalence 12%, very large population",
"Expected prevalence 12%, population N=2,000"),
required_complete = c(
prop_n(p = 0.50, margin = 0.03),
prop_n(p = 0.12, margin = 0.03),
prop_n(p = 0.12, margin = 0.03, population = 2000)
)
)
planningUsing maximizes when prevalence is unknown. The formula assumes an unweighted simple random sample and a Wald-style planning approximation; rare proportions or strict coverage requirements may need exact or simulation-based work.
An approximate equal-allocation, two-sided calculation is
p0 <- 0.20
p1 <- 0.15
n_per_group <- two_prop_n(p0, p1, power = 0.80, alpha = 0.05)
c(per_group = n_per_group, total_complete = 2 * n_per_group)## per_group total_complete
## 906 1812
# Compare with R's built-in approximation (returns per-group n).
power.prop.test(p1 = p0, p2 = p1, power = 0.80,
sig.level = 0.05, alternative = "two.sided")##
## Two-sample comparison of proportions power calculation
##
## n = 905.4
## p1 = 0.2
## p2 = 0.15
## sig.level = 0.05
## power = 0.8
## alternative = two.sided
##
## NOTE: n is number in *each* group
The target difference must be scientifically meaningful, not chosen only to make the study affordable. Unequal allocation, covariate adjustment, noninferiority, repeated measures, survival outcomes, and multiplicity need design-specific calculations.
For equal cluster size and intracluster correlation coefficient , a common approximation is:
independent_total <- 600
m <- 25
icc <- 0.03
deff <- design_effect(m, icc)
cluster_adjusted <- ceiling(independent_total * deff)
n_clusters <- ceiling(cluster_adjusted / m)
c(design_effect = deff,
people_before_response_adjustment = cluster_adjusted,
clusters_rounded_up = n_clusters,
people_after_cluster_rounding = n_clusters * m)## design_effect people_before_response_adjustment
## 1.72 1032.00
## clusters_rounded_up people_after_cluster_rounding
## 42.00 1050.00
This shortcut does not replace a cluster-trial calculation. Unequal cluster sizes, few clusters, stratification, repeated periods, cluster-level attrition, treatment effect heterogeneity, and the intended analysis all matter. Explore a plausible ICC range and prioritize enough independent clusters.
Let be the SRS-equivalent number of completed observations required for the analysis, , , and the design effect. A transparent first inflation is
If the primary sample-size calculation already incorporated clustering, set here to avoid applying the design effect twice. For cluster trials, round to whole clusters and plan for cluster-level loss separately; losing one cluster is not the same as losing the same number of independent participants.
complete_needed <- 450
deff <- 1.35
response <- 0.70
retention <- 0.88
approach_needed <- ceiling(complete_needed * deff / (response * retention))
c(complete_needed = complete_needed,
approach_or_invite = approach_needed,
anticipated_complete = approach_needed * response * retention / deff)## complete_needed approach_or_invite anticipated_complete
## 450.0 987.0 450.4
Inflation restores expected count, not validity. If response or dropout depends on unmeasured prognosis, recruiting more people does not remove selection bias.
For survival studies, information is driven largely by the number and timing of events, not only enrollment. For prediction models, needed sample size depends on outcome frequency, number and form of candidate parameters, anticipated model fit, shrinkage, and desired precision. A universal rule such as “10 events per variable” is not a sufficient planning strategy. Plan the complete modeling and validation process, and use external validation when transportability matters.
Simulation is useful for stepped-wedge, longitudinal, adaptive, clustered, recurrent- event, or complex missing-data settings. Simulate the data-generating process, apply the exact planned analysis, and calculate operating characteristics across uncertain parameters.
simulate_two_group_power <- function(n_per_group, p0, p1,
repetitions = 2000, alpha = 0.05) {
stopifnot(length(n_per_group) == 1L, is.finite(n_per_group),
n_per_group >= 2, n_per_group == floor(n_per_group),
length(p0) == 1L, is.finite(p0), p0 > 0, p0 < 1,
length(p1) == 1L, is.finite(p1), p1 > 0, p1 < 1,
length(repetitions) == 1L, is.finite(repetitions),
repetitions >= 100, repetitions == floor(repetitions),
length(alpha) == 1L, is.finite(alpha), alpha > 0, alpha < 1)
p_values <- replicate(repetitions, {
y0 <- rbinom(n_per_group, 1, p0)
y1 <- rbinom(n_per_group, 1, p1)
tab <- matrix(c(sum(y1), n_per_group - sum(y1),
sum(y0), n_per_group - sum(y0)), nrow = 2, byrow = TRUE)
prop.test(tab, correct = FALSE)$p.value
})
estimated_power <- mean(p_values < alpha)
c(simulated_power = estimated_power,
repetitions = repetitions,
monte_carlo_se = sqrt(estimated_power * (1 - estimated_power) / repetitions))
}
simulation_result <- simulate_two_group_power(n_per_group = n_per_group,
p0 = p0, p1 = p1)
c(analytic_target = 0.80, simulation_result)## analytic_target simulated_power repetitions monte_carlo_se
## 0.800000 0.795500 2000.000000 0.009019
Monte Carlo error means repeated runs differ slightly. Increase repetitions for final planning and report the seed, code, scenarios, and uncertainty around simulated operating characteristics.
For every exposure, outcome, covariate, and process measure, specify:
Use validated instruments in a population and language close to the intended setting; validation is not a permanent property of a questionnaire. Pilot the workflow, not only the wording.
Reliability concerns consistency; validity concerns whether the measure captures the intended construct. A consistently miscalibrated device can be reliable but invalid. Measure agreement across the relevant range; correlation alone does not establish agreement. If feasible, embed a blinded validation substudy and preserve enough information for correction or sensitivity analysis.
Maintain a data dictionary, source-to-analysis map, edit-check rules, derivation log, instrument version, query history, and immutable raw-data layer. Store identifiers separately from analytic data, apply least-privilege access, encrypt transfers and storage as appropriate, and define retention and destruction schedules.
A pilot should test recruitment, measurement, randomization, delivery, data flow, and retention. Its progression criteria should be prespecified. A small pilot generally does not estimate the definitive treatment effect precisely and should not be treated as a miniature underpowered efficacy trial. If pilot participants enter the main analysis, plan and disclose that integration in advance.
Write and version the SAP before outcome-driven analytic choices are possible. Include:
# Built-in data illustrate a transparent sequence; this is not a causal analysis.
analysis_data <- within(mtcars, {
high_mpg <- as.integer(mpg >= 22)
automatic <- as.integer(am == 0)
weight_1000lb <- wt
})
# 1. Freeze the analytic cohort and count missingness.
n_analyzed <- nrow(analysis_data)
missing_by_variable <- colSums(is.na(analysis_data[c("high_mpg", "automatic",
"weight_1000lb")]))
# 2. Produce prespecified descriptive summaries.
descriptive <- aggregate(cbind(mpg, weight_1000lb) ~ automatic,
data = analysis_data, FUN = mean)
# 3. Fit the prespecified illustrative model and report effect with uncertainty.
fit <- glm(high_mpg ~ automatic + weight_1000lb,
family = binomial(), data = analysis_data)
estimate_table <- cbind(estimate = coef(fit), confint.default(fit))
list(n_analyzed = n_analyzed,
missing = missing_by_variable,
descriptive = descriptive,
model_estimates = estimate_table)## $n_analyzed
## [1] 32
##
## $missing
## high_mpg automatic weight_1000lb
## 0 0 0
##
## $descriptive
## automatic mpg weight_1000lb
## 1 0 24.39 2.411
## 2 1 17.15 3.769
##
## $model_estimates
## estimate 2.5 % 97.5 %
## (Intercept) 12.248 2.335 22.161
## automatic 2.213 -1.887 6.313
## weight_1000lb -4.963 -8.931 -0.995
The example separates cohort construction, missingness, description,
and modeling. Because mtcars is a small historical
convenience dataset and the model is not tied to a causal design,
coefficients should not be interpreted as transportable causal
effects.
Ethics is part of design quality, not a form attached afterward.
Seek the appropriate research ethics board or institutional review board determination before recruitment or access to identifiable data. Publicly available data are not automatically risk-free; small cells, linkage, stigmatizing labels, and community-level inference can cause harm.
Choose a reporting guideline during protocol development, not only at manuscript submission.
| Study/output | Common guideline |
|---|---|
| Observational cohort, case-control, cross-sectional | STROBE |
| Routinely collected health data | RECORD extension to STROBE |
| Randomized trial | CONSORT 2025 and relevant extension |
| Trial protocol | SPIRIT 2025 |
| Nonrandomized behavioral or public-health intervention evaluation | TREND or design-specific EQUATOR guidance |
| Diagnostic accuracy | STARD 2015; add STARD-AI 2025 when applicable |
| Regression or machine-learning prediction model | TRIPOD+AI 2024; add TRIPOD-Cluster when applicable |
| Systematic review or meta-analysis of prediction-model studies | TRIPOD-SRMA |
| Study developing, tuning, prompting, or evaluating an LLM | TRIPOD-LLM 2025 |
| Systematic review | PRISMA 2020 |
| Qualitative interviews/focus groups | COREQ |
Reporting guidelines improve completeness; they do not by themselves certify design quality. Report participant flow and dates, missingness by variable, outcome events, both absolute and relative estimates with uncertainty, protocol deviations, harms, sampling and weighting, subgroup denominators, and every material change from the protocol or SAP.
For reproducibility, preserve human-readable code, data dictionaries, software versions, computational environment, seeds where meaningful, and checksums or version identifiers. Share deidentified data only when consent, governance, law, and reidentification risk permit; otherwise share metadata, code, synthetic examples, and a governed access route.
Title and decision: What decision will this study inform?
Question/estimand: Population; strategies/exposure; comparator; outcome; time zero; horizon; effect scale; intercurrent events.
Design/setting: Design label plus operational description, sites, dates, and unit.
Population flow: Target -> source -> frame -> eligible -> sampled -> enrolled -> followed -> analyzed.
Sampling/recruitment: Frame, selection stages and probabilities, contact modes, consent, incentives, oversampling, retention, and weights.
Measurement: Definitions, sources, timing, blinding, validation, and quality checks.
Bias plan: Main causal diagram; top five threats; prevention and sensitivity plans.
Size/feasibility: Primary calculation, inputs and sources, design effect, response, attrition, events, sensitivity range, and operational capacity.
Analysis: Primary model and effect measure; missing data; clustering/weights; subgroups; multiplicity; sensitivity analyses.
Ethics/governance: Risks, benefits, equity, privacy, community engagement, return of results, access, retention, and oversight.
Transparency: Registration, protocol/SAP version, reporting guideline, authorship, dissemination, data/code availability.
Keep a short record of rejected alternatives:
| Candidate design | What it would estimate | Advantage | Fatal/major limitation | Decision |
|---|---|---|---|---|
| Design A | … | … | … | retain/reject |
| Design B | … | … | … | retain/reject |
| Design C | … | … | … | retain/reject |
This makes compromises visible and prevents a convenient dataset from silently redefining the scientific question.
Question. What proportion of adults in a city experienced food insecurity in the past 12 months in autumn 2026, overall and by neighborhood deprivation?
Design. Address-based, stratified, two-stage probability survey: sample small areas within deprivation strata, addresses within areas, then one adult per household using a reproducible random rule. Offer web, telephone, and interviewer modes in major local languages.
Sampling. Oversample the highest-deprivation stratum for subgroup precision. Compute inverse stage probabilities, adjust for nonresponse using frame variables, then calibrate age-group, sex/gender where appropriate, and region margins. Preserve strata and clusters for variance estimation.
Size. Base precision on the smallest priority domain, not only the city total; inflate for clustering and response. Assess sensitivity across anticipated prevalence, ICC, and response.
Bias plan. Address undercoverage of people without conventional housing, mode effects, stigma-related nonresponse, proxy responses, seasonal specificity, and the 12-month recall window. Consider a parallel venue-based component for excluded groups and report its estimand separately rather than merging incompatible frames casually.
Question. Among adults initiating treatment for hypertension, what is the 2-year risk difference for acute kidney injury under drug A versus drug B?
Design. Active-comparator, new-user cohort from linked pharmacy and hospital data. Time zero is the dispensing date after a washout period; eligibility and baseline covariates are assessed before that date. Both groups are initiators for the same indication.
Analysis. Define intention-to-treat-like and per-protocol strategies separately. Estimate standardized 2-year risks and their difference. Plan for treatment switching, death as a competing event, informative loss of insurance coverage, missing laboratory data, and positivity. Use a DAG to select baseline confounders and quantitative sensitivity analysis for residual confounding.
Why not a prevalent-user cohort? Including long-term survivors after treatment initiation can miss early harms and condition entry on prior response and persistence.
Question. Did a province-wide naloxone policy change monthly opioid-overdose death rates beyond contemporaneous change in a comparable province?
Design. Prespecify intervention date, sufficient monthly pre/post observations, population offsets, seasonality, autocorrelation, lag or ramp-up, and an unaffected or differently exposed comparison series. Document other policies, toxic-drug-supply changes, coding changes, and pandemic disruptions.
Sensitivity. Falsification dates, alternate intervention lags, control regions, count versus rate models, exclusion of transition months, and negative-control outcomes. Do not interpret a post-policy level change alone as proof that the policy caused it.
Investigators identify all new tuberculosis cases during 2026 and, for each case, sample four people still at risk on the case’s diagnosis date from the same population. They compare prior workplace exposure. What is the design and natural association measure?
This is a nested incidence-density (risk-set) case-control study. With appropriate risk-set sampling and analysis, the exposure odds ratio estimates the incidence rate ratio without requiring tuberculosis to be rare.
A university emails a mental-health survey to currently enrolled students and claims to estimate anxiety prevalence among all young adults in the city. Name two gaps.
The frame excludes young adults who are not currently enrolled, so the source/frame does not cover the stated target population. Within students, opening email, volunteering, and completing a sensitive survey can depend on mental health and other characteristics, creating nonresponse/selection. A large number of replies does not repair either gap.
A health authority needs regional vaccination coverage and reliable estimates for a small remote region representing 4% of the population. What design feature helps?
Stratify by region and oversample the remote stratum to reach its precision target. Use inverse inclusion-probability weights, possibly followed by planned nonresponse and calibration adjustments, for total-population estimates. Incorporate strata, clusters if any, and weights in variance estimation.
A simple-random-sample calculation requires 500 completed participants. The planned cluster design has , expected response is 65%, and expected retention is 90%. Approximately how many people should be approached?
Use 1197. This preserves the expected information count only under the planning assumptions; it does not fix bias from selective response or dropout.
In a cohort of hospital admissions, patients are labeled “treated” if they receive a drug at any point in the first seven days. Mortality follow-up begins at admission. Why is this problematic?
Patients must remain alive and in care long enough to receive the drug, so the period between admission and treatment is guaranteed survival time used to define the treated group. A defensible design could align time zero at treatment eligibility/assignment, use a landmark with appropriate eligibility, or model treatment as time varying, depending on the target strategy and assumptions.
In a study of exercise and cardiovascular disease , baseline health causes both exercise and disease. Weight loss occurs after exercise and affects disease. Which variable is a confounder for the total effect, and which is a mediator?
Baseline health is a confounder and should be measured and appropriately controlled to block . Weight loss is a mediator on a causal path from exercise to disease; controlling it would remove part of the total effect and may introduce additional bias if mediator-outcome confounding is mishandled.
Rewrite “Does green space improve health?” as a PECO(T) question and name one major selection threat and one measurement threat.
Among adults residing in the city on January 1, 2027 (P), what is the 5-year difference in incident depression risk under residence with at least 30% versus less than 10% tree-canopy cover within 500 meters (E/C), with depression defined by a prespecified validated algorithm and follow-up through December 31, 2031 (O/T)? Residential self-selection related to income, preferences, and health is a confounding/selection threat. Satellite canopy near an address may not represent actual access, use, quality, or time spent in green space, creating exposure measurement error.
Which general guideline fits (a) an observational cohort using routinely collected hospital data, and (b) a randomized trial protocol?
| Term | Brief definition |
|---|---|
| Allocation concealment | Preventing foreknowledge of the next randomized assignment |
| Analytic population | Records contributing to a specified analysis |
| Base weight | Inverse of a unit’s sampling inclusion probability |
| Calibration | Adjusting weights so selected weighted totals match reliable known totals |
| Case-control study | Outcome-anchored sampling of cases and source-population controls |
| Cluster | Group sampled or assigned as a unit, with potentially correlated observations |
| Cohort study | Longitudinal follow-up from defined eligibility/time zero to incident outcome |
| Collider | Common effect of two variables; conditioning can induce a noncausal association |
| Confounder | Common cause of exposure and outcome that can distort their causal contrast |
| Coverage error | Mismatch between sampling frame and eligible target/source population |
| Design effect | Variance under a complex design divided by variance under a reference SRS |
| Estimand | Precisely defined quantity the study aims to estimate |
| Finite-population correction | Precision adjustment when sampling a substantial fraction without replacement |
| Inclusion probability | Probability a unit enters the sample under the design |
| Intention-to-treat effect | Effect of randomized assignment regardless of subsequent adherence |
| Intracluster correlation | Similarity of outcomes among units within the same cluster |
| Nonresponse bias | Distortion when response relates to variables relevant to the estimand |
| Positivity | Nonzero probability of each compared strategy in relevant covariate patterns |
| Probability sample | Sample selected with known, nonzero inclusion probabilities |
| Recruitment | Processes used to contact, invite, consent, and enroll selected/eligible units |
| Sampling frame | Operational list or mechanism from which units are selected |
| Source population | Population giving rise to observed study participants or cases |
| Stratification | Dividing a population into groups and sampling/assigning within them |
| Target population | Population to which the intended inference applies |
| Time zero | Aligned start of eligibility, assignment/exposure, and follow-up |