SMDs are central; baseline p-values are not
For a continuous variable, the standardized mean difference can be
written as:
where
is a predefined standardization scale. For the ATT design on this page,
MatchIt uses the standard deviation in the original
intervention group as the SMD denominator and retains the same
denominator before and after matching. The reference scale can differ
for ATE or ATC designs. Categorical variables are usually expanded into
indicator variables and checked separately. An SMD does not shrink
mechanically with sample size, so it is useful for describing design
balance. A baseline significance test asks whether population means
could be equal, not whether an observed difference is large enough to
confound the analysis.
An absolute SMD below 0.1 is a common heuristic, not an automatic
certificate. Important variables may require a stricter standard, and
mean balance does not guarantee balance in tails, nonlinear terms, or
interactions.
balance_terms <- setdiff(rownames(match_summary$sum.all), "distance")
balance_labels <- c(
age = "Age",
bmi = "BMI",
smokerNo = "Nonsmoker",
smokerYes = "Smoker",
severityMild = "Mild",
severityModerate = "Moderate",
severitySevere = "Severe",
biomarker = "Biomarker"
)
balance_table <- data.frame(
Covariate = unname(balance_labels[balance_terms]),
`SMD before` = match_summary$sum.all[balance_terms, "Std. Mean Diff."],
`SMD after` = match_summary$sum.matched[
balance_terms, "Std. Mean Diff."
],
`Variance ratio before` = match_summary$sum.all[balance_terms, "Var. Ratio"],
`Variance ratio after` = match_summary$sum.matched[
balance_terms, "Var. Ratio"
],
`Maximum eCDF difference before` = match_summary$sum.all[
balance_terms, "eCDF Max"
],
`Maximum eCDF difference after` = match_summary$sum.matched[
balance_terms, "eCDF Max"
],
check.names = FALSE
)
max_smd_before <- max(abs(balance_table$`SMD before`), na.rm = TRUE)
max_smd_after <- max(abs(balance_table$`SMD after`), na.rm = TRUE)
knitr::kable(
balance_table,
digits = 3,
caption = "Balance diagnostics for pretreatment covariates before and after matching"
)
Balance diagnostics for pretreatment covariates before and
after matching
| age |
Age |
0.566 |
-0.017 |
0.992 |
0.958 |
0.237 |
0.035 |
| bmi |
BMI |
0.371 |
-0.009 |
1.049 |
1.052 |
0.148 |
0.040 |
| smokerNo |
Nonsmoker |
-0.312 |
0.000 |
NA |
NA |
0.151 |
0.000 |
| smokerYes |
Smoker |
0.312 |
0.000 |
NA |
NA |
0.151 |
0.000 |
| severityMild |
Mild |
-0.296 |
-0.066 |
NA |
NA |
0.144 |
0.032 |
| severityModerate |
Moderate |
0.076 |
0.033 |
NA |
NA |
0.037 |
0.016 |
| severitySevere |
Severe |
0.252 |
0.038 |
NA |
NA |
0.107 |
0.016 |
| biomarker |
Biomarker |
0.250 |
-0.031 |
0.991 |
1.052 |
0.131 |
0.040 |
The largest absolute SMD among pretreatment covariates falls from
0.566 to 0.066. The distance row is intentionally excluded
from this “largest covariate” calculation: balancing the PS itself
cannot replace diagnostics for every covariate.
love_before <- abs(balance_table$`SMD before`)
love_after <- abs(balance_table$`SMD after`)
love_order <- order(love_before)
love_y <- seq_along(love_order)
plot(
love_before[love_order],
love_y,
pch = 16,
col = palette_psm["orange"],
xlim = c(0, max(love_before) * 1.08),
yaxt = "n",
xlab = "Absolute standardized mean difference",
ylab = "",
main = "Covariate balance before and after matching",
las = 1
)
points(
love_after[love_order],
love_y,
pch = 17,
col = palette_psm["teal"]
)
axis(2, at = love_y, labels = balance_table$Covariate[love_order], las = 1)
abline(v = 0.10, lty = 2, col = palette_psm["vermillion"])
legend(
"bottomright",
legend = c("Before matching", "After matching", "|SMD|=0.10"),
col = c(
palette_psm["orange"],
palette_psm["teal"],
palette_psm["vermillion"]
),
pch = c(16, 17, NA),
lty = c(NA, NA, 2),
bty = "n"
)
Variance ratios and eCDFs supplement mean balance
The variance ratio (VR) compares dispersion between groups. Here
MatchIt reports the intervention-group variance divided by
the control-group variance and applies the corresponding matching
weights after matching. The ideal value for a continuous variable is
near 1. Ranges such as 0.5–2, or the stricter 0.8–1.25, can flag
concerns but should not determine causal validity mechanically. A binary
indicator’s variance is determined by its mean, so its VR may be blank
in the table; that is not a computational failure.
Empirical cumulative distribution functions (eCDFs) compare
distributions across the entire observed range. eCDF Max is
the largest vertical separation between the two eCDFs; values closer to
zero are better. It can detect settings in which means agree but
distributions do not.
old_par <- par(mfrow = c(1, 2), mar = c(4.2, 4.2, 3, 1))
plot(
ecdf(psm_design_data$age[psm_design_data$treatment_num == 0]),
col = palette_psm["orange"],
lwd = 2.2,
verticals = TRUE,
do.points = FALSE,
main = "Before matching",
xlab = "Age",
ylab = "Empirical cumulative probability",
las = 1
)
lines(
ecdf(psm_design_data$age[psm_design_data$treatment_num == 1]),
col = palette_psm["teal"],
lwd = 2.2,
verticals = TRUE,
do.points = FALSE
)
plot(
ecdf(matched_design$age[matched_design$treatment_num == 0]),
col = palette_psm["orange"],
lwd = 2.2,
verticals = TRUE,
do.points = FALSE,
main = "After matching",
xlab = "Age",
ylab = "Empirical cumulative probability",
las = 1
)
lines(
ecdf(matched_design$age[matched_design$treatment_num == 1]),
col = palette_psm["teal"],
lwd = 2.2,
verticals = TRUE,
do.points = FALSE
)
legend(
"bottomright",
legend = c("Control", "Intervention"),
col = c(palette_psm["orange"], palette_psm["teal"]),
lwd = 2.2,
bty = "n"
)
continuous_vr_before <- balance_table$`Variance ratio before`[
is.finite(balance_table$`Variance ratio before`)
]
continuous_vr_after <- balance_table$`Variance ratio after`[
is.finite(balance_table$`Variance ratio after`)
]
balance_summary <- data.frame(
Diagnostic = c(
"Largest covariate |SMD|",
"Continuous-variable variance-ratio range",
"Largest covariate eCDF Max"
),
Before = c(
sprintf("%.3f", max_smd_before),
sprintf(
"%.3f–%.3f",
min(continuous_vr_before),
max(continuous_vr_before)
),
sprintf("%.3f", max(balance_table$`Maximum eCDF difference before`))
),
After = c(
sprintf("%.3f", max_smd_after),
sprintf(
"%.3f–%.3f",
min(continuous_vr_after),
max(continuous_vr_after)
),
sprintf("%.3f", max(balance_table$`Maximum eCDF difference after`))
),
check.names = FALSE
)
knitr::kable(balance_summary, caption = "Overall balance summary for the primary specification")
Overall balance summary for the primary specification
| Largest covariate |SMD| |
0.566 |
0.066 |
| Continuous-variable variance-ratio range |
0.991–1.049 |
0.958–1.052 |
| Largest covariate eCDF Max |
0.237 |
0.040 |
PS
distributions after matching
matched_probability <- plogis(matched_design$distance)
matched_control_density <- density(
matched_probability[matched_design$treatment_num == 0],
from = 0,
to = 1
)
matched_treated_density <- density(
matched_probability[matched_design$treatment_num == 1],
from = 0,
to = 1
)
plot(
matched_control_density,
col = palette_psm["orange"],
lwd = 2.4,
xlim = c(0, 1),
ylim = c(0, max(matched_control_density$y, matched_treated_density$y)),
xlab = "Estimated propensity score",
ylab = "Density",
main = "Region of support after matching",
las = 1
)
lines(
matched_treated_density,
col = palette_psm["teal"],
lwd = 2.4
)
legend(
"topright",
legend = c("Matched controls", "Retained intervention recipients"),
col = c(palette_psm["orange"], palette_psm["teal"]),
lwd = 2.4,
bty = "n"
)
Check your understanding: If the PS is balanced, can the original
covariates be ignored?
No. Different covariate combinations can share a similar PS. Every
prespecified pretreatment covariate must be checked individually, along
with nonlinear terms, interactions, tails, and clinically important
strata when appropriate.