Note: Shared synthetic release 2.0. Real hospital names
are reference labels only; all patient records and outcomes are
simulated.
Overview
Research
question
Among the same within-hospital matched patients as in A07, what is
the standardized difference in admission risk between asthma and
non-asthma patients once outcome regression is added? This version uses
unified data release 2.0-2026-10-09; it does not
rematch and does not create a second study cohort.
Simulation only. This memo estimates differences under
a simulated data-generating mechanism. Real hospital names are used only
as a structural reference; the results do not represent those hospitals’
quality of care or actual admission rates.
The primary model uses 7,700 pairs from 30
hospitals. After standardization, the risk is
31.8% for the asthma group and 22.2%
for the control scenario; the risk difference is 9.54 percentage
points (95% CI 8.47 to 10.61) and the risk ratio is
1.43 (95% CI 1.37 to 1.49).
Cohort and target
A07 selects patients aged ≥18 from the main NACRS file, identifies
asthma by main diagnosis, and then performs within-hospital 1:1
matching. This memo reads those pair IDs directly. Each prediction
scenario targets the successfully matched asthma
patients: their background and hospital are kept, the current
asthma indicator is set to 1 and then to 0, and the predicted
probabilities are averaged.
This keeps A07’s hypothetical disease-scenario comparison and its
limitations. It is not a randomized intervention on a real disease. A08
is also not an independent replication, so agreement with A07 should not
be read as two independent pieces of evidence.
Outcome models
Primary
specification
admitted_flag ~ asthma_visit + age10 + sex + smoking + obesity
+ cardiometabolic_history + asthma_history + prior_ed_visits
+ patient_residence + season + hospital_id
+ asthma_visit:hospital_region
Binomial logistic regression is used. Hospital fixed effects absorb
baseline differences between hospitals; the asthma-by-region interaction
lets the asthma–control difference vary by region. The region main
effect is absorbed by the hospital fixed effects and cannot also be
identified separately. CTAS, arrival severity, length of stay,
disposition, and the true values are not used as adjustment variables,
so that the simulated severity pathway of asthma is preserved.
Flexible sensitivity
model
Age is entered as a natural spline with 3 degrees of freedom, and
asthma is allowed to interact with every background variable; hospital
fixed effects and the asthma-by-region interaction are retained. The
model was fixed before results were inspected, and variables were not
selected by significance.
Standardization and
uncertainty
For each target patient, probabilities are predicted under both
scenarios. The RD is the difference in mean predicted probabilities and
the RR is the ratio of the two mean probabilities; an exponentiated
logistic coefficient is not mislabeled as an RR.
Standard errors use a hospital-level HC1 sandwich
covariance that includes the scores of all matched pairs in the
same hospital; intervals for RD and log(RR) are computed by the delta
method using a t critical value with the number of hospitals minus one
degrees of freedom. Hospital fixed effects and hospital-clustered
standard errors play different roles.
The intervals are conditional on the existing matches and target
composition; they do not fully propagate the uncertainty from the PS and
match selection, and they do not automatically form a doubly robust
estimator. Repeated simulation is used to check empirical bias and
coverage.
Results
The true simulated risk difference for the same target population is
10.38 percentage points. In this sample, the primary
model’s estimation error is -0.84 percentage points; adding more model
terms does not guarantee that every random sample comes closer to the
truth.
Repeated-simulation
comparison
A07’s 100 positive-effect and 100 null-effect simulations are
replayed with the same seeds and sample sizes, and each run is checked
against the pair count, the A07 estimate, and the true value. Both A08
models are fitted in each run, and failures are recorded as well.
The total number of failed models is 0. Coverage and
the false-positive share under the null scenario both carry Monte Carlo
uncertainty; with only 30 hospitals, the precision of the cluster-robust
approximation is also limited.
Interpretation and
limitations
This memo addresses the asthma–control difference in the matched
population as a whole. Whether hospital regions show different admission
patterns is compared directly in A09. GLM adjustment, within-hospital
matching, and sandwich standard errors cannot remove unmeasured
confounding in real data; the current results serve only to validate the
simulated analysis workflow.
Reproducibility
The inputs are A07’s result bundle and the same main NACRS cache. The
versions of the two Excel workbooks are checked at A07’s entry point,
and A08 additionally checks the A07 results and analysis code.
Mathematical checks cover pairing and within-hospital constraints, the
hospital-clustered covariance, numerical gradients, consistency of
spline predictions, and exclusion of downstream variables from the
model.
Back to project home · Matching A07 · Regional analysis A09