Note: Shared synthetic release 2.0. Real hospital names
are reference labels only; all patient records and outcomes are
simulated.
Overview
Research
question
Among patients aged ≥18 whose main diagnosis for this ED visit is
asthma, how much does their probability of admission differ from that of
other patients at the same hospital with comparable background
characteristics? This memo keeps the A00–A06 format and uses unified
data release 2.0-2026-10-09.
Simulation only. Hospital names and the ED service
directory come from public sources. All cases, counts, risks, and
hospital effects are simulated and cannot be used to evaluate the
corresponding real hospitals.
Main result
From the 30,000 records in the main NACRS file,
24,926 adult visits were selected, of which
8,170 have a main diagnosis of asthma. Within-hospital
1:1 matching retained 7,700 pairs from 30
hospitals.
After matching, the admission rate is 31.8% in the
asthma group and 21.8% in the control group; the risk
difference is 10.00 percentage points (95% CI 7.59 to
12.41) and the risk ratio is 1.46 (95% CI 1.34 to
1.59). The simulated true value is 10.38 percentage points.
Data and
provenance
A00 creates the analytic caches from pseudo_NACRS.xlsx
and pseudo_DAD.xlsx. A07 selects only the adults in
analytic_nacrs.rds, A08 uses the same pairs as this memo,
and A09 uses the adult cohort from the same main data. There is
no longer a separate 12,000-person A07 main dataset. The old
version has been archived, and old results should not be mixed with the
figures in this release.
Each simulated patient has exactly one ED visit. Hospital IDs
H001–H030 are project-internal codes, not official facility codes. The
30 hospitals are a preselected sample of ED hospitals, 10 each for
Calgary, Edmonton, and Other; the first two groups include surrounding
towns, and Other contains the selected Central, North, and South
hospitals. This is not a province-wide hospital census.
hospital_city is the municipality,
hospital_region is the study region, and
patient_residence is the patient’s area of residence; the
three are distinct.
The main exposure is defined as main_diagnosis_icd10
beginning with J45; the current simulation uses J45.9 throughout. A
history of asthma, or asthma as a secondary diagnosis, does not mean the
visit was mainly for asthma. CIHI’s related ED indicator likewise
defines asthma cases by J45 in the Main Problem. CIHI
definition
The outcome is defined by whether the ED disposition is
Admitted inpatient; whether a DAD record is successfully
linked does not change this outcome. A small share of non-admitted
records are simulated as LWBS and keep the diagnosis label assigned at
generation; this is a teaching simplification.
Design and generating
assumptions
Within a fixed ED population, a hypothetical “asthma attack” is
compared with a prespecified mix of non-asthma reasons for the visit.
The target is the average risk difference and risk ratio among
successfully matched asthma visits. This is not a
disease-switching intervention that could be carried out in practice,
and the process by which patients decide to come to the ED is not
simulated.
Hospital assignment precedes the current attack and is determined by
area of residence and an independent random draw. Shared background
factors are age, sex, smoking, obesity, cardiovascular/metabolic
history, prior asthma, number of ED visits in the past 12 months, area
of residence, and season. Arrival severity is affected by the current
attack, and CTAS is its graded measure, so neither enters the PS model
for the total effect. Admission, length of stay, and the simulated truth
do not enter matching either.
Individual risk is
logistic(background + 0.60 × severity + region baseline + hospital effect + asthma direct/interaction).
Hospital random intercepts come from a normal distribution with mean 0;
the full background equation is kept in
code/unified_data.R. Real hospital names provide reference
labels only; no real case volumes, bed counts, or clinical outcomes were
used to set these parameters.
Propensity-score
matching
asthma_visit ~ age10 + sex + smoking + obesity + cardiometabolic_history
+ asthma_history + prior_ed_visits + patient_residence + season
After fitting a logistic PS model, each group is restricted,
within each hospital, to the common support of
logit(PS). Nearest-neighbor 1:1 matching without replacement proceeds in
descending order of the asthma group’s scores; the caliper is 0.20 times
the standard deviation of logit(PS) across all adults, with 0.10 as a
sensitivity analysis. Ties are broken by original row order; the
admission outcome is never used to form pairs.
Matching retained 94.2% of the original asthma group. Hospital
distributions are identical after matching, but the other background
factors still need to be checked.
The maximum absolute SMD is 0.302 before matching and
0.025 after matching. The denominator is the standard
deviation in the asthma group before matching; the 0.10 threshold is
only a diagnostic guide and cannot rule out unmeasured confounding.
Admission risk
comparison
The crude risk difference is 16.24 percentage points. The crude and
matched comparisons describe different populations, so the change in the
difference cannot be attributed entirely to removing confounding. A
narrower caliper can also change the target population.
Hospital clustering. Influence-function values for
the risk difference and risk ratio are first formed for each pair, then
summed by hospital to compute cluster-robust standard errors; intervals
use the t critical value with the number of hospitals minus one degrees
of freedom. Every matched pair lies within one hospital, so hospital
clustering also covers the within-pair correlation. This is approximate
inference conditional on the given pairs and the empirical target
population; it does not fully propagate the uncertainty from PS
estimation and match selection.
Repeated-simulation
validation
Using the same generating function, 100 runs were made under each of
a positive-effect scenario and an asthma null-effect scenario, each with
4,000 adults and 30 hospitals; seeds are 20261101–20261200. Each run
refits the PS, matches within hospitals, and compares the estimate with
the true value for its target population. The null scenario removes
asthma’s severity shift, direct effect, and region interaction, while
keeping baseline region and hospital differences.
These repeated samples are used only to check the method; they are
not written to the main NACRS/DAD files and do not count toward the
project sample size. Coverage from 100 repetitions still has Monte Carlo
error and cannot guarantee confidence-interval coverage in any real
study.
Limitations and next
memo
This is a methods demonstration under a known generating mechanism.
Real data may have diagnosis-timing bias, unmeasured severity, selection
into the ED, inter-hospital transfers, and more missing data; matching
and small p values cannot verify causal assumptions. Only selected
hospital names were used, and the simulated case count at each hospital
does not reflect its real size.
A08 applies outcome
regression and standardization to these fixed pairs. A09 studies differences between
hospital regions using the full adult cohort and hospital random
intercepts.
Reproducibility
This memo reads precomputed results and verifies input checksums. The
main data come from the two Excel workbooks; the RDS files are analysis
caches, and the simulated counterfactual truth is kept separately in
data/simulation_truth.rds. The one-step pipeline is
code/rebuild_unified_project.R, and historical snapshots
are kept in outputs/unification/before_v2/. By project
convention, working code, patient-level synthetic files, and analysis
results are not committed to the public Git repository.
Back to project home · Data definitions A00