MemoA07
ClientDr. ABC
Client OrganizationABC Department
Version2.0
StatusDraft
ClassificationInternal use only
Note: Shared synthetic release 2.0. Real hospital names are reference labels only; all patient records and outcomes are simulated.

1 Overview

1.1 Research question

Among patients aged ≥18 whose main diagnosis for this ED visit is asthma, how much does their probability of admission differ from that of other patients at the same hospital with comparable background characteristics? This memo keeps the A00–A06 format and uses unified data release 2.0-2026-10-09.

Simulation only. Hospital names and the ED service directory come from public sources. All cases, counts, risks, and hospital effects are simulated and cannot be used to evaluate the corresponding real hospitals.

1.2 Main result

From the 30,000 records in the main NACRS file, 24,926 adult visits were selected, of which 8,170 have a main diagnosis of asthma. Within-hospital 1:1 matching retained 7,700 pairs from 30 hospitals.

After matching, the admission rate is 31.8% in the asthma group and 21.8% in the control group; the risk difference is 10.00 percentage points (95% CI 7.59 to 12.41) and the risk ratio is 1.46 (95% CI 1.34 to 1.59). The simulated true value is 10.38 percentage points.

2 Data and provenance

A00 creates the analytic caches from pseudo_NACRS.xlsx and pseudo_DAD.xlsx. A07 selects only the adults in analytic_nacrs.rds, A08 uses the same pairs as this memo, and A09 uses the adult cohort from the same main data. There is no longer a separate 12,000-person A07 main dataset. The old version has been archived, and old results should not be mixed with the figures in this release.

Each simulated patient has exactly one ED visit. Hospital IDs H001–H030 are project-internal codes, not official facility codes. The 30 hospitals are a preselected sample of ED hospitals, 10 each for Calgary, Edmonton, and Other; the first two groups include surrounding towns, and Other contains the selected Central, North, and South hospitals. This is not a province-wide hospital census. hospital_city is the municipality, hospital_region is the study region, and patient_residence is the patient’s area of residence; the three are distinct.

The main exposure is defined as main_diagnosis_icd10 beginning with J45; the current simulation uses J45.9 throughout. A history of asthma, or asthma as a secondary diagnosis, does not mean the visit was mainly for asthma. CIHI’s related ED indicator likewise defines asthma cases by J45 in the Main Problem. CIHI definition

The outcome is defined by whether the ED disposition is Admitted inpatient; whether a DAD record is successfully linked does not change this outcome. A small share of non-admitted records are simulated as LWBS and keep the diagnosis label assigned at generation; this is a teaching simplification.

3 Design and generating assumptions

Within a fixed ED population, a hypothetical “asthma attack” is compared with a prespecified mix of non-asthma reasons for the visit. The target is the average risk difference and risk ratio among successfully matched asthma visits. This is not a disease-switching intervention that could be carried out in practice, and the process by which patients decide to come to the ED is not simulated.

Hospital assignment precedes the current attack and is determined by area of residence and an independent random draw. Shared background factors are age, sex, smoking, obesity, cardiovascular/metabolic history, prior asthma, number of ED visits in the past 12 months, area of residence, and season. Arrival severity is affected by the current attack, and CTAS is its graded measure, so neither enters the PS model for the total effect. Admission, length of stay, and the simulated truth do not enter matching either.

Prespecified simulation parameters; not derived from real hospital statistics
Parameter Value
Main seed 20261009
Hospital random-intercept SD 0.35
Asthma severity shift 0.65
Direct asthma log-odds 0.25
Region baseline: Other / Calgary / Edmonton 0 / -0.10 / 0.12
Asthma × region: Other / Calgary / Edmonton 0 / -0.35 / 0.35

Individual risk is logistic(background + 0.60 × severity + region baseline + hospital effect + asthma direct/interaction). Hospital random intercepts come from a normal distribution with mean 0; the full background equation is kept in code/unified_data.R. Real hospital names provide reference labels only; no real case volumes, bed counts, or clinical outcomes were used to set these parameters.

4 Propensity-score matching

asthma_visit ~ age10 + sex + smoking + obesity + cardiometabolic_history
  + asthma_history + prior_ed_visits + patient_residence + season

After fitting a logistic PS model, each group is restricted, within each hospital, to the common support of logit(PS). Nearest-neighbor 1:1 matching without replacement proceeds in descending order of the asthma group’s scores; the caliper is 0.20 times the standard deviation of logit(PS) across all adults, with 0.10 as a sensitivity analysis. Ties are broken by original row order; the admission outcome is never used to form pairs.

Cohort flow
group status Freq
Asthma Eligible, not matched 373
Other Eligible, not matched 8903
Asthma Matched 7700
Other Matched 7700
Asthma Outside common support 97
Other Outside common support 153

Matching retained 94.2% of the original asthma group. Hospital distributions are identical after matching, but the other background factors still need to be checked.

Propensity-score distributions in the unified adult NACRS cohort.

Propensity-score distributions in the unified adult NACRS cohort.

Baseline balance; hospital identity is exact-matched.

Baseline balance; hospital identity is exact-matched.

The maximum absolute SMD is 0.302 before matching and 0.025 after matching. The denominator is the standard deviation in the asthma group before matching; the 0.10 threshold is only a diagnostic guide and cannot rule out unmeasured confounding.

5 Admission risk comparison

Analysis Pairs Asthma risk Control risk RD (pp) 95% CI (pp) RR RR 95% CI
Primary: caliper 0.20 7,700 31.8% 21.8% 10.00 7.59 to 12.41 1.46 1.34 to 1.59
Sensitivity: caliper 0.10 7,449 31.1% 21.5% 9.59 7.20 to 11.97 1.45 1.33 to 1.57

The crude risk difference is 16.24 percentage points. The crude and matched comparisons describe different populations, so the change in the difference cannot be attributed entirely to removing confounding. A narrower caliper can also change the target population.

Hospital clustering. Influence-function values for the risk difference and risk ratio are first formed for each pair, then summed by hospital to compute cluster-robust standard errors; intervals use the t critical value with the number of hospitals minus one degrees of freedom. Every matched pair lies within one hospital, so hospital clustering also covers the within-pair correlation. This is approximate inference conditional on the given pairs and the empirical target population; it does not fully propagate the uncertainty from PS estimation and match selection.

6 Repeated-simulation validation

Using the same generating function, 100 runs were made under each of a positive-effect scenario and an asthma null-effect scenario, each with 4,000 adults and 30 hospitals; seeds are 20261101–20261200. Each run refits the PS, matches within hospitals, and compares the estimate with the true value for its target population. The null scenario removes asthma’s severity shift, direct effect, and region interaction, while keeping baseline region and hospital differences.

Scenario Runs Mean bias (pp) MC SE of bias (pp) Coverage CI excludes zero
positive 100 0.49 0.18 98.0% 100.0%
null 100 0.61 0.18 97.0% 3.0%

These repeated samples are used only to check the method; they are not written to the main NACRS/DAD files and do not count toward the project sample size. Coverage from 100 repetitions still has Monte Carlo error and cannot guarantee confidence-interval coverage in any real study.

7 Limitations and next memo

This is a methods demonstration under a known generating mechanism. Real data may have diagnosis-timing bias, unmeasured severity, selection into the ED, inter-hospital transfers, and more missing data; matching and small p values cannot verify causal assumptions. Only selected hospital names were used, and the simulated case count at each hospital does not reflect its real size.

A08 applies outcome regression and standardization to these fixed pairs. A09 studies differences between hospital regions using the full adult cohort and hospital random intercepts.

8 Reproducibility

This memo reads precomputed results and verifies input checksums. The main data come from the two Excel workbooks; the RDS files are analysis caches, and the simulated counterfactual truth is kept separately in data/simulation_truth.rds. The one-step pipeline is code/rebuild_unified_project.R, and historical snapshots are kept in outputs/unification/before_v2/. By project convention, working code, patient-level synthetic files, and analysis results are not committed to the public Git repository.

Back to project home · Data definitions A00