Deciding Before Seeing: A Statistical Analysis Plan for Evaluating Neighborhood Blood Pressure Support
Student Name
Doctor of Public Health Program, Aspen University
DPH 860: Advanced Biostatistics
Instructor Name
Month Day, Year
Deciding Before Seeing: A Statistical Analysis Plan for Evaluating Neighborhood Blood Pressure Support
Writing the statistical analysis plan before looking at outcomes protects a study from choosing methods that produce favorable results. It also gives the committee a clear picture of what will be done and why. This paper presents the analysis plan for a DPH project evaluating whether a community health worker program improved blood pressure outcomes among east-side adults compared with matched adults.
Aims and Hypotheses
The first aim estimates how program participation affects blood pressure control below 140/90 at 12 months; the hypothesis is that participants will have higher control. The second aim estimates its effect on systolic pressure change. The third aim explores whether effects differ by age, sex, insurance and baseline pressure.
Design and Matching
The project uses a matched comparison design. Each participant is paired with a nonenrolled adult from the same clinics, using propensity scores that estimate the probability of enrollment from age, sex, race and ethnicity, insurance, baseline systolic pressure, diabetes, number of medications and neighborhood. Austin (2011) describes how propensity score matching can reduce confounding in observational studies by creating groups similar on measured characteristics.
Checking Balance
After matching, balance will be checked using standardized differences for each covariate, with values below 0.1 indicating adequate balance. Significance tests are not used for balance, since they depend on sample size. Covariates remaining imbalanced will be included in outcome models. A balance table showing standardized differences before and after matching will appear in the results chapter.
Analysis Map
The table links each aim to its outcome, method, covariates and effect measure.
| Aim | Outcome | Method | Adjustment | Effect measure |
|---|---|---|---|---|
| 1 | Blood pressure controlled at 12 months | Conditional logistic regression on matched pairs | Residual imbalanced covariates | Odds ratio; risk difference with 95% CI |
| 2 | Change in systolic pressure | Linear mixed model with pair as random effect | Baseline systolic pressure | Mean difference with 95% CI |
| 3 | Aims 1 and 2 outcomes | Interaction terms by subgroup | As above | Subgroup estimates with 95% CIs |
Primary Analysis
For Aim 1, conditional logistic regression accounts for matched pairs and estimates the odds ratio for control among participants compared with their matches. Because odds ratios can overstate effects when outcomes are common, the risk difference will also be reported, estimated from the matched data with a confidence interval.
Secondary Analysis
For Aim 2, a linear mixed model with pair as a random effect estimates the difference in systolic change, adjusting for baseline systolic pressure to account for regression to the mean. Residuals will be examined for normality and outliers.
Missing Data
About 10% of outcome values are expected to be missing. The primary analysis will use multiple imputation with 20 imputed datasets, including the outcome, exposure, all matching variables and predictors of missingness in the imputation model. Sterne et al. (2009) caution that imputation models must include the outcome and be compatible with the analysis model; the plan follows these recommendations. A complete-case analysis will be reported as a sensitivity analysis. Imputed results will be combined using standard rules that account for uncertainty from imputation.
Sensitivity Analyses
Sensitivity analyses will vary the matching caliper, use inverse probability weighting instead of matching, define control using only systolic pressure and calculate the minimum strength an unmeasured confounder would require to account for the whole result. Consistent findings across these analyses would strengthen confidence. The unmeasured confounding analysis will report an E-value, the minimum strength of association such a confounder would need with both exposure and outcome.
Subgroup Analyses
Subgroup analyses under Aim 3 are exploratory. Interaction terms will test whether effects differ, and subgroup estimates will be reported with intervals in a forest plot. Given limited power, nonsignificant interactions will not be read as evidence of equal effects.
Software and Reproducibility
Analyses will be conducted in R, with code documented and archived. The codebook, cleaning scripts and analysis scripts will allow the analysis to be reproduced.
Reporting Standards
Results will follow the STROBE guidelines, which ask authors to report the design and setting, who took part, how variables and bias were handled, study size, methods, missing data and sensitivity analyses (von Elm et al., 2007).
Timeline
Matching and balance checks will be completed in month one, imputation and primary analyses in months two and three, sensitivity and subgroup analyses in month four, and drafting of results in month five.
Defining Exposure
Exposure is defined as enrollment in the program, regardless of the number of visits received, an intention-to-treat approach that avoids bias from differences between those who engaged more or less. A secondary analysis will examine the relationship between visit count and outcomes among participants, interpreted cautiously.
Handling Regression to the Mean
Adults were eligible because their blood pressure was high, so some improvement is expected by chance as later readings move toward typical values. Comparing participants with matched adults who met the same criteria, and adjusting for baseline pressure, addresses this. Without a comparison group, regression to the mean could be mistaken for a program effect.
Deviations From the Plan
Any deviation from the plan, such as a change in model because assumptions fail, will be documented with reasons and reported. Prespecified and post hoc analyses will be clearly distinguished in the results chapter.
Committee Review
The plan was reviewed by the committee chair and methodologist before outcome data were accessed. Their suggestions, including adding the E-value sensitivity analysis and specifying the time window for the 12-month reading, were incorporated and the plan was dated and archived.
Interpretation Rules
The plan also states how results will be interpreted. A benefit will be described as supported if the primary analysis interval excludes no effect and sensitivity analyses point in the same direction. Results will be discussed in terms of size and uncertainty, and the absence of statistical significance in subgroups will not be treated as evidence of no effect.
Prespecified Covariates
Covariates are chosen based on prior knowledge of factors that influence both enrollment and blood pressure outcomes, not on statistical significance in the data. Selecting covariates by significance testing can introduce bias and overfitting. The plan lists all covariates in advance and justifies each.
Conclusion
The analysis plan fixes aims, matching, outcome models, missing data methods, sensitivity and subgroup analyses and reporting standards before outcomes are examined. Linking each aim to a specific method and effect measure gives the DPH project a transparent, defensible path from data to conclusions.
References
Austin, P. C. (2011). An introduction to propensity score methods for reducing the effects of confounding in observational studies. Multivariate Behavioral Research, 46(3), 399-424. https://doi.org/10.1080/00273171.2011.568786
Sterne, J. A. C., White, I. R., Carlin, J. B., Spratt, M., Royston, P., Kenward, M. G., Wood, A. M., & Carpenter, J. R. (2009). Multiple imputation for missing data in epidemiological and clinical research: Potential and pitfalls. BMJ, 338, Article b2393. https://doi.org/10.1136/bmj.b2393
von Elm, E., Altman, D. G., Egger, M., Pocock, S. J., Gøtzsche, P. C., & Vandenbroucke, J. P. (2007). The Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) statement: Guidelines for reporting observational studies. The Lancet, 370(9596), 1453-1457. https://doi.org/10.1016/S0140-6736(07)61602-X
How this DPH 860 Module 7 example is structured
Check the DPH 860 prompt in your Aspen classroom before using this example. It states aims, describes matching and balance, tables the analysis map, then covers primary and secondary analyses, missing data, sensitivity and subgroup analyses, software, reporting and timeline.
DPH 860 Module 7 questions, answered
What does DPH 860 Module 7 usually ask for?
Aspen's DPH 860 builds toward the doctoral project, so a statistical analysis plan is a typical assignment. Confirm with your classroom prompt.
What is propensity score matching?
Matching participants to nonparticipants with similar estimated probabilities of participation based on measured characteristics.
Why write an analysis plan before seeing outcomes?
To prevent choosing methods that favor particular results and to make the analysis transparent.
Write yours, or have the desk draft it
This paper is an original model document written by our desk, not a submitted student paper and not an official Aspen University document. Read it for the moves, then write your own to the instructions in your classroom. If you want one built to your exact prompt and rubric, the first custom sample is free and arrives in 24 to 48 hours.