Enough to Know: Power and Sample Size for a Community Health Worker Program Evaluation
Student Name
Doctor of Public Health Program, Aspen University
DPH 860: Advanced Biostatistics
Instructor Name
Month Day, Year
Enough to Know: Power and Sample Size for a Community Health Worker Program Evaluation
A study that is too small may miss a real effect; one that is larger than needed wastes resources. Power and sample size calculations decide, before data collection, how many participants are needed to detect an effect of a given size with reasonable confidence. This paper calculates power and sample size for a DPH project evaluating a community health worker initiative for high blood pressure.
The Ingredients
A sample size calculation needs five inputs: the significance level, usually 0.05; the desired power, usually 80% or 90%; the expected outcome in the comparison group; the smallest difference worth detecting; and, for continuous outcomes, the expected standard deviation. Whitley and Ball (2002) stress that the smallest important difference should be chosen on clinical or practical grounds, not to make the numbers convenient.
Choosing the Inputs
Blood pressure control among comparable adults in the county is about 48%, based on clinic data. A 10-point improvement, to 58%, was judged the smallest difference that would justify expanding the program, given its cost. For systolic change, a standard deviation of about 16.5 mmHg comes from prior clinic data, and a 4 mmHg greater reduction was judged clinically meaningful.
Scenarios
The table shows required sample sizes per group for several scenarios, with alpha set at 0.05, two-sided.
| Outcome and difference | Power | Required per group | With 15% missing |
|---|---|---|---|
| Control, 47.6% vs 57.6% | 80% | 391 | 460 |
| Control, 47.6% vs 57.6% | 90% | 522 | 615 |
| Control, 47.6% vs 52.6% | 80% | 1,568 | 1,845 |
| Systolic change, 4 mmHg difference, SD 16.5 | 80% | 268 | 316 |
| Systolic change, 4 mmHg difference, SD 16.5 | 90% | 358 | 422 |
Working the Main Calculation
For two proportions, the required sample size per group equals the square of the sum of two terms, divided by the squared difference. With proportions of 0.476 and 0.576, the average proportion is 0.526. The first term is 1.96 times the square root of two times 0.526 times 0.474, about 1.384. The second is 0.842 times the square root of the sum of 0.476 times 0.524 and 0.576 times 0.424, about 0.591. Their sum, 1.975, squared is 3.902, and dividing by 0.01 gives about 391 per group.
Allowing for Missing Data
About 15% of participants may lack 12-month blood pressure readings. Dividing 391 by 0.85 gives 460 per group to be enrolled. With 700 per group available, the project exceeds this requirement for its primary outcome.
Power for Smaller Effects
If the true difference were only 5 points, about 1,568 per group would be needed for 80% power. With 700 per group, power to detect a 5-point difference is roughly 45%. The project could therefore miss a small but real benefit, a limitation to state openly.
Continuous Outcome
For systolic change, 268 per group gives 80% power to detect a 4 mmHg difference when the standard deviation is 16.5. The continuous outcome is more efficient than the binary one, reflecting the information lost when blood pressure is reduced to controlled or not.
Subgroups
Subgroup analyses have much less power. With about 118 adults over 70 in each group, power to detect a 10-point difference in control is only about 34%. Subgroup results will therefore be treated as exploratory, and nonsignificant findings will not be read as evidence of no effect (Altman & Bland, 1995). Stating this in advance prevents over-interpreting subgroup patterns after the fact.
Effect of Matching
Matching on factors related to the outcome can increase precision, slightly reducing the sample size needed. Because the gain is uncertain, the calculation conservatively assumes independent groups; analysis that accounts for matching will likely have somewhat more power than calculated.
Checking With Software
Hand calculations were checked with power analysis software. G*Power, a widely used free program, performs calculations for many common tests and produces power curves showing how power changes with sample size (Faul et al., 2007). Its results for the main scenarios matched the hand calculations within one or two participants.
Power After the Fact
Calculating power after a study using the observed effect adds little information, since it is determined by the P value. The project reports its planned power and relies on confidence intervals to show what effects the data can and cannot rule out.
Implications for the Project
The project is well powered to detect a 10-point difference in control and a 4 mmHg difference in systolic change, but not to detect small effects or subgroup differences reliably. These limits will be stated in the methods chapter and considered when interpreting results.
Where the Inputs Came From
The comparison group's expected control rate came from clinic data on similar adults in the year before the program began. The standard deviation for systolic change came from the same data. Using local rather than published values makes the calculation more realistic, since published trials often enroll more selected populations with different variability.
Sensitivity of the Calculation
Sample size is sensitive to inputs. If the comparison group's control rate were 40% rather than 47.6%, the required sample for a 10-point difference would change only slightly, but if the standard deviation for systolic change were 20 rather than 16.5, the required sample for a 4 mmHg difference would rise by nearly half. Reporting how conclusions change with inputs shows the committee the calculation's robustness.
Explaining Power to Decision Makers
Health department leaders may ask why the study cannot answer every question. Explaining that 700 per group can reliably detect a meaningful overall benefit but not small differences among subgroups helps set realistic expectations and guides decisions about whether a larger or longer evaluation is needed.
Precision Rather Than Power
An alternative approach sets the sample size needed to estimate an effect with a given precision. With 700 per group, the 95% interval for the difference in control has a half-width of about 5 percentage points, precise enough to distinguish a meaningful benefit from none. Framing sample size in terms of precision fits the project's emphasis on estimation.
Documenting the Calculation
The methods chapter will report every input, its source, the formula or software used and the resulting sample size, so that readers and reviewers can reproduce the calculation. Omitting inputs is one of the most common weaknesses reviewers identify in sample size sections.
Conclusion
Power and sample size calculations show that the DPH project's 700 participants and 700 matched adults provide ample power for its primary outcomes, even allowing for missing data, while subgroup and small-effect analyses remain underpowered. Documenting inputs, working calculations by hand and checking them with software make the project's design transparent and its conclusions appropriately bounded.
References
Altman, D. G., & Bland, J. M. (1995). Absence of evidence is not evidence of absence. BMJ, 311(7003), 485. https://doi.org/10.1136/bmj.311.7003.485
Faul, F., Erdfelder, E., Lang, A.-G., & Buchner, A. (2007). G*Power 3: A flexible statistical power analysis program for the social, behavioral, and biomedical sciences. Behavior Research Methods, 39(2), 175-191. https://doi.org/10.3758/BF03193146
Whitley, E., & Ball, J. (2002). Statistics review 4: Sample size calculations. Critical Care, 6(4), 335-341. https://doi.org/10.1186/cc1521
How this DPH 860 Module 4 example is structured
Check the DPH 860 prompt in your Aspen classroom before using this example. It lists and justifies inputs, tables scenarios, works the main calculation, then covers missing data, smaller effects, the continuous outcome, subgroups, matching, software, post hoc power and implications.
DPH 860 Module 4 questions, answered
What does DPH 860 Module 4 usually ask for?
Aspen's DPH 860 covers power and sample size, so a calculation for the student's own project is typical. Follow your classroom prompt.
What inputs does a sample size calculation need?
Significance level, power, the expected comparison value, the smallest important difference and, for continuous outcomes, the standard deviation.
Why is post hoc power not useful?
Because it is determined by the observed P value and adds no new information; confidence intervals are more informative.
Write yours, or have the desk draft it
This paper is an original model document written by our desk, not a submitted student paper and not an official Aspen University document. Read it for the moves, then write your own to the instructions in your classroom. If you want one built to your exact prompt and rubric, the first custom sample is free and arrives in 24 to 48 hours.