DPH 860 Module 3 assignment: doctoral paper on hypothesis testing for group comparisons, a full sample

Reviewed by Douglas Renshaw, MBA Aspen University True APA form Annotated

A complete DPH 860 Module 3 example in true APA form: hypothesis testing for group comparisons in a DPH project on an east-side program in which community health workers support blood pressure care, with hypotheses, choosing tests, a table of comparisons, tests and results, worked t and chi-square tests, paired analysis for matching, a nonparametric test, assumptions, nonsignificant subgroups and multiple comparisons.

1

Chance or Program Effect? Hypothesis Testing for Group Comparisons in a DPH Project

Student Name

Doctor of Public Health Program, Aspen University

DPH 860: Advanced Biostatistics

Instructor Name

Month Day, Year

What this page is doingThe title poses the question hypothesis testing is designed to address. APA 7 student title page.
2

Chance or Program Effect? Hypothesis Testing for Group Comparisons in a DPH Project

Hypothesis testing asks whether observed differences between groups are larger than would be expected by chance if there were no true difference. Used well, it complements estimation; used poorly, it becomes a ritual of declaring results significant or not. This paper applies hypothesis tests to a DPH project that sets 700 enrollees of an east-side program, in which community health workers support blood pressure care, against an equal number of matched nonparticipants.

Hypotheses

For the continuous outcome, the null hypothesis holds that mean systolic change is equal across groups, and the alternative holds that it is not. For the binary outcome, the null states that the proportion with controlled blood pressure is the same; the alternative states that it differs. Two-sided tests are used, since the program could in principle make outcomes worse.

What this page is doingStating both hypotheses before any test shows the grader the logic is set in advance.
3

Choosing a Test

The appropriate test depends on the outcome type, the number of groups and whether observations are independent or paired. The table matches the project's comparisons to tests and results.

ComparisonOutcome typeTestResult
Systolic change, participants vs comparisonContinuousTwo-sample t testt = -6.44, P < .001
Blood pressure control, participants vs comparisonBinaryChi-square testChi-square = 15.5, P < .001
Systolic change within matched pairsContinuous, pairedPaired t testConsistent with unpaired result
Community health worker visits by insurance typeSkewed countMann-Whitney U testP = .21

The t Test for Systolic Change

The difference in mean systolic change was 5.6 mmHg, with a standard error of 0.87. The t statistic is the difference divided by its standard error, about 6.44. With nearly 1,400 degrees of freedom, this corresponds to P < .001. The data are highly unlikely under the null hypothesis of no difference, consistent with the confidence interval that excluded zero.

The Chi-Square Test for Control

Control was achieved by 58.1% of participants and 47.6% of comparison adults. Under the null hypothesis, the pooled proportion is 52.9%, and the difference's standard error is roughly 2.7 points. The resulting z statistic of about 3.94 corresponds to a chi-square value of about 15.5 with one degree of freedom and P < .001.

Accounting for Matching

Because comparison adults were matched to participants, the samples are not independent. Paired analyses, such as a paired t test on differences within pairs or McNemar's test for binary outcomes, account for this structure. In the project, paired results were consistent with unpaired ones, but the final analysis will use methods that respect matching.

Nonparametric Tests

Visit counts are skewed, so the comparison by insurance type uses the Mann-Whitney U test, a rank-based method that does not rely on means. The result, P = .21, provides little evidence of a difference, though this does not show that visits were equal across insurance groups.

Assumptions

The t test assumes approximately normal sampling distributions and, in its standard form, equal variances; with large samples and similar standard deviations, both are reasonable here. The chi-square test requires expected counts of at least five in each cell, easily met. Whitley and Ball (2002) emphasize checking such assumptions before interpreting results.

Nonsignificant Subgroup Results

Among adults over 70, the difference in control was 6 points with P = .34, based on only 118 participants and 118 matched comparison adults. That result cannot be read as proof the program does not help older adults. Altman and Bland (1995) caution that failing to find an effect is different from finding that there is none; the confidence interval, from about -7 to 19 points, is wide and includes both no effect and a benefit similar to the overall result.

Multiple Comparisons

Testing many subgroups raises the chance that some differences appear significant by chance. The project prespecifies two primary comparisons and treats subgroup analyses as exploratory, reporting them with intervals and without claims of confirmation.

Interpreting P Values

A P value indicates how well the data fit a particular statistical model, with the null hypothesis among its assumptions. It is not the probability that the null is true, and a small P value does not by itself show a large or important effect (Wasserstein & Lazar, 2016). Both main results are reported with estimates and intervals, not P values alone.

Reporting

Results will report the test used, the test statistic, degrees of freedom where relevant, exact P values except when below .001, the estimate of effect and its confidence interval. Stating tests in the methods section before results prevents choosing tests after seeing data.

Type I and Type II Errors

A type I error occurs when a test rejects a true null hypothesis; its probability is set by the significance level, here 0.05. A type II error occurs when a test fails to reject a false null; its probability is one minus power. With the project's large samples, the risk of a type II error for the main outcomes is low, but it is substantial for small subgroups.

One-Sided or Two-Sided

One-sided tests have more power to detect an effect in the expected direction but cannot detect harm. Because a program could in principle raise blood pressure, for instance by causing participants to stop medications they misunderstood, the project uses two-sided tests throughout. The choice was made before data were examined.

Tests and Estimates Together

Each test result is paired with its estimate and interval. The t test shows that the difference in systolic change is unlikely under the null hypothesis; the interval of 3.9 to 7.3 mmHg shows how large the difference plausibly is. Presenting both avoids the trap of treating statistical significance as the finding.

Software Output and Hand Checks

Tests were run in statistical software and checked by hand for the main comparisons. Hand calculations using the reported means, standard deviations and counts reproduced the software results closely, a simple step that catches data or coding errors before they reach the manuscript.

Choosing Between Paired and Unpaired Analysis

When matching is strong and the matching variables predict the outcome, paired analysis gains precision; when matching is weak, the gain is small. Because matched pairs share clinics and baseline pressure, the project expects a modest gain. Reporting both analyses in a supplement lets readers see that conclusions do not depend on the choice.

Communicating Test Results

For county leaders, test results are translated into plain statements: the difference in control between participants and matched adults is larger than would be expected by chance, and the likely size of the benefit is between 5 and 16 percentage points. Technical details remain available in the manuscript for readers who want them.

Conclusion

Hypothesis tests confirm that differences in systolic change and blood pressure control between program participants and matched adults would be very surprising if chance were the only explanation. Choosing tests that fit the outcome and design, accounting for matching, checking assumptions and interpreting nonsignificant and multiple results cautiously keep testing in service of estimation rather than replacing it.

References

Altman, D. G., & Bland, J. M. (1995). Absence of evidence is not evidence of absence. BMJ, 311(7003), 485. https://doi.org/10.1136/bmj.311.7003.485

Wasserstein, R. L., & Lazar, N. A. (2016). The ASA statement on p-values: Context, process, and purpose. The American Statistician, 70(2), 129-133. https://doi.org/10.1080/00031305.2016.1154108

Whitley, E., & Ball, J. (2002). Statistics review 3: Hypothesis testing and P values. Critical Care, 6(3), 222-225. https://doi.org/10.1186/cc1493

How this DPH 860 Module 3 example is structured

Check the DPH 860 prompt in your Aspen classroom before using this example. It states hypotheses, tables comparisons and tests, works the t and chi-square tests, then covers matching, nonparametric tests, assumptions, nonsignificant subgroups, multiple comparisons, P values and reporting.

DPH 860 Module 3 questions, answered

What does DPH 860 Module 3 usually ask for?

Aspen's DPH 860 covers hypothesis testing, so a paper applying tests to group comparisons is typical. Confirm with your classroom prompt.

Which test compares two proportions?

A chi-square test or a z test for two proportions; for matched pairs, McNemar's test.

Does a nonsignificant result mean no effect?

No; it may reflect limited data, so the confidence interval should be examined.

Write yours, or have the desk draft it

This paper is an original model document written by our desk, not a submitted student paper and not an official Aspen University document. Read it for the moves, then write your own to the instructions in your classroom. If you want one built to your exact prompt and rubric, the first custom sample is free and arrives in 24 to 48 hours.