Knowing the Data Before Testing It: Data Types and Descriptive Statistics for a DPH Project
Student Name
Doctor of Public Health Program, Aspen University
DPH 860: Advanced Biostatistics
Instructor Name
Month Day, Year
Knowing the Data Before Testing It: Data Types and Descriptive Statistics for a DPH Project
Every statistical analysis begins with understanding the data: what each variable measures, what kind of values it takes and how those values are distributed. Mistakes at this stage, such as averaging a category or ignoring skew, carry through every later test. This paper classifies and describes the variables in a DPH project evaluating a county blood pressure outreach program staffed by community health workers for 700 participants compared with 700 matched adults.
The Dataset
The project uses clinical records from a hospital system and community health centers linked to program enrollment data. Each record includes demographic characteristics, insurance, baseline and 12-month blood pressure, the number of community health worker visits and whether blood pressure was controlled below 140/90 at 12 months. Each comparison adult resembles a participant in age, sex, starting blood pressure, insurance and neighborhood.
Types of Data
Nominal variables, such as sex or insurance type, name categories without order. Ordinal variables, such as education level, have ordered categories with uneven spacing. Discrete numerical variables, such as the number of visits, take whole-number counts. Continuous variables, such as systolic blood pressure or age, can take any value within a range. The type determines which summaries and tests are appropriate.
Variables and Summaries
The table lists key variables, their type and the appropriate summary statistic.
| Variable | Type | Appropriate summary |
|---|---|---|
| Age | Continuous | Mean and standard deviation |
| Sex | Nominal | Count and percentage |
| Insurance type | Nominal | Count and percentage |
| Education level | Ordinal | Count and percentage; median category |
| Baseline systolic pressure | Continuous | Mean and standard deviation |
| Community health worker visits | Discrete, skewed | Median and interquartile range |
| Blood pressure controlled at 12 months | Binary | Count and percentage |
Center and Spread
For roughly symmetric continuous data, the mean and standard deviation summarize center and spread. Participants had a mean age of 56.4 years with a standard deviation of 11.2, and a mean baseline systolic pressure of 152.3 mmHg with a standard deviation of 14.8. Whitley and Ball (2002) note that these summaries are informative only when the distribution is not heavily skewed.
Skewed Data
The number of community health worker visits is right-skewed: most participants had between six and 12 visits, but a few had more than 25. The mean of 9.8 is pulled upward by these few, while the median of 9 better represents a typical participant. The interquartile range of 6 to 12 describes the middle half of participants without being distorted by extremes.
Categorical Variables
Categorical variables are summarized as counts and percentages. Among participants, 58% were women, 64% had Medicaid, 21% had private insurance and 15% were uninsured. Percentages should always be reported with the denominator, especially when data are missing, so readers know what the percentage represents.
Continuous Versus Dichotomized Outcomes
The project's primary outcome, blood pressure controlled below 140/90, turns a continuous measure into a yes or no. Altman and Royston (2006) caution that dichotomizing continuous variables discards information, reduces statistical power and can misrepresent risk, since a person at 139 mmHg is treated as very different from one at 141. The project therefore analyzes systolic pressure change as a continuous secondary outcome.
Describing Both Groups
A descriptive table comparing participants and matched comparison adults shows whether matching worked. Mean age was 56.4 and 56.1 years, women made up 58% and 57%, and mean baseline systolic pressure was 152.3 and 151.9 mmHg. Such similarity supports comparing outcomes, though unmeasured differences, such as motivation, may remain.
Outcome Description
At 12 months, 407 of 700 participants (58.1%) had controlled blood pressure, compared with 333 of 700 comparison adults (47.6%). Mean systolic pressure fell by 14.2 mmHg among participants and 8.6 mmHg among comparison adults. These descriptive results set up the inferential questions in later modules.
Missing Data
Twelve-month blood pressure was missing for 9% of participants and 11% of comparison adults. Describing who has missing data, for example whether they differ in age or baseline pressure, is part of describing the data and informs how missingness will be handled in analysis.
Reporting Guidance
The SAMPL guidelines recommend reporting means with standard deviations for approximately normal data, medians with interquartile ranges for skewed data, and numerators with denominators for percentages, and they advise against reporting standard errors as measures of spread (Lang & Altman, 2015).
Common Errors
Common errors include reporting means for ordinal scales, using standard errors instead of standard deviations to describe spread, omitting denominators and failing to check distributions before choosing summaries. Examining histograms and box plots for each continuous variable prevents most of these problems.
Checking Distributions
Before choosing summaries, each continuous variable was plotted. Age and baseline systolic pressure were roughly symmetric, with a few readings above 200 mmHg that were checked against source records and confirmed as real. Visit counts showed a long right tail. Histograms and box plots take minutes to produce and prevent many later errors, such as reporting a misleading mean.
Ordinal Data
Education was recorded in five ordered categories, from less than high school to graduate degree. Although the categories are numbered one to five, averaging them would assume equal spacing between levels, which does not hold. The project reports the percentage in each category and, where a single summary is needed, the median category.
Describing Change
Change in systolic pressure is calculated for each person as the 12-month reading minus the baseline reading, so negative values indicate improvement. Its distribution was roughly symmetric, with a mean and median both close to -14 mmHg among participants, supporting use of the mean. Describing change scores separately from baseline and follow-up values helps readers see the size of improvement directly.
Precision in Reporting
Summary statistics should be reported with sensible precision: blood pressure to one decimal place, percentages to one decimal place and ages to one decimal place. Excessive decimals imply more precision than the measurements support, while too few can hide meaningful differences.
Why Description Matters for Policy
Health department leaders will read the descriptive table before any model. Knowing that participants were mostly women with Medicaid, averaging 56 years old, tells them who the program reached and who it may have missed, such as younger men or uninsured residents. Description is therefore a finding in its own right, not only a preliminary step.
Conclusion
Classifying variables correctly and choosing summaries that fit their distributions lay the foundation for the DPH project's analysis. Means and standard deviations suit age and blood pressure, medians and interquartile ranges suit skewed visit counts, and counts with percentages suit categorical data. Careful description also shows that matching produced comparable groups and that the continuous outcome should not be lost to dichotomization.
References
Altman, D. G., & Royston, P. (2006). The cost of dichotomising continuous variables. BMJ, 332(7549), 1080. https://doi.org/10.1136/bmj.332.7549.1080
Lang, T. A., & Altman, D. G. (2015). Basic statistical reporting for articles published in biomedical journals: The "Statistical Analyses and Methods in the Published Literature" or the SAMPL guidelines. International Journal of Nursing Studies, 52(1), 5-9. https://doi.org/10.1016/j.ijnurstu.2014.09.006
Whitley, E., & Ball, J. (2002). Statistics review 1: Presenting and summarising data. Critical Care, 6(1), 66-71. https://doi.org/10.1186/cc1455
How this DPH 860 Module 1 example is structured
Check the DPH 860 prompt in your Aspen classroom before using this example. It describes the dataset and data types, tables variables and summaries, then covers center and spread, skew, categorical data, dichotomization, group description, outcomes, missing data, reporting guidance and common errors.
DPH 860 Module 1 questions, answered
What does DPH 860 Module 1 usually ask for?
Aspen's DPH 860 covers data types and describing data, so a paper classifying and summarizing project variables is typical. Follow your classroom prompt.
When should you report a median instead of a mean?
When data are skewed or contain extreme values, since the median is not pulled by outliers.
Why avoid dichotomizing continuous variables?
It discards information, reduces power and can misrepresent differences near the cut point.
Write yours, or have the desk draft it
This paper is an original model document written by our desk, not a submitted student paper and not an official Aspen University document. Read it for the moves, then write your own to the instructions in your classroom. If you want one built to your exact prompt and rubric, the first custom sample is free and arrives in 24 to 48 hours.