| Course | MPH 560 Applied Biostatistics for Public Health |
|---|---|
| Module | Module 7 |
| Paper type | Correlation and regression paper |
| Length | About 1,060 words, 6 pages |
| Format | APA 7 student paper |
| School | Aspen University |
| Program | Master of Public Health |
| Updated | September 2026 |
Free sample paper for MPH 560 Module 7
Measuring Relationships: Correlation and Linear Regression With Age and Blood Pressure
Student Name
Master of Public Health Program, Aspen University
MPH 560: Applied Biostatistics for Public Health
Instructor Name
Month Day, Year
Measuring Relationships: Correlation and Linear Regression With Age and Blood Pressure
Public health often asks how two numerical variables are related: does blood pressure rise with age, does air pollution relate to asthma visits, does income relate to life expectancy? Correlation measures the strength of a linear relationship, and regression describes it with an equation that can be used for prediction. This paper explains both through composite data on age and systolic blood pressure.
The Composite Data
Ten adults aged 25 to 70, in five-year steps, had systolic readings of 114, 117, 121, 119, 126, 130, 133, 138, 136 and 145 mm Hg. A scatter plot would show points rising from lower left to upper right, close to a straight line.
Correlation
Pearson's correlation coefficient, r, ranges from -1 to 1. Values near 1 indicate a strong positive linear relationship, near -1 a strong negative one and near 0 little linear relationship. For these data, r equals 0.98, a very strong positive correlation (Bewick et al., 2003). Real population data are usually much noisier; this composite example was built to show the calculations clearly. Correlation is symmetric: the correlation of age with pressure is the same as that of pressure with age. Regression, by contrast, treats one variable as the outcome and the other as the predictor.
Regression
Simple linear regression fits the line that minimizes the squared vertical distances between points and the line. The table summarizes the results.
| Quantity | Value | Meaning |
|---|---|---|
| Mean age | 47.5 years | Center of x |
| Mean systolic pressure | 127.9 mm Hg | Center of y |
| Slope | 0.66 mm Hg per year | Average change in pressure per year of age |
| Intercept | 96.6 mm Hg | Predicted pressure at age 0 (not meaningful) |
| Correlation, r | 0.98 | Strength of linear relationship |
| R-squared | 0.96 | Share of variation in pressure explained by age |
Interpreting the Slope and Intercept
The slope indicates that, on average, systolic pressure was 0.66 mm Hg higher for each additional year of age in these data, or about 6.6 mm Hg per decade. The intercept, 96.6, is the predicted value at age zero, which lies far outside the data and has no practical meaning; it simply anchors the line. The slope should be reported with its confidence interval so readers can judge its precision, as statisticians have long recommended for all estimates (Gardner & Altman, 1986).
Prediction
The equation predicted pressure = 96.6 + 0.66 times age allows prediction within the range of the data. For a 50-year-old, the prediction is 96.6 plus 33.0, or about 129.6 mm Hg. Predicting outside the observed range, for a 90-year-old or a child, is extrapolation and can be badly wrong, because the relationship may not stay linear.
R-Squared
R-squared, the square of r, is 0.96, meaning 96% of the variation in systolic pressure in these 10 people is accounted for by age. In real data, age typically explains far less, because diet, genetics, medication and other factors also matter.
Correlation Is Not Causation
A strong correlation does not prove that one variable causes another. Both may be driven by a third factor, the relationship may run in the opposite direction or the association may arise by chance or bias. Causal conclusions require study designs and reasoning that address these alternatives.
Assumptions and Diagnostics
A straight-line model needs the relationship to be roughly linear, the observations independent, the scatter around the line about even and, for tests and intervals, residuals that are close to normal. Plotting residuals, the differences between observed and predicted values, against predicted values reveals curves, funnel shapes or outliers that signal problems (Bewick et al., 2003).
Outliers and Influence
A single unusual point can change a regression line substantially, especially in small samples. Analysts should examine outliers, check for data errors and report how results change if influential points are removed. In skewed data, transforming variables or using rank correlation may be more appropriate (Whitley & Ball, 2002). A sensitivity analysis that reruns the regression without a suspect point shows readers whether the conclusion depends on it.
Toward Multiple Regression
Public health questions usually involve several variables. Multiple regression estimates the relationship between an outcome and one variable while holding others constant, which helps address confounding. For example, the association between income and blood pressure can be examined while adjusting for age and sex.
Confidence Intervals and Tests for the Slope
The slope has its own standard error, confidence interval and test of whether it differs from zero. A slope whose interval excludes zero indicates an association unlikely to be due to chance alone. Reporting the interval conveys how precisely the relationship is estimated.
Ecological Correlations
Correlations calculated from group data, such as county averages, can differ from those within individuals, a problem known as the ecological fallacy. A strong correlation between county income and life expectancy does not mean every wealthier person lives longer. Conclusions should match the level at which data were collected.
Using Regression in Public Health Practice
Health departments use regression to adjust rates for age, to estimate trends over time and to identify factors associated with outcomes, such as neighborhood characteristics linked to asthma visits. Clear reporting of the model, variables and assumptions allows others to evaluate the results.
Nonlinear Relationships
Not all relationships are straight lines. Blood pressure may rise faster at older ages, and some exposures have thresholds or plateaus. Plotting the data first shows whether a straight line fits. Curves can be modeled by adding squared terms or using other methods.
Reporting Regression Results
A clear report states the outcome and predictor, the slope with its confidence interval, the intercept if meaningful, R-squared, the sample size and any checks of assumptions. Graphs of the fitted line with the data help readers see the relationship.
Correlation Versus Agreement
A high correlation does not mean two measures agree. Two blood pressure devices might correlate at 0.95 while one reads consistently 10 mm Hg higher. Agreement between methods is assessed with other approaches, such as plotting differences against averages, not with correlation alone.
Conclusion
Correlation and regression describe and quantify relationships between numerical variables. In the composite data, age and systolic pressure had a correlation of 0.98 and a slope of 0.66 mm Hg per year, with 96% of variation explained. Interpreting slopes carefully, avoiding extrapolation, checking assumptions and remembering that correlation is not causation allow these tools to inform public health responsibly.
References
Bewick, V., Cheek, L., & Ball, J. (2003). Statistics review 7: Correlation and regression. Critical Care, 7(6), 451-459. https://doi.org/10.1186/cc2401
Gardner, M. J., & Altman, D. G. (1986). Confidence intervals rather than P values: Estimation rather than hypothesis testing. BMJ, 292(6522), 746-750. https://doi.org/10.1136/bmj.292.6522.746
Whitley, E., & Ball, J. (2002). Statistics review 1: Presenting and summarising data. Critical Care, 6(1), 66-71. https://doi.org/10.1186/cc1455
What the MPH 560 Module 7 instructions ask for
Aspen describes MPH 560 as preparing students to use quantitative methods applied by public health professionals, and because the seventh module's prompt is available only in the classroom, this example covers correlation and regression. Assignments on this topic usually give paired data and ask for a correlation, a regression equation and interpretation. Plot the data first. Report r, the slope, the intercept and R-squared. Interpret the slope in the variables' units. Say whether the intercept is meaningful. Predict only within the data range. Discuss causation and assumptions. Give the equation of the fitted line in words and symbols. Mention any unusual points and how you handled them, including whether results changed without them.
How this MPH 560 Module 7 example is built
Seventeen headings organize a little over 1,000 words here, and a three-column regression table reports six quantities with their meanings. The paper presents the data and explains correlation, then gives the table and interprets slope, intercept, prediction and R-squared. Causation, assumptions, outliers and multiple regression follow, along with the slope's interval, ecological correlations, practical uses, nonlinear relationships, reporting and correlation versus agreement. A margin comment notes that listing the raw data lets readers check each calculation. The last paragraph summarizes responsible use of these tools. The prediction for a 50-year-old is calculated in full so readers can check it. Each quantity in the table is explained in the text that follows, from slope to R-squared.
Reading the MPH 560 Module 7 grading rubric
Correlation and regression papers earn credit for correct calculation, accurate interpretation in context, attention to assumptions and caution about causation and extrapolation. This paper cites a statistics review on correlation and regression, a review on presenting data and a classic article on confidence intervals, formatted in APA. The table's values can be reproduced from the listed data. The intercept is correctly described as not meaningful. Residual diagnostics and outliers are addressed. Graders reward the distinction between correlation and agreement. Clear warnings about extrapolation and ecological reasoning, and a residual plot or its description, show understanding beyond calculation. Explaining that R-squared depends on the range of the data adds depth.
MPH 560 Module 7 help: mistakes that cost marks
Students often interpret a high correlation as causation or extrapolate far beyond the data. Others report r without a plot or interpret the intercept literally. Show a scatter plot or describe its shape. Interpret the slope with units. Predict only within range. Add a sentence on confounding. If your regression output is confusing, our tutors can explain each line of a typical software table. Finish with what other variables you would add in a multiple regression. Report the sample size and the range of each variable. Explain what a one-unit change in the predictor means in real terms. Mention at least one confounder that multiple regression could address.
Write yours, or have the desk draft it
This paper is an original model document written by our desk, not a submitted student paper and not an official Aspen University document. Read it for the moves, then write your own to the instructions in your classroom. If you want one built to your exact prompt and rubric, the first custom sample is free and arrives in 24 to 48 hours.
More MPH 560 and Master of Public Health sample papers
- MPH 560 Module 1: Descriptive Statistics
- MPH 560 Module 2: Probability in Public Health
- MPH 560 Module 3: Sampling and Confidence Intervals
- MPH 560 Module 4: Hypothesis Testing and P Values
- MPH 560 Module 5: Comparing Groups
- MPH 560 Module 6: Nonparametric Tests
- MPH 560 Module 8: Appraising Published Statistics
- MPH 570 Module 1: Evidence-Based Public Health
- MPH 530 Module 5: Climate, Weather and Health
- MPH 510 Module 3: Case-Control Studies and Odds Ratios
- MPH 590 Module 6: Capstone Discussion
MPH 560 Module 7 questions, answered
What does MPH 560 Module 7 usually ask for?
Aspen's MPH 560 covers quantitative methods used by public health professionals, so correlation and regression analysis is a typical assignment. Check the prompt in your classroom.
What does the regression slope mean?
The average change in the outcome for each one-unit increase in the predictor.
What does R-squared tell you?
The proportion of variation in the outcome explained by the predictor or predictors in the model.
Where can I find a free MPH 560 Module 7 sample paper?
The age and blood pressure regression paper is published on this page, with a table of slope, intercept, r and R-squared.
What is the difference between correlation and regression in MPH 560 Module 7?
Correlation measures the strength of a linear relationship; regression describes it with an equation used to predict one variable from another.