The Dataset's Instruction Manual: Building and Using a Codebook for a DPH Project
Student Name
Doctor of Public Health Program, Aspen University
DPH 860: Advanced Biostatistics
Instructor Name
Month Day, Year
The Dataset's Instruction Manual: Building and Using a Codebook for a DPH Project
A dataset without a codebook is a puzzle. Variable names may be cryptic, categories unexplained and missing values indistinguishable from real ones. A codebook documents every variable so that the analyst, the committee and future researchers can understand and reproduce the work. This paper builds and applies a codebook for a DPH project evaluating a blood pressure program delivered by community health workers.
What a Codebook Contains
For each variable, a codebook records its name, a descriptive label, its type, allowed values or ranges, units, the meaning of each code, how missing values are coded and, for derived variables, how they were calculated. It also records the data source and the date of extraction.
Selected Codebook Entries
The table shows selected entries.
| Name | Label | Type | Values | Missing code |
|---|---|---|---|---|
| age_yrs | Age at enrollment in years | Continuous | 18-90 | -9 |
| sex | Sex recorded in clinic record | Nominal | 1 = female, 2 = male, 3 = other | -9 |
| ins_type | Insurance at enrollment | Nominal | 1 = Medicaid, 2 = private, 3 = Medicare, 4 = none | -9 |
| sbp_base | Baseline systolic pressure, mmHg | Continuous | 90-240 | -9 |
| sbp_12m | Systolic pressure at 12 months, mmHg | Continuous | 90-240 | -9 |
| chw_visits | Community health worker visits | Count | 0-60 | -9 |
| bp_ctrl_12m | Controlled below 140/90 at 12 months | Binary, derived | 0 = no, 1 = yes | -9 |
Naming Rules
Variable names use lowercase letters with no spaces, begin with a letter and mark time points consistently, as in sbp_base and sbp_12m. Consistent names reduce errors in code and make the dataset easier to read. Broman and Woo (2018) recommend such consistency and a rectangular layout in which each column holds a single variable and each row a single record.
Coding Categorical Variables
Categorical variables are stored as numeric codes with labels defined in the codebook. Codes should be consistent across variables, so that, for example, 1 always means yes where yes or no is recorded. Free-text entries, such as other insurance types, are reviewed and recoded into defined categories where possible.
Derived Variables
Blood pressure control at 12 months is derived from systolic and diastolic readings: it equals 1 if systolic pressure is below 140 and diastolic pressure below 90, and 0 otherwise. The codebook records this rule exactly, along with how the reading was selected when several were taken, so that anyone can reproduce the variable. When readings were taken on the same day, the lowest valid reading after rest is used, following clinic protocol.
Missing Data Codes
Missing values are coded as -9 with a companion variable noting the reason, such as no visit, refused or not recorded. Using a distinct code prevents missing values from being mistaken for real ones, and recording reasons helps judge whether data are missing at random.
Data Organization
Broman and Woo (2018) offer principles for organizing data: be consistent, write dates in a standard format, leave no cells empty, put one thing in a cell, make the data rectangular, create a data dictionary, avoid calculations in raw data files and do not use color to encode information. The project follows these principles in its data files.
Checking Data Against the Codebook
Before analysis, every variable is checked against its codebook entry: values outside allowed ranges, impossible combinations such as a 12-month reading before enrollment and unexpected codes are listed and resolved with the data source. In the project, 23 systolic readings above 240 mmHg were traced to data entry errors and corrected. Every correction is logged with its reason and source, so the audit trail can be reviewed by the committee.
Planning for Missing Data
The codebook's missing data information feeds the analysis plan. Sterne et al. (2009) describe multiple imputation as a way to handle missing data under the assumption that missingness depends on observed variables, and they caution that leaving the outcome out of the imputation model is a frequent error. The project records predictors of missingness so that imputation can be done properly.
Version Control
Raw data are stored unchanged, and all cleaning is done through documented code that produces a new analysis file. Each version of the codebook and dataset is dated. This allows any step to be traced and repeated and protects against accidental changes.
Reproducible Analysis
With a complete codebook, raw data and cleaning code, another analyst could reproduce every result. Reporting guidelines encourage authors to describe how variables were defined and handled so that readers can judge and replicate analyses (Lang & Altman, 2015).
Sharing the Codebook
The codebook will be included as an appendix to the doctoral project and shared with the health department. It also helps committee members understand results without needing to query the analyst about variable meanings.
Linking Files
Clinical and program records come from different systems. The codebook documents the linkage variable, the method used and the match rate, which was 96% for participants. Records that could not be linked are described so that readers can judge whether their exclusion might bias results.
Dates and Time Windows
The 12-month outcome uses the blood pressure reading closest to 365 days after enrollment, within a window of 300 to 430 days. The codebook records this rule, the date variables used and what happens when two readings fall on the same day. Clear time windows prevent inconsistent outcome definitions across participants.
Privacy in the Codebook
The codebook describes variables without revealing individual data. Identifiers such as names and addresses are removed after linkage, and the codebook notes which variables were removed and why. Dates are shifted consistently for each person to protect privacy while preserving intervals between events.
Codebooks for Collaborators
The committee's methodologist reviewed the codebook before analysis and suggested adding a variable recording which clinic took each blood pressure reading. That addition later allowed a sensitivity analysis examining differences in measurement practice across clinics.
Common Codebook Errors
Frequent errors include labels that do not match values, derived variables without their rules, missing codes that overlap real values, such as 0 for missing visit counts, and codebooks that fall out of date when the data change. Updating the codebook each time the dataset changes avoids these problems.
From Codebook to Analysis File
The analysis file is produced by a script that reads raw data, applies codebook rules, derives variables, flags out-of-range values and writes a clean file with a date stamp. Running the script again on the same raw data produces an identical file, which is the practical test of reproducibility.
Training Others
Research assistants who help with data checks are trained using the codebook. They learn each variable's definition and allowed values before touching the data, which reduces errors and ensures consistent decisions when ambiguous values appear.
Conclusion
A codebook turns a collection of numbers into an understandable, reproducible dataset. By documenting names, labels, types, values, missing codes and derived variables, following sound data organization principles and checking data against the codebook, the DPH project builds a foundation for trustworthy analysis and reporting.
References
Broman, K. W., & Woo, K. H. (2018). Data organization in spreadsheets. The American Statistician, 72(1), 2-10. https://doi.org/10.1080/00031305.2017.1375989
Lang, T. A., & Altman, D. G. (2015). Basic statistical reporting for articles published in biomedical journals: The "Statistical Analyses and Methods in the Published Literature" or the SAMPL guidelines. International Journal of Nursing Studies, 52(1), 5-9. https://doi.org/10.1016/j.ijnurstu.2014.09.006
Sterne, J. A. C., White, I. R., Carlin, J. B., Spratt, M., Royston, P., Kenward, M. G., Wood, A. M., & Carpenter, J. R. (2009). Multiple imputation for missing data in epidemiological and clinical research: Potential and pitfalls. BMJ, 338, Article b2393. https://doi.org/10.1136/bmj.b2393
How this DPH 860 Module 5 example is structured
Check the DPH 860 prompt in your Aspen classroom before using this example. It explains what a codebook contains, tables selected entries, then covers naming, categorical and derived variables, missing codes, organization, data checks, missing data planning, version control, reproducibility and sharing.
DPH 860 Module 5 questions, answered
What does DPH 860 Module 5 usually ask for?
Aspen's DPH 860 covers codebooks, so a paper building or using a codebook for the project is typical. Confirm with your classroom prompt.
What should a codebook include?
Variable names, labels, types, allowed values, units, code meanings, missing value codes, derivation rules and data sources.
Why use a distinct missing value code?
So missing values are not mistaken for real data and reasons for missingness can be tracked.
Write yours, or have the desk draft it
This paper is an original model document written by our desk, not a submitted student paper and not an official Aspen University document. Read it for the moves, then write your own to the instructions in your classroom. If you want one built to your exact prompt and rubric, the first custom sample is free and arrives in 24 to 48 hours.