Six Stations, Two Raters: Designing and Evaluating an Objective Structured Clinical Examination for a Health Assessment Course
Student Name
Master of Science in Nursing Program, Aspen University
N582: Teaching Strategies in Nursing Education
Instructor Name
Month Day, Year
Six Stations, Two Raters: Designing and Evaluating an Objective Structured Clinical Examination for a Health Assessment Course
In a composite baccalaureate program, students in the health assessment course have traditionally been graded on a single recorded head-to-toe examination of a classmate, scored by one faculty member with a long checklist. Faculty have noticed problems. Students memorize a sequence and perform it without interpreting findings, scores vary with the rater, and a student who is anxious on the day of the single examination can fail a course they otherwise understand. The course committee has proposed replacing the single examination with an objective structured clinical examination, known as an OSCE.
This paper designs that examination. It explains what an OSCE is and what the evidence says about its reliability and validity, sets out a blueprint of six stations linked to the course outcomes, describes the scoring tools and rater preparation, and explains how the results will be analyzed to improve both the examination and the course.
What an OSCE Is and What the Evidence Shows
Harden and Gleeson (1979) introduced the OSCE as a way to assess clinical competence through a series of timed stations, each testing a defined skill with a standardized patient, a model or a clinical scenario, and each scored against predetermined criteria. Because every student meets the same tasks under the same conditions, the format was designed to be more objective than a single observed encounter. It has since spread from medicine to nursing and other health professions.
The evidence supports the OSCE but warns against treating it as automatically reliable. In a narrative review focused on nursing, Rushforth (2007) concluded that OSCEs can make a meaningful contribution to assessment if new examinations and marking tools are carefully prepared and piloted, and if the length, number and interdependence of stations are balanced to protect both validity and reliability, while cautioning against using the OSCE as the only means of assessment. A meta-analysis of 188 reliability estimates from 39 studies found an average reliability across stations of only about 0.66, higher within stations, and better reliability when examinations used more stations and more than one examiner per station; communication skills were assessed less reliably across stations than clinical skills (Brannick et al., 2011). These findings shape the design that follows.
The Blueprint
A blueprint links each station to a course outcome so that the examination samples the course as a whole rather than one routine. The six stations, each ten minutes with two minutes between, are shown in the table.
| Station | Task | Course outcome | Scoring |
|---|---|---|---|
| 1 | Focused health history with a standardized patient reporting chest discomfort | Obtain a relevant history using therapeutic communication | Checklist plus communication global rating |
| 2 | Cardiac and peripheral vascular examination | Perform a systematic examination using correct technique | Checklist plus global rating |
| 3 | Respiratory examination with interpretation of recorded breath sounds | Distinguish normal from abnormal findings | Checklist plus short written interpretation |
| 4 | Abdominal examination on a standardized patient | Perform techniques in the correct sequence | Checklist plus global rating |
| 5 | Neurological screening of an older adult with new unsteadiness | Adapt assessment to age and presenting concern | Checklist plus global rating |
| 6 | Written documentation and SBAR report of findings from station 5 | Document and communicate findings accurately | Rubric |
Scoring Tools and Rater Preparation
Each station uses a short checklist of critical actions, no more than twelve items, combined with a global rating of overall performance on a five-point scale. Long checklists reward memorized sequences, while a global rating allows experienced raters to judge whether the student is interpreting findings and adapting to the patient, which addresses the first weakness of the current examination. Critical safety items, such as hand hygiene and patient identification, are flagged so that omitting them triggers a remediation requirement regardless of total score.
To address rater variation, every station except the written station has two raters: a faculty member and a trained clinical partner or teaching assistant. Raters attend a 90-minute calibration session in which they score recorded performances together and discuss disagreements until their scores converge. Standardized patients are trained with scripts so that each student receives the same information. The examination is piloted with a small group of volunteers from the previous cohort to test timing, clarity and scoring before it counts toward grades.
Standard Setting and Fairness
The passing standard is set before the examination using a panel method in which faculty judge how a borderline student would perform at each station, rather than using an arbitrary percentage. Students must pass four of the six stations and all critical safety items. Spreading the assessment across six stations also addresses the third weakness of the current exam: a single anxious performance no longer determines the outcome. Students with approved accommodations, such as extended time, receive them at every station, and all students complete a practice OSCE in week eight so that the format itself is not a source of anxiety.
Analyzing the Results
After the examination, the course committee analyzes results at three levels. For each station, it examines the distribution of scores, the agreement between the two raters and the correlation between checklist and global scores; low agreement signals a station or rater that needs attention. Across the examination, it estimates reliability, recognizing from the meta-analysis that values around 0.66 are common and that adding stations or examiners is the most reliable way to improve them. For the course, it looks for patterns: if many students perform the respiratory examination correctly but misinterpret the recorded sounds, the problem lies in teaching, not in the students, and the course should add practice with auscultation. Results are shared with faculty each semester and the blueprint is revised annually.
Conclusion
Replacing a single observed examination with a six-station OSCE addresses the weaknesses the faculty identified: memorized performance, rater variation and the weight placed on one attempt. The evidence shows that OSCEs are useful but not automatically reliable, and that more stations, more than one examiner, careful preparation and piloting all matter. Blueprinted to the course outcomes, scored with brief checklists and global ratings, and analyzed to improve both the examination and the teaching, the OSCE gives the health assessment course a fairer and more informative way to judge whether students are ready to assess real patients.
References
Brannick, M. T., Erol-Korkmaz, H. T., & Prewett, M. (2011). A systematic review of the reliability of objective structured clinical examination scores. Medical Education, 45(12), 1181-1189. https://doi.org/10.1111/j.1365-2923.2011.04075.x
Harden, R. M., & Gleeson, F. A. (1979). Assessment of clinical competence using an objective structured clinical examination (OSCE). Medical Education, 13(1), 39-54. https://doi.org/10.1111/j.1365-2923.1979.tb00918.x
Rushforth, H. E. (2007). Objective structured clinical examination (OSCE): Review of literature and implications for nursing education. Nurse Education Today, 27(5), 481-490. https://doi.org/10.1016/j.nedt.2006.08.009
How this N 582 Module 8 example is structured
Aspen does not publish N582 module prompts, so check your classroom for the exact instructions. This example identifies the weaknesses of a current assessment, defines the OSCE and appraises its evidence, presents a blueprint in a table, describes scoring tools and rater calibration, sets the standard fairly and ends with a three-level plan for analyzing results.
N582 Module 8 questions, answered
What does N582 Module 8 usually ask for?
Aspen's N582 description closes with designing, conducting and analyzing assessments and evaluations of learning outcomes, so a paper designing and evaluating an assessment is a typical final assignment. Check your classroom for the assessment type required.
What is an OSCE blueprint?
A table linking each station to a course outcome and a scoring method, showing that the examination samples the intended content and skills rather than a single routine.
How can OSCE reliability be improved?
By using more stations, more than one trained examiner per station, calibrated raters, piloted checklists and standardized patients who follow scripts.
Write yours, or have the desk draft it
This paper is an original model document written by our desk, not a submitted student paper and not an official Aspen University document. Read it for the moves, then write your own to the instructions in your classroom. If you want one built to your exact prompt and rubric, the first custom sample is free and arrives in 24 to 48 hours.