When Well-Done Is Not the Same: Dissecting Communication Skill Assessment in Virtual versus In-Person OSCEs Among Medicine Residents

Highlight

This study evaluates communication skill assessment across virtual and in-person Objective Structured Clinical Examinations (OSCEs) among internal medicine residents using Item Response Theory (IRT). While overall communication ratings were similar across modalities, the analysis revealed task-specific differences in proficiency thresholds, notably in the ability to explain medical jargon. Such findings have implications for the design, interpretation, and comparability of OSCE assessments across modalities.

Study Background

Assessment of clinical communication forms a critical component of medical training and certification. The Objective Structured Clinical Examination (OSCE) is a widely used tool that measures specific clinical and communication competencies using standardized patients (SPs) and behaviorally anchored rating scales. Traditionally conducted in person, virtual OSCEs have gained prominence, especially accelerated by the COVID-19 pandemic, offering flexibility and access advantages. However, prior research has mainly compared overall performance scores between virtual and in-person OSCEs, without dissecting whether individual communication behaviors are evaluated equivalently across modalities. This gap risks masking subtle variations in skill demonstration or rating, which may impact the validity and fairness of OSCE-based assessments.

Study Design

This investigation analyzed data from 126 postgraduate year-1 internal medicine residents assessed through six case-based OSCEs conducted between 2019 and 2023. Of these, 82 encounters were in-person and 54 virtual. For consistency, the same internal medicine cases were used in both formats. Standardized patients rated resident communication across three domains: information gathering, relationship development, and patient education, using a behaviorally anchored scale with three ordinal ratings: “not done,” “partially done,” or “well done.” Outcomes included mean domain scores, rating distributions, and, crucially, item-level proficiency thresholds estimated with a graded response Item Response Theory (IRT) model. This approach quantifies the latent communication proficiency needed to transition between rating categories per item. Comparing thresholds between modalities with Wald tests allowed identification of items functioning differently across virtual and in-person encounters.

Key Findings

Overall communication scores and rating distributions did not significantly differ between virtual and in-person OSCEs, suggesting comparable aggregate performance in both formats. Most residents demonstrated adequate communication skill to earn “partly done” or “well done” ratings on the majority of checklist items.

However, item-level IRT threshold analysis highlighted that one specific communication behavior—”using words the patient understood and explaining medical jargon”—exhibited a significant difference in proficiency requirement between modalities. For this item, residents needed a lower level of underlying communication proficiency to achieve a “partly done” rating in virtual encounters compared to in-person (z = 3.92, adjusted p = 0.001). This signifies that either the virtual format facilitates easier attainment of basic recognition for this skill or that raters apply different standards or perceive this behavior differently when assessed virtually. Other communication items did not display statistically significant threshold differences.

These findings illustrate that despite similar overall score profiles, the meaning of specific ratings—particularly for nuanced communication behaviors—may not be uniform across OSCE modalities. This is an important consideration for educators and program directors interpreting longitudinal or comparative assessment results when shifting between or combining virtual and in-person OSCE formats.

Expert Commentary

The application of Item Response Theory in clinical skills assessment provides a sophisticated framework to parse item functioning beyond aggregate scores. This study adeptly demonstrates how IRT threshold parameters can reveal latent differences in communication proficiency standards tied to OSCE modality.

Differences observed in explaining medical jargon could stem from challenges unique to virtual encounters—such as altered patient cues, audio-visual limitations, or modified interaction dynamics—which may either exaggerate or obscure resident ability in this domain. Alternatively, standardized patients might rate this behavior differently due to modality-specific expectations or perceptions. Future research incorporating rater training standardization and qualitative rater feedback would help clarify these mechanisms.

Importantly, this work alerts medical educators to the risk of assuming direct equivalency in OSCE scores across modalities without granular psychometric evaluation. As virtual assessments become increasingly prevalent, tools like IRT threshold analysis can support validity arguments and guide tailored formative feedback, remediation, and summative judgment strategies.

Conclusion

This study provides meaningful evidence that while overall communication competency scores in internal medicine OSCEs appear similar between virtual and in-person modalities, individual communication behaviors may be rated at different levels of proficiency depending on context. The demonstrated difference in explaining jargon underscores potential modality-specific biases or interaction effects in clinical communication assessment. Incorporating Item Response Theory threshold analyses in OSCE validation extends understanding of scoring equivalency and helps ensure fair, accurate appraisal of resident communication skills regardless of testing format. Medical education programs should consider these nuances when designing, interpreting, and transitioning between virtual and in-person OSCE formats to maintain the integrity and meaningfulness of communication skill evaluations.

Funding and ClinicalTrials.gov

The original publication did not specify funding sources or clinical trial registration details.

References

Beltran CP, Nallamaddi S, Wilhite JA, Hardowar K, Hanley K, Altshuler L, Zabar SR, Gillespie C. When “Well-Done” Is Not the Same: Item Response Theory Analysis of Medicine Residents’ OSCE Communication Ratings Across Modalities. J Gen Intern Med. 2026 Sep 22. doi: 10.1007/s11606-026-4277-3. PMID: 42773393.

Cook DA, Holmboe ES. Translating OSCE scores into meaningful judgments: the promise of item response theory models. Med Educ. 2015 Feb;49(2):115-7. doi:10.1111/medu.12634.

Bokken L, Rethans JJ, van Heurn L, et al. Virtual patients in medical education: design and dimensions for research. Adv Health Sci Educ Theory Pract. 2010 Aug;15(3):495-502. doi:10.1007/s10459-009-9190-1.

Hawkins RE, Perfetto J, Brandt B, et al. Exploring best practices for equating scores from multiple assessment methods. Acad Med. 2015 Jul;90(7):932-8. doi:10.1097/ACM.0000000000000711.

Comments

No comments yet. Why don’t you start the discussion?

Leave a Reply