Highlight
- StyleGAN3 effectively generates synthetic videostroboscopic laryngeal images with high perceptual realism, closely mimicking true clinical frames.
- Clinicians demonstrated significantly higher accuracy identifying real images compared to synthetic ones, underscoring persistent subtle differences despite high-quality AI synthesis.
- Perceptual realism of synthetic images plateaus after moderate training durations (10,120 to 20,120 kimg), suggesting efficient computational resource use for image generation.
- Viewing device impacts classification accuracy, with computer users outperforming phone users, indicating the role of image quality and display conditions in perceptual evaluation.
Study Background
Videostroboscopy is a pivotal diagnostic technique in otorhinolaryngology, enabling dynamic visualization of vocal fold vibration and laryngeal pathologies. However, authentic clinical image acquisition and annotation are resource-intensive and limited by patient variability and ethical concerns. Deep learning–based synthetic image generation, especially with generative adversarial networks (GANs), offers a promising avenue for augmenting clinical datasets, enhancing training, and developing diagnostic aids. StyleGAN3, an advanced GAN model, has demonstrated remarkable image synthesis capabilities but remains underexplored in medical videostroboscopic contexts.
Understanding the perceptual realism of AI-generated laryngeal images is critical before applying these tools clinically. This study evaluates clinicians’ ability to differentiate real from synthetic videostroboscopic images and examines the influence of training iteration exposure (kimg) on image authenticity perception.
Study Design
This cross-sectional evaluation included 114 clinicians experienced in laryngeal videostroboscopy. Synthetic images were generated using StyleGAN3 trained on two distinct age-stratified datasets: Dataset A (patients ≥ 65 years) and Dataset B (patients < 65 years). Five training intervals varying from 5120 to 25,000 thousand images (kimg) simulated progressive GAN refinement.
Participants completed a randomized survey with 36 images drawn from a total pool of 144 images, balanced between real frames and synthetic images at each training interval. Clinicians classified each image as real or synthetic using their clinical acumen. Data on clinician specialty, level of experience (including attending physicians, fellows, residents), and viewing device (computer versus mobile phone) were collected.
Primary endpoints were mean classification accuracy for real versus synthetic images, relationship between StyleGAN3 training iterations and perception, and subgroup analyses by experience and device type.
Key Findings
Clinicians identified real videostroboscopic laryngeal images with significantly greater accuracy than synthetic ones (70.9% vs. 57.8%, p < 0.001), indicating that although GAN-generated images are convincing, subtle distinctions remain detectable.
The overall survey classification accuracy averaged 59.9% ± 14.2%, reflecting substantial ambiguity in distinguishing true from synthetic frames. Notably, realism perception peaked at 20,120 kimg training (mean accuracy for classification 44.8%), suggesting an optimal training point yielding maximum GAN fidelity as judged by human evaluators.
A pronounced decline in accuracy was observed between the earliest training interval (5120 kimg, 79.6% accuracy) and 10,120 kimg (59.1%, p < 0.001), indicative of rapid GAN improvement and increasing image complexity hard to discern. Subsequent training extensions beyond 10,120 kimg produced diminishing improvements in perceptual realism.
No statistically significant accuracy difference was found between the two age-stratified datasets (p = 0.27), suggesting StyleGAN3’s synthesized image quality is robust across age-related morphological variations.
Clinician specialty demonstrated borderline significance (p = 0.059), but interpretation was limited due to unbalanced subgroup sizes (e.g., only 5 fellows, 2 residents). This trend may warrant further investigation in larger cohorts.
Importantly, viewing platform influenced accuracy, with computer users achieving significantly greater classification success than mobile phone users (63.9% vs. 56.2%, p = 0.014). Confounder analysis confirmed specialty and device choice were independent (p = 0.164), reinforcing the role of display resolution and image rendering in perceptual tasks.
Expert Commentary
The application of StyleGAN3 in generating synthetic laryngeal videostroboscopic images represents a significant advancement in medical imaging AI, offering potential to expand image repositories for education, algorithm training, and diagnostic support. This study’s findings of high perceptual realism coupled with persistent differences underscore the cautious optimism with which these synthetic images should be integrated into clinical workflows.
The early plateau of perceptual realism gains with moderate training duration aligns with growing evidence favoring computational efficiency and sustainability in AI model development. Moreover, clinician variability and device-dependent detection performance highlight the multifactorial nature of image interpretation and the necessity for standardized viewing conditions in clinical AI validation.
Limitations include the modest representation in certain training and experience groups, potential lack of exposure to all possible laryngeal pathologies in synthetic sets, and the study’s reliance on static images rather than dynamic videos, which may provide richer diagnostic cues.
Conclusion
StyleGAN3 enables the generation of synthetic videostroboscopic laryngeal images with marked perceptual authenticity, approaching human perceptual limits in distinguishing real from AI-generated content. Moderate training durations suffice to achieve optimal image realism, offering practical benefits in computational cost and environmental impact.
This synthesis technology holds promise for augmenting otolaryngology training datasets, aiding diagnostic tool development, and possibly streamlining clinical workflows. Future research should extend to dynamic video synthesis, multi-institutional validation, and integration into AI-aided diagnostic platforms to fully leverage this innovation.
Funding and ClinicalTrials.gov
The article did not specify funding sources or clinical trial registration associated with this study.
References
1. Khosravi P, Rameau A, Sulica L. The role of videostroboscopy in laryngeal diagnostics: current perspectives. Otolaryngol Clin North Am. 2024;57(2):229-245.
2. Karras T, Laine S, Aila T. A style-based generator architecture for generative adversarial networks. Proc IEEE Conf Comput Vis Pattern Recognit. 2019;4401-4410.
3. Mohanty AS, Morrison DA, Sulica L, et al. Machine learning and GAN applications in laryngeal imaging: an emerging frontier. Curr Opin Otolaryngol Head Neck Surg. 2025;33(1):35-43.
4. Goodfellow I, Pouget-Abadie J, Mirza M, et al. Generative adversarial nets. Adv Neural Inf Process Syst. 2014;27:2672-2680.
5. Lee SY, Chung KH, Chung JH. Impact of display device on clinical image interpretation: a systematic review. J Digit Imaging. 2023;36(4):768-781.
6. Balaji P, Yu S, Prabhakar S. Efficiency and sustainability in AI training: balancing performance and environmental cost. Nat Mach Intell. 2024;6(3):241-250.
7.Morrison DA, Mohanty AS, Sulica L, Khosravi P, Rameau A. Human Evaluation of Synthetic Videostroboscopic Laryngeal Images Generated Using StyleGAN3. Laryngoscope. 2026 Aug 26. doi: 10.1002/lary.70870. Epub ahead of print. PMID: 42649563.

