Highlight
• In a cross-sectional evaluation of offline FRCOphth Part 2 preparation questions, seven foundation models (FMs) showed strong performance on textual multiple-choice items; the best-performing FM (Claude 3.5 Sonnet) achieved 77.7% accuracy, comparable with expert ophthalmologists.
• Multimodal performance (questions that included images or other non-text inputs) remained substantially lower: the top multimodal FM (GPT-4o) scored 57.5%, underperforming expert clinicians and trainees.
