Assessing the clinical readiness of foundation models in ophthalmology
Foundation models, a class of generative artificial intelligence (AI) systems that includes large language models (LLMs), are increasingly being evaluated for applications in medicine (1-4). Survey data suggest that physicians are already using LLMs at substantial rates (5). As evidence supporting these tools expands, clinicians may begin to incorporate them into educational and clinical workflows. Because these systems are evolving rapidly, ongoing benchmarking is essential to understand their capabilities and limitations in ophthalmology (6).
Many investigations evaluating LLM capabilities in medicine have relied on standardized question banks or board-style examination questions (7-9). These approaches allow reproducible comparisons across models and over time. Other study formats have evaluated LLM responses to clinical questions qualitatively by having physicians grade responses for characteristics such as subjective accuracy, completeness, readability, and overall quality (10,11). Although such evaluations provide insight into the quality of model-generated explanations, they are inherently subjective. Quantitative benchmarking using standardized examination questions therefore remains an important method for evaluating the medical knowledge and reasoning performance of these models. However, recent reports of near-perfect LLM performance on certain medical questions have raised concerns about potential contamination of training data when examination questions have previously appeared online or may have been encountered by models during training (12).
In this context, the study “Performance of Foundation Models vs Physicians in Textual and Multimodal Ophthalmological Questions” by Rocha and colleagues provides an important evaluation of foundation model performance using a multiple choice question set derived from a physical examination-preparation textbook that, according to the authors, was not readily available online at the time of the analysis (13). The authors assessed seven widely used foundation models, including GPT-4o (OpenAI), Gemini 1.5 Pro (Google), Claude 3.5 Sonnet (Anthropic), Llama-3.2-11B (Meta), DeepSeek V3 (High-Flyer), Qwen2.5-Max (Alibaba Cloud), and Qwen2.5-VL-72B (Alibaba Cloud). Model performance was compared with that of 10 physicians with varying levels of expertise, including junior physicians, ophthalmology trainees, and specialist ophthalmologists. The evaluation used questions derived from a preparation resource for the Fellowship of the Royal College of Ophthalmologists part 2 written examination, a standardized question set.
A key finding of Rocha et al.’s study is the strong performance of previous generation leading foundation models on ophthalmology questions. Claude 3.5 Sonnet achieved the highest overall accuracy and significantly outperformed junior physicians and ophthalmology trainees, with performance similar to specialist ophthalmologists. GPT-4o also demonstrated strong performance and exceeded that of earlier OpenAI models, including GPT-4 and GPT-3.5. These findings illustrate the rapid progress of foundation models in clinical reasoning and suggest that modern models may approximate expert-level performance on structured ophthalmic knowledge tasks (14).
However, an important limitation highlighted by Rocha et al.’s study is the persistent weakness of multimodal reasoning among foundation models. Although GPT-4o achieved the highest performance among the evaluated models on image-containing questions, its accuracy remained lower than that of both trainees and specialist ophthalmologists. All specialists in the study outperformed each of the foundation models on multimodal questions requiring image interpretation. This finding is particularly relevant for ophthalmology, where diagnostic reasoning frequently relies on interpretation of imaging modalities such as fundus photography, optical coherence tomography, and visual field reports.
Several methodological considerations should be acknowledged when interpreting these findings. First, only 40 questions in the dataset contained image-based components, 27 of which were generated by an ophthalmologist specifically for this study rather than taken from the Fellowship of the Royal College of Ophthalmologists part 2 written examination question set. This represents a relatively small portion of the total 385-question set and limits the strength of conclusions regarding multimodal reasoning. Second, the construction of the multimodal question set may introduce selection bias. Two image-based questions from the original textbook were excluded, and additional questions were incorporated to form the final dataset. As a result, the question set may not fully reflect the content and difficulty of the source examination. Third, the evaluation period of the models spanned several months in an environment where foundation models evolve rapidly. During such periods, models may undergo updates, architectural changes, or system refinements that are not publicly documented. Reporting the exact timing of model evaluations or repeating analyses at multiple time points may help clarify how model performance evolves over time.
Another limitation common to this area of research is the uncertain relationship between performance on board-style questions and real-world clinical performance. Standardized examination questions evaluate knowledge recall and structured reasoning under controlled conditions. Clinical decision-making requires interpretation of incomplete information, contextual judgment, and integration of patient-specific factors. Importantly, not all incorrect answers carry equal clinical significance. Clinicians are trained to prioritize the exclusion of vision- or life-threatening conditions and may use clinical reasoning to guide safer decision-making. In contrast, evaluation based on multiple-choice accuracy treats all incorrect responses equivalently and does not account for the potential clinical consequences of different types of errors. In addition, LLMs generate responses through internal processes that are often not visible to the user. This lack of transparency, often referred to as the “black box” problem, remains a barrier to physician trust and clinical adoption. Foundation models may also produce hallucinated outputs, which can reduce clinician trust, particularly when such errors are presented with high confidence or maintained over multiple queries. Concerns that these models may replace, rather than augment, the role of the physician may also contribute to reluctance among clinicians to adopt these tools.
Since the time period in which the models were evaluated in this study, newer generations of foundation models have been released. Recent foundation models have demonstrated improved performance on multimodal tasks, including image-based questions, although performance in these domains remains lower than on text-based tasks (15). The relative performance of different models continues to shift rapidly as capabilities improve (15,16). Because clinicians are increasingly using these systems to obtain medical information, continual benchmarking of newer models will remain important. Such evaluations can help inform clinicians about the strengths and limitations of available tools and guide responsible use in clinical and educational settings. Future benchmarking efforts may benefit from standardized evaluation datasets that incorporate multimodal inputs and higher-order questions designed to assess clinical reasoning and critical thinking rather than simple factual recall.
The increasing use of LLMs in medicine also raises important considerations regarding data privacy and security. Entering patient information into publicly available models may pose privacy risks, particularly when those models lack formal safeguards or regulatory protections. Development of secure clinical AI systems that comply with privacy regulations will be essential if these technologies are to be safely integrated into patient care.
Further research is needed to evaluate how foundation models interact with clinicians in real-world workflows. These systems may assist with tasks such as clinical documentation, synthesis of medical information, or educational support for trainees. Continued benchmarking studies similar to the study by Rocha and colleagues will be essential for understanding how different models perform across clinical domains and for informing their responsible integration into clinical and educational settings. Specifically, future studies should incorporate larger and more diverse multimodal datasets, curated offline datasets designed to minimize potential training data contamination, standardized benchmarking evaluations to enable longitudinal comparisons of model performance with testing conducted under conditions in which model training on input data is disabled, and evaluation against clinicians at different stages of training. Such efforts will be critical to ensuring that foundation models are evaluated rigorously for safe and effective integration into clinical practice.
Acknowledgments
The content is solely the responsibility of the authors and does not necessarily represent the official views of the National Institutes of Health.
Footnote
Provenance and Peer Review: This article was commissioned by the editorial office, Annals of Eye Science. The article has undergone external peer review.
Peer Review File: Available at https://aes.amegroups.com/article/view/10.21037/aes-2026-0022/prf
Funding: This work was supported by
Conflicts of Interest: Both authors have completed the ICMJE uniform disclosure form (available at https://aes.amegroups.com/article/view/10.21037/aes-2026-0022/coif). R.S.S. received grants from Research to Prevent Blindness Medical Student Eye Research Fellowship and Dean’s Research Scholarship, Keck School of Medicine of the University of Southern California. B.Y.X. received grants from Heidelberg Engineering, ArcScan, Ocular Therapeutix, Movu and Topcon. B.Y.X. received consulting fees from Movu and AbbVie. The authors have no other conflicts of interest to declare.
Ethical Statement: The authors are accountable for all aspects of the work in ensuring that questions related to the accuracy or integrity of any part of the work are appropriately investigated and resolved.
Open Access Statement: This is an Open Access article distributed in accordance with the Creative Commons Attribution-NonCommercial-NoDerivs 4.0 International License (CC BY-NC-ND 4.0), which permits the non-commercial replication and distribution of the article with the strict proviso that no changes or edits are made and the original work is properly cited (including links to both the formal publication through the relevant DOI and the license). See: https://creativecommons.org/licenses/by-nc-nd/4.0/.
References
- Betzler BK, Chen H, Cheng CY, et al. Large language models and their impact in ophthalmology. Lancet Digit Health 2023;5:e917-24. [Crossref] [PubMed]
- Wu LL, Hong AT, Davuluru SS, et al. Utility of ChatGPT-4o in Creating Patient Handouts in Ophthalmology: A Comparison With American Academy of Ophthalmology Educational Materials. Transl Vis Sci Technol 2026;15:14. [Crossref] [PubMed]
- Shool S, Adimi S, Saboori Amleshi R, et al. A systematic review of large language model (LLM) evaluations in clinical medicine. BMC Med Inform Decis Mak 2025;25:117. [Crossref] [PubMed]
- Zhang Z, Zhang H, Pan Z, et al. Evaluating Large Language Models in Ophthalmology: Systematic Review. J Med Internet Res 2025;27:e76947. [Crossref] [PubMed]
- Hong HJ, Shah NH, Pfeffer MA, et al. Physician Perspectives on Large Language Models in Health Care: A Cross-Sectional Survey Study. Appl Clin Inform 2025;16:1738-48. [Crossref] [PubMed]
- Cascella M, Semeraro F, Montomoli J, et al. The Breakthrough of Large Language Models Release for Medical Applications: 1-Year Timeline and Perspectives. J Med Syst 2024;48:22. [Crossref] [PubMed]
- Srinivasan S, Ai X, Lo TWS, et al. BEnchmarking Large Language Models for Ophthalmology (BELO): An Expert-Curated Data Set and Evaluation Framework for Knowledge and Reasoning. Ophthalmol Sci 2026;6:101050. [Crossref] [PubMed]
- Shean R, Shah T, Sobhani S, et al. OpenAI o1 Large Language Model Outperforms GPT-4o, Gemini 1.5 Flash, and Human Test Takers on Ophthalmology Board-Style Questions. Ophthalmol Sci 2025;5:100844. [Crossref] [PubMed]
- Shean R, Shah T, Pandiarajan A, et al. A comparative analysis of DeepSeek R1, DeepSeek-R1-Lite, OpenAi o1 Pro, and Grok 3 performance on ophthalmology board-style questions. Sci Rep 2025;15:23101. [Crossref] [PubMed]
- Yan Z, Liu J, Fan Y, et al. Ability of ChatGPT to Replace Doctors in Patient Education: Cross-Sectional Comparative Analysis of Inflammatory Bowel Disease. J Med Internet Res 2025;27:e62857. [Crossref] [PubMed]
- Pushpanathan K, Zou M, Srinivasan S, et al. Can OpenAI's New o1 Model Outperform Its Predecessors in Common Eye Care Queries?. Ophthalmol Sci 2025;5:100745. [Crossref] [PubMed]
- Alohali KI, Almusaeeb LA, Almubarak AA, et al. Reasoning-based LLMs surpass average human performance on medical social skills. Sci Rep 2025;15:36453. [Crossref] [PubMed]
- Rocha H, Chong YJ, Thirunavukarasu AJ, et al. Performance of Foundation Models vs Physicians in Textual and Multimodal Ophthalmological Questions. JAMA Ophthalmol 2026;144:5-13. [Crossref] [PubMed]
- Zhang J, Ma Y, Zhang R, et al. A comparative study of GPT-4o and human ophthalmologists in glaucoma diagnosis. Sci Rep 2024;14:30385. [Crossref] [PubMed]
- Shean RS, Mallapu JK, Shah T, et al. Comparative Performance of Gemini 3 Pro and GPT-5 Family Models on Ophthalmology Board-Style Questions. Ophthalmol Sci 2026;6:101145. [Crossref] [PubMed]
- Zhang P, Wang J, Hu X, et al. Comparative performance of GPT-4, GPT-o3, GPT-5, Gemini-3-Flash, and DeepSeek-R1 in ophthalmology question answering. Front Cell Dev Biol 2026;14:1744389. [Crossref] [PubMed]
Cite this article as: Shean RS, Xu BY. Assessing the clinical readiness of foundation models in ophthalmology. Ann Eye Sci 2026;11:39.

