Artificial Intelligence In Radiology Reporting: The Role Of Natural Language Processing And Large Language Models
Özet
This study examines the role of natural language processing and large language models in radiology reporting within the context of current literature. Radiology reports are essential communication tools that include the clinical indication, imaging technique, radiological findings, and interpretation. However, free-text reports may pose limitations in data extraction, quality control, research applications, and clinical decision-support processes. In this context, natural language processing offers the potential to convert unstructured report narratives into analyzable and reusable data. The study outlines the evolution of this field from rule-based approaches to machine learning and deep learning, followed by Transformer architectures and large language models. It also discusses the potential applications of large language models in areas such as error detection in reports, impression generation, structured report creation, patient-friendly summarization, education, decision support, and workflow management. At the same time, key limitations including hallucination, data privacy, bias, explainability, ethical responsibility, and regulatory challenges are addressed. Overall, it is emphasized that the safest and most appropriate use of these technologies in radiology is as assistive tools under the supervision of radiologists rather than as autonomous diagnostic systems.
Referanslar
Casey A, Davidson E, Poon M, et al. A systematic review of natural language processing applied to radiology reports. BMC Medical Informatics and Decision Making. 2021;21:179. doi:10.1186/s12911-021-01533-7
Moezzi SAR, Ghaedi A, Rahmanian M, et al. Application of deep learning in generating structured radiology reports: A transformer-based technique. Journal of Digital Imaging. 2023;36(1):80-90. doi:10.1007/s10278-022-00692-x
Kreimeyer K, Foster M, Pandey A, et al. Natural language processing systems for capturing and standardizing unstructured clinical information: A systematic review. Journal of Biomedical Informatics. 2017;73:14-29. doi:10.1016/j.jbi.2017.07.012
Akinci D’Antonoli T, Stanzione A, Bluethgen C, et al. Large language models in radiology: Fundamentals, applications, ethical considerations, risks, and future directions. Diagnostic and Interventional Radiology. 2024;30(2):80-90. doi:10.4274/dir.2023.232417
Stephan D, Bertsch AS, Schumacher S, et al. Improving patient communication by simplifying AI-generated dental radiology reports with ChatGPT: Comparative study. Journal of Medical Internet Research. 2025;27:e73337. doi:10.2196/73337
Busch F, Hoffmann L, dos Santos DP, et al. Large language models for structured reporting in radiology: Past, present, and future. European Radiology. 2025;35(5):2589-2602. doi:10.1007/s00330-024-11107-6
Gupta A, Hussain M, Nikhileshwar K, et al. Integrating large language models into radiology workflow: Impact of generating personalized report templates from summary. European Journal of Radiology. 2025;189:112198. doi:10.1016/j.ejrad.2025.112198
Bhayana R. Chatbots and large language models in radiology: A practical primer for clinical and research applications. Radiology. 2024;310(1):e232756. doi:10.1148/radiol.232756
Thirunavukarasu AJ, Ting DSJ, Elangovan K, et al. Large language models in medicine. Nature Medicine. 2023;29(8):1930-1940. doi:10.1038/s41591-023-02448-8
Fanni SC, Tumminello L, Formica V, et al. The journey from natural language processing to large language models: Key insights for radiologists. Journal of Medical Imaging and Interventional Radiology. 2024;11:43. doi:10.1007/s44326-024-00043-w
Friedman C, Shagina L, Lussier Y, et al. Automated encoding of clinical documents based on natural language processing. Journal of the American Medical Informatics Association. 2004;11(5):392-402. doi:10.1197/jamia.M1552
Chapman WW, Bridewell W, Hanbury P, et al. A simple algorithm for identifying negated findings and diseases in discharge summaries. Journal of Biomedical Informatics. 2001;34(5):301-310. doi:10.1006/jbin.2001.1029
Aronson AR. Effective mapping of biomedical text to the UMLS Metathesaurus: The MetaMap program. Proceedings of the AMIA Symposium. 2001:17-21.
Savova GK, Masanz JJ, Ogren PV, et al. Mayo clinical Text Analysis and Knowledge Extraction System (cTAKES): Architecture, component evaluation and applications. Journal of the American Medical Informatics Association. 2010;17(5):507-513. doi:10.1136/jamia.2009.001560
Goldberg Y. A primer on neural network models for natural language processing. Journal of Artificial Intelligence Research. 2016;57:345-420. doi:10.1613/jair.4992
Hochreiter S, Schmidhuber J. Long short-term memory. Neural Computation. 1997;9(8):1735-1780. doi:10.1162/neco.1997.9.8.1735
Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need. Advances in Neural Information Processing Systems. 2017;30.
Devlin J, Chang MW, Lee K, et al. BERT: Pre-training of deep bidirectional transformers for language understanding. Proceedings of NAACL-HLT 2019. 2019:4171-4186.
Brown TB, Mann B, Ryder N, et al. Language models are few-shot learners. Advances in Neural Information Processing Systems. 2020;33:1877-1901.
Moor M, Banerjee O, Abad ZSH, et al. Foundation models for generalist medical artificial intelligence. Nature. 2023;616:259-265. doi:10.1038/s41586-023-05881-4
O’Sullivan JW, Palepu A, Saab K, et al. A large language model for complex cardiology care. Nature Medicine. 2026;32:616-623. doi:10.1038/s41591-025-04190-9
Gertz RJ, Dratsch T, Bunck AC, et al. Potential of GPT-4 for detecting errors in radiology reports. Radiology. 2024;311(1):e232714. doi:10.1148/radiol.232714
Sun C, Teichman K, Zhou Y, et al. Generative large language models trained for detecting errors in radiology reports. Radiology. 2025;315(2):e242575. doi:10.1148/radiol.242575
Sun Z, Ong H, Kennedy P, et al. Evaluating GPT-4 on impressions generation in radiology reports. Radiology. 2023;307(5):e231259. doi:10.1148/radiol.231259
Ziegelmayer S, Marka AW, Lenhart N, et al. Evaluation of GPT-4’s chest X-ray impression generation: A reader study on performance and perception. Journal of Medical Internet Research. 2023;25:e50865. doi:10.2196/50865
Li C, Wong C, Zhang S, et al. LLaVA-Med: Training a large language-and-vision assistant for biomedicine in one day. arXiv [Preprint]. 2023. doi:10.48550/arXiv.2306.00890
Tu T, Azizi S, Driess D, et al. Towards generalist biomedical AI. NEJM AI. 2024;1(3):AIoa2300138. doi:10.1056/AIoa2300138.
Nakaura T, Ito R, Ueda D, et al. The impact of large language models on radiology: A guide for radiologists on the latest innovations in AI. Japanese Journal of Radiology. 2024;42:1143-1156. doi:10.1007/s11604-024-01552-0
Kim SH, Wihl J, Schramm S, et al. Human-AI collaboration in large language model-assisted brain MRI differential diagnosis: A usability study. European Radiology. 2025;35(9):5252-5263. doi:10.1007/s00330-025-11484-6
Kim SH, Schramm S, Wihl J, et al. Boosting LLM-assisted diagnosis: 10-minute LLM tutorial elevates radiology residents’ performance in brain MRI interpretation. Neuroradiology. 2025;67(8):2069-2081. doi:10.1007/s00234-025-03664-4
Zaki HA, Aoun J, Abdullah KG, et al. The application of large language models for radiologic decision-making. Academic Radiology. 2024;31(5):2142-2147. doi:10.1016/j.acra.2024.01.007
Tan JR, Lim DYZ, Le Q, et al. ChatGPT performance in assessing musculoskeletal MRI scan appropriateness based on ACR appropriateness criteria. Scientific Reports. 2025;15:7140. doi:10.1038/s41598-025-88925-1
Prasad, S., Rao, A. S., Russo, M. V., Ghoshal, S., Roux, E., Vo, C., Kim, J., Hirsch, J. A., Lev, M. H., Gupta, R., Marks, W. H., Korchi, A., Landman, A., Bizzo, B. C., Raja, A. S., Dreyer, K. J., & Succi, M. D. (2026). Large language models as cost-conscious decision aids in emergency medicine: protocol support for imaging in lower back pain. Emergency radiology, 33(1), 53–59. https://doi.org/10.1007/s10140-025-02412-8
Gertz RJ, Bunck AC, Lennartz S, et al. GPT-4 for automated determination of radiologic study and protocol based on radiology request forms: A feasibility study. Radiology. 2023;307(5):e230877. doi:10.1148/radiol.230877
Benary M, Wang XD, Schmidt M, et al. Leveraging large language models for decision support in personalized oncology. JAMA Network Open. 2023;6(11):e2343689. doi:10.1001/jamanetworkopen.2023.43689
Contaldo MT, Pasceri G, Vignati G, et al. AI in radiology: Navigating medical responsibility. Diagnostics. 2024;14(14):1506. doi:10.3390/diagnostics14141506
U.S. Food and Drug Administration. Artificial intelligence-enabled device software functions: Lifecycle management and marketing submission recommendations. Draft guidance for industry and Food and Drug Administration staff. January 2025. Available from: https://www.fda.gov/regulatory-information/search-fda-guidance-documents/artificial-intelligence-enabled-device-software-functions-lifecycle-management-and-marketing. Accessed June 25, 2026.
Cross JL, Choma MA, Onofrey JA. Bias in medical AI: Implications for clinical decision-making. PLOS Digital Health. 2024;3(11):e0000651. doi:10.1371/journal.pdig.0000651
Koçak B, Ponsiglione A, Stanzione A, et al. Bias in artificial intelligence for medical imaging: Fundamentals, detection, avoidance, mitigation, challenges, ethics, and prospects. Diagnostic and Interventional Radiology. 2025;31(2):75-88. doi:10.4274/dir.2024.242854
Norori N, Hu Q, Aellen FM, et al. Addressing bias in big data and AI for health care: A call for open science. Patterns. 2021;2(10):100347. doi:10.1016/j.patter.2021.100347
U.S. Food and Drug Administration. Total product lifecycle considerations for generative AI-enabled devices. Executive summary for the Digital Health Advisory Committee meeting. November 20-21, 2024. Available from: https://www.fda.gov/media/182871/download. Accessed June 25, 2026.