The Concordance Between Radiologists and AI in Identifying Nipple Shadows on Chest X-rays During Health Checkups at a Private Hospital in Thailand
Main Article Content
Abstract
OBJECTIVES: To evaluate the diagnostic performance of an AI system in differentiating nipple shadows from non-nipple abnormalities, identify imaging characteristics associated with AI misclassification, and compare its performance according to nipple-marker status among annual health examination participants with discordant initial AI–radiologist interpretations of chest radiographs.
MATERIALS AND METHODS: This retrospective study included adults undergoing annual health examinations who had discordant initial chest radiograph interpretations between the AI system and radiologists. Of 54,809 screening examinations, 1,564 showed discordant interpretations and underwent repeat chest radiography with nipple markers followed by final radiologist review. A simple random sample of 902 discordant examinations was analyzed. The final radiologist diagnosis served as the reference standard for distinguishing nipple shadows from non-nipple abnormalities. Within this discordant subset, diagnostic performance was assessed using sensitivity, specificity, overall accuracy, and the area under the receiver operating characteristic curve (AUC), each with a 95% confidence interval, and was stratified by nipple-marker status on the initial radiograph.
RESULTS: Among 54,809 screening chest radiographs, initial AI and radiologist interpretations were concordant in 53,245 examinations (97.1%) and discordant in 1,564 (2.9%). Of the 902 randomly selected discordant examinations, 418 (46.3%) had nipple markers on the initial radiograph, whereas 484 (53.7%) did not. Within this discordant subset, AI sensitivity, specificity, accuracy, and AUC were 66.3%, 51.1%, 52.7%, and 0.587, respectively. Compared with examinations performed with nipple markers, those performed without markers showed higher specificity (74.2% vs. 25.2%), accuracy (72.1% vs. 29.9%), and AUC (0.658 vs. 0.538), but lower sensitivity (57.4% vs. 82.4%).
CONCLUSION: AI demonstrated high initial concordance with radiologist interpretation in a large health-screening population, supporting its potential use as an adjunct to chest radiograph screening and workflow prioritization. These findings suggest that routine nipple-marker placement may not provide additional benefit for AI-based interpretation. A selective approach, reserving repeat radiography with nipple markers for cases of persistent radiologic uncertainty requiring radiologist confirmation, may be more appropriate.
Article Details

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.
This is an open access article distributed under the terms of the Creative Commons Attribution Licence, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.
References
Broder J. Imaging the chest: the chest radiograph. In: Broder J, editor. Diagnostic imaging for the emergency physician. Philadelphia: Elsevier Saunders; 2011. p. 185-296. doi: 10.1016/B978-1-4160-6113-7.10005-5.
Jones CM, Buchlak QD, Oakden-Rayner L, et al. Chest radiographs and machine learning—past, present and future. J Med Imaging Radiat Oncol. 2021;65(5):538-44. doi: 10.1111/1754-9485.13274.
Ferris RA, White AF. The round nipple shadow. Radiology. 1976;121(2):293-4. doi: 10.1148/121.2.293.
Knipe H. Nipple shadows. Radiopaedia.org [Internet]. [cited 2026 Jul 25]. Available from: https://radiopaedia.org/articles/nipple-shadows doi: 10.53347/rID-29788.
Miller WT, Aronchick JM, Epstein DM, et al. The troublesome nipple shadow. AJR Am J Roentgenol. 1985;145(3):521-3. doi: 10.2214/ajr.145.3.521.
Huang HY, Huang YH, Lin CH, et al. AI-assisted chest radiograph interpretation enhances diagnostic confidence and standardizes diagnostic accuracy across radiologists: a multi-reader study. Clin Imaging. 2026;130:110694. doi: 10.1016/j.clinimag.2025.110694.
Yoo H, Kim EY, Kim H, et al. Artificial intelligence-based identification of normal chest radiographs: a simulation study in a multicenter health screening cohort. Korean J Radiol. 2022;23(10):1009-18. doi: 10.3348/kjr.2022.0189.
Yoon SH, Park S, Jang S, et al. Use of artificial intelligence in triaging of chest radiographs to reduce radiologists’ workload. Eur Radiol. 2024;34(2):1094-103. doi: 10.1007/s00330-023-10124-1.
Plesner LL, Müller FC, Brejnebøl MW, et al. Using AI to identify unremarkable chest radiographs for automatic reporting. Radiology. 2024;312(2):e240272. doi: 10.1148/radiol.240272.
Lee E, Sayyouh M, Aslam A, et al. Utility of nipple markers in the era of digital imaging. J Thorac Imaging. 2023;38(1): 4-9. doi: 10.1097/RTI.0000000000000628.
Hunink MGM, Gazelle GS. CT screening: a trade-off of risks, benefits, and costs. J Clin Invest. 2003;111(11):1612-9. doi: 10.1172/JCI18842.
Yoshida K, Takamatsu A, Tanaka R, et al. Can AI substitute the first reader in chest radiograph screening? A retrospective non-inferiority evaluation. Jpn J Radiol. 2026;44(7):1168-76. doi: 10.1007/s11604-026-01973-z.
Tugwell-Allsup JR, England A, Owen BW. Educational insights of artificial intelligence-assisted chest radiograph interpretation in lung cancer detection—a pictorial review. BJR Artif Intell. 2025;2(1):ubaf002. doi: 10.1093/bjrai/ubaf002.
Vasilev YA, Vladzymyrskyy AV, Arzamasov KM, et al. Limitations of using artificial intelligence services to analyze chest X-ray imaging. Digit Diagn. 2024;5(3):407-20. doi: 10.17816/DD626310.
Plesner LL, Müller FC, Brejnebøl MW, et al. Commercially available chest radiograph AI tools for detecting airspace disease, pneumothorax, and pleural effusion. Radiology. 2023;308(3):e231236. doi: 10.1148/radiol.231236.
Yanagawa M, Tomiyama N. Clinical performance of current-generation AI tools for chest radiographs. Radiology. 2023;308(3):e232139. doi: 10.1148/radiol.232139.