Posted in

AI Models Match Resident Axillary Ultrasound Diagnostics

A young Indian doctor reviewing radiology scans on a digital screen, symbolising the integration of clinical diagnostics with advanced imaging technology in 2025

Accurate axillary lymph node ultrasound evaluation remains a crucial component of breast cancer management and staging. Modern individualized breast cancer therapy relies heavily on accurate nodal characterization before surgical intervention. Recently, advanced vision-language models have shown potential as image-only decision-support tools in diagnostic radiology.

Evaluating AI Performance in Axillary Lymph Node Ultrasound

A retrospective study evaluated 718 primary invasive breast cancer patients who underwent node biopsy. Researchers collected grayscale and Doppler ultrasound images for each suspicious target node. Then, two inexperienced residents and two expert breast radiologists independently evaluated these images. Furthermore, three advanced vision-language models—GPT-5.2, Gemini-3 Pro, and Claude-Opus 4.5—classified the same images as metastatic or nonmetastatic.

Comparing Diagnostic Accuracy and Specificity

Experienced radiologists achieved the highest overall diagnostic accuracy at 83 percent. In contrast, inexperienced radiologists reached 76 percent accuracy with a high sensitivity of 92 percent. Among the artificial intelligence systems, GPT-5.2 performed best with 78 percent accuracy, 76 percent sensitivity, and 80 percent specificity. Additionally, Gemini-3 Pro achieved 73 percent accuracy, while Claude-Opus 4.5 achieved 67 percent accuracy. Consequently, experienced radiologists significantly outperformed all three artificial intelligence models and inexperienced readers in overall accuracy.

Clinical Applications for AI Decision Support

GPT-5.2 showed comparable accuracy to resident radiologists without a statistically significant difference. However, GPT-5.2 demonstrated higher specificity than inexperienced clinicians. This specific strength indicates that artificial intelligence can serve as a valuable supervised second reader. Therefore, implementing vision-language models may effectively reduce false-positive assessments made by junior radiologists in clinical practice.

Frequently Asked Questions

Q1: How did GPT-5.2 compare with inexperienced radiologists in axillary nodal assessment?

GPT-5.2 demonstrated overall accuracy similar to inexperienced radiologists, achieving higher specificity but lower sensitivity.

Q2: Did any AI model surpass experienced breast radiologists in diagnostic accuracy?

No AI model outperformed experienced radiologists, who maintained significantly superior diagnostic accuracy and specificity overall.

Q3: What is the main clinical value of using vision-language models in axillary ultrasound?

These models can serve as supervised decision-support tools to help junior radiologists reduce unnecessary biopsies and false-positive findings.

References

  1. He P et al. Vision Language Models for Ultrasound Assessment of Suspicious Axillary Lymph Nodes in Breast Cancer. AJR Am J Roentgenol. 2026 Jul 22. doi: 10.2214/AJR.26.35227. PMID: 42485405.

Leave a Reply

Your email address will not be published. Required fields are marked *