Background: Axillary management in patients with breast cancer is becoming increasingly individualized, requiring accurate nodal characterization. Objective: To evaluate the performance of vision-language models (VLMs) with image-only diagnostic input for detecting malignancy in biopsied axillary lymph nodes in patients with breast cancer and to compare performance with radiologists of different experience levels. Methods: This retrospective study included 718 patients (mean age, 53.1±11.4 years) with primary invasive breast cancer who underwent ultrasound-guided fine-needle aspiration or core-needle biopsy (FNA/CNB) of a suspicious target axillary lymph node from January 2025 through June 2025. A grayscale image and corresponding Doppler image, if available, were saved for a single biopsied node in each patient. Two inexperienced radiologists (both with only residency exposure to breast ultrasound) and two experienced radiologists (both breast imaging radiologists) reviewed images to classify nodes as metastatic or nonmetastatic; each pair reached consensus. The images were also inputted to three VLMs (GPT-5.2, Gemini-3-Pro, and Claude-Opus-4.5), with a text prompt requesting binary outputs (metastatic or nonmetastatic). Performance was compared between the reader groups and between the models and reader groups using Bonferroni adjustment and FNA/CNB as reference. Results: Accuracy, sensitivity, and specificity for inexperienced radiologists were 76%, 92%, and 64%; for experienced radiologists were 83%, 75%, and 89%; for GPT-5.2 were 78%, 76%, and 80%; for Gemini-3-Pro were 73%, 79%, and 68%; and for Claude-Opus-4.5 were 67%, 63%, and 70%, respectively. Accuracy was significantly higher for experienced radiologists compared with inexperienced radiologists and all three models; and for inexperienced radiologists compared with Claude-Opus-4.5. Sensitivity was significantly higher for inexperienced radiologists compared with experienced radiologists and all three models and for experienced radiologists compared with Claude-Opus-4.5. Specificity was significantly higher for experienced radiologists compared with inexperienced radiologists and all three models and for GPT-5.2 compared with inexperienced radiologists. Other comparisons were not significant. Conclusion: GPT-5.2 (the VLM with highest accuracy) showed no significant difference in accuracy versus inexperienced radiologists, albeit had lower sensitivity and greater specificity. GPT-5.2 had lower accuracy than experienced radiologists. Clinical Impact: The results support investigation of VLMs as a supervised decision-support tool to help reduce false-positive assessments by inexperienced radiologists.