Artificial Intelligence Performance Under Different Conditions in Answering China's Standardized Training Examination for Resident Physician in Radiology: A Comparative Analysis
{{output}}
Background: The capabilities of general-purpose large language models (LLMs) on specialized medical examinations have not been systematically compared. To evaluate the performance differences among three LLMs-DeepSeek-R1, ChatGPT... ...