General-Purpose LLMs Outperform Specialized Clinical AI on Medical Benchmarks
A study published in Nature Medicine has found that general-purpose large language models (LLMs) such as GPT-4 and Claude outperform specialized clinical AI tools across a range of medical benchmarks.
The research compared general-purpose LLMs against domain-specific clinical AI systems on various medical evaluation tasks. The results showed that the