OpenAI’s o1 model correctly diagnosed 78.3% of cases in NEJM clinicopathologic conferences, outperforming human physicians. The model also maintained high accuracy on real-world, unstructured data from a major emergency department. #MedSky
Performance of a large language model on the reasoning tasks of a physician
More than 65 years ago, complex clinical diagnostic reasoning cases were introduced as the gold standard for the evaluation of expert medical computing systems, a standard that has held ever since.…
science.org