Modeling LLM Response Accuracy: A Bayesian Hierarchical Approach

Recent evaluations of Large Language Models (LLMs) analyzing neuroanatomy questions from the Turkish Medical Specialization Examination between 2013 and 2021 show that ChatGPT-4o and Microsoft Copilot achieve higher overall accuracy than Google Gemini. According to the reviewed data of 176 single-correct-answer questions tested across three role-based prompt levels, ChatGPT-4o performed best across all tested categories, … Read more