AI paper index

Can ChatGPT pass the polish national medical specialization examination in orthopedics and traumatology?

2026-08-10 · Archives of Orthopaedic and Trauma Surgery

One-line summary

An AI research paper on Can ChatGPT pass the polish national medical specialization examination in orthopedics and traumatology?.

Engineering notes

Engineering notes will be added by the aipentium editorial team.

Chinese explanation / 中文解读

中文解读待补充:本站会优先为大语言模型、生成式AI、ChatGPT相关技术、计算机视觉、深度学习等高价值论文补充中文说明。

Original abstract

Abstract Introduction Artificial intelligence (AI) has evolved rapidly in recent years and is becoming increasingly integrated into many areas of medicine. In the medical field, these systems have attracted considerable attention because of their potential applications in clinical decision support, medical education, scientific communication, and postgraduate training. The aim of the present study was to evaluate whether ChatGPT could achieve a passing score on the Polish National Medical Specialization Examination (Panstwowy Egzamin Specjalizacyjny, PES) in orthopedics and traumatology and to determine how different prompting strategies influenced its performance. Materials and methods Authors systematically assessed the performance of ChatGPT-4 across five consecutive official PES exams (from Autumn 2023 to Spring 2025) in orthopedics and traumatology. Each exam was administered using three distinct prompting strategies: Professor prompt, Specialist prompt, Resident prompt. For each exam session (e.g., Spring 2024, Autumn 2023), the prompts were submitted to ChatGPT sequentially. The model’s responses were evaluated against the official answer key published by the Polish Center of Medical Exams (Centrum Egzaminów Medycznych, CEM). Results The results were averaged across all prompts. The overall accuracy was 81% (range: 68.33%–85.83%). The highest score (85.83%) was recorded multiple times across different prompt types, indicating that the model could approach or exceed the minimum passing threshold depending on prompt structure. Agreement between prompting strategies was moderate to substantial (Cohen’s κ 0.518–0.632; Fleiss’ κ = 0.571). Radiology-related questions demonstrated the highest error rate among all question categories. Conclusions Large language models such as ChatGPT can perform well on knowledge-based orthopedic examinations and may serve as useful tools for examination preparation and educational support. However, examination success should not be interpreted as evidence of clinical competence or readiness for independent surgical practice. Specialist certification and patient care remain dependent on human expertise, practical skills, and professional responsibility.

5.0Engineering value
7.0Research novelty
4.0Business relevance

Links and sources

Need this topic turned into a technical roadmap?

aipentium can prepare a custom AI literature review, code map, dataset map, and B2B technology assessment.

Request B2B AI research

Comments

No comments yet. Be the first to share your thoughts on this paper.
Login or register to leave a comment