AI Models Near Perfect Scores on German Medical Licensing Exams
August 31, 2026


Frontier AI models now score close to 100% on the German medical licensing exams, but questions that include images still knock their accuracy down far more than they do human candidates. That is the central finding of a benchmark published in npj Digital Medicine, run in partnership with Seoul National University Bundang Hospital.
Researchers from Marburg University and the German Institute for State Examinations in Medicine, Pharmacy, Dentistry and Psychotherapy (IMPP) tested proprietary and open-weight foundation models on 7,485 official examination items from 24 exam sittings between 2019 and 2024. The dataset also carried student response data from 119,878 sittings, which let the team compare model difficulty with human difficulty item by item. Thirteen models took the shared text-only subset; eight vision-capable models took the full benchmark, images included.
Near ceiling on German medical licensing exams
On text-only items, the strongest models have effectively maxed out the test. Gemini 3.1 Pro scored 99.63% on the first exam (M1) and 98.86% on the second (M2), ahead of GPT-5.4 and Claude Opus 4.6. On the full benchmark, Gemini 3.1 Pro reached 99.31% on M1 and 98.37% on M2.
The finding with the most direct deployment relevance is how close the open-weight models came. GLM-5 hit 99.10% on M1 text items, DeepSeek V3.2-Thinking 98.73%, and Kimi K2.5 98.76%. Even compact models cleared the average student, who managed 71.21% on M1 text-only items and 74.66% on M2.
The image gap
The picture changed when items included images. Image questions were harder for everyone, but the penalty fell far more heavily on the models. Mean student accuracy dropped from 71.21% to 64.27% on M1 image items. Claude Opus 4.6 fell from 99.26% on text to 82.50% on images, and Qwen3-VL fell from 98.42% to 69.81%. Gemini 3.1 Pro held up best, slipping only 2.32 percentage points on M1.
Expressed as an error-rate ratio, student errors rose about 1.2-fold on image items, while most models’ errors rose several-fold. The authors note this ratio normalizes each solver against its own text baseline, which points to an extra, model-specific difficulty rather than just a generally harder question.
Human and model difficulty diverge
Model difficulty and student difficulty aligned only partially. Item-level correlation was moderate: ρ = 0.318 on M1 and ρ = 0.333 on M2. Items that were hardest for students were often easy for the models, and the reverse. Two difficulty-enriched subsets (human-hard versus model-hard) showed limited overlap, with the model-hard set doing a better job of separating the top systems.
What this means for deployment
Two implications stand out for health systems. First, aggregate text accuracy is no longer a useful discriminator at the frontier, so future benchmarks need to be modality-stratified and difficulty-enriched. Second, the strong open-weight results matter for Europe, where data protection and sovereignty rules often block cloud APIs. Local, privacy-preserving deployment of open-weight models for medical education and assessment is now realistic rather than aspirational, the authors argue.
The evaluation itself was deliberately austere: zero-shot prompting in German, single-pass inference, no retrieval, no external tools, and no ensembling. The authors also ran two contamination probes based on training cutoffs and exam chronology and found no evidence of substantial direct contamination, though they caution that indirect exposure through paraphrased study materials cannot be ruled out.
Let Google know we are your trusted source.
Add our editorial as a preferred source in your search results.
Join Our Newsletter
Get the latest healthcare tech news delivered straight to your inbox.




