AdAdvertisement
← Back to News

Multi-Agent AI Improves AI-Generated Medical Exams

J

By João L. Carapinha

September 14, 2026

Artificial intelligence and machine learning
AI-generated medical exams

An open-source framework that divides medical exam-item writing across specialist AI agents has outperformed a single model in a blinded evaluation, suggesting the limits of AI-generated medical exams can be pushed well beyond current assumptions. The framework, called MAID, won 57.7% of expert preferences against 42.3% for a single-model baseline.

The single-model boundary

The work responds to a study by Qian et al., which tested DeepSeek’s ability to write in-training examination questions for radiology resident education. That team found single-model AI matched human experts on fact-based recall items, the A1-type questions, but fell short on higher-order clinical reasoning questions, the A3/A4-type items. They traced the gap to a lack of genuine clinical experience and tacit knowledge, which left generated scenarios feeling superficial, without narrative depth or authentic decisional tension.

How the MAID framework works

Zhehan Jiang’s answer is architectural. MAID, short for Multi-Agent Item Development, breaks item development into specialized roles that mirror how high-stakes exam programs such as the USMLE work. One Author Agent drafts the initial item. Three Reviewer Agents then critique it from separate angles: one checks factual and scientific accuracy, another clinical plausibility and reasoning quality, and a third item-writing quality, including distractor construction. An Editor Agent reconciles the feedback into a single revision directive and mediates the final rewrite.

Models are deliberately not tied to roles. GPT-4o, Claude 3.5, DeepSeek-R1, Qwen, Gemini, and MedSeek are randomly assigned across tasks so no single model’s biases or knowledge gaps propagate unchecked. Optional image generation and verification agents handle multimodal content, such as radiographs or CT scans.

Expert evaluation favors the multi-agent output

To test the approach, Jiang ran a cross-sectional, double-blind, two-alternative forced-choice comparison with 14 experts across seven disciplines: neurosurgery, gastroenterology, physiology, general surgery, hematology, cardiology, and neurology. Each expert compared 25 matched pairs of questions, giving 350 item-level observations.

The multi-agent items won 57.7% of preferences (95% CI [52.5%, 62.8%]) against 42.3% for the single-model baseline (95% CI [37.2%, 47.5%]; χ²(1) = 8.33, p = 0.004). A paired t-test confirmed the pattern held at the individual level: mean difference = 3.86, SD = 5.12, t(13) = 2.81, p = 0.015. The advantage showed up across all seven disciplines.

Implications for AI-generated medical exams

The findings suggest the ceiling for AI-generated medical exams written by a single large language model can be raised substantially through multi-agent design. Jiang stresses that human expertise remains essential, but the role shifts: instead of correcting basic structural problems, expert reviewers can focus on finer clinical judgment and pedagogical alignment. The study adds to a broader conversation about artificial intelligence in healthcare, where specialist architectures are increasingly tested against general-purpose models. The framework is available on GitHub at https://github.com/zjiang4/MAID.

The work was supported by the National Natural Science Foundation of China (Grant No. 72474004) and the Health Human Resources Development Center, National Health Commission in China (Grant No. 2026000164 & 2026000166).

Source: Jiang, Z. “Multi-Agent collaboration as a complementary architecture for AI-generated medical examination items.” npj Digital Medicine 9, 701 (2026). https://doi.org/10.1038/s41746-026-03187-z

Let Google know we are your trusted source.

Add our editorial as a preferred source in your search results.

Trust this Source