AdAdvertisement
← Back to News

AI preoperative communication cuts anxiety and surgeon workload in prostate cancer

H

By HEOR Staff Writer

September 28, 2026

Asia
Stylised painting of patients in conversation before prostate cancer surgery, for Syenza News coverage of a randomized trial of AI preoperative communication

A randomized phase II trial published in npj Digital Medicine found that AI preoperative communication lowered anxiety scores and reduced surgeon consultation time for men scheduled for radical prostatectomy. Conducted at Fudan University Shanghai Cancer Center, the study compared an AI system using a locally deployed large language model against standard physician-led communication in 268 men with newly diagnosed prostate cancer.

Men in the AI-assisted group read answers to their own questions before the routine face-to-face conversation. They reported lower anxiety and better understanding of their illness than the control group, while the three surgeons who ran the consultations, all blinded to group allocation, reported less workload and spent close to half the time with each patient.

How the trial was run

Between July and December 2025, researchers screened 371 patients at FUSCC. Ninety-five were ineligible and 276 were randomized, but eight did not complete the scheduled preoperative communication. The final analysis covers 138 patients in the AI-assisted group and 130 in the control group. Because inpatients share rooms of three or four beds, randomization was done by room rather than by patient, to limit conversations between men in different arms. The trial is registered as NCT07082049 and was approved by the FUSCC institutional review board (No. 2505-Exp193).

Eligible patients were 18 to 80 years old, had histologically confirmed prostate cancer and were scheduled for radical prostatectomy, either laparoscopic radical prostatectomy (LAP-RP) or robot-assisted radical prostatectomy (RARP). They also needed a smartphone and enough education to complete the questionnaires, which were written in Mandarin. The sample size assumed a clinically meaningful 2.0 point difference in the final GAD-7 score, a standard deviation of 4.5 points, an intracluster correlation of 0.05 and an average cluster of about 3.8 patients, giving 116 evaluable patients per group for 85% power.

Both groups submitted questions through the same system. Only the AI-assisted group could log back in an hour later to read the generated answers, each of which had been approved by a content reviewer. Everyone then had the usual one-to-one preoperative conversation with one of three attending-level surgeons, who were instructed not to ask or infer whether a patient had used the AI answers. Baseline characteristics were balanced, including age (67.4 years, SD 6.0, in controls against 68.2 years, SD 6.2, in the AI group), median PSA at diagnosis (14.8 against 15.1 ng/mL) and baseline anxiety (GAD-7 9.8, SD 3.7, against 9.5, SD 3.5).

Anxiety fell furthest in the AI-assisted group

After the final preoperative conversation, the control group scored 5.7 (SD 3.1) on the GAD-7 against 3.2 (SD 2.7) in the AI-assisted group. A linear mixed-effects model adjusting for baseline GAD-7 score and physician, with a random intercept for room, estimated an adjusted mean difference of -2.2 points (95% CI -3.4 to -1.1; P = 0.001). The intracluster correlation coefficient was 0.03 (95% CI 0.01 to 0.06).

The visual analogue scale for anxiety moved the same way, ending at 3.8 (SD 1.8) in controls against 2.0 (SD 1.5) in the AI group (P < 0.001), from baseline scores of 4.8 and 5.1. Anxiety severity also shifted: the proportion of patients with severe or moderate anxiety was similar in both groups at baseline, but fewer AI-assisted patients remained in those categories at the final assessment.

Endpoint Group N Baseline mean (SD) Final mean (SD) Adjusted mean difference (AI minus control) 95% CI P
GAD-7 score Control 130 9.8 (3.7) 5.7 (3.1)
GAD-7 score AI-assisted 138 9.5 (3.5) 3.2 (2.7) -2.2 -3.4 to -1.1 0.001
NASA-TLX score Control 130 not available 56.8 (19.5)
NASA-TLX score AI-assisted 138 not available 39.9 (16.9) -10.6 -14.3 to -6.9 <0.001

Source: Table 3, primary outcomes in patients and physicians. GAD-7, Generalized Anxiety Disorder 7-item scale; NASA-TLX, NASA Task Load Index. Treatment effects are adjusted mean differences from linear mixed-effects models.

Surgeon workload and the near halving of consultation time

Surgeon workload fell alongside patient anxiety. NASA-TLX totals were 56.8 (SD 19.5) for consultations in the control group against 39.9 (SD 16.9) in the AI-assisted group, an adjusted mean difference of -10.6 points (95% CI -14.3 to -6.9; P < 0.001). Mental demand, temporal demand, effort and frustration all fell. Physical demand and perceived performance did not differ significantly between groups.

Recorded communication time fell from 19.9 (SD 6.1) minutes to 11.3 (SD 4.5) minutes overall, and the reduction held for each surgeon individually: 19.2 to 11.2 minutes for surgeon A, 20.2 to 11.3 minutes for surgeon B and 20.8 to 12.0 minutes for surgeon C, all at P < 0.001.

Secondary outcomes pointed the same way. Positive affect on the I-PANAS-SF rose from 9.2 (SD 2.7) to 11.3 (SD 3.0) and negative affect fell from 12.3 (SD 4.1) to 10.2 (SD 3.6). Illness perception scores on the B-IPQ dropped from 38.2 (SD 14.8) to 27.5 (SD 13.2), and a lower score here means better understanding of the condition. Satisfaction measured by the PSQ-18 rose from 46.8 (SD 9.5) to 57.0 (SD 9.0). All of these comparisons reached P < 0.001.

What AI preoperative communication did and did not replace

The system was a retrieval-augmented pipeline built around DeepSeek-R1, deployed on a local server cluster inside the hospital network and reached through an internal interface. A curated knowledge base of urology and oncologic surgery references, reviewed by experts, supplied the material. Retrieval combined dense semantic embeddings (bge-m3, indexed with FAISS) with sparse TF-IDF search, prioritised guideline-level content and ran a second pass through an outbound-only secure gateway only when nothing local was relevant enough. Only de-identified query text left the intranet.

Answers were capped at 500 words and written for patients with a lower secondary education, under a fixed clinical context of prostate cancer scheduled for radical prostatectomy. Every draft response went to a secure human review backend before release, and access was verified through system logs. The team plans to automate the triage of unsuitable questions in a phase III study, so the standard reply no longer needs a reviewer.

The design kept the physician in the loop. AI-assisted communication was a preparation step rather than a substitute, and the face-to-face discussion remained the point at which individual decisions were made. Against prerecorded videos or printed materials, the authors argue, a question-and-answer system can respond to what a particular patient is worried about, and its knowledge base can be updated as guidance changes, where fixed content has to be recorded again and again. This is closer to collaborative intelligence than to automation, and it matches the transparency and human oversight that payers and health technology assessment bodies now require when AI is used in clinical care. Earlier work in the field has ranged from clinical trial screening with GPT-4 to safety evaluation of LLM judges in multilingual health settings.

Source: Liu Z, Xu H, Lin G-w, et al. Large language model-assisted preoperative communication reduces patient anxiety and physician workload in prostate cancer: a prospective randomized phase II trial. npj Digital Medicine, article in press, 2026. Trial registration NCT07082049. The study code is available at github.com/SHMCLz/RAG-LLM, and de-identified datasets are available from the corresponding author on request. Analysis used R 4.5.2, PASS 11.0.7 and GraphPad Prism 10.2.1.

Let Google know we are your trusted source.

Add our editorial as a preferred source in your search results.

Trust this Source