AdAdvertisement
← Back to News

EpiScreen epilepsy detection model beats specialists on routine clinical notes

J

By João L. Carapinha

October 9, 2026

Clinical Practice
Abstract cubist artwork of a clinician writing notes beside a monitor showing brain wave traces, with dancers in a village street scene

The EpiScreen epilepsy detection model reads the notes a neurologist writes after referral and estimates whether seizure-like events are epileptic or psychogenic without inpatient video monitoring. Researchers at the University of Minnesota developed the system, which they reported in npj Digital Medicine. It achieved an AUC of up to 0.875 on the public MIMIC-IV dataset and 0.980 on a private University of Minnesota cohort. When practicing neurologists used the model, their accuracy rose by as much as 14.3 percentage points compared with unaided decisions.

Epilepsy affects roughly 50 million people worldwide, including about 2.9 million with active epilepsy in the United States. Psychogenic non-epileptic seizures (PNES), also called functional or dissociative seizures, look similar when they happen but have different causes, treatments and outlooks, so separating them early matters.

Why epilepsy and PNES are hard to tell apart

Between 20% and 30% of patients referred to tertiary epilepsy centers for what looks like drug-resistant epilepsy are ultimately diagnosed with PNES. Neurologists can weigh a range of semiological and historical features, among them ictal duration, eye closure and tongue biting, but none of these is pathognomonic. Some PNES patients display signs usually linked to epileptic seizures, and some epilepsy patients lack the classical markers, so rule-based and feature-dependent models perform poorly on the task. The narrative itself adds difficulty: descriptions of seizures often come from patients, family members or eyewitnesses, and phrases such as “loss of consciousness” or “unresponsiveness” appear in both conditions.

Prolonged video-electroencephalography (vEEG) remains the diagnostic gold standard because it correlates clinical events with electrical activity in the brain. A single admission in the United States commonly exceeds $15,000, typical wait times run 8 to 12 weeks, and monitoring itself lasts 1 to 5 days.

Detection accuracy in two health systems

On MIMIC-IV, the phenotype-based FSLS method reached an AUC of 0.695 and ClinicalBERT reached 0.769. EpiScreen performed better than every baseline on both datasets, with the Med42-70B version reaching 0.891 on MIMIC-IV. On the Minnesota cohort, ClinicalBERT reached 0.960 against 0.944 for FSLS, while the Qwen-2.5-32B and DeepSeek-R1-32B versions of EpiScreen both reached 0.992. Fine-tuning mattered more than prompting: relative to off-the-shelf language models, it produced average absolute AUC gains of 13.2% on MIMIC-IV and 11.6% on the Minnesota data.

Model MIMIC-IV AUC UMN AUC
FSLS (phenotype-based) 0.695 0.944
BERT 0.756 0.952
ClinicalBERT 0.769 0.960
EpiScreen (Qwen-2.5-32B) 0.887 0.992
EpiScreen (DeepSeek-R1-32B) 0.882 0.992
EpiScreen (Llama-3.3-70B) 0.889 0.974
EpiScreen (Med42-70B) 0.891 0.976

Source: Zhou et al., npj Digital Medicine, 2026 (Table 2). Training and testing used data from the same institution; values carry 95% confidence intervals.

Clinicians plus AI beat either alone

To test whether the output helps in practice, the researchers drew 300 cases from the MIMIC-IV test cohort by stratified random sampling and asked eight clinicians to classify them. Groups A, B and C each contained two neurology residents with an average of four years of training; Group D contained two senior neurologists with more than ten years of experience each. Decisions were made by within-group consensus, and clinicians could use the internet as they would in normal practice.

Working unaided, the three resident groups scored 0.793, 0.763 and 0.767, and the attendings scored 0.817, none of which differed significantly from the standalone model at 0.810. With model predictions and sentence-level attributions in front of them, the resident groups rose to 0.880, 0.890 and 0.877 and the attendings to 0.893. The errors moved in complementary directions: assistance cut false negatives among residents from 36 to 14, and cut false positives among attendings from 37 to 19. The overall gain was up to 14.3% over unaided clinicians and up to 10.2% over the model working on its own.

What comes next

EpiScreen is designed for the specialty-care stage, using notes generated during and after a neurology referral to support decisions before confirmatory video EEG. On the authors’ estimate that could shorten diagnostic delays by roughly 8 to 12 weeks, and it could help guide the urgency and allocation of vEEG monitoring or reduce unnecessary exposure to antiseizure medications in patients who are likely to have PNES. It is not a replacement for definitive testing.

Source: Early epilepsy detection from electronic health records with large language models, npj Digital Medicine, 8 October 2026.

Let Google know we are your trusted source.

Add our editorial as a preferred source in your search results.

Trust this Source