Skip to content
All projects

LUMINA

Automating the manual screening of medical reviews

first author · NUS research · 2024–25

Why it exists

A medical systematic review takes a team of five about 15 months (Borah et al., 2017 (opens in new tab)). The first step is the grind: reading thousands of abstracts by hand to find the 3 in 100 that belong. Miss one and the evidence is wrong.

What I built

I built LUMINA to do that screening. Four agents work in two passes: a generous first cut, a strict check against the review’s criteria, and a reasoning model that challenges every call before it stands. As first author, I designed it, built it and ran the evaluation.

Impact

  • Missed 8 of 378 included studies across 15 published reviews (98.2% sensitivity).
  • Cut what a team reads in full to under 9% of the candidates.
  • 35× fewer misses than a reimplementation of Li et al. (2024), and every study found on the four reviews Tran et al. reported.
  • Under a cent per citation, across about 150,000 abstracts screened.
LUMINA · screening, review, and decision traces

Scroll across to follow the full system

Citation titles and abstracts pass through a lenient classifier. Relevant or uncertain candidates advance to PICOS screening with the review's objectives and methods. Each tier uses an o3-mini reviewer and a gpt-4o-mini improver, with at most three reviews per tier. At the limit, the last improver output is used. Included citations proceed to human full-text review. Completed candidates' decisions, feedback, revisions, and costs are saved as JSONL.Candidate citationsCSV / RIS · title + abstractReview protocolTitle / abstract · objectives + methodsINPUT01TriageTitle + abstractrevise, then recheckdisagreeClassifiergpt-4o-miniReviewero3-miniImprovergpt-4o-miniExcludeLikely irrelevant02PICOS screeningCriteria + methodsrevise, then recheckdisagreeDetailed screenergpt-4o-miniReviewero3-miniImprovergpt-4o-miniExcludeDecision recordedrelevant / uncertainFull-text reviewHuman eligibility checksJSONL traceDecisions + justificationsReview feedback + revisionsToken cost per completed citationPython library + CLI · versioned promptsAt most 3 reviews per tier. On the iteration limit, the last improver output is used.

Method and results

Two screening tiers

The classifier sees the review and candidate titles and abstracts. It keeps potentially relevant and uncertain citations for the detailed screener, which also receives the review’s objectives and methods. The second tier works through Population, Intervention, Comparison, Outcome, and Study design (PICOS) before returning an inclusion or exclusion. Read the prompts (opens in new tab).

Review, revision, and termination

The classifier, screener, and improver use gpt-4o-mini; the reviewer uses o3-mini. At each tier, the reviewer checks the decision and its justification. Disagreement sends the original task and reviewer feedback to the improver, then back for review. The implementation caps each tier at three reviewer calls; if agreement has not been reached, it uses the last improver output. Follow the control flow (opens in new tab).

Decisions you can inspect

Each completed citation produces a typed AgentTrace: classifier and screener outputs, review feedback, revisions, final decision, and cost. The CLI accepts CSV or RIS candidates and writes one JSONL record per completed citation. Prompts are versioned alongside the Python code. Decision markers are parsed and validated, with bounded retries for invalid output.

Evaluation

LUMINA achieved 98.2% mean sensitivity and 87.9% mean specificity across 15 published systematic reviews, missing 8 of the 378 studies their authors included. Metrics are calculated per review, then averaged. Against a reimplementation of Li et al. (2024) on the same reviews, its miss rate was 35× lower (0.018 against 0.630). In all, the pipeline has screened about 150,000 candidate abstracts, at under a cent each. Sensitivity comes first: the goal is to retain relevant studies while reducing the material that needs full-text review. Inspect the results and evaluation data (opens in new tab).

The baseline and the design change

Tran et al., Annals of Internal Medicine (2024) (opens in new tab) tested GPT-3.5 Turbo with separate eligibility prompts, combining their answers under a balanced rule or a rule favoring sensitivity. LUMINA adds a lenient first pass, contextual PICOS screening, and explicit review and revision at both tiers. On the four reviews Tran reported, LUMINA found every included study; Tran’s pipeline found 81% to 97%.

Scope

Retrospective screening of published reviews; included citations still go to human full-text review.

Python · gpt-4o-mini · o3-mini · PICOS · JSONL

zanwenfu/agentic-reviewers-for-SRMA