LUMINA
Automating the manual screening of medical reviews
first author · NUS research · 2024–25
Why it exists
A medical systematic review takes a team of five about 15 months (Borah et al., 2017 (opens in new tab)). The first step is the grind: reading thousands of abstracts by hand to find the 3 in 100 that belong. Miss one and the evidence is wrong.
What I built
I built LUMINA to do that screening. Four agents work in two passes: a generous first cut, a strict check against the review’s criteria, and a reasoning model that challenges every call before it stands. As first author, I designed it, built it and ran the evaluation.
Source code (opens in new tab)Results & data (opens in new tab)
Impact
- Missed 8 of 378 included studies across 15 published reviews (98.2% sensitivity).
- Cut what a team reads in full to under 9% of the candidates.
- 35× fewer misses than a reimplementation of Li et al. (2024), and every study found on the four reviews Tran et al. reported.
- Under a cent per citation, across about 150,000 abstracts screened.
Scroll across to follow the full system
Method and results
Two screening tiers
The classifier sees the review and candidate titles and abstracts. It keeps potentially relevant and uncertain citations for the detailed screener, which also receives the review’s objectives and methods. The second tier works through Population, Intervention, Comparison, Outcome, and Study design (PICOS) before returning an inclusion or exclusion. Read the prompts (opens in new tab).
Review, revision, and termination
The classifier, screener, and improver use gpt-4o-mini; the reviewer uses o3-mini. At each tier, the reviewer checks the decision and its justification. Disagreement sends the original task and reviewer feedback to the improver, then back for review. The implementation caps each tier at three reviewer calls; if agreement has not been reached, it uses the last improver output. Follow the control flow (opens in new tab).
Decisions you can inspect
Each completed citation produces a typed AgentTrace: classifier and screener outputs, review feedback, revisions, final decision, and cost. The CLI accepts CSV or RIS candidates and writes one JSONL record per completed citation. Prompts are versioned alongside the Python code. Decision markers are parsed and validated, with bounded retries for invalid output.
Evaluation
LUMINA achieved 98.2% mean sensitivity and 87.9% mean specificity across 15 published systematic reviews, missing 8 of the 378 studies their authors included. Metrics are calculated per review, then averaged. Against a reimplementation of Li et al. (2024) on the same reviews, its miss rate was 35× lower (0.018 against 0.630). In all, the pipeline has screened about 150,000 candidate abstracts, at under a cent each. Sensitivity comes first: the goal is to retain relevant studies while reducing the material that needs full-text review. Inspect the results and evaluation data (opens in new tab).
The baseline and the design change
Tran et al., Annals of Internal Medicine (2024) (opens in new tab) tested GPT-3.5 Turbo with separate eligibility prompts, combining their answers under a balanced rule or a rule favoring sensitivity. LUMINA adds a lenient first pass, contextual PICOS screening, and explicit review and revision at both tiers. On the four reviews Tran reported, LUMINA found every included study; Tran’s pipeline found 81% to 97%.
Scope
Retrospective screening of published reviews; included citations still go to human full-text review.
Python · gpt-4o-mini · o3-mini · PICOS · JSONL