pubextractor

Starting pubextractor

Loading the workspace. The first open takes a few seconds.

Demo: short papers only (up to 3 files, 20 pages or 10 MB each). Upload only papers that may be shared; files are deleted when the page closes. Run locally for full use.
PDF, DOCX, or image · several at once
Papers

What to extract from each paper: one row per field. Edit here or import your Excel sheet.

Double-click a cell to edit it; click a row to select it.
Model
License
Keys load from .env and are never shown.
Review
Designssuggested from each Method; confirm or change
Extracted values
Extracted tables
Click a cell to correct it; an edited value is recorded with layer human. Status takes extracted, review, not_found, or refused.
Image tables
Type each value with its arm, outcome, time, and stat; group is read from the arm label when left blank. Stat takes mean, sd, mean_sd, n, se, adj_mean, adj_mean_se, percent, r, t, f, p.
Reading with a vision model sends the page crop under the license check in the Model step. Local Ollama needs a vision-capable model: gemma3:4b reads images, tinyllama does not.
Effect sizes
Calibrate: click two x-axis ticks, then two y-axis ticks, and type their values. Edit points: click to add, click a point to remove.
Axes
x (session)
tick 1
tick 2
y (outcome)
tick 1
tick 2
Study
n per box from the text layer; none means hand entry.
Kept with its caption; no values read.
Export
Long table
One row per value, with page, section, and quote.
Wide table
One row per paper, one column per field.
Effect rows
g, variance, method, and checks per comparison.
metafor
yi and vi per effect, for metafor::rma.
scdbayes
One row per graph point, for scdbayes.
Correlation matrices
One CSV per study over the constructs, and n.csv, zipped.
Trace (JSON)
Steps behind each value, per paper.
Methods text
PRISMA 2020 items 9 and 11, to edit.
Agreement
Coding sheet
agreement.csv