Nothing you paste here is sent anywhere: it is all computed in your browser. That is why you can use it on an unpublished manuscript, or on one you are reviewing.
The PDF is read in your browser: it is not uploaded anywhere. What comes out is the text, not the tables nor formulas that are images.
What this screen has NOT looked at
- Results inside tables: pasting as plain text loses the structure, so there is no way to tell which number goes with which.
- Incomplete tests: missing degrees of freedom, or missing p.
- Notation outside APA: square brackets instead of parentheses, a semicolon instead of a comma.
- Formulas and results inserted as an image.
- p-values corrected for multiplicity or sphericity: they are flagged "uncheckable" on purpose, because they need not follow from the test statistic.
- Intervals with an open endpoint (
95 % CI 73 to not reached), which are the norm in survival studies: there is no number to measure.
A clean screen does not mean the manuscript is clean.
On an in-house corpus of 19 hand-labelled test results, this recogniser finds 100 %, gets the verdict right on 100 % of those it finds, and its precision is 100 %. It is a regression net, not a measure of real-world accuracy: the corpus was written by whoever built the tool, so it only contains the cases they thought of. And against 151 test results from 34 published articles (Europe PMC, CC-BY) reviewed by hand, it extracts exactly the ones it should and leaves alone the 15 candidates that are not its business. In those same articles it also finds 97 confidence intervals, of which it can judge 65 without accusing a single one: those it cannot judge say why.