Eleven stages, explained twice
How the model works
Eleven stages, from a PDF file to a matrix a nurse can use. Each stage is explained twice: once in everyday language, once in its technical terms.
- 2.1b
Domain relevance filter
Is this article actually about palliative care and drug side effects? If not, it stops here.
Technical
A relevance score over six term groups; threshold 8. Of 405 articles, 263 passed.
- 2.2
PDF text extraction
The PDF becomes plain text, page by page.
Technical
PyMuPDF, `page.get_text('text')` joined across pages.
- 2.3
Preprocessing
Tidying the text: words broken at a line end are rejoined, citation numbers are lifted out of the sentence.
Technical
NFKC, Unicode dash folding (U+2013 appears in 8.6% of sentences), superscript citation-marker separation.
- 2.4
Sentence segmentation
The text is cut into sentences. Very short ones are usually layout debris; very long ones are usually a table that turned into prose.
Technical
NLTK punkt, then word-break repair against a corpus vocabulary, then a 30–2,000 character length gate. 137,906 sentences, 92,713 into NER.
- 3.1
Entity tagging
Two AI models read each sentence and highlight which words are drug names and which are side effects.
Technical
BioBERT chemical + diseases side by side, batch 8, 512 tokens, plus dictionary matching over 38/32 terms.
- 3.2
Term normalisation
One thing can be written many ways. All of them collapse to a single standard term.
Technical
Lemmatisation, British→American spelling, RapidFuzz with two guards: refusing a mapping to a more specific term, and refusing drug pairs that merely look alike.
- 3.2b
Strict entity filter
Fragments that got through as entities but are not medical terms are dropped here — one- and two-letter abbreviations, numbers, and ordinary words.
Technical
A letter-count threshold, a list of 32 disallowed terms, and an exemption for the suffixes typical of drug names. Its counters report in, out, and dropped.
- 3.3
Co-occurrence (baseline path)
Every drug is paired with every side effect appearing in the same sentence. No meaning has been checked yet.
Technical
Token distance ≤30. Negation is detected over the side-effect span (NegEx via negspaCy), not the whole sentence; negated pairs are weighted 0.65 rather than discarded.
- 3.4
Association scoring
The more often a pair appears together, and the rarer each is alone, the stronger the relation is taken to be.
Technical
PMI over weighted frequency; Final Score = PMI · log(1+frequency); Confidence = sigmoid.
- 4.1
Advanced relations and meaning filters
This is where the difference lies. The sentence structure is examined: is the drug really the cause, does the sentence deny it, is the drug in fact being used to treat that effect, and does the effect actually belong to a different drug in the same sentence.
Technical
spaCy dependency parsing with a causal co-occurrence fallback. Recorded rejections: NO_CAUSAL_SIGNAL, NON_ADVERSE_CONTEXT, DISTANCE_EXCEEDED, attribution to another drug.
- 4.1b
Frequency gate
A pair only enters the operational matrix if it appears at least twice across the whole corpus.
Technical
`advanced_min_pair_frequency = 2`. Lowering it to 1 raises recall 0.800→0.933 but drops precision 0.632→0.538.
What this site cannot do
Three limitations that have to be stated
1. The frequency gate is corpus-level
The operational matrix requires a pair to appear at least twice across all 405 articles. A document you upload almost never satisfies that on its own, so the “passed the gate” figure in the simulation is not a judgement on whether that pair belongs in the matrix.
2. The word-repair vocabulary comes from the corpus
Rejoining broken words (“consti pation” → “constipation”) uses a word list built from the 137,906 sentences of the research corpus. A document from outside the corpus therefore behaves slightly differently than it would had it been part of it. This is a structural limit, not something that can be measured away.
3. Two numeric regimes, measured at zero difference
The article's numbers were computed on GPU fp16; this site runs on CPU fp32. The two were tested on three separate axes — extraction fidelity (GPU), precision (CPU), and batch composition — and all three came back at zero difference across 211 advanced pairs and 333 baseline pairs. So what you see here does not approximate the article's numbers; it is the same numbers.