Evidence Knowledge Base
Entity × function evidence mined from the literature. Each cell shows supporting papers (▲ evidence for / ▼ evidence against the therapeutic direction), recomputed live for the selected context and confidence. Click a cell for the underlying papers.
How this knowledge base was built — and important caveats
- Define functions. Four function domains — fibrosis, inflammation, angiogenesis, apoptosis — were seeded, each with keyword patterns (e.g.
collagen,fibro*). - Retrieve literature. For each target entity (miRNA or gene) PubMed was searched in a heart / muscle / Duchenne context, and matching papers were collected.
- Get the text. Full text was pulled from PubMed Central where available; otherwise the abstract was used.
- Find passages. Papers were split into sections, keeping only passages that co-mention a target entity and a domain keyword.
- AI extraction. A large language model read each passage and produced a structured claim: entity → domain → direction (promotes / inhibits / no effect), a 0–100 confidence, and context tags (organ/tissue, species, disease). The model did not write the quote — to avoid fabricated text, the supporting quote is a real sentence lifted verbatim from the source passage afterward by a deterministic heuristic (the sentence that best evidences the entity). Each passage was read several times and the majority answer kept.
- Deduplicate. Claims were collapsed into assertions (entity × domain × direction), with one evidence row per (paper, quote).
- Present. This matrix reframes direction into therapeutic for / against (e.g. inhibiting fibrosis = anti-fibrotic = "for"), with live filters for context, confidence, and source.
- AI-generated & not peer-reviewed. Every direction, confidence, and quote is produced by a language model and can be wrong. Treat this as a triage / hypothesis-generation tool, not ground truth.
- Verify at the source. Always open the linked PMID and read the paper before relying on a claim.
- Direction errors happen. Notably "agent inversion" (a drug or antagomir that inhibits the entity and reduces fibrosis can be mis-coded), and correlation / biomarker statements ("X is elevated in fibrosis") can be read as causal.
- Confidence ≠ correctness. The score reflects how clearly a sentence states its claim, not whether the claim is causal or true. Filtering to ≥ 80 removes vague text but not clearly-worded associations.
- Context tags are noisy. Organ / species / disease are free-text from the model and may be over-tagged or imprecise.
- Quotes are real, but heuristically chosen. Quotes are never AI-written (lifted verbatim from the source), but an algorithm picks which sentence to show and can occasionally surface one adjacent to — rather than exactly — the sentence that drove the model's call.
- Counts can be inflated. Numbers are distinct papers; reviews repeating the same claim (citation propagation) can add several papers for one underlying finding.
- Coverage is bounded. Limited to the seeded domains, the searched context, and the target entity list — absence of evidence here is not evidence of absence.
Summary statistics
| Metric | miRNAs | Genes |
|---|---|---|
| Retrieval | ||
| Discovery runs (search batches) | — | — |
| Papers analyzed | — | — |
| — with full text retrieved | — | — |
| Sectioning | ||
| Text sections analyzed | — | — |
| — from full text | — | — |
| — from abstracts | — | — |
| AI extraction | ||
| Passage extractions run | — | — |
| Results | ||
| Entities in KB | — | — |
| Assertions (entity × function × direction) | — | — |
| Evidence records | — | — |
| — from full text | — | — |
| — from abstracts | — | — |
| Distinct papers with evidence | — | — |
Retrieval / sectioning / extraction counts are per analysis run; entity / assertion / evidence counts are per entity type. "Evidence records" are raw AI extractions (before the matrix's per-(paper, quote) dedup), so they exceed the deduplicated evidence shown in cells.
Each row is one PubMed query session behind the KB. The context is what was actually searched — organ / species / disease filters in the matrix reflect only what papers happened to mention, not exhaustive coverage of that context.
| Type | Entities | Context | Domains | Papers | Runs | Evidence |
|---|