Evidence Knowledge Base

Entity × function evidence mined from the literature. Each cell shows supporting papers (▲ evidence for / ▼ evidence against the therapeutic direction), recomputed live for the selected context and confidence. Click a cell for the underlying papers.

How this knowledge base was built — and important caveats
Build pipeline
  1. Define functions. Four function domains — fibrosis, inflammation, angiogenesis, apoptosis — were seeded, each with keyword patterns (e.g. collagen, fibro*).
  2. Retrieve literature. For each target entity (miRNA or gene) PubMed was searched in a heart / muscle / Duchenne context, and matching papers were collected.
  3. Get the text. Full text was pulled from PubMed Central where available; otherwise the abstract was used.
  4. Find passages. Papers were split into sections, keeping only passages that co-mention a target entity and a domain keyword.
  5. AI extraction. A large language model read each passage and produced a structured claim: entity → domain → direction (promotes / inhibits / no effect), a 0–100 confidence, and context tags (organ/tissue, species, disease). The model did not write the quote — to avoid fabricated text, the supporting quote is a real sentence lifted verbatim from the source passage afterward by a deterministic heuristic (the sentence that best evidences the entity). Each passage was read several times and the majority answer kept.
  6. Deduplicate. Claims were collapsed into assertions (entity × domain × direction), with one evidence row per (paper, quote).
  7. Present. This matrix reframes direction into therapeutic for / against (e.g. inhibiting fibrosis = anti-fibrotic = "for"), with live filters for context, confidence, and source.
⚠ Caveats — this is AI-generated
  • AI-generated & not peer-reviewed. Every direction, confidence, and quote is produced by a language model and can be wrong. Treat this as a triage / hypothesis-generation tool, not ground truth.
  • Verify at the source. Always open the linked PMID and read the paper before relying on a claim.
  • Direction errors happen. Notably "agent inversion" (a drug or antagomir that inhibits the entity and reduces fibrosis can be mis-coded), and correlation / biomarker statements ("X is elevated in fibrosis") can be read as causal.
  • Confidence ≠ correctness. The score reflects how clearly a sentence states its claim, not whether the claim is causal or true. Filtering to ≥ 80 removes vague text but not clearly-worded associations.
  • Context tags are noisy. Organ / species / disease are free-text from the model and may be over-tagged or imprecise.
  • Quotes are real, but heuristically chosen. Quotes are never AI-written (lifted verbatim from the source), but an algorithm picks which sentence to show and can occasionally surface one adjacent to — rather than exactly — the sentence that drove the model's call.
  • Counts can be inflated. Numbers are distinct papers; reviews repeating the same claim (citation propagation) can add several papers for one underlying finding.
  • Coverage is bounded. Limited to the seeded domains, the searched context, and the target entity list — absence of evidence here is not evidence of absence.
Summary statistics
MetricmiRNAsGenesY RNA
Retrieval
Discovery runs (search batches)———
Papers analyzed———
— with full text retrieved———
Sectioning
Text sections analyzed———
— from full text———
— from abstracts———
AI extraction
Passage extractions run———
Results
Entities in KB———
Assertions (entity × function × direction)———
Evidence records———
— from full text———
— from abstracts———
Distinct papers with evidence———

Retrieval / sectioning / extraction counts are per analysis run; entity / assertion / evidence counts are per entity type. "Evidence records" are raw AI extractions (before the matrix's per-(paper, quote) dedup), so they exceed the deduplicated evidence shown in cells.

Retrieval sessions & scope

Each row is one PubMed query session behind the KB. The context is what was actually searched — organ / species / disease filters in the matrix reflect only what papers happened to mention, not exhaustive coverage of that context.

TypeEntitiesContextDomainsPapersRunsEvidence
entity
search
only these miRNAs
order by
organ
species
disease
source
min confidence: 80
evidence for evidence against mixed / no-effectcell = ▲for ▼against papers · ○ no-effect
0 miRNAs with evidence (of 0) · 0 evidence