Evidence Knowledge Base

Entity × function evidence mined from the literature. Each cell shows supporting papers (▲ evidence for / ▼ evidence against the therapeutic direction), recomputed live for the selected context and confidence. Click a cell for the underlying papers.

How this knowledge base was built — and important caveats
Build pipeline
  1. Define functions. Four function domains — fibrosis, inflammation, angiogenesis, apoptosis — were seeded, each with keyword patterns (e.g. collagen, fibro*).
  2. Retrieve literature. For each target entity (miRNA or gene) PubMed was searched in a heart / muscle / Duchenne context, and matching papers were collected.
  3. Get the text. Full text was pulled from PubMed Central where available; otherwise the abstract was used.
  4. Find passages. Papers were split into sections, keeping only passages that co-mention a target entity and a domain keyword.
  5. AI extraction. A large language model read each passage and produced a structured claim: entity → domain → direction (promotes / inhibits / no effect), a 0–100 confidence, and context tags (organ/tissue, species, disease). The model did not write the quote — to avoid fabricated text, the supporting quote is a real sentence lifted verbatim from the source passage afterward by a deterministic heuristic (the sentence that best evidences the entity). Each passage was read several times and the majority answer kept.
  6. Deduplicate. Claims were collapsed into assertions (entity × domain × direction), with one evidence row per (paper, quote).
  7. Present. This matrix reframes direction into therapeutic for / against (e.g. inhibiting fibrosis = anti-fibrotic = "for"), with live filters for context, confidence, and source.
⚠ Caveats — this is AI-generated
  • AI-generated & not peer-reviewed. Every direction, confidence, and quote is produced by a language model and can be wrong. Treat this as a triage / hypothesis-generation tool, not ground truth.
  • Verify at the source. Always open the linked PMID and read the paper before relying on a claim.
  • Direction errors happen. Notably "agent inversion" (a drug or antagomir that inhibits the entity and reduces fibrosis can be mis-coded), and correlation / biomarker statements ("X is elevated in fibrosis") can be read as causal.
  • Confidence ≠ correctness. The score reflects how clearly a sentence states its claim, not whether the claim is causal or true. Filtering to ≥ 80 removes vague text but not clearly-worded associations.
  • Context tags are noisy. Organ / species / disease are free-text from the model and may be over-tagged or imprecise.
  • Quotes are real, but heuristically chosen. Quotes are never AI-written (lifted verbatim from the source), but an algorithm picks which sentence to show and can occasionally surface one adjacent to — rather than exactly — the sentence that drove the model's call.
  • Counts can be inflated. Numbers are distinct papers; reviews repeating the same claim (citation propagation) can add several papers for one underlying finding.
  • Coverage is bounded. Limited to the seeded domains, the searched context, and the target entity list — absence of evidence here is not evidence of absence.
Summary statistics
MetricmiRNAsGenes
Retrieval
Discovery runs (search batches)
Papers analyzed
— with full text retrieved
Sectioning
Text sections analyzed
— from full text
— from abstracts
AI extraction
Passage extractions run
Results
Entities in KB
Assertions (entity × function × direction)
Evidence records
— from full text
— from abstracts
Distinct papers with evidence

Retrieval / sectioning / extraction counts are per analysis run; entity / assertion / evidence counts are per entity type. "Evidence records" are raw AI extractions (before the matrix's per-(paper, quote) dedup), so they exceed the deduplicated evidence shown in cells.

Retrieval sessions & scope

Each row is one PubMed query session behind the KB. The context is what was actually searched — organ / species / disease filters in the matrix reflect only what papers happened to mention, not exhaustive coverage of that context.

TypeEntitiesContextDomainsPapersRunsEvidence
entity
search
order by
organ
species
disease
source
min confidence: 80
evidence for evidence against mixed / no-effectcell = ▲for ▼against papers · ○ no-effect
0 miRNAs with evidence (of 0) · 0 evidence