Skill for analyzing and improving compliance in the dismech knowledge base...
Analyze and improve the completeness of disorder YAML files in the dismech knowledge base. The compliance system checks for recommended fields (ontology terms, evidence items, descriptions) and generates scores to identify priority curation targets.
just compliance kb/disorders/Asthma.yaml
Output includes:
just compliance-all
Multi-file report showing:
just compliance-weighted
Uses conf/qc_config.yaml to:
# CSV format for spreadsheet analysis
just compliance-csv
# JSON format for programmatic processing
just compliance-report
just gen-dashboard
Creates dashboard/index.html with:
Recommended-slot compliance only measures whether fields are populated on an
object. It cannot express cross-object graph properties ā most importantly,
whether a phenotype is actually wired into the causal pathograph. A
phenotype can have a perfect HPO term, evidence, and description (full
compliance credit) yet still float as a disconnected node, because the edge that
connects it lives on a different object's downstream list.
dismech.qc_plugins fills this gap with a graph-derived QC metric computed from
build_causal_graph() and emitted as an AggregatedPathScore (path
phenotypes[].causal_inlink), so it composes with weighted compliance and
conf/qc_config.yaml weights/thresholds like any other field. It is graded
coverage, not a binary gate: a file with 9/12 phenotypes wired in scores 75%.
# Per-file connectivity coverage across the KB (lists files with gaps)
just compliance-connectivity
# Show the floating phenotype node names
just compliance-connectivity --list-unconnected
# Fail CI if aggregate coverage drops below a percentage
just compliance-connectivity --fail-under 30
just compliance-connectivity --genes-fail-under 40
just compliance-connectivity --activity-fail-under 20
A phenotype counts as connected when at least one causal edge (causes,
leads_to, triggers, exacerbates ā see qc_plugins.CAUSAL_PREDICATES)
targets it. To fix a floating phenotype, add the phenotype's name as a
downstream: [{target: <phenotype name>}] on the upstream pathophysiology node.
Three metrics come out of the same command, and the two gene ones are a pair:
| path | asks |
|---|---|
phenotypes[].causal_inlink |
is the phenotype reached by a causal edge? |
genetic[].mechanism_outlink |
does the gene reach the mechanism graph at all? |
genetic[].mechanism_activity_grounding |
does where it lands name a molecular function? |
Gene wiring. A genetic[] entry reaches the pathograph only when some
pathophysiology node carries the same gene in its gene:/genes: descriptor
ā that shared CURIE is the entire edge. A causal gene with no such node floats
in the genetic block and never appears in the graph. Non-causal items
(BIOMARKER, PROTECTIVE, MODIFIER, DISPUTED, UNKNOWN) are excluded from the
denominator.
Gene activity grounding. GO puts a level between a gene and a process ā
gene ā molecular function ā biological process ā and an edge running from a gene
straight to a node annotated only with biological_processes: skips it: the
graph says what the cell can no longer do without saying what the protein
can no longer do. A wired gene counts as grounded when at least one node it
reaches carries molecular_functions:. The denominator is the wired genes, so
an unwired gene is charged once (against mechanism_outlink), not twice.
To fix a flagged gene, add molecular_functions: to the node it lands on:
- name: SLC25A20 transporter molecular function deficiency
molecular_functions:
- preferred_term: acyl carnitine transmembrane transporter activity
term:
id: GO:0015227
label: O-acyl-L-carnitine transmembrane transporter activity
biological_processes:
- preferred_term: carnitine transport
term: {id: GO:0015879, label: carnitine transport}
Two cases where the term is not the fix. When the landing node collects many
genes with unrelated activities (Primary_Ciliary_Dyskinesia / "Ciliary
Dysfunction" carries 21 ā dynein motors, radial-spoke constituents, and
transcription factors), no single MF term is true of it and the node wants
splitting. And some classes genuinely have no shared molecular function; say so
in the node's description rather than binding a term that overstates it.
just list-disconnected-phenotypesThe recipe above is the compliance view of the metric, and a gate ā both
halves, the aggregate percentages, and the min_compliance floor it enforces
over the corpus. For per-entry triage of the phenotype half, use the report
added for issue #11935, which calls the same causal_inlink_coverage function
and so cannot disagree with it:
just list-disconnected-phenotypes # census + ranked worklist
just list-disconnected-phenotypes --format tsv # one row per phenotype
just list-disconnected-phenotypes --zero-only # only 0-connected entries
just list-disconnected-phenotypes kb/disorders/Asthma.yaml
It adds three things compliance-connectivity does not have:
ISOLATED, TREATED,
READOUT, NONCAUSAL_INBOUND, SEQUELA_SOURCE ā distinguishing "nothing in
the graph knows this node exists" from "a treatment or an observational
readout points at it, but no mechanism does". Neither counts as connected;
they are different curation jobs.--format tsv / --format json for a census.Report-only by design (exit 0), with --strict / --fail-under opt-in ā that
is this view, not the metric, whose aggregate is gated by
compliance-connectivity. Do not treat the percentage as a target: the edge
asserts which mechanism produces which clinical feature, and some phenotypes
legitimately have no upstream node in the entry. An edge added to clear a report
is worse than no edge.
The QCMetricPlugin protocol in src/dismech/qc_plugins.py is the generic seam
for further graph-derived metrics (orphan-target rate, dead-end nodes).
| Metric | Description |
|---|---|
| Global Compliance | Simple percentage: populated fields / total recommended fields |
| Weighted Compliance | Adjusted by field importance from conf/qc_config.yaml |
| Field | Weight | Min Threshold | Why |
|---|---|---|---|
disease_term.term |
5.0 | 95% | Root disease identity - always required |
phenotypes[].phenotype_term.term |
3.0 | 90% | Core clinical data |
pathophysiology[].cell_types[].term |
3.0 | 85% | Mechanistic understanding |
treatments[].treatment_term.term |
2.5 | 80% | Clinical relevance |
term (general) |
2.0 | 80% | All ontology bindings |
pathophysiology[].evidence |
2.0 | 80% | Scientific backing |
evidence (general) |
1.5 | - | Valuable but not always required |
description |
0.5 | - | Nice-to-have context |
| Status | Meaning |
|---|---|
| OK | Field is populated |
| MISSING | Recommended field is empty/absent |
Address fields in this priority order based on weights:
disease_term:
preferred_term: Asthma
term:
id: MONDO:0004979
label: asthma
Look up: uv run runoak -i sqlite:obo:mondo info "asthma"
phenotypes:
- name: Wheezing
phenotype_term:
preferred_term: Wheezing
term:
id: HP:0030828
label: Wheezing
Look up: uv run runoak -i sqlite:obo:hp info "l~wheezing"
cell_types:
- preferred_term: Mast cells
term:
id: CL:0000097
label: mast cell
Look up: uv run runoak -i sqlite:obo:cl info "l~mast cell"
treatments:
- name: Inhaled corticosteroids
treatment_term:
preferred_term: corticosteroid therapy
term:
id: NCIT:C15986
label: Pharmacotherapy
Look up: uv run runoak -i sqlite:obo:ncit search "corticosteroid"
evidence:
- reference: PMID:12345678
supports: SUPPORT
snippet: "Exact quote from abstract"
explanation: "Why this supports the claim"
just gen-dashboard
# Check dashboard/index.html for "Priority Curation Targets"
Or use compliance-all and sort:
just compliance-report | jq -r '.files | sort_by(.weighted_compliance) | .[:10] | .[].file'
just compliance-weighted 2>&1 | grep "VIOLATION"
For systematically missing fields across many files:
import yaml
import glob
# Example: Find files missing disease_term.term
for f in glob.glob("kb/disorders/*.yaml"):
with open(f) as file:
data = yaml.safe_load(file)
dt = data.get('disease_term', {})
if not dt.get('term'):
print(f"{f}: missing disease_term.term")
# Schema validation
just validate kb/disorders/MyDisease.yaml
# Term validation (labels match ontology)
just validate-terms kb/disorders/MyDisease.yaml
# Re-check compliance
just compliance kb/disorders/MyDisease.yaml
# Default for unconfigured fields
default_weight: 1.0
default_min_compliance: null
# Per-slot config (applies everywhere that slot appears)
slots:
term:
weight: 2.0
min_compliance: 80.0
# Per-path config (overrides slot config for specific locations)
paths:
"phenotypes[].phenotype_term.term":
weight: 3.0
min_compliance: 90.0
Edit conf/qc_config.yaml to:
just qc after improvements for full validationThis indicates your important fields (high weight) have different coverage than low-priority fields. Focus on improving high-weight fields first.
Descriptions have low weight (0.5) and no minimum threshold. Address these last, or not at all if not needed.
Check conf/qc_config.yaml for min_compliance settings. Either:
Ensure the dashboard directory exists and you have write permissions:
mkdir -p dashboard
just gen-dashboard