📢 Accepted at ECML PKDD 2026
Welcome to the official repository for:
PathogenKG: Cross-Species Drug Repurposing via Heterogeneous Knowledge Graph Link Prediction
European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases
📍 Naples, Italy | 🗓️ September 7-11, 2026
PathogenKG is a heterogeneous knowledge graph embedding framework for cross-species antimicrobial drug repurposing. It integrates protein–protein interaction (PPI) networks from 31 bacterial pathogens (STRING), COG orthology, Gene Ontology annotations, and validated drug–target associations (DrugBank) into a unified graph of over 3 million triples. Heterogeneous GNN encoders (R-GCN, CompGCN) with a DistMult decoder are trained to predict new drug–target interactions (DTIs) via link prediction on the TARGET relation.
PathogenKG/
├── train_and_eval.py # Main training & evaluation script (Section 3.2)
├── drug_eval.py # Compound-centric repurposing evaluation (Section 3.3)
├── drug_eval_results.py # Process and summarise drug evaluation outputs
├── tuning_hyperparameter.py # Bayesian HPO via Weights & Biases sweeps (Section 3.2)
├── tuning_dataset.py # HPO on PathogenKG dataset
├── tuning_dataset_drkg.py # HPO on DRKG benchmark (Section 3.2, pipeline validation)
├── model_eval.py # Standard model evaluation utilities
├── evaluate_ranking.py # Ranking metric computation
├── kg_stats_visualization.py # KG statistics and figures (Table 2, Figure 2)
├── build_pathogenkg.py # Build organism-specific KGs from raw data (Section 3.1)
├── merge_pathogen_kgs.py # Merge organism KGs into PathogenKG (Section 3.1)
├── generate_all_pathogenkg.py # Batch-build all 31 pathogen KGs
├── get_pathogenkg.py # Download pre-built datasets from HuggingFace
│
├── src/
│ ├── utils.py # Data loading, splitting, negative sampling, metrics
│ ├── evaluation_metrics_filtered.py # Type-constrained filtered evaluation (Section 3.2)
│ ├── hetero_compgcn.py # CompGCN encoder (Section 3.2)
│ ├── hetero_rgcn.py # R-GCN encoder (Section 3.2)
│ ├── hetero_rgat.py # R-GAT encoder (experimental)
│ ├── models_params.json # Best hyperparameter configurations
│ ├── bio_utils.py # Biological ID mapping utilities
│ └── downloaders/ # Scripts to fetch raw data from STRING, DrugBank, UniProt, GO, COG
│
├── dataset/
│ ├── PathogenKG_n31_core.tsv.zip # Pre-built PathogenKG (31 species, 3M triples)
│ └── DRUGBANK/ # Auxiliary DrugBank mappings
│
├── figures/ # Figures used in the paper
├── requirements.txt
└── environment.yml
Requirements: Python 3.10, CUDA-capable GPU (≥16 GB VRAM recommended for CompGCN with hidden dim 200).
# Create conda environment and install dependencies
conda create -n pathogenkg python=3.10 -y
conda activate pathogenkg
pip install -r requirements.txtIf torch-sparse installation fails (common on some platforms), install PyTorch and PyG extensions separately:
pip install --index-url https://download.pytorch.org/whl/cu128 torch torchvision torchaudio
pip install --no-cache-dir --only-binary=:all: \
pyg_lib torch-geometric torch-scatter torch-sparse torch-cluster torch-spline-conv \
termcolor torcheval \
-f https://data.pyg.org/whl/torch-2.7.1+cu128.htmlAdjust the CUDA version (
cu128) to match your local installation.
The default dataset PathogenKG_n31_core.tsv.zip (31 core STRING species, ~3M triples) is included in dataset/. No further action is needed.
Additional datasets can be downloaded from HuggingFace:
python get_pathogenkg.pyThis downloads both PathogenKG variants and the raw source data to dataset/.
To reproduce the KG construction pipeline from raw biological databases:
# 1. Download source data (STRING PPI, DrugBank, COG, GO)
# Scripts in src/downloaders/ fetch from public APIs.
# 2. Build organism-specific KGs for all 31 pathogens
python generate_all_pathogenkg.py
# 3. Merge into a single multi-organism graph
python merge_pathogen_kgs.pyPathogenKG is stored as a TSV of triples: head, interaction, tail.
- Entity types:
ExtGene(pathogen proteins, UniProt accession) andCompound(drugs, PubChem CID) - Relation types (13): 8 PPI types (ASSOCIATION, GENE_BIND, ACTIVATION, REACTION, CATALYSIS, PTMOD, INHIBITION, EXPRESSION), ORTHOLOGY, 3 GO types (BiologicalProcess, MolecularFunction, CellularComponent), and TARGET
Train a CompGCN model on PathogenKG with the best hyperparameters (stored in src/models_params.json):
# Single run with default settings (CompGCN, TARGET task, PathogenKG_n31_core)
python train_and_eval.py
# 12-run evaluation for statistical comparison (as in Table 4)
python train_and_eval.py --model compgcn --runs 12 --epochs 400 --task TARGET
# R-GCN comparison
python train_and_eval.py --model rgcn --runs 12 --epochs 400 --task TARGETKey arguments:
| Argument | Default | Description |
|---|---|---|
--model |
compgcn |
GNN encoder: compgcn, rgcn, or rgat |
--tsv |
dataset/PathogenKG_n31_core.tsv.zip |
Path to dataset |
--task |
TARGET |
Relation type(s) to predict |
--runs |
1 |
Number of independent runs |
--epochs |
400 |
Training epochs |
--patience |
20 |
Early stopping patience (with --early_stopping) |
--alpha |
0.25 |
Focal loss alpha |
--gamma |
3.0 |
Focal loss gamma |
--alpha_adv |
2.0 |
Adversarial negative weighting temperature |
--oversample_rate |
5 |
TARGET triple repetition factor |
--undersample_rate |
0.5 |
Fraction of non-TARGET edges to keep |
--negative_rate |
1 |
Negatives per positive |
--pretrain_epochs |
0 |
Multi-relational pretraining epochs |
The script produces:
- Console output: AUROC, AUPRC, MRR, Hits@1/3/10 per run and aggregated
- Saved model:
models/target_<dataset>_<timestamp>/(weights, config, entity/relation mappings)
To validate the training pipeline on DRKG before applying it to PathogenKG:
# Download DRKG
python get_pathogenkg.py
# Run HPO or train directly
python tuning_dataset_drkg.pyBayesian HPO via Weights & Biases:
# Edit ENTITY in tuning_hyperparameter.py with your W&B username
python tuning_hyperparameter.pyThe sweep optimises the composite metric M = 0.2·AUROC + 0.4·AUPRC + 0.4·MRR. Search space and best configurations are documented in src/models_params.json.
After training, run the compound-centric evaluation to rank all pathogen proteins for each drug:
# Evaluate all compounds (produces top-k ranked lists per compound)
python drug_eval.py --model_folder models/<your_model_folder>
# Evaluate a single compound
python drug_eval.py --model_folder models/<your_model_folder> \
--compound "Compound::Pubchem:2764" --topk 20
# Process results into summary tables
python drug_eval_results.pyKey arguments for drug_eval.py:
| Argument | Default | Description |
|---|---|---|
--model_folder |
(required) | Path to trained model directory |
--compound |
all |
Compound ID or all |
--topk |
20 |
Number of top predictions to return |
--batch_size |
4096 |
Scoring batch size |
This produces per-compound ranked lists of predicted protein targets across all 31 pathogen species, corresponding to the biological plausibility analysis in Section 5 and Table 5 of the paper.
To reproduce the dataset statistics (Table 2) and schema figures:
python kg_stats_visualization.pyOutput figures are saved to figures/.
📌 Citation information will be available upon publication. The paper has been accepted at ECML PKDD 2026 (September 7-11, Naples, Italy). BibTeX and additional citation formats will be provided once the proceedings are published.
This project is released under the MIT License.
