Skip to content

Latest commit

 

History

6 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

PathogenKG Logo

Hugging Face PyTorch PyG Python License

📢 Accepted at ECML PKDD 2026

Welcome to the official repository for:

PathogenKG: Cross-Species Drug Repurposing via Heterogeneous Knowledge Graph Link Prediction

European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases

📍 Naples, Italy | 🗓️ September 7-11, 2026



PathogenKG is a heterogeneous knowledge graph embedding framework for cross-species antimicrobial drug repurposing. It integrates protein–protein interaction (PPI) networks from 31 bacterial pathogens (STRING), COG orthology, Gene Ontology annotations, and validated drug–target associations (DrugBank) into a unified graph of over 3 million triples. Heterogeneous GNN encoders (R-GCN, CompGCN) with a DistMult decoder are trained to predict new drug–target interactions (DTIs) via link prediction on the TARGET relation.

PathogenKG schema


Repository Structure

PathogenKG/
├── train_and_eval.py           # Main training & evaluation script (Section 3.2)
├── drug_eval.py                # Compound-centric repurposing evaluation (Section 3.3)
├── drug_eval_results.py        # Process and summarise drug evaluation outputs
├── tuning_hyperparameter.py    # Bayesian HPO via Weights & Biases sweeps (Section 3.2)
├── tuning_dataset.py           # HPO on PathogenKG dataset
├── tuning_dataset_drkg.py      # HPO on DRKG benchmark (Section 3.2, pipeline validation)
├── model_eval.py               # Standard model evaluation utilities
├── evaluate_ranking.py         # Ranking metric computation
├── kg_stats_visualization.py   # KG statistics and figures (Table 2, Figure 2)
├── build_pathogenkg.py         # Build organism-specific KGs from raw data (Section 3.1)
├── merge_pathogen_kgs.py       # Merge organism KGs into PathogenKG (Section 3.1)
├── generate_all_pathogenkg.py  # Batch-build all 31 pathogen KGs
├── get_pathogenkg.py           # Download pre-built datasets from HuggingFace
│
├── src/
│   ├── utils.py                # Data loading, splitting, negative sampling, metrics
│   ├── evaluation_metrics_filtered.py  # Type-constrained filtered evaluation (Section 3.2)
│   ├── hetero_compgcn.py       # CompGCN encoder (Section 3.2)
│   ├── hetero_rgcn.py          # R-GCN encoder (Section 3.2)
│   ├── hetero_rgat.py          # R-GAT encoder (experimental)
│   ├── models_params.json      # Best hyperparameter configurations
│   ├── bio_utils.py            # Biological ID mapping utilities
│   └── downloaders/            # Scripts to fetch raw data from STRING, DrugBank, UniProt, GO, COG
│
├── dataset/
│   ├── PathogenKG_n31_core.tsv.zip   # Pre-built PathogenKG (31 species, 3M triples)
│   └── DRUGBANK/                     # Auxiliary DrugBank mappings
│
├── figures/                    # Figures used in the paper
├── requirements.txt
└── environment.yml

1. Environment Setup

Requirements: Python 3.10, CUDA-capable GPU (≥16 GB VRAM recommended for CompGCN with hidden dim 200).

# Create conda environment and install dependencies
conda create -n pathogenkg python=3.10 -y
conda activate pathogenkg
pip install -r requirements.txt

If torch-sparse installation fails (common on some platforms), install PyTorch and PyG extensions separately:

pip install --index-url https://download.pytorch.org/whl/cu128 torch torchvision torchaudio
pip install --no-cache-dir --only-binary=:all: \
  pyg_lib torch-geometric torch-scatter torch-sparse torch-cluster torch-spline-conv \
  termcolor torcheval \
  -f https://data.pyg.org/whl/torch-2.7.1+cu128.html

Adjust the CUDA version (cu128) to match your local installation.


2. Dataset

Option A: Use the pre-built dataset (recommended)

The default dataset PathogenKG_n31_core.tsv.zip (31 core STRING species, ~3M triples) is included in dataset/. No further action is needed.

Additional datasets can be downloaded from HuggingFace:

python get_pathogenkg.py

This downloads both PathogenKG variants and the raw source data to dataset/.

Option B: Build from scratch (Section 3.1)

To reproduce the KG construction pipeline from raw biological databases:

# 1. Download source data (STRING PPI, DrugBank, COG, GO)
#    Scripts in src/downloaders/ fetch from public APIs.

# 2. Build organism-specific KGs for all 31 pathogens
python generate_all_pathogenkg.py

# 3. Merge into a single multi-organism graph
python merge_pathogen_kgs.py

Dataset format

PathogenKG is stored as a TSV of triples: head, interaction, tail.

  • Entity types: ExtGene (pathogen proteins, UniProt accession) and Compound (drugs, PubChem CID)
  • Relation types (13): 8 PPI types (ASSOCIATION, GENE_BIND, ACTIVATION, REACTION, CATALYSIS, PTMOD, INHIBITION, EXPRESSION), ORTHOLOGY, 3 GO types (BiologicalProcess, MolecularFunction, CellularComponent), and TARGET

3. Reproducing the Results

3.1 Training and Evaluation (Table 4)

Train a CompGCN model on PathogenKG with the best hyperparameters (stored in src/models_params.json):

# Single run with default settings (CompGCN, TARGET task, PathogenKG_n31_core)
python train_and_eval.py

# 12-run evaluation for statistical comparison (as in Table 4)
python train_and_eval.py --model compgcn --runs 12 --epochs 400 --task TARGET

# R-GCN comparison
python train_and_eval.py --model rgcn --runs 12 --epochs 400 --task TARGET

Key arguments:

Argument Default Description
--model compgcn GNN encoder: compgcn, rgcn, or rgat
--tsv dataset/PathogenKG_n31_core.tsv.zip Path to dataset
--task TARGET Relation type(s) to predict
--runs 1 Number of independent runs
--epochs 400 Training epochs
--patience 20 Early stopping patience (with --early_stopping)
--alpha 0.25 Focal loss alpha
--gamma 3.0 Focal loss gamma
--alpha_adv 2.0 Adversarial negative weighting temperature
--oversample_rate 5 TARGET triple repetition factor
--undersample_rate 0.5 Fraction of non-TARGET edges to keep
--negative_rate 1 Negatives per positive
--pretrain_epochs 0 Multi-relational pretraining epochs

The script produces:

  • Console output: AUROC, AUPRC, MRR, Hits@1/3/10 per run and aggregated
  • Saved model: models/target_<dataset>_<timestamp>/ (weights, config, entity/relation mappings)

3.2 DRKG Pipeline Validation (Section 3.2)

To validate the training pipeline on DRKG before applying it to PathogenKG:

# Download DRKG
python get_pathogenkg.py

# Run HPO or train directly
python tuning_dataset_drkg.py

3.3 Hyperparameter Optimisation (Section 3.2)

Bayesian HPO via Weights & Biases:

# Edit ENTITY in tuning_hyperparameter.py with your W&B username
python tuning_hyperparameter.py

The sweep optimises the composite metric M = 0.2·AUROC + 0.4·AUPRC + 0.4·MRR. Search space and best configurations are documented in src/models_params.json.

3.4 Compound-Centric Drug Repurposing (Section 3.3, Table 5)

After training, run the compound-centric evaluation to rank all pathogen proteins for each drug:

# Evaluate all compounds (produces top-k ranked lists per compound)
python drug_eval.py --model_folder models/<your_model_folder>

# Evaluate a single compound
python drug_eval.py --model_folder models/<your_model_folder> \
    --compound "Compound::Pubchem:2764"  --topk 20

# Process results into summary tables
python drug_eval_results.py

Key arguments for drug_eval.py:

Argument Default Description
--model_folder (required) Path to trained model directory
--compound all Compound ID or all
--topk 20 Number of top predictions to return
--batch_size 4096 Scoring batch size

This produces per-compound ranked lists of predicted protein targets across all 31 pathogen species, corresponding to the biological plausibility analysis in Section 5 and Table 5 of the paper.


4. KG Statistics and Visualisation

To reproduce the dataset statistics (Table 2) and schema figures:

python kg_stats_visualization.py

Output figures are saved to figures/.


Citation

📌 Citation information will be available upon publication. The paper has been accepted at ECML PKDD 2026 (September 7-11, Naples, Italy). BibTeX and additional citation formats will be provided once the proceedings are published.

License

This project is released under the MIT License.

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages