This repository contains a deep learning model for predicting protein functions, including EC numbers and Gene Ontology (GO) terms, from amino acid sequences.
The model is designed for accurate functional annotation in computational biology research, drug discovery, and protein engineering.
- Multi-label prediction of protein functions
- Embeddings for protein sequences
- Supports GPU acceleration with PyTorch
- Pre-trained model available for use
- Clone the repository:
git clone https://github.com/yaan-jang/emo.git
cd eimo- Create a Python virtual environment:
conda create -n eimo python=3.9 -y
conda activate eimo- Install dependencies:
conda install pytorch torchvision torchaudio pytorch-cuda=11.8 -c pytorch -c nvidia
conda install numpy scikit-learn biopython tqdm -c conda-forgeFor example, predicting enzyme number for given proteins
python predict.py --config=default_config.yml --kind=ECN --params=checkpoints/ECN.pt --input_list=list_of_protein_ids --input_path=path_features - Embedding Layer: Encodes amino acids
- Convolutional / Transformer Layers: Capture sequence motifs
- Fully Connected Layers: Map features to function classes
- Output Layer: Multi-label probabilities via sigmoid activation
- Python >= 3.9
- PyTorch >= 2.0
- NumPy, Scikit-learn
- Biopython
This repository is licensed under the Apache License 2.0. See LICENSE for details.