You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Classifying tumors by supervised network propagation
Software overview
We develop a general algorithmic framework by adapting the Supervised Random Walk (SRW) algorithm (Backstrom and Leskovec, 2010) with a novel loss function designed specifically for cancer subtype classification. The package is called Network-Based Supervised Stratification (NBS2).
The package
SRW_v044 This software package contains all the functions of NBS2.
data_processing_BRCA Script to processes PathwayCommons interaction features and Breast Cancer mutation profiles.
SRW_cookbook_BRCA Run the NBS2 package to classify breast cancer subtypes.
Equations
equations_v044 This document contains equations of the algorithm.
Figures
Fig. 1. NBS2 workflow. NBS2 takes three input data sets, represented by red arrows: 1) a molecular network where each edge is annotated by a set of features x, and each feature is assigned a initial weight w; 2) a tumor-by-gene matrix P(0) representing the mutation profile of a cohort; and 3) the defined subtype of each tumor. In each iteration, NBS2 compute an activation score a for each edge (Eq. 3), calculate a transition matrix Q (Eq. 2), perform a random walk (Eq. 1), and compute the value of the cost function J(w) (Eq. 4). Training the classifier is conducted iteratively using gradient descent. To minimize J(w), the algorithm calculates the partial derivative of J(w) with respect to the edge feature weights w using the chain rule (Eqs. 7-11), and updates w accordingly. Upon convergence, the algorithm outputs the final feature weights w, transition matrix Q and propagated mutation profiles P, which together defines the classification model.
Fig. 2. Experiments on simulated data. (A) Simulated mutation dataset including characteristic genes of two subtypes (gene 1-10) and an Frequently Mutated Gene (FMG) (gene 16). Mutated genes are shown in dark red and non-mutated genes are shown in white. A reduced set of tumors and genes (10 x 20) is shown as an example; the full simulation is (100 ✕ 1000). An edge-by-feature matrix is also used as a input for the supervised random walk. (B) Unsupervised and (C) supervised random walk of mutations over a simulated gene interaction network. Shades of red show propagated mutation values for tumor sample #7. (D, E) Propagated mutation profiles following (D) unsupervised and (E) supervised random walk. (F, G) Principal components analysis (PCA) of the full simulation (100✕1000) between (F) unsupervised random walk-based tumor stratification and (G) supervised random walk-based tumor classification.
Fig. 3. Performance of breast cancer subtypes classification. Values of cost function (left y-axis) and classification accuracy (right y-axis) are plotted against the number of iterations of NBS2 on the (A) training data and (B) validation data. Dash line indicates the accuracy of tumor stratification based on unsupervised random walk and non-propagated mutation profiles on the validation data.
Fig. 4. Subnetworks of breast cancer subtypes. Subnetworks characterizing breast cancer subtypes extracted from Pathway Commons by NBS2, defined as a set of genes with significantly different network-transformed scores between different subtypes (ANOVA FDR < 0.05) and their molecular interactions with higher than average activation scores (> 0.01) learned from NBS2. The pie chart represents the relative proportions of the average propagated mutation score for the four subtypes. For example, the large light blue pie slice on ERBB2 represents that its average propagation score is much higher in HER2 tumors than other subtypes.