WEBVTT
Kind: captions
Language: en

00:00:00.400 --> 00:00:05.920
Hello, everyone my name is Ho-Joon Lee from&nbsp;
Yale School of Medicine and I am going to talk&nbsp;&nbsp;

00:00:05.920 --> 00:00:11.920
about An interactome landscape of SARS-CoV-2&nbsp;
virus-human protein-protein interactions&nbsp;&nbsp;

00:00:11.920 --> 00:00:16.800
by machine learning.
There are two objectives.&nbsp;&nbsp;

00:00:19.280 --> 00:00:25.440
The first is to develop the protein sequence based&nbsp;
multi-class machine learning or deep learning&nbsp;&nbsp;

00:00:25.440 --> 00:00:33.040
classifiers for evidence or confidence level&nbsp;
prediction using the Viruses.STRING database.&nbsp;&nbsp;

00:00:34.000 --> 00:00:41.040
The second is to - using those classifiers we&nbsp;
want to create a draft interactive landscape&nbsp;&nbsp;

00:00:41.040 --> 00:00:44.800
of cytoskeletal virus human&nbsp;
protein-protein interactions.&nbsp;

00:00:46.320 --> 00:00:53.200
So here is an overview of our machine&nbsp;
learning and deep learning workflow. So we use&nbsp;&nbsp;

00:00:53.760 --> 00:00:59.280
the Viruses.STRING database, which did not&nbsp;
include the SARS-CoV-2 at the time of the&nbsp;&nbsp;

00:00:59.280 --> 00:01:06.560
analysis. This is the network of PPI&nbsp;
virus-human PPIs which contain more than 80,000&nbsp;&nbsp;

00:01:07.760 --> 00:01:18.240
interactions between about 1,200 virus proteins&nbsp;
from 102 virus species and about 8,500 human&nbsp;&nbsp;

00:01:18.240 --> 00:01:27.040
proteins. And each interaction has a combined&nbsp;
score ranging from zero to one thousand which&nbsp;&nbsp;

00:01:27.040 --> 00:01:33.520
we convert into five evidence classes pieces.&nbsp;
And this is the distribution of the number PPIs&nbsp;&nbsp;

00:01:34.480 --> 00:01:42.720
for evidence classes. And we are going to focus&nbsp;
on the experimental PPIs which belong to evidence&nbsp;&nbsp;

00:01:42.720 --> 00:01:53.840
class 3 or 2 based on the zero index here. And&nbsp;
based on the data, we first extract node features,&nbsp;&nbsp;

00:01:54.400 --> 00:02:01.200
another protein features which are fractional&nbsp;
compositions of 20 amino acids. And at this point&nbsp;&nbsp;

00:02:01.200 --> 00:02:08.800
we are developing two different models&nbsp;
- one is more canonical [inaudible]&nbsp;&nbsp;

00:02:09.440 --> 00:02:14.960
models like Random Forests and XGBoost, in this&nbsp;
case. And another one is based on deep learning.&nbsp;&nbsp;

00:02:14.960 --> 00:02:21.040
We specifically use graph neural networks like&nbsp;
GraphSAGE or datalized version of HinSAGE..&nbsp;&nbsp;

00:02:23.120 --> 00:02:30.320
For connected machine learning, we also extract&nbsp;
Edge features which are 72 distance or similarity&nbsp;&nbsp;

00:02:30.320 --> 00:02:36.560
measures between amino acid composition profiles&nbsp;
between virus proteins and human proteins.&nbsp;&nbsp;

00:02:37.360 --> 00:02:40.880
And based on the features, we developed&nbsp;
the Random Forests and XGBoost.&nbsp;&nbsp;

00:02:41.600 --> 00:02:47.920
For Random Forests, we optimize 36 models&nbsp;
by research with temp request regulation&nbsp;&nbsp;

00:02:48.480 --> 00:02:55.520
and 432 models for executive space which has the&nbsp;
same contemporary transplantation. And, in short,&nbsp;&nbsp;

00:02:55.520 --> 00:03:02.640
we obtain up to 67% accuracy and 37%&nbsp;
accuracy for Random Forests cases&nbsp;&nbsp;

00:03:02.640 --> 00:03:10.240
and 74% accuracy and 67% accuracy for XGBoost&nbsp;
cases. And this work, this part, has been&nbsp;&nbsp;

00:03:10.240 --> 00:03:16.000
published as a preprint recently. So&nbsp;
you can refer to the paper in detail&nbsp;&nbsp;

00:03:16.000 --> 00:03:19.840
[https://www.biorxiv.org/content/10.1101/2021.11.07.467640v2].&nbsp;
And for GraphSAGE here, still in&nbsp;&nbsp;

00:03:19.840 --> 00:03:24.800
advanced reading and preparation, but I'm going&nbsp;
to show you, briefly show you, the results from&nbsp;&nbsp;

00:03:24.800 --> 00:03:31.840
GraphSAGE as well. Because this shows more than&nbsp;
70% accuracy which is plenty promising as well.&nbsp;

00:03:34.720 --> 00:03:41.040
And here I'm going to show you a performance&nbsp;
example for the best models for 20% of [inaudible]&nbsp;&nbsp;

00:03:41.040 --> 00:03:48.960
with this random seed. We see, in this case,&nbsp;
when the Forest shows 60% accuracy, XBG was&nbsp;&nbsp;

00:03:48.960 --> 00:03:56.160
67.7% accuracy. And if you look at computer&nbsp;
metrics, again, I'm going to focus on this [EC3?]&nbsp;&nbsp;

00:03:56.160 --> 00:04:03.520
which implies mostly expanded PPIs. And&nbsp;
if we look at the individual classes,&nbsp;&nbsp;

00:04:03.520 --> 00:04:09.440
focusing on the f1-score, the extra booster shows&nbsp;
higher f1 scores across four individual classes.&nbsp;

00:04:12.640 --> 00:04:19.920
Based on, based on this to [inaudible]&nbsp;
boost model. The important features&nbsp;&nbsp;

00:04:19.920 --> 00:04:25.760
were identified using two alternating methods&nbsp;
here. One by Gini index and the other by SHAP&nbsp;&nbsp;

00:04:26.320 --> 00:04:30.000
analysis, which based on SHAP,&nbsp;
came through at the SHAP values.&nbsp;&nbsp;

00:04:30.800 --> 00:04:38.800
And interestingly enough, we see that cysteine and&nbsp;
histidine are most - two most important features.&nbsp;&nbsp;

00:04:39.360 --> 00:04:42.480
Where this minus [C_minus and&nbsp;
H_minus] means that the fraction&nbsp;&nbsp;

00:04:42.480 --> 00:04:49.200
of cysteine between virus and human. And the ratio&nbsp;
means the ratio between the fractions cysteine and&nbsp;&nbsp;

00:04:49.200 --> 00:04:58.640
histidine reactions between virus and humans.
One control experiment we performed is to&nbsp;&nbsp;

00:04:58.640 --> 00:05:04.880
compare prediction of experimental PPIs&nbsp;
and - with a prediction of text mining&nbsp;&nbsp;

00:05:04.880 --> 00:05:11.760
PPIs in the virus [inaudible]. Because the&nbsp;
data size, the difference is pretty big here,&nbsp;&nbsp;

00:05:11.760 --> 00:05:18.640
six squared difference, but what we observed here&nbsp;
is that XGBoost, in fact, shows higher accuracy.&nbsp;&nbsp;

00:05:18.640 --> 00:05:26.240
With 94% accuracy compared to 90% accuracy for&nbsp;
the text mining case. So despite the data size&nbsp;&nbsp;

00:05:26.240 --> 00:05:34.160
difference, XGBoost it shows a good prediction&nbsp;
performance. And this is the agreement between&nbsp;&nbsp;

00:05:34.160 --> 00:05:43.280
random force activities for ec3 and test&nbsp;
binding as we expect shows mostly ec1 or ec2.&nbsp;

00:05:46.880 --> 00:05:54.560
So based on those encouraging results, we&nbsp;
applied those classifiers to SARS-CoV-2&nbsp;&nbsp;

00:05:55.280 --> 00:06:02.960
for our second objective in two ways. So&nbsp;
first, we apply that to IntAct database,&nbsp;&nbsp;

00:06:02.960 --> 00:06:10.960
which are a collection of experimental&nbsp;
PPIs. And here I’m showing you the network&nbsp;&nbsp;

00:06:11.680 --> 00:06:18.080
by XGBoost with [inaudible] predicted&nbsp;
evidence. So ec3 for blue, ec4, red.&nbsp;&nbsp;

00:06:18.720 --> 00:06:26.800
So this can be viewed as prioritizing a network.&nbsp;
So although these links would be about 2,000 links&nbsp;&nbsp;

00:06:26.800 --> 00:06:35.040
from experimental data are equally meaningful,&nbsp;
we can also prioritize those links based on this&nbsp;&nbsp;

00:06:35.680 --> 00:06:42.400
[inaudible] class predicted XGBoost in&nbsp;
this case. Secondly, we also apply that&nbsp;&nbsp;

00:06:42.400 --> 00:06:48.160
to protein-wide interaction [inaudible]&nbsp;
the old pairs of more than half a million&nbsp;&nbsp;

00:06:49.360 --> 00:06:58.000
between 27 SARS-CoV-2 proteins and about more&nbsp;
than 20,000 human proteins. And here I'm showing&nbsp;&nbsp;

00:06:58.000 --> 00:07:09.120
you the subset of 22,000 PPIs with evidence class&nbsp;
of at least 2. I either actually use XGBoost or&nbsp;&nbsp;

00:07:09.120 --> 00:07:17.440
Random Forest. And this is the another subset -&nbsp;
140 PPIs with the highest evidence class, 5, by&nbsp;&nbsp;

00:07:18.480 --> 00:07:27.520
XGBoost. And based on this interaction network we&nbsp;
observed that many human proteins are enriching&nbsp;&nbsp;

00:07:27.520 --> 00:07:32.720
vascular smooth muscle contraction and&nbsp;
the targets and also H2A components.&nbsp;

00:07:36.800 --> 00:07:42.320
There are a few more applications of this&nbsp;
work that have been found in the past month,&nbsp;&nbsp;

00:07:42.320 --> 00:07:50.640
actually. So Giuseppe Novelli, who is the renowned&nbsp;
geneticist in Rome, in Italy, he reached out to me&nbsp;&nbsp;

00:07:50.640 --> 00:07:56.000
by email and by surprise, last month. He had&nbsp;
read my preprint telling me about his quality&nbsp;&nbsp;

00:07:56.000 --> 00:08:06.160
therapeutics publication for HECT E3 ligases and&nbsp;
his idea of using the results from this important&nbsp;&nbsp;

00:08:06.160 --> 00:08:12.960
work through this ongoing research. And we&nbsp;
immediately realized that we can help each other&nbsp;&nbsp;

00:08:15.200 --> 00:08:23.600
based on my results on interactive network&nbsp;
results. And we found that HECT-domain protein&nbsp;&nbsp;

00:08:23.600 --> 00:08:28.400
tend to interact with the SARS-CoV-2 proteins&nbsp;
with an evidence class of greater than 2 with&nbsp;&nbsp;

00:08:28.400 --> 00:08:33.760
statistical significance. In other words,&nbsp;
HECT-domain proteins are favored by SARS-CoV-2.&nbsp;&nbsp;

00:08:35.040 --> 00:08:41.200
Based on that observation you're asking whether&nbsp;
there are other protein families favored by&nbsp;&nbsp;

00:08:41.200 --> 00:08:49.600
SARS-CoV-2. In addition we can also extend that&nbsp;
to other virus species like human metapneumovirus,&nbsp;&nbsp;

00:08:49.600 --> 00:08:53.840
which Dr. Novelli is also working on as well.&nbsp;

00:08:55.680 --> 00:08:59.280
So finally, I'm going to show you&nbsp;
briefly about the graph neural networks&nbsp;&nbsp;

00:08:59.280 --> 00:09:07.840
using GraphSAGE and HinSAGE architecture. On&nbsp;
the left are the accuracies by 15 different&nbsp;&nbsp;

00:09:07.840 --> 00:09:13.840
models using three different Java weights in the&nbsp;
columns and five different edge embedding methods.&nbsp;&nbsp;

00:09:15.040 --> 00:09:22.000
As you see, without dropout rates, in fact, we&nbsp;
see more than 70% accuracy values and accuracies&nbsp;&nbsp;

00:09:22.000 --> 00:09:27.440
which are very promising. This is a based on&nbsp;
Viruses.STRING and if we apply that to SARS-CoV-2&nbsp;&nbsp;

00:09:27.440 --> 00:09:34.080
IntAct database, you see the prediction is&nbsp;
enriched with evidence class 2, or 3, in fact,&nbsp;&nbsp;

00:09:34.720 --> 00:09:41.520
which are mostly [inaudible] PPIs. And&nbsp;
this consensus is number of agreements&nbsp;&nbsp;

00:09:42.080 --> 00:09:50.720
among these 15 different models. We see&nbsp;
more consensus of, like, 8-9 for against&nbsp;&nbsp;

00:09:50.720 --> 00:09:58.320
two compared to 6-7 against 1 but I think this&nbsp;
is also very important because - we will see.&nbsp;

00:09:59.840 --> 00:10:03.600
Okay, with that I'd like to thank&nbsp;
my collaborators for very helpful&nbsp;&nbsp;

00:10:03.600 --> 00:10:09.920
discussions and feedback and support. And the Yale&nbsp;
Center for Research Computing for computational&nbsp;&nbsp;

00:10:09.920 --> 00:10:15.760
resources. And COVID HASTE community from Yale’s&nbsp;
School of Engineering &amp; Applied Science. And&nbsp;&nbsp;

00:10:15.760 --> 00:10:23.280
finally the Northeast Big Data Hub Seed Fund&nbsp;
for supporting this work. Thank you so much.

