WEBVTT
Kind: captions
Language: en

00:00:01.120 --> 00:00:08.160
So hello everyone. I'll be speaking today&nbsp;
about computational work on the problem of&nbsp;&nbsp;

00:00:08.160 --> 00:00:16.000
peptide binding and this is a collaboration&nbsp;
with George Vasmatzis from the Mayo Clinic.

00:00:16.000 --> 00:00:27.440
So with COVID-19, we've all heard that the disease&nbsp;
is- it behaves like any disease. There is a range&nbsp;&nbsp;

00:00:27.440 --> 00:00:33.120
in severity of symptoms that people experience,&nbsp;
but it's perhaps more pronounced with COVID-19&nbsp;&nbsp;

00:00:33.120 --> 00:00:38.240
than other diseases. A good fraction of people who&nbsp;
contract the virus show no symptoms whatsoever.&nbsp;&nbsp;

00:00:38.880 --> 00:00:44.560
Most show either no or mild symptoms, but a&nbsp;
significant fraction shows severe symptoms.&nbsp;&nbsp;

00:00:44.560 --> 00:00:51.680
And of course a small fraction show critical&nbsp;
symptoms or even experience death as a result. So&nbsp;&nbsp;

00:00:52.880 --> 00:00:56.160
there are many factors influencing&nbsp;
the disease severity and experts&nbsp;&nbsp;

00:00:56.960 --> 00:01:02.880
can speak better about this than I can. Age,&nbsp;
sex, and in particular comorbidities are very&nbsp;&nbsp;

00:01:02.880 --> 00:01:10.880
significant and past exposure to similar viruses&nbsp;
plays a role, but a part of the equation here in&nbsp;&nbsp;

00:01:10.880 --> 00:01:17.120
terms of disease severity is also the innate&nbsp;
differences in our immune systems. Essentially&nbsp;&nbsp;

00:01:17.120 --> 00:01:23.360
the genes we've inherited from our parents play&nbsp;
a role in how well our bodies fight off this&nbsp;&nbsp;

00:01:23.360 --> 00:01:32.480
disease or most diseases. So our topic of study&nbsp;
here is to computationally predict one aspect&nbsp;&nbsp;

00:01:32.480 --> 00:01:36.480
of the innate difference in an individual's&nbsp;
immune system's so-called cellular response.

00:01:37.840 --> 00:01:41.440
So the cellular response is&nbsp;
the first line of defense that&nbsp;&nbsp;

00:01:41.440 --> 00:01:48.240
our bodies have in response to any viral&nbsp;
infection. It is the mechanism by which&nbsp;&nbsp;

00:01:49.440 --> 00:01:55.760
foreign peptides introduced into cells are&nbsp;
chopped up into small fragments called peptides.&nbsp;&nbsp;

00:01:56.320 --> 00:02:00.320
These then can be transported to the cell&nbsp;
surface where they can bind to cell surface&nbsp;&nbsp;

00:02:00.320 --> 00:02:04.720
receptors called MHC [Major Histocompatibility&nbsp;
Complex] Class 1 molecules. If bound this way,&nbsp;&nbsp;

00:02:05.280 --> 00:02:11.120
the infected cells become targets for&nbsp;
killer T cells which can come off and kill&nbsp;&nbsp;

00:02:11.120 --> 00:02:16.880
the infected cells- either killed from a viral&nbsp;
infection or as it turns out, this is the most&nbsp;&nbsp;

00:02:16.880 --> 00:02:20.720
effective defense that we have against cancer.&nbsp;
Cancerous cells are also killed off this way.&nbsp;&nbsp;

00:02:21.360 --> 00:02:28.880
This is the first line of defense because if it&nbsp;
kicks in once we're infected, the immune system-&nbsp;&nbsp;

00:02:28.880 --> 00:02:33.680
the killer T cells can kill off all the infected&nbsp;
cells before they have a chance to get going. But&nbsp;&nbsp;

00:02:33.680 --> 00:02:38.800
if the cellular response fails, then the infected&nbsp;
cells become factories. They start churning&nbsp;&nbsp;

00:02:38.800 --> 00:02:44.240
out many many copies of the virus. and other&nbsp;
aspects of the immune system have to take over.

00:02:45.200 --> 00:02:53.760
So the cellular response is- as with all aspects&nbsp;
of immunology complex. Understanding it is very&nbsp;&nbsp;

00:02:53.760 --> 00:02:59.280
important. It's critical for understanding and&nbsp;
predicting the severity of novel viruses such as&nbsp;&nbsp;

00:02:59.280 --> 00:03:05.840
SARS-CoV-2, for vaccine development targeting&nbsp;
viruses, also understanding the impacts of&nbsp;&nbsp;

00:03:05.840 --> 00:03:13.440
viral mutations. How different viruses will affect&nbsp;
different people through the their innate cellular&nbsp;&nbsp;

00:03:13.440 --> 00:03:18.080
immune response, and as I keep returning to&nbsp;
the topic for cancer immunotherapy which is&nbsp;&nbsp;

00:03:18.080 --> 00:03:23.280
the area of expertise of my collaborator-&nbsp;
like many people we turned our attention to&nbsp;&nbsp;

00:03:23.280 --> 00:03:30.400
COVID-19 and repurposed the skill sets. So in this&nbsp;
case, the computational predictions are originally&nbsp;&nbsp;

00:03:30.400 --> 00:03:38.960
targeting cancer immunotherapy. So and of course&nbsp;
all of this same theory applies to autoimmune&nbsp;&nbsp;

00:03:38.960 --> 00:03:48.000
diseases. So the cellular immune response at its&nbsp;
core is a computational problem. If we have the&nbsp;&nbsp;

00:03:48.000 --> 00:03:54.560
blue peptides in the figure on the right this is&nbsp;
the protein fragment associated with the virus.&nbsp;&nbsp;

00:03:54.560 --> 00:04:01.680
The question is will that blue peptide bind&nbsp;
inside a groove or a cleft inside the yellow&nbsp;&nbsp;

00:04:01.680 --> 00:04:10.080
cell surface molecule the MHC 1 molecule? And&nbsp;
the immune response then will be dependent upon&nbsp;&nbsp;

00:04:10.080 --> 00:04:16.560
whether this protein fragment binds and whether&nbsp;
it binds well enough for the killer T cells to&nbsp;&nbsp;

00:04:16.560 --> 00:04:20.800
recognize it. the blue fragments that&nbsp;
come from the virus- those are novel.&nbsp;&nbsp;

00:04:21.680 --> 00:04:29.200
So given a new, virus we will have a completely&nbsp;
new set of peptides. The yellow cell surface&nbsp;&nbsp;

00:04:29.200 --> 00:04:37.520
molecules- those are determined by our genes&nbsp;
and each individual has up to six different&nbsp;&nbsp;

00:04:39.760 --> 00:04:47.600
MHC 1 molecules determined by our inheritance-&nbsp;
three from each parent. And the cell surface&nbsp;&nbsp;

00:04:47.600 --> 00:04:52.880
molecules, the MHC molecules, are among the most&nbsp;
diverse in our genome there are approximately&nbsp;&nbsp;

00:04:52.880 --> 00:05:00.480
21000 variants in the human population. This is no&nbsp;
accident. Evolution has ensured that this aspect&nbsp;&nbsp;

00:05:00.480 --> 00:05:09.280
of our immune system is very diverse so that&nbsp;
we've been able to survive past viral infections.

00:05:10.640 --> 00:05:14.960
But this also poses a very significant&nbsp;
computational challenge as I'll describe.&nbsp;&nbsp;

00:05:14.960 --> 00:05:22.880
So the core problem that we're tackling is how to&nbsp;
apply computer science to predict whether the blue&nbsp;&nbsp;

00:05:22.880 --> 00:05:29.520
peptide will bind inside the yellow cell surface&nbsp;
molecule and to do this for a very wide range of&nbsp;&nbsp;

00:05:29.520 --> 00:05:34.880
peptides all of those associated with the virus&nbsp;
and for the very wide range of cell surface&nbsp;&nbsp;

00:05:34.880 --> 00:05:41.200
receptors the MHC 1 molecules. Now there's been&nbsp;
a lot of prior work- computational prior work.&nbsp;&nbsp;

00:05:41.200 --> 00:05:47.520
There's been experimental work that of course has&nbsp;
provided molecular information on the structure&nbsp;&nbsp;

00:05:47.520 --> 00:05:52.080
of these molecules, and the computational work&nbsp;
that's been very successful has been to apply&nbsp;&nbsp;

00:05:52.080 --> 00:05:55.840
a machine learning- no surprise there for&nbsp;
those who have a background in computer science&nbsp;&nbsp;

00:05:55.840 --> 00:06:02.080
and neural networks. So there's a package called&nbsp;
NetMHC that has been trained extensively on the&nbsp;&nbsp;

00:06:02.080 --> 00:06:07.520
experimental data with the binding strength&nbsp;
for known pairs of the blue yellow molecules-&nbsp;&nbsp;

00:06:07.520 --> 00:06:14.080
the MHC 1 peptide pairs. And based upon&nbsp;
a neural network structure, this program&nbsp;&nbsp;

00:06:14.080 --> 00:06:21.760
can predict, given a new peptide how strong it&nbsp;
will predict. And this inference is based upon&nbsp;&nbsp;

00:06:22.480 --> 00:06:26.960
simply peptide sequence, so the sequence&nbsp;
of letters these are the amino acids in the&nbsp;&nbsp;

00:06:26.960 --> 00:06:33.760
peptide are paired with a label that corresponds&nbsp;
to the cell surface molecule the, MHC 1 molecule.&nbsp;&nbsp;

00:06:33.760 --> 00:06:38.720
There's a strength. That's all the experimental&nbsp;
data. And so given many many thousands such pairs,&nbsp;&nbsp;

00:06:40.160 --> 00:06:44.320
the neural network can be trained. And once&nbsp;
it's trained, given a novel peptide sequence,&nbsp;&nbsp;

00:06:44.320 --> 00:06:50.640
a novel sequence of amino acids, it can predict&nbsp;
how well that peptide will bind to a given&nbsp;&nbsp;

00:06:51.280 --> 00:06:57.680
MHC 1 molecule. So here the peptide sequence&nbsp;
would be the blue peptide molecule in the drawing,&nbsp;&nbsp;

00:06:57.680 --> 00:07:02.800
and the label would correspond to the yellow&nbsp;
cell surface receptor, the MHC 1 molecule.&nbsp;&nbsp;

00:07:04.000 --> 00:07:07.520
This is great and neural networks&nbsp;
are powerful in the sense that&nbsp;&nbsp;

00:07:08.960 --> 00:07:14.960
they can be easily trained and they can very&nbsp;
effectively make predictions based upon the&nbsp;&nbsp;

00:07:14.960 --> 00:07:20.880
data they're given. But the limitation here is&nbsp;
there is absolutely no molecular data whatsoever,&nbsp;&nbsp;

00:07:21.920 --> 00:07:29.760
we're simply training labels and letters, and also&nbsp;
the training data is taken for a very diverse set&nbsp;&nbsp;

00:07:29.760 --> 00:07:35.120
of experimental data. A lot of it for instance&nbsp;
comes from HIV, a different virus, and the&nbsp;&nbsp;

00:07:35.120 --> 00:07:40.880
peptides associated with HIV, and the problem&nbsp;
is given a brand new virus like SARS-CoV-2,&nbsp;&nbsp;

00:07:42.080 --> 00:07:47.600
most of the peptides have never been seen before,&nbsp;
and a neural network will make predictions that&nbsp;&nbsp;

00:07:47.600 --> 00:07:52.640
are spurious because it's making inferences from&nbsp;
a very different region of the peptide space.

00:07:54.240 --> 00:07:58.720
And so our approach would- is to do&nbsp;
molecular-level simulations and of&nbsp;&nbsp;

00:07:58.720 --> 00:08:04.000
course there's been a lot of prior work on this&nbsp;
topic. Very sophisticated molecular simulation&nbsp;&nbsp;

00:08:04.000 --> 00:08:09.440
techniques are known and widely used. One is&nbsp;
called Molecular Dynamics. There's also Monte&nbsp;&nbsp;

00:08:09.440 --> 00:08:15.840
Carlo based simulations. They use a technique&nbsp;
called Simulated Annealing, and Molecular Docking&nbsp;&nbsp;

00:08:15.840 --> 00:08:21.920
is another approach. Software available for such&nbsp;
molecular level simulations are widely used, but&nbsp;&nbsp;

00:08:21.920 --> 00:08:27.120
they're special. They're general purpose. They've&nbsp;
been developed for broad classes of molecules&nbsp;&nbsp;

00:08:28.160 --> 00:08:34.880
binding and they're extremely computationally&nbsp;
intensive. To take a peptide, an MHC 1 pair, and&nbsp;&nbsp;

00:08:34.880 --> 00:08:42.160
to use existing software to simulate it, it takes&nbsp;
days, sometimes weeks to simulate a single binding&nbsp;&nbsp;

00:08:43.200 --> 00:08:50.560
event. So weeks of actually super computing&nbsp;
time to make a single prediction. And the scope&nbsp;&nbsp;

00:08:50.560 --> 00:08:57.360
of the problem we're confronting is we have 21000&nbsp;
variants of the yellow cell surface molecules, the&nbsp;&nbsp;

00:08:58.640 --> 00:09:06.720
MHC1 molecules, and for SARS-CoV-2, if we focus&nbsp;
just on the spike protein and we chop that up&nbsp;&nbsp;

00:09:06.720 --> 00:09:10.800
into little bits for the peptides, we have&nbsp;
about 38000 of those. So we're talking about&nbsp;&nbsp;

00:09:10.800 --> 00:09:16.880
1 billion combinations that we want to simulate&nbsp;
in terms of the strain, and if it takes a week&nbsp;&nbsp;

00:09:16.880 --> 00:09:21.040
of supercomputing time each we obviously&nbsp;
don't have a billion weeks to study this.

00:09:21.760 --> 00:09:28.640
So our approach is twofold. On the one hand,&nbsp;
we're creating highly customized software for&nbsp;&nbsp;

00:09:28.640 --> 00:09:37.680
molecular simulation and we're using the details&nbsp;
of the domain we're working in. So we're starting&nbsp;&nbsp;

00:09:37.680 --> 00:09:41.760
with the peptide. The peptides don't vary so&nbsp;
much in terms of their length or their shape.&nbsp;&nbsp;

00:09:41.760 --> 00:09:46.400
We're starting with the peptides correctly aligned&nbsp;
inside the cleft of the cell surface molecule,&nbsp;&nbsp;

00:09:46.400 --> 00:09:51.120
the MHC 1 molecule, so we don't spend a lot&nbsp;
of time just rotating the entire peptide in&nbsp;&nbsp;

00:09:51.120 --> 00:09:55.920
space. We place it exactly where it should&nbsp;
be and we perform the entire search in the&nbsp;&nbsp;

00:09:55.920 --> 00:09:59.440
torsional space. So instead of moving the&nbsp;
molecule around three-dimensional space we&nbsp;&nbsp;

00:09:59.440 --> 00:10:06.160
just twist and turn its bonds to try to find the&nbsp;
optimal configuration. The other contribution is&nbsp;&nbsp;

00:10:06.160 --> 00:10:11.840
we're deploying this at scale. So we're using GPUs&nbsp;
and then eventually cloud computing infrastructure&nbsp;&nbsp;

00:10:11.840 --> 00:10:17.120
to really throw computing power at the problem.&nbsp;
And our goal as I stated in the title is to turn&nbsp;&nbsp;

00:10:17.120 --> 00:10:21.600
a billion days into a million minutes or&nbsp;
perhaps one month of cloud computing time.

00:10:22.560 --> 00:10:27.440
So the challenge is here and I'll gloss&nbsp;
over this- I'm almost out of time.&nbsp;&nbsp;

00:10:28.480 --> 00:10:35.840
One thing the experimental data is not complete.&nbsp;
We don't have full molecular models, not only of&nbsp;&nbsp;

00:10:35.840 --> 00:10:41.600
all 21000 variants, but we don't even have a good&nbsp;
geographic spreads. This ties in with some of the&nbsp;&nbsp;

00:10:41.600 --> 00:10:44.720
previous talks. Most of the&nbsp;
experimental data is for&nbsp;&nbsp;

00:10:45.280 --> 00:10:52.160
the variance of the MHC 1 molecules from&nbsp;
Western Caucasian demographics. So the other&nbsp;&nbsp;

00:10:52.160 --> 00:10:55.600
approach is we must infer the structure&nbsp;
of the molecules and then simulate them.

00:10:56.800 --> 00:11:03.360
And so final slide: the impact of this work.&nbsp;
Well it's relatively straightforward to&nbsp;&nbsp;

00:11:04.240 --> 00:11:09.040
determine for an individual what genes they&nbsp;
have- that code for the relevant molecules,&nbsp;&nbsp;

00:11:09.040 --> 00:11:14.160
the MHC 1 molecules. It can be done through HLA&nbsp;
typing, which is done for paternity testing.&nbsp;&nbsp;

00:11:14.160 --> 00:11:19.040
Given that information and given our computational&nbsp;
infrastructure, we'll be able to predict for an&nbsp;&nbsp;

00:11:19.040 --> 00:11:25.520
individual how that individual will respond to a&nbsp;
new pathogen. to a virus, to variants of the virus&nbsp;&nbsp;

00:11:26.480 --> 00:11:29.840
for different individuals, and for the&nbsp;
effect of different vaccines on different&nbsp;&nbsp;

00:11:29.840 --> 00:11:33.920
viruses for different individuals. And&nbsp;
of course as I stated at the outset,&nbsp;&nbsp;

00:11:34.560 --> 00:11:38.320
this work will apply not only to viruses but&nbsp;
also potentially the cancer immunotherapy&nbsp;&nbsp;

00:11:38.320 --> 00:11:41.840
and autoimmune diseases. So I'll&nbsp;
stop there. Thank you very much.

