WEBVTT
Kind: captions
Language: en

00:00:01.980 --> 00:00:07.620
Hi, my name is Hong Qin and I'm&nbsp;
faculty at the University of Tennessee,&nbsp;&nbsp;

00:00:07.620 --> 00:00:17.760
Chattanooga. This talk is about a project&nbsp;
I was fortunate enough to work on with a&nbsp;&nbsp;

00:00:17.760 --> 00:00:26.940
large team of colleagues from North Carolina&nbsp;
A&amp;T, Spelman College, Catholic University of&nbsp;&nbsp;

00:00:26.940 --> 00:00:33.780
America. These are the collaborators on paper,&nbsp;
but in reality, there are many many more. In&nbsp;&nbsp;

00:00:34.440 --> 00:00:41.100
our collaboration we probably had more than a&nbsp;
dozen team members that have been working on this.&nbsp;

00:00:41.940 --> 00:00:50.340
So this is a National Science Foundation supported&nbsp;
project. It's called PIPP Phase One. My goal is&nbsp;&nbsp;

00:00:50.340 --> 00:00:59.820
to develop an AI based framework to predict&nbsp;
and prevent future coronavirus pandemics.&nbsp;&nbsp;

00:00:59.820 --> 00:01:07.560
Although we say it's coronavirus, the model&nbsp;
now has the possibility to be generalizable to&nbsp;&nbsp;

00:01:09.240 --> 00:01:27.420
presumably many other viral pandemics.
So this is the - ok, it's automatically&nbsp;&nbsp;

00:01:27.420 --> 00:01:36.240
forwarding - my apologies. Ok, so this is the&nbsp;
overview of we proposed in the beginning, although&nbsp;&nbsp;

00:01:36.240 --> 00:01:46.080
we we have made a lot of modifications right now.&nbsp;
The idea is still the same. We propose how do we&nbsp;&nbsp;

00:01:46.080 --> 00:01:57.780
predict the virus - a new virus pandemic? Our idea&nbsp;
is first we generate all the possible SARS-CoV-X&nbsp;&nbsp;

00:01:57.780 --> 00:02:03.480
from that are predictable based on recombination&nbsp;
patterns or habitat change. We then&nbsp;&nbsp;

00:02:04.560 --> 00:02:10.860
predict those potential SARS-CoV-X reactions.&nbsp;
Here's the challenge - it's easy to generate&nbsp;&nbsp;

00:02:10.860 --> 00:02:17.340
those sequences, but how do we know which one&nbsp;
actually will become virulent or transmit in&nbsp;&nbsp;

00:02:17.340 --> 00:02:27.000
a human population? That's really the key here.&nbsp;
So we turn to AI. When I say AI is a black box,&nbsp;&nbsp;

00:02:27.000 --> 00:02:32.460
well in this case that's probably also a&nbsp;
blessing because we really don't know how&nbsp;&nbsp;

00:02:33.060 --> 00:02:41.580
a virus jumps from host, from an animal, to a&nbsp;
human to become virulent. So AI is a black box but&nbsp;&nbsp;

00:02:41.580 --> 00:02:47.280
it's probably a great tool we can use even though&nbsp;
it may be a black box. That's my argument here.&nbsp;

00:02:47.280 --> 00:02:54.060
And my apology, I had to black out certain things&nbsp;
because my University is applying for a patent&nbsp;&nbsp;

00:02:54.060 --> 00:03:03.900
based on the function of this. So the goal is here&nbsp;
we have input sequences - this is just primary&nbsp;&nbsp;

00:03:03.900 --> 00:03:10.500
nucleotide sequences. From this sequence, we will&nbsp;
do feature engineering to generate some helpful&nbsp;&nbsp;

00:03:10.500 --> 00:03:15.900
features. Then we use the sequence feature&nbsp;
to feed into our AI model. Based on our AI&nbsp;&nbsp;

00:03:15.900 --> 00:03:25.260
model we will predict the finish. Given the recent&nbsp;
advance of AI, we can also use the [inaudible]&nbsp;&nbsp;

00:03:25.260 --> 00:03:32.100
dynamic simulation or even experimental methods&nbsp;
to verify some candidate genes as the potential&nbsp;&nbsp;

00:03:32.100 --> 00:03:42.840
drivers for the virulence change. Given the AI&nbsp;
model, we also generalize this - basically use&nbsp;&nbsp;

00:03:42.840 --> 00:03:52.380
transfer learning from coronaviruses, presumably&nbsp;
to other viruses such as influenza or HIV,&nbsp;&nbsp;

00:03:55.680 --> 00:04:03.840
HPV. There's many viruses. I'm not actually really&nbsp;
not expert in biology, I'm just computer science.&nbsp;

00:04:03.840 --> 00:04:11.040
And so the challenges we are really dealing with -&nbsp;
my apologies, it's automatically forwarding again.&nbsp;&nbsp;

00:04:13.620 --> 00:04:19.920
The challenge is really to predict&nbsp;
pathogenicity from sequences. We can&nbsp;&nbsp;

00:04:19.920 --> 00:04:24.540
also found a potential rule and test the&nbsp;
general ability by transforming predict&nbsp;&nbsp;

00:04:24.540 --> 00:04:32.940
the - future viruses and predict the outcome.&nbsp;
It will be AI enabled, the early warning system.&nbsp;&nbsp;

00:04:36.540 --> 00:04:43.380
We want to use AI to predict vaccine development.&nbsp;
Given that we know how the virulence variance&nbsp;&nbsp;

00:04:43.380 --> 00:04:50.940
changes. So that's the big picture of this.
So here are some results we have&nbsp;&nbsp;

00:04:51.660 --> 00:05:03.960
developed - that we have accomplished. This&nbsp;
is an estimation of the viral fitness based&nbsp;&nbsp;

00:05:03.960 --> 00:05:09.720
on the information - the data we have. This&nbsp;
is actually for the Omicron sub-variant&nbsp;&nbsp;

00:05:11.940 --> 00:05:19.200
and we can compare the Omicron variant using&nbsp;
a method with development pairwise comparison.&nbsp;&nbsp;

00:05:19.200 --> 00:05:23.700
Then we can estimate the relative&nbsp;
difference between those variants.&nbsp;

00:05:24.360 --> 00:05:29.280
Based on this data we can generate&nbsp;
a so-called viral fitness landscape.&nbsp;&nbsp;

00:05:30.240 --> 00:05:36.480
For those aware of the theory of evolution - in&nbsp;
evolution there is a so-called fitness landscape&nbsp;&nbsp;

00:05:36.480 --> 00:05:41.700
and there's really a lot of arguments about that,&nbsp;
but there's also a lot of theory based on the&nbsp;&nbsp;

00:05:41.700 --> 00:05:52.140
evolution landscape. With the fitness landscape of&nbsp;
a potential virus, we could in theory predict its&nbsp;&nbsp;

00:05:52.140 --> 00:06:04.680
evolutionary trajectory. This - we can - we have&nbsp;
a method to do this. This is the estimated result&nbsp;&nbsp;

00:06:04.680 --> 00:06:13.320
for the alpha beta delta omicron subvariants in&nbsp;
the USA. We are also in the process of applying&nbsp;&nbsp;

00:06:13.320 --> 00:06:21.720
this to influenza or smallpox. We don't have&nbsp;
enough data for smallpox or other viruses,&nbsp;&nbsp;

00:06:21.720 --> 00:06:32.460
but influenza seem to be promising. So this is a&nbsp;
significant finding of this - our current project.&nbsp;

00:06:35.100 --> 00:06:43.560
Here's some model results. Without divulging too&nbsp;
much detail into the AI model we have developed.&nbsp;&nbsp;

00:06:44.160 --> 00:06:50.400
There's really a lot of parameters, but the&nbsp;
accuracy of current model can achieve quite a&nbsp;&nbsp;

00:06:50.400 --> 00:07:00.360
high accuracy. Of course, this is based on what we&nbsp;
know. The challenge is really about the unknown.&nbsp;&nbsp;

00:07:03.720 --> 00:07:10.500
Even though we are predicting the unknown, the&nbsp;
data is already known. In machine learning we&nbsp;&nbsp;

00:07:10.500 --> 00:07:16.500
train it with input training and testing.&nbsp;
The challenge about predicting the unknown&nbsp;&nbsp;

00:07:19.440 --> 00:07:31.740
still remains an open question.
Some people may argue, well why do we have to&nbsp;&nbsp;

00:07:31.740 --> 00:07:37.140
do deep learning? There is a lot of statistical,&nbsp;
genetics, genomics, quantitative treatment - there&nbsp;&nbsp;

00:07:37.140 --> 00:07:44.340
are so many other methods and why we do we even&nbsp;
have to turn into deep learning? Here, we have&nbsp;&nbsp;

00:07:44.340 --> 00:07:50.340
some evidence that show deep learning, at the&nbsp;
minimum, it detects different signatures from the&nbsp;&nbsp;

00:07:50.340 --> 00:07:56.700
standard genome-wide association studies. On the&nbsp;
left, is a result from our deep learning result.&nbsp;&nbsp;

00:07:58.020 --> 00:08:09.120
This is actually based on the WHO labeled&nbsp;
variances Alpha, Beta, Delta, Omicron,&nbsp;&nbsp;

00:08:09.120 --> 00:08:16.320
and others. We want to see what signature in the&nbsp;
SARS-CoV-2 genome are important to contribute to&nbsp;&nbsp;

00:08:16.320 --> 00:08:27.780
this increase of virulence. Those high signals&nbsp;
you can see at the end of the SARS-CoV-2 genome.&nbsp;&nbsp;

00:08:29.220 --> 00:08:36.060
And around this region I'm highlighting at&nbsp;
about 20,000 to 25,000, that's the spike&nbsp;&nbsp;

00:08:36.060 --> 00:08:45.900
gene which WHO [has identified] and most of the&nbsp;
immune response will be reacted to. Based on the&nbsp;&nbsp;

00:08:45.900 --> 00:08:52.260
conventional knowledge, that should be in fact,&nbsp;
that's how WHO classified the SARS-CoV-2 variants.&nbsp;&nbsp;

00:08:54.300 --> 00:09:00.600
Surprisingly AI models pick up those signals,&nbsp;
but those are not the strongest signal,&nbsp;&nbsp;

00:09:00.600 --> 00:09:05.400
and they pick up many other signals. Some&nbsp;
of the stronger signals are not there.&nbsp;&nbsp;

00:09:06.960 --> 00:09:15.300
By using conventional statistical genomic&nbsp;
methods like Genome Wide Association Study - that&nbsp;&nbsp;

00:09:17.460 --> 00:09:22.620
still picks the spike gene as the&nbsp;
stronger signal. There are some others,&nbsp;&nbsp;

00:09:22.620 --> 00:09:30.840
but not very strong. So in this case, the&nbsp;
AI and the conventional statistical method&nbsp;&nbsp;

00:09:30.840 --> 00:09:35.280
pick, or at minimum, give a&nbsp;
different weight to those signals.&nbsp;

00:09:37.440 --> 00:09:48.780
This is surprising and also reassuring, in&nbsp;
a way. So we picked the conventional wisdom,&nbsp;&nbsp;

00:09:48.780 --> 00:09:53.820
the spiking, but we also picked some&nbsp;
other signals which may or may not be&nbsp;&nbsp;

00:09:55.320 --> 00:10:01.620
verified by the experimental method. But how do&nbsp;
we predict and how do we verify this, right? So we&nbsp;&nbsp;

00:10:01.620 --> 00:10:08.700
use AI to predict many things and we know AI can&nbsp;
generate a lot of false predictions. In this case&nbsp;&nbsp;

00:10:08.700 --> 00:10:18.060
how can we verify it? That's quite challenging.
We are trying, in general, two different ways.&nbsp;&nbsp;

00:10:18.060 --> 00:10:29.160
One is to generate a model experimental system&nbsp;
to verify the findings. Our team is using a&nbsp;&nbsp;

00:10:32.280 --> 00:10:39.240
biological model this is the [inaudible]&nbsp;
we also have a cell line system to&nbsp;&nbsp;

00:10:40.320 --> 00:10:46.620
convert those genes into the life cycle&nbsp;
and measure their relative activities.&nbsp;

00:10:49.320 --> 00:10:56.700
Then we also have a computational person to&nbsp;
perform molecular dynamics to simulate how&nbsp;&nbsp;

00:10:56.700 --> 00:11:02.700
those mutations affect [inaudible]&nbsp;
of their activity. In this case,&nbsp;&nbsp;

00:11:02.700 --> 00:11:09.720
ACE2 and RBD, but we we also have other&nbsp;
[inaudible] in which predict in different&nbsp;&nbsp;

00:11:09.720 --> 00:11:17.520
regions how they react with the human genes -&nbsp;
potential targeting the human immune systems.&nbsp;

00:11:19.680 --> 00:11:32.220
We also organize a workshop and expert panels to&nbsp;
discuss how to design trustworthy AI to promote&nbsp;&nbsp;

00:11:32.220 --> 00:11:45.240
social trust in the AI - equitable AI. The COVID&nbsp;
pandemic has shown there is high disparity in&nbsp;&nbsp;

00:11:45.240 --> 00:11:53.880
our current system so if we use AI to predict the&nbsp;
future pandemics, it can it can easily amplify the&nbsp;&nbsp;

00:11:53.880 --> 00:12:03.660
hidden disparities in our current system, probably&nbsp;
in our current data as well. If we recognize this&nbsp;&nbsp;

00:12:03.660 --> 00:12:11.820
and we organize a quite diverse panel, and&nbsp;
including people from South Africa, Kenya,&nbsp;&nbsp;

00:12:13.500 --> 00:12:22.560
there's a person from Europe and across the - with&nbsp;
different background like lawyers, physicians,&nbsp;&nbsp;

00:12:23.280 --> 00:12:35.100
policy makers, governmental - NIST - governmental&nbsp;
agencies to have all kinds of perspective on how&nbsp;&nbsp;

00:12:35.100 --> 00:12:42.480
to promote AI how and to promote trust in&nbsp;
AI in historically marginalized communities.&nbsp;

00:12:45.360 --> 00:12:53.280
We also emphasize a lot of workforce&nbsp;
development given our collaborations with&nbsp;&nbsp;

00:12:55.560 --> 00:12:59.940
Minority Serving Institutions, including&nbsp;
Spelman College, North Carolina A&amp;T,&nbsp;&nbsp;

00:12:59.940 --> 00:13:08.400
and Catholic University of America, which&nbsp;
contain a lot of Hispanic students, and my&nbsp;&nbsp;

00:13:08.400 --> 00:13:12.960
home institution University of Tennessee,&nbsp;
Chattanooga. We also collaborate with a&nbsp;&nbsp;

00:13:12.960 --> 00:13:18.180
hospital system Global South.
Lastly I want to thank NSF,&nbsp;&nbsp;

00:13:18.180 --> 00:13:29.820
UTC and the AI Tennessee Initiative for the&nbsp;
support of this. And lastly, I didn't put in&nbsp;&nbsp;

00:13:29.820 --> 00:13:36.720
the slides, NSF has released call for the phase&nbsp;
2 application for a National Center, to be at&nbsp;&nbsp;

00:13:36.720 --> 00:13:43.980
a scale probably 10 times bigger, if not a 100&nbsp;
times bigger so I'm looking for collaborators,&nbsp;&nbsp;

00:13:43.980 --> 00:13:49.560
especially with community engagement, public&nbsp;
policy, I heard someone working policy. So&nbsp;&nbsp;

00:13:49.560 --> 00:13:59.760
I wish to connect after this.
So I'll just stop here, thank you.

