WEBVTT
Kind: captions
Language: en

00:00:00.480 --> 00:00:05.460
Ioannis Paschalidis:
Thank you so much and thank you for inviting me&nbsp;&nbsp;

00:00:05.460 --> 00:00:14.160
to be part of this exciting afternoon with a great&nbsp;
set of speakers. I will talk about predictive&nbsp;&nbsp;

00:00:14.160 --> 00:00:20.100
models of COVID-19 severity and patient outcomes&nbsp;
that we have developed at Boston University.&nbsp;

00:00:21.540 --> 00:00:26.340
Just to give a little bit of a context - we've&nbsp;
been working for quite some time before the&nbsp;&nbsp;

00:00:26.340 --> 00:00:31.920
pandemic on a variety of models that predict&nbsp;
disease and important key events, for instance,&nbsp;&nbsp;

00:00:31.920 --> 00:00:39.120
hospitalizations, and also models that prescribe&nbsp;
treatments. So when the pandemic started,&nbsp;&nbsp;

00:00:39.120 --> 00:00:45.480
we mobilized, like the rest of the scientific&nbsp;
community, and we also mobilized a relatively&nbsp;&nbsp;

00:00:45.480 --> 00:00:50.640
large network of collaborators so we obtained&nbsp;
access to a variety of different data sets.&nbsp;

00:00:51.420 --> 00:00:57.660
You can sort of see in this slide the&nbsp;
various data sets we had. Data sets&nbsp;&nbsp;

00:00:57.660 --> 00:01:05.040
that were local from Massachusetts - from two&nbsp;
different hospital networks in Massachusetts.&nbsp;&nbsp;

00:01:05.040 --> 00:01:13.500
One was from Mass General Brigham network,&nbsp;
from five different hospitals about 2,500 cases&nbsp;&nbsp;

00:01:13.500 --> 00:01:20.100
and then another relatively large cohort&nbsp;
from the Boston Medical Center of about&nbsp;&nbsp;

00:01:20.880 --> 00:01:29.100
7,000 cases. We also got some cases from&nbsp;
Wuhan, China which was the origin, obviously,&nbsp;&nbsp;

00:01:29.100 --> 00:01:35.640
of the epidemic. And finally we were able to get&nbsp;
access to some large national datasets - so a data&nbsp;&nbsp;

00:01:35.640 --> 00:01:44.940
set from Brazil and another data set from Mexico.
In this very short presentation I will focus more&nbsp;&nbsp;

00:01:44.940 --> 00:01:52.500
on the most recent work that considered the&nbsp;
dataset - the series of patients from the&nbsp;&nbsp;

00:01:52.500 --> 00:02:00.120
Boston Medical Center. Boston Medical Center is&nbsp;
the teaching hospital affiliated with BU Medical&nbsp;&nbsp;

00:02:00.120 --> 00:02:06.960
School and is also a safetynet hospital. As you&nbsp;
will see, this has some interesting implications&nbsp;&nbsp;

00:02:06.960 --> 00:02:16.080
in the findings that we were able to get.
So we got access to the entire 2020 BMC&nbsp;&nbsp;

00:02:16.080 --> 00:02:21.840
cohort of about 7,000 patients. Those were&nbsp;
patients who tested positive for COVID-19.&nbsp;&nbsp;

00:02:23.340 --> 00:02:31.980
Just some rough statistics - about 20 percent were&nbsp;
admitted. From those admitted, about 23 or so were&nbsp;&nbsp;

00:02:31.980 --> 00:02:40.740
admitted to the ICU. From those admitted to the&nbsp;
ICU, about 60 or 58.7 percent were intubated. From&nbsp;&nbsp;

00:02:40.740 --> 00:02:48.180
those intubated, about 70 percent unfortunately&nbsp;
did not make it. We had lots of information&nbsp;&nbsp;

00:02:48.180 --> 00:02:53.220
about these patients, including demographics,&nbsp;
their vitals throughout the hospital stay,&nbsp;&nbsp;

00:02:53.220 --> 00:03:00.660
radiology reports, their medical history, any&nbsp;
symptoms, any lab results, any medications,&nbsp;&nbsp;

00:03:00.660 --> 00:03:07.320
and even information about depression status,&nbsp;
zip code, and also we had information about&nbsp;&nbsp;

00:03:07.320 --> 00:03:12.840
the occupancy of the hospital at the time&nbsp;
that each one of these patients were seen.&nbsp;&nbsp;

00:03:13.380 --> 00:03:21.840
We also had information on social determinants&nbsp;
of Health. BMC runs a program called the Thrive&nbsp;&nbsp;

00:03:21.840 --> 00:03:29.580
Program that everyone who has an encounter with&nbsp;
the hospital is given a survey where we asked&nbsp;&nbsp;

00:03:30.720 --> 00:03:38.580
them about needs in a variety of different areas,&nbsp;
including housing, food, transportation, help with&nbsp;&nbsp;

00:03:38.580 --> 00:03:46.200
caretaking, medications, helping getting access to&nbsp;
medications and paying for medications, education,&nbsp;&nbsp;

00:03:46.200 --> 00:03:56.280
and employment. A lot of the information was&nbsp;
in tabular format that is more easily handled&nbsp;&nbsp;

00:03:56.280 --> 00:04:01.860
by AI machine learning methods, but also there&nbsp;
was quite a bit of information, particularly&nbsp;&nbsp;

00:04:01.860 --> 00:04:10.140
radiology reports and other reports during the&nbsp;
hospital stay, that were just narratives - reports&nbsp;&nbsp;

00:04:10.140 --> 00:04:16.260
by clinicians. We had to use quite a bit of&nbsp;
natural language processing in order to extract&nbsp;&nbsp;

00:04:16.260 --> 00:04:24.780
appropriate findings from the text. What we did&nbsp;
is we developed a set of robust interpretable&nbsp;&nbsp;

00:04:24.780 --> 00:04:30.240
models that can predict hospitalization,&nbsp;
ICU admission, mechanical ventilation,&nbsp;&nbsp;

00:04:30.240 --> 00:04:36.660
and death. I'll show you some examples. I will&nbsp;
not be exhaustive, again in the interest of time.&nbsp;

00:04:37.620 --> 00:04:48.300
First, what we did is for every patient we&nbsp;
constructed a timeline from the time that we were&nbsp;&nbsp;

00:04:48.300 --> 00:04:57.060
aware of the patient testing positive and having&nbsp;
information about this patient up to the point of&nbsp;&nbsp;

00:04:57.060 --> 00:05:05.460
the event of interest whether that was, let's say,&nbsp;
an ICU admission or mechanical ventilation. Here&nbsp;&nbsp;

00:05:05.460 --> 00:05:14.880
is the outcome of interest, and then looking back&nbsp;
in time, we created these data buckets. We dropped&nbsp;&nbsp;

00:05:14.880 --> 00:05:25.020
any information that was available just before&nbsp;
the time of the event of interest becuase we&nbsp;&nbsp;

00:05:25.020 --> 00:05:30.900
wanted models that could predict what will happen&nbsp;
into the future. Perhaps even a first-year medical&nbsp;&nbsp;

00:05:30.900 --> 00:05:39.120
student can identify a patient heading to the ICU,&nbsp;
so we wanted the models to make that prediction&nbsp;&nbsp;

00:05:39.120 --> 00:05:46.680
with earlier information. You will see that we&nbsp;
obtained different versions of the models that had&nbsp;&nbsp;

00:05:46.680 --> 00:05:52.140
different cutoffs in terms of the information that&nbsp;
the model used in order to make the prediction.&nbsp;&nbsp;

00:05:52.140 --> 00:06:01.200
Another reason we created these time buckets is&nbsp;
we wanted to capture the dynamic evolution of&nbsp;&nbsp;

00:06:01.200 --> 00:06:06.840
the progress of the patient while in the hospital,&nbsp;
for instance, the dynamic evolution of the vitals.&nbsp;&nbsp;

00:06:07.620 --> 00:06:13.440
We understood that this was rather important.&nbsp;
So rather than just looking at the snapshot of&nbsp;&nbsp;

00:06:13.440 --> 00:06:19.440
vitals, let's say at some specific time and using&nbsp;
that information for making a forward prediction,&nbsp;&nbsp;

00:06:21.360 --> 00:06:27.840
the exact values are important but&nbsp;
trends are important as well. Physicians,&nbsp;&nbsp;

00:06:27.840 --> 00:06:33.180
when they look at patients, they sort of look&nbsp;
at the trends of the vitals in the patient. So&nbsp;&nbsp;

00:06:33.180 --> 00:06:38.940
we developed a model that used some fairly&nbsp;
sophisticated deep learning methodologies,&nbsp;&nbsp;

00:06:38.940 --> 00:06:50.040
including LSTM type of networks and a transformer&nbsp;
architecture that took as input the vitals - six&nbsp;&nbsp;

00:06:50.040 --> 00:06:58.320
vital signals at different points in time - and&nbsp;
produced a score. That score captured the dynamic&nbsp;&nbsp;

00:06:58.320 --> 00:07:05.460
evolution of the vitals and that vital score was&nbsp;
then used in an ensemble model that was attempting&nbsp;&nbsp;

00:07:05.460 --> 00:07:12.780
to make a prediction for the outcome of interest.
For instance, hospitalization predictions. You&nbsp;&nbsp;

00:07:12.780 --> 00:07:19.800
can see that these are fairly accurate. Some of&nbsp;
the best models give you a 92 percent - this is&nbsp;&nbsp;

00:07:19.800 --> 00:07:26.640
in area under the curve. You can think of this&nbsp;
as a measure of accuracy of the model. The best&nbsp;&nbsp;

00:07:26.640 --> 00:07:35.040
is 100 percent, a random guess will give you 50&nbsp;
percent. So 92 percent is quite good performance.&nbsp;&nbsp;

00:07:36.240 --> 00:07:42.720
You can see here that from some linear models&nbsp;
that we developed we also found some of the&nbsp;&nbsp;

00:07:42.720 --> 00:07:50.880
factors that were important in making the&nbsp;
hospitalization prediction. In blue, you&nbsp;&nbsp;

00:07:50.880 --> 00:07:59.400
see some of the variables that are associated with&nbsp;
some earlier health conditions that, for instance,&nbsp;&nbsp;

00:07:59.400 --> 00:08:06.780
are highly correlated with hospitalization. You&nbsp;
will also see that the occupancy of the hospital,&nbsp;&nbsp;

00:08:06.780 --> 00:08:15.780
if it was high, reduced the likelihood that&nbsp;
the patient is going to be hospitalized. Also,&nbsp;&nbsp;

00:08:15.780 --> 00:08:20.520
you will see two social determinants&nbsp;
of health: need for food and need for&nbsp;&nbsp;

00:08:20.520 --> 00:08:28.200
transportation. These were both contributing to&nbsp;
a hospitalization decision. Patients with those&nbsp;&nbsp;

00:08:28.200 --> 00:08:34.380
needs were more likely to be hospitalized.&nbsp;
I would like to emphasize the role of&nbsp;&nbsp;

00:08:34.380 --> 00:08:39.900
these social determinants of health. This is&nbsp;
something that we also saw in other datasets,&nbsp;&nbsp;

00:08:39.900 --> 00:08:45.240
particularly in the Brazilian dataset,&nbsp;
that was a national data set and we found&nbsp;&nbsp;

00:08:45.240 --> 00:08:52.020
that socio-demographic factors were&nbsp;
impacting hospitalization decisions.&nbsp;

00:08:53.220 --> 00:09:01.860
What we also found was that the model - the naive&nbsp;
model that one is able to produce - is actually&nbsp;&nbsp;

00:09:01.860 --> 00:09:11.400
rather biased. You could see here how the model&nbsp;
performs out of samples for Black individuals and&nbsp;&nbsp;

00:09:11.400 --> 00:09:19.620
White individuals. The false positive rate of the&nbsp;
model for Black individuals was twice as much as&nbsp;&nbsp;

00:09:19.620 --> 00:09:26.940
the model for - the false positive rate for white&nbsp;
individuals - even though we were controlling for&nbsp;&nbsp;

00:09:26.940 --> 00:09:33.180
race, we were controlling for socio-demographic&nbsp;
factors, we were controlling for social&nbsp;&nbsp;

00:09:33.180 --> 00:09:41.400
determinants of health. Despite that, the model&nbsp;
was much more eager to make the prediction that&nbsp;&nbsp;

00:09:41.400 --> 00:09:47.820
the Black individual was going to be hospitalized&nbsp;
compared to a White individual. Correspondingly,&nbsp;&nbsp;

00:09:47.820 --> 00:09:55.380
it was most likely to make a false negative&nbsp;
prediction for a White individual compared to&nbsp;&nbsp;

00:09:55.380 --> 00:10:03.060
a Black individual, which suggests that they are&nbsp;
apparently hidden features in the data that are&nbsp;&nbsp;

00:10:03.060 --> 00:10:13.320
not visible to us, perhaps reflecting structural&nbsp;
bias and other factors that make the model make&nbsp;&nbsp;

00:10:13.320 --> 00:10:21.000
that biased prediction. There are ways and we have&nbsp;
addressed them in a paper we published on how one&nbsp;&nbsp;

00:10:21.000 --> 00:10:26.160
can correct for these factors and produce&nbsp;
models that do not have this sort of bias.&nbsp;

00:10:28.740 --> 00:10:36.240
These are some results on ICU prediction -&nbsp;
predicting an ICU admission -roughly the median&nbsp;&nbsp;

00:10:36.240 --> 00:10:45.420
gap between an admission to the hospital and&nbsp;
admission to the ICU at least in our data set was&nbsp;&nbsp;

00:10:45.420 --> 00:10:55.260
about four hours. We will use different cutoffs.&nbsp;
If you use the latest information, then you get&nbsp;&nbsp;

00:10:55.260 --> 00:11:04.440
quite accurate models with AUCs on the order&nbsp;
of 93-95 percent. If you start cutting off the&nbsp;&nbsp;

00:11:04.440 --> 00:11:09.720
information that you were going to use so 12 hours&nbsp;
in advance the performance of the model drops&nbsp;&nbsp;

00:11:09.720 --> 00:11:18.060
to about 86 percent. 24 hours in advance, the&nbsp;
performance of the model drops to about roughly&nbsp;&nbsp;

00:11:18.060 --> 00:11:26.880
80 percent. What I would like to emphasize is that&nbsp;
we compared these models that we developed to some&nbsp;&nbsp;

00:11:26.880 --> 00:11:35.580
standard models that predict ICU admissions, there&nbsp;
are some well-known sepsis models, NEWS, qSOFA,&nbsp;&nbsp;

00:11:35.580 --> 00:11:42.120
they're called, and these are fairly inaccurate&nbsp;
in this case, indicating that standard models&nbsp;&nbsp;

00:11:42.120 --> 00:11:49.140
for ICU prediction, at least in the COVID cases,&nbsp;
fail to predict an ICU admission. This indicates&nbsp;&nbsp;

00:11:49.140 --> 00:11:57.180
a rather unique signature of the disease. Here,&nbsp;
you find some of the variables that were again&nbsp;&nbsp;

00:11:57.180 --> 00:12:06.180
highly correlated with the outcome with an ICU&nbsp;
admission. What was interesting to us was that&nbsp;&nbsp;

00:12:06.180 --> 00:12:13.320
this vital score that we produced that captured&nbsp;
the dynamic evolution of the vitals pretty&nbsp;&nbsp;

00:12:13.320 --> 00:12:20.400
much tells the entire story. There are some&nbsp;
other variables or some lab variables (LDH,&nbsp;&nbsp;

00:12:20.400 --> 00:12:29.580
CRP) that have been identified by other studies&nbsp;
that also are contributing, but if one just takes&nbsp;&nbsp;

00:12:29.580 --> 00:12:36.480
the dynamic evolution of the vitals, that&nbsp;
pretty much tells the ICU admission story.&nbsp;

00:12:39.900 --> 00:12:49.560
Finally we produced a number of calculators&nbsp;
that we made available on on the web. This,&nbsp;&nbsp;

00:12:49.560 --> 00:12:56.220
I understand, were used by our colleagues and&nbsp;
collaborators at Mass General Hospital in the&nbsp;&nbsp;

00:12:56.220 --> 00:13:03.120
early stages of the epidemic. It was very easy&nbsp;
to input some of the key variables and get a&nbsp;&nbsp;

00:13:03.120 --> 00:13:09.660
prediction, for instance, for an ICU admission,&nbsp;
or for a mechanical intubation need of a patient,&nbsp;&nbsp;

00:13:09.660 --> 00:13:17.940
and we found that, you know, we were challenged&nbsp;
to find cases - what is the model telling me that&nbsp;&nbsp;

00:13:17.940 --> 00:13:26.040
the very experienced clinician cannot potentially&nbsp;
predict by just looking at the patient. Here are&nbsp;&nbsp;

00:13:26.040 --> 00:13:32.400
some cases, and there are many others, where&nbsp;
patients were admitted, they were stable for&nbsp;&nbsp;

00:13:32.400 --> 00:13:39.540
a couple of days, nothing in their clinical&nbsp;
outlook suggested that the condition of this&nbsp;&nbsp;

00:13:39.540 --> 00:13:46.380
patient was going to deteriorate, but then the&nbsp;
model was able - upon admission - based on some&nbsp;&nbsp;

00:13:46.380 --> 00:13:54.660
special laboratory results, predict that the&nbsp;
patient was going to need the ICU care. These&nbsp;&nbsp;

00:13:54.660 --> 00:14:04.380
are two cases with the details of those cases.
So that brings me to a conclusion. Of course, I'm&nbsp;&nbsp;

00:14:04.380 --> 00:14:11.760
just presenting. There have been many people that&nbsp;
have contributed to this work and I would like to&nbsp;&nbsp;

00:14:11.760 --> 00:14:18.480
thank them, including the students in my group,&nbsp;
but also our collaborators at Mass General Brigham&nbsp;&nbsp;

00:14:18.480 --> 00:14:25.380
Boston Medical Center and some of the other&nbsp;
areas where we've been able to get data from.&nbsp;&nbsp;

00:14:25.380 --> 00:14:30.900
Thank you so much for your attention and looking&nbsp;
forward to questions at the end of the session.

