Abstract-Guannan Wang- College of William and Mary

Title: Comparing and Integrating US COVID-19 Data from Multiple Sources with Anomaly Detection and Repairing

Abstract:

Over the past few months, the outbreak of COVID-19 has been expanding over the world. A reliable and accurate dataset of the cases is vital for scientists to conduct related research and for policy-makers to make better decisions. We collect the United States COVID-19 daily reported data from four open sources: the New York Times, the COVID-19 Data Repository by Johns Hopkins University, the COVID Tracking Project at the Atlantic, and the USAFacts, then compare the similarities and differences among them. To obtain reliable data for further analysis, we first examine the cyclical pattern and the following anomalies which frequently occur in the reported cases: (1) the order dependencies violation, (2) abnormal data point or data period, and (3) the delay-reported issue. Based on the issues detected, we propose the corresponding repairing methods and procedures if corrections are necessary. In addition, we integrate the COVID-19 reported cases with the county-level auxiliary information of the local features from official sources, such as health infrastructure, demographic, socioeconomic, and environment information, which are also essential for understanding the spread of the virus.