Data Cleaning
Before we began our analysis of the data, we had to perform some data cleaning. The dataset was imported into OpenRefine, an open-source data wrangling tool, which we used to help clean the data. Since the records were manually compiled together by a team of volunteers, there were many inconsistencies in the syntax and format in which the data was recorded. A lot of the data cleaning work revolved around standardizing these formats, like capitalizing state abbreviations, fixing spelling errors in country names, and making a consistent format for "remarks". These changes allowed us to more easily and accurately analyze the data once we had a standard format for each of the columns.
Isolating Data & Creating Visualizations
We based our visualizations and research questions on the data we had available from the MNHS dataset itself as well as readily accessible public data, namely the 1910 US Census. This led us to focus on the county of origin from decedents (Visualization 1), population of each county (Visualization 1), cause of death (Visualizations 2 & 3), and age of decedents (Visualization 3). From there, we developed visualizations that best suited the data we explored.