9/18/2014
Problem:
Businesses are often interested in expanding their customer bases, and many have likely been diligent about maintaining location records of their current customers. It seems reasonable to consider targeting unrepresented areas to drive their expansion. In order to understand which areas are under-represented and improve the efficiencies of advertisement campaigns, they would likely benefit from an analysis of their current data base against locational reference data to identify where their current customers reside and to provide analyses on the best strategy for their advertisement campaigns. The data used for these analyses were provided by Dr. Perver Baran, GIS instructor at NCSU, and the most current versions were downloadable from the Wake County, NC website at: http://www.wakegov.com/gis/services/pages/data.aspx. The files included information about Wake County, NC zip code boundaries, County boundaries, and streets, the U.S. Postal Service zip code look-up tool (https://tools.usps.com/go/ZipLookupAction!input.action), and a customer database. To analyze these data, the customer-reported locations provided by the customer database were geocoded as referenced against the official records kept by Wake County, NC.
Analysis Procedure:
The first step was to geocode addresses using the Wake County Zip Code layer. The second step was to geocode addresses using the Wake County Street Addresses layer. All analyses were performed in ArcMap using the geocoding tools. In ArcCatalog, create an Address Locator using the Create Address Locator tool, paying particular attention to the chosen style, as this will determine how to format the Geocode Addresses tool. This address locator should be deposited into the same folder as the data that will be analyzed. Following this, in ArcMap, geocode addresses, where the address table contains the “unknown” data and is compared against the reference data generated in the ArcCatalog process. This produces matched, tied, and unmatched entries. One objective is to minimize the number of unmatched entries, so the next step is to assess results of the geocode address process and fix and rematch where feasible using the Review/Rematch Addresses tool. In Part 1, this involved looking up Zip Codes via the US Postal Service on-line tool. In Part 2, this involved concatenating the address numbers with the street names, as dictated by the US Addresses – Dual Ranges style, then manually changing components of the addresses to fix them so they could be identified.
Figure 1. Workflow diagram for Geocoding Tabular Data
Results:
When assessing the customer distribution by zip code (Figure 2), the zip codes are fairly evenly distributed across Wake County; there are no distinct patterns in the customer distribution, so using this strategy for this analysis may be less preferable. When assessing customer distribution by street address (Figure 3), there appears to be a stronger pattern to the customer distribution indicating areas with fewer current customers that may be beneficial for targeted advertisement.
Figure 2. Customers Geocoded by Zip Code
Figure 3. Customers Geocoded by Streets
Application and Reflection:
Most of my work is on small plot agricultural research where each field has known GPS coordinates, and each plot within the field is the same size and planted in the same layout. The centroids of each plot might be assigned a “known” address, and therefore essentially become the reference points. When we collect notes from our plots, we tend to collect them by plot name, which is akin to the general single field concept as described in the ESRI geocoding material. It seems this concept is essentially an address locator style, and thus can be used in a matching process. In the process of evaluating our field data, aligning it within the whole field, and then visualizing it against their position within the whole field, it could be useful to run a matching process to identify and fix mis-labeled plot information.