10/23/14
Problem:
Pattern analysis allows us to evaluate characteristics about our datasets. Sometimes, it is challenging to determine if the patterns are real or if they are occurring by chance. Statistical analysis provides tools with which we can answer these questions. Some available tools conduct nearest neighbor analysis - including but not limited to clustering by location, clustering by value, multi-distance clustering, spatial auto-correlation, cluster/outlier analysis, and hotspot analysis. This exercise examined some of these tools. The data used for these analyses were acquired from Dr. Perver Baran at NCSU, as adopted from Chapter 8: Analyzing Patterns from GIS Tutorial II - Spatial Analysis Workbook by David W. Allen (2011). The coordinate system used for these data was NAD_1983_StatePlane_ Texas_North_Central_FIPS_4202_Feet. The data were specific to Battalion 2 of the Fort Worth, Texas, Fire Department, in January and February, 2007. Information was provided about EMT calls for service and fire alarm distributions.
Analysis Procedure:
After opening the map, define the properties of interest, in this case if fire alarms cluster. Open the Average Nearest Neighbor tool to perform the analyses, accessing the values via the report from the Geoprocessing tab and observing the histogram to identify if the data were occurring by chance. The report values include the nearest neighbor ratio and Z-score. The Fire Department would like to know if the calls' ranks are clustering. The null hypothesis for this analysis is: The Priority Ranking values for Calls for Service are randomly distributed across the study area. After opening the map, calculate the distance band with the Neighborhood Count tool to understand the distance to enter into the General G tool. Run the High/Low Clustering tool on several distance bands to identify the distance band with the highest Z-score. The response call data set will be used to analyze for a set of distances to capture the maximum z-score, confidence level, and distance band for clustering. Null hypothesis: The Calls for Service data set, when weighted with priority rankings, are randomly distributed across the study area. After opening the map, run the Get Multi-distance Spatial Cluster Analysis tool to determine at which distance there is the highest significant clustering with both no confidence envelope and another with confidence envelope defined. Scenario: The fire department wants to know at what distance calls for service per block are clustering. Null Hypothesis: Calls for service are randomly distributed among city blocks. After opening the map, aggregate the data by turning on the grid and joining the grid to the layer containing the number of service calls. Define the properties of interest (Count > 0), then open the Spatial Auto-correlation tool to find the peak Z-score. Run tool for some number of iterations (in this case, 8).
Figure 1. Workflow diagram for Spatial Statistics
Results:
With the Nearest Neighbor index, a ratio of less than one and a Z-score highly positive or negative indicates clustering is not due to random chance. For this exercise, at a confidence level of 90%, the Nearest Neighbor Index was less than one and the Z-score was highly negative, indicating that clustering was not due to random chance (Figure 2).
The High/Low Clustering tool performed on several distance bands identified the distance band with the highest Z-score, indicating the most significant clustering occurred at 200 feet (Figure 3). Output produced with the multi-distance spatial cluster analysis indicate that the value exceeding the confidence interval with the highest margin and thus with the most significant clustering was at 900 feet (Figure 4). Based on these results, a high Z-score at 550 feet with a confidence greater than 99% indicates that this data is not clustered by random chance (Figure 5).
Figure 2. Nearest Neighbor Index
Figure 3. High/Low Clustering tool to identify the distance band with the highest Z-score
Figure 4. The Multi-distance Spatial Cluster Analysis to determine at which distance there is the highest significant clustering with both no confidence envelope and another with confidence envelope defined
Figure 5. Spatial Auto-correlation tool to find the peak Z-score run with eight iterations
Application and Reflection:
Spatial statistics appears to be a good means of evaluating how data are distributed and may relate to one another in the study area. After performing the Linear Referencing operations on the mushroom sightings (see Our Mushroom Hunting), it seems that the next step would be to conduct some spatial statistics. I would like to know, for example, if certain mushroom varieties can be affiliated with specific environments, which genera or families cluster in which conditions and whether these might be predictable. With spatial statistics, we may be able to assess whether or not there is clustering and where the hot spots are with respect to specific mushroom groups. The data I would like to have for these analyses are soil moisture content, light level, temperature, relative humidity, barometric pressure, and GPS coordinates collected at the time of the mushroom sightings. I would also like to know the trail segment location, slope, aspect, elevation, and the information produced through linear referencing. I envision using the procedures demonstrated in this segment to begin the process of statistically analyzing mushroom distribution within Eno River State Park.