I recently completed a project in which I was given 500,000 lines of greyhound racing data and was tasked with predicting scores in upcoming races. Completing the modeling in Python, I was able to use a mix of clustering algorithms and ridge Regression in order to create a high-performing solution. I used many different approaches and models before finalizing my results, which are detailed below.
The goal of this project was to minimize the mean squared error of our final model when predicting times in upcoming races. Each line of data was provided in CSV format, as illustrated below:
Each entry describes a forthcoming race for a particular greyhound, along with information about its previous race. The date1, time1, distance1, trap1, and comment1 columns contain details of the dog's last race, while date2, distance2, and trap2 describe the upcoming race. The time2 column records the time run in the upcoming race and appears only in the training data. Using this training set, we aim to build a model that can predict time2 for similar datasets where that column is missing.
In my initial approach, I began by engineering several features from the dataset, including greyhound age, a consistency metric based on the standard deviation of recent times, and the difference in distance between the previous and upcoming race. I then examined a correlation matrix to identify any relationships or patterns that could inform the model.
As expected, several factors exhibited strong correlations. The time and distance of the upcoming race were both highly correlated with those of the previous race. I also observed a correlation between trap1 and trap2, which may suggest that greyhounds tend to start from familiar trap positions. Age showed an interesting relationship with both the dog's consistency and the sentiment of the race comments.
Guided by these findings, I next examined scatterplots to look for clearer patterns in the data. The plot of previous race times versus upcoming race times seemed of importance:
This scatterplot reveals several distinct linear groupings. Within each of these lines, there also appear to be smaller clusters of points. I set out to explicitly model these group structures so that the final model could better capture and leverage these underlying patterns.
Given the large number of features I engineered, it was impractical to inspect every pairwise combination for patterns. Instead, I applied clustering algorithms to automatically separate the observations into groups and then examined which features were most influential in distinguishing those clusters.
At first, I struggled to isolate only the linear patterns. Although the earlier clustering approaches were useful later in the workflow, my immediate goal was to identify the distinguishing features between the main linear groupings. Ultimately, I achieved this by using a modified k-means clustering procedure.
In brief, the method begins by randomly assigning data points to a fixed number of clusters and fitting a separate regression line within each cluster. Cluster quality is then evaluated using the residuals from these regressions, and points are iteratively reassigned to the cluster that yields the smallest residual until convergence. Based on the data, which exhibited between three and five distinct linear trends, I initialized the algorithm with the same number of clusters. This approach produced substantially clearer and more meaningful groupings.
The next step in my process was to examine how feature distributions varied across the clusters. I generated distribution plots for each engineered feature, comparing their behavior among the three groups. This revealed clear differences in the distributions of race distance and the distance differential between consecutive races.
When I colored the original plot of time1 versus time2 by distance differential, the clusters became much clearer. This visualization confirmed that the difference in distance between the previous and upcoming races was a key factor separating the groups.
I was able to further separate the elliptical groupings within the linear patterns by incorporating the distance2 feature. In combination, clustering models with 9–12 clusters, conditioned on both distance differential and distance2, produced the best separation and achieved silhouette scores greater than 0.8. With these robust clusters established, I then focused on developing predictive models to accurately assign new observations to the appropriate group based on their features.
One of the first feature families I explored was the racing comments. My goal was to classify dogs (identified by their birthdates) according to the sentiment of their race descriptions. I initially used the TextBlob and VADER libraries in Python to label each comment as positive, neutral, or negative and to generate sentiment scores. From these, I engineered additional features such as overall sentiment, rolling sentiment scores, and related aggregates for inclusion in the model. However, TextBlob struggled with domain‑specific language in greyhound racing, so I decided to develop a custom sentiment analysis approach. As a starting point, I examined the distribution of words appearing in the comments.
Building on this, I researched the contextual meaning of each word appearing in the comments and constructed a custom sentiment dictionary of 157 terms. Each word was labeled as positive, negative, or neutral with scores of 1, −1, or 0, respectively. I then tokenized the comments and applied this dictionary to assign sentiment scores, normalizing the total by summing the frequencies of each word weighted by its sentiment value. Using these scores, I engineered rolling and average sentiment features for each dog. In practice, a combination of the custom dictionary and VADER-based sentiment features performed best in feature selection, capturing more domain-specific nuances than the off-the-shelf models alone.
With the clusters identified and additional features engineered, I began the feature selection process. Although the clusters exhibited relatively clear linear patterns, I had created roughly 34 distinct features for each observation. Correlation matrices revealed strong relationships among many variables, and an initial linear regression showed variance inflation factors in the thousands for several feature pairs, indicating severe multicollinearity if all features were included.
To address this, I adopted ridge regression, which applies an L2 penalty to the model coefficients and shrinks them in proportion to their contribution. This penalization discourages the model from assigning excessively large coefficients to highly correlated features, allowing me to retain important predictors while mitigating multicollinearity. I tuned the regularization parameter to balance underfitting and overfitting and then fit separate ridge models within each cluster. This approach substantially reduced the effects of multicollinearity and improved the overall predictive performance.
Bringing all of the engineered features together, I built a final model. The workflow proceeded as follows:
Read in a dataset of race entries without the time2 variable and append it to the existing dataset with known time2 values.
Engineer additional features such as consistency metrics, age, sentiment scores, and related variables, and add them to the combined dataset.
Cluster the entries into distinct linear groups and elliptical subgroups using distance-based metrics.
Fit a ridge regression model within each cluster.
Use these cluster-specific models to predict time2 for races in the new dataset.
Using this approach, I was able to predict upcoming racing times with high accuracy, achieving a mean squared error below 0.18 seconds squared, corresponding to an average absolute error of about 0.42 seconds. Standard measures of predictive performance, including R2R^2R2 and AIC, indicated that the models were both effective and robust, and in all cases they outperformed simpler baselines such as basic linear regression and other conventional techniques.
For anyone interested in implementation details—such as code snippets, methodological choices, or additional evaluation metrics—I would be happy to share more information.