Adapting Karun Singh's xT Model to Professional Lacrosse
I have, like many others I would assume, spent a lot of this summer watching the FIFA Men's World Cup. Some games were on in the background of my room as I worked, others a few feet in front of my face as I watched the entire 90+ minute match. Maybe my X algorithm was sharp then in recently transitioning a lot of my feed to (soccer) news and thankfully, a lot more visualization using soccer (I will continue to refer to the "Beautiful Game", as soccer for the entirity of this project, sorry Karun) data.
I recently saw a post from Karun Singh, a research engineer for Arsenal, that highlighted a job opportunity with the team. The post -- to my initial surprise has over 2.8 million views. But the more I look into this role, and Karun's work, I am not too shocked. The ability to use tracking data and implement AI models for sport research has been taking off over the past few years and it is truly interesting work.
Since I am not a huge soccer guy (sorry Mom), and as I have been recently looking into the Premier Lacrosse League (not to get confused with the EPL in this setting), I decided to look at one of Karun's ideas in a lacrosse context. A few years ago, he created what he calls "xT" or "Expected Threat", which basically values every position on the field by how likely a possession starting there is to eventually produce a goal, learned entirely from what actually happens next — whether the ball is passed, carried, or shot from that spot. In a lacrosse setting ,it lets you credit the pass or dodge that sets up a chance, not just the shot itself, by working backward from the goal through however many actions typically come before it.
It is a really cool idea and one that I think has some very interesting overlap in a game like lacrosse. Since lacrosse is still in its "infancy" as a professional league, and thus lacks robust technology that would make something like this very concrete, I decided to engineer a way to bring xT to lacrosse. I think the xT premise has incredible carry over in lacrosse, and one day I hope that lacrosse sees tracking data in their workflows. I believe that the adaptation and insights gained are boundless.
Lacrosse Background
For those that are not as directly aware of how lacrosse is played, here is a brief overview and why I think a model like this has so much benefit. Firstly, lacrosse, a lot like soccer is very fluid. Space is constantly being contested, and off-ball play has a lot more value than many other sports. Lacrosse is also highly possessional/territorial in nature and the team that dominates possession in a game is likely to see similar results on the scoreboard. The dynamic of the game is something that draws people in, but it is something that is incredibly difficult to quantify when determining individual value. The eye test will tell you which players are the "Alphas", and this year it has been C.J. Kirst's world. Yet, I am more interested in further exploring how players play together and think the frameworks established in data science from soccer have good carryover.
The catch: soccer has years of large-scale public tracking and event data to build xT on other models that track player and ball movement. Lacrosse has essentially none. So this is what building toward it looks like when you have to generate that data yourself, one game at a time, by hand.
I want to be upfront that this is a small, rough first pass, and nowhere near a finished model. I'm sharing it because the process itself — and where it genuinely breaks down against lacrosse's structure — seems to show some of the promise that exists in the game going forward.
WHAT I DID
"The zone system — coarser than xT's original grid, built for hand-charting in real time"
Since I don't have player-tracking data for lacrosse, I split the attacking half of the field into named zones (Crease, Wing Box Side, X, Up Top Far Side, and so on — not arbitrary grid cells) and charted an entire game from the Archers and Whipsnakes this past weekend in Denver, by narrating every meaningful touch: which zone the ball started in, what happened (a pass, a carry, a shot, a turnover), and where it ended up. I built a small tool to speed this up, and eventually a voice-dictation workflow so I could narrate a live game and transcribe it after.
From that charted data, I built the same kind of value model xT uses: a zone's value is a function of how often shots from it score, plus the value of wherever it tends to lead — solved iteratively, the same way the original model propagates value backward from the goal.
"Zone values from the full game — color indicates how many shots the value is actually based on"
What the Visualizations Actually Showed:
Convergence By Zone
For this convergence plot, each line is one (arbitrary by my determination) zone on the field. The x-axis is how many actions of "lookahead" the model was given — iteration 1 only knows about an immediate shot, iteration 20 accounts for long chains of passes before a shot happens. As you move right, a zone's true value gets revealed. Crease (red) shoots up almost instantly and flattens — meaning it's valuable because of direct shot-taking, which makes sense to those who watch games. The other zones climb slowly and together, meaning the model needed to "look ahead" much further to realize they're valuable — not because shots go in from there, but because the ball usually ends up somewhere dangerous after passing through them.
Delta Through Iterations
This heatmap shows how much each zone's value changed between two specific lookahead depths — iteration 2 versus iteration 6. A small number means the model already had that zone figured out early (like we saw with the crease in the line plot above); a big number means it took longer for the zone's real value to show up. The wing and GLE zones changed two to three times as much — the model only realized how valuable they are once it was allowed to think several steps further ahead, which tells us their value increases as play moves closer towards those zones (defenses becoming more condensed and attackers getting better positioning as the possession moves along).
Ball Movement Transitions Between Zones
This is another adaptation from soccer that I think has massive potential carryover to lacrosse. This diagram shows where the ball actually travels (not the ball tracking in space, but from Point A to Point B). Every arc is a real pass or carry that happened in the charted games; thicker and brighter lines mean that specific move happened more often. Look at Crease: even though it's the single most valuable zone on the field (per the earlier chart), it barely has any thick lines connecting to it — the ball rarely moves "through" the crease. Instead, all the thick, busy traffic is out at "Up Top", Wing, and GLE, passing or carrying among themselves. Players building possession with frequent "alley" dodges or dodges out on the wings, find the windows of opening in the crease for quick shots. With a more precise tracking system, I think this tool can be incredibly valuable as teams see the flow of their possessions that have the most positive outcome.
My Many! Limitations in the Model:
I ran into a ton of issues as one could imagine of trying to "hand-track" all of the events in a game. I began by trying to do it in a Excel sheet with player names and events, yet I could not keep up with the speed of the game. I decided to narrate the game through positions of the ball on the field and use a Python function to translate my voice memos. Here is a snippet of one of my transcriptions:
"new possession archers possession wing far side carry top center pass up top box side pass up top box side pass wing box side carry goal line extended box side pass crease shot save whipsnakes ball".
I obviously run into the issue of sample size. One game, even a full one, is a small dataset by the standards xT was built on — the original model drew on years of tracking data across entire leagues of professional soccer. Several zones in my data still have single-digit shot counts. Any specific number in this model should be treated as *very* provisional.
I arbitrarily made zones on the field, rather than using x/y coordinates that other leagues are able to pull. My model buckets the field into about eleven named regions. The real xT operates on a much finer grid (or, more recently, on literal continuous player coordinates). A shot two yards inside the arc and a shot two yards outside it may land in the same zone in my data, even though they're meaningfully different shots. This is a real loss of resolution and something that will prove to be invaluable in future iterations. My denotion of box side and far side is also majorly flawed as teams flip attacking ends. Left-handed players for example would be on the box side when a team is attacking "from right to left" and then on far side when that team is going "left to right". That degrades the insights gathered on the team level for tendencies on where attacks begin from.
I found that hand-charting introduces human error and subjectivity. I charted this myself, narrating live and reconciling messy speech-to-text output afterward. There are specific plays in my underlying data where I had to guess at an ambiguous zone, assume a same-zone fallback for an unstated pass destination, or interpret a garbled word. I've kept notes on these judgment calls, but they exist, and a different charter watching the same game might draw slightly different boundaries. I also fell victim to fatigue and inconsistency in passes and dodges within the same zones.
This is a team-level, not player-level outlook. My data says "Archers had the ball in the Crease area," not which specific Archers player. Real per-player attribution needs jersey-number recognition or reliable player tracking, neither of which I have working yet. Everything here describes team tendencies, not individual value.
The 2-point arc isn't modeled. Arguably professional lacrosse's most rule-distinctive feature — shots from beyond a certain distance are worth two points instead of one — isn't reflected in my value calculation at all right now. A zone's value here only reflects whether a shot went in, not what it was worth if it did.
Man-up and man-down situations aren't separated out. A zone's value should look very different on a man-up possession than at even strength, and lacrosse has penalties and extra-man situations far more often than soccer has anything comparable. My current model pools all of this together.
"Restarts" reset the game constantly, unlike soccer's continuous flow. Lacrosse stops for a faceoff after every single goal and end of quarter. Soccer's xT seems to be built for a sport where play flows continuously for 45-minute halves. I don't yet know how much this structural difference should change the model's assumptions, only that it's a real difference I haven't fully reasoned through, especially with how shots out of bounds are inserted where the team "chases the ball." This means that a missed shot from the "box side wing" can get to the "far side behind" area nearly instantly with play resuming from the opposite corner of the field.
Only the attacking half is modeled at all. My zones don't cover the clearing game, the ride, or general defensive-half play. Real value is unquestionably created and destroyed in that part of the field — a bad clear can end a possession before it ever reaches my zones — and none of it is currently captured.
This is one specific matchup. Every number in this write-up describes how two specific teams played against each other in one specific game. None of it should be read as a general lacrosse truth yet. That requires charting many more games, across many more teams.
No opponent or context adjustment. The model doesn't yet know if a defense was elite or average, whether a team was playing from ahead or behind, or how much time was left. All of that shapes real decision-making and isn't in here.
There are likely many more that I have not even realized yet, but that seems like a good start.
Where to From Here?
I have a bunch of iterations off this but the realistic next step is volume: charting more games until zone values stop moving much from game to game. Hand charting games takes a lot of time, so the PLL investing in some tracking technology would be a gamechanger (probably on multiple levels). Real player-tracking coordinates from broadcast video or systems — even through a better computer-vision pipeline would be much better to feed a model like this.
If nothing else, I hope this is a useful data point on how far the core idea behind xT travels beyond the sport it was built for and where I see the future of lacrosse analytics being able to reach if the investment from the league comes in time. I see many parallels of this game with soccer on the analytics side and am excited about what is able to be deduced from both going forward.
Thanks for reading! Would love to hear your thoughts -- email me @ djduffy4@gmail.com or DM me on X @DigestDuffy !