Showing posts with label stats. Show all posts
Showing posts with label stats. Show all posts

Saturday, February 16, 2013

The Effect of Overtime Games

Notre Dame just played an OT game against DePaul, just four days after a marathon 4 OT game against Louisville, and their 3rd OT in the last five games.  This prompted the question of whether playing OT games "tires out" teams and affects their performance in subsequent games.

There are various ways you might look at this question, but an easy one is to look at scoring averages for teams after they play an OT game and see if they differ significantly from the scoring average when they haven't just played an OT game.  Ignoring neutral court games, I looked at how many points home teams scored and gave up for both cases:

No OTAfter OT
Home Score70.370.2
Away Score65.366

As you can see, there's a mild effect, particularly on the defensive side of the ball, which totals to about 1 point difference in MOV.  There are about 1200 games in my database that meet the "home team's last game went to OT" criteria, so this isn't a huge sample.

For what it's worth, since 2009 there have been 11 Tournament games where the home team's previous game went to OT.  In those games, the home team averaged 69.5 points and gave up 65.5 points.

Thursday, January 24, 2013

Oddball Statistic

I've recently been modifying my data collection to include the score by halves, and in looking at the data, I discovered this oddball statistic: Only 11 NCAA teams average more points in the first half than in the second half:

Loyola (MD)  (#2 MAAC)
Charleston Sou. (#1 Big South, South)
New Mexico St. (#3 WAC)
Indiana (#2 Big Ten)
La Salle (#5 Atlantic Ten)
Oregon (#1 Pac-12)
Oregon St. (#12 Pac-12)
Holy Cross (#3 Patriot League)
Kansas (#1 Big-12)
Lehigh (#1 Patriot League)
Saint Louis (#10 Atlantic Ten)

With the notable exceptions of Oregon State and St. Louis, these are all good to excellent teams.  What's the connection?  Why do almost all teams score more in the second half than in the first half, and why would the few that don't be generally better than average?

Thursday, February 23, 2012

3PT Attempt Percentage

Ken Pomeroy recently made a couple of blog postings concerning defense, and specifically a statistic he calls the "3 Point Attempt Percentage" (3PA%).  He defines this statistic as the "percentage of field-goal attempts that are from three-point range."  Ken Pomeroy thinks this is a better measure of defense than 3PT%.  His reasoning is that most teams only take 3 point shots when they are relatively unguarded; the effect of defense is not to make these shots harder, but to cut down on the number of opportunities.  Hence the claim that it's really how many 3 pointers your opponent takes that reveals the quality of your 3PT defense.  Near the end of the second posting he says:
People that are unaware of 3PA% (which is to say nearly everyone) are missing a very telling statistic that explains a lot of how defense works.
This is a strong statement and worthy of a little research to see whether it is true (at least so far as predicting outcomes is concerned).

3PA% is similar to Effective Field Goal Percentage, one of Dean Oliver's Four Factors.  I have previously considered the Four Factors and concluded that they didn't add any predictive value to my models, but 3PA% captures a slightly different slice of information.

When I recently looked at derived statistics, one of the derived statistics was pretty close to 3PA%:

(Ave. number of 3PT attempts by the opposing team)
------------------------------------------------------
(Ave. number of FG attempts by the opposing team)
This isn't quite the same statistic, because it is using game averages rather than cumulatives, but it is close.  This statistic turned out to have no predictive value, but a couple of statistics based upon 3PT attempts did have value:

(Ave. number of 3PT attempts by the opposing team)
-----------------------------------------------------
               (Ave. number of turnovers)

(Ave. number of 3PT attempts by the opposing team)
-----------------------------------------------------
               (Ave. number of rebounds)
Note that these statistics are relating the number of 3PT attempts by the opponent to a statistic for the defending team.  I'm not entirely sure what these statistics are capturing, but I don't think it is 3PT defense.  (The latter might be indirectly saying something about how a team defends against the three pointer, from how it is positioned to rebound effectively or not after a taken three pointer.)

That aside, I modified my models to generate four new statistics:  the 3PA% for the home team in previous games, the 3PA% for the away team in previous games, the 3PA% for the home team's opponents in previous games, and the 3PA% for the away team's opponents in previous games.  I then tested the model both with and without these statistics:

  Model   Error      %Correct  
Base Statistical model  11.0672.7%
Base Statistical model + 3PA% statistics 11.0572.7%

There's a very small improvement in RMSE with the added 3PA% statistics.  So at least for my model, the 3PA% statistics don't seem to add any significant new information.

Friday, February 10, 2012

The Continued (Slow) Pursuit of Statistical Prediction (Part III)

As promised last time, we'll now look at a different type of derived statistic. We're going to look at statistics which are the ratio between the two teams of the same base statistic, e.g.,

(Ave # of offensive rebounds for the home team / Ave # of offensive rebounds for the away team)

The idea here is that it may be more predictive to look at the relative strengths of the teams rather than the absolute strengths. 

The first statistics I want to try this upon are the strength measures like TrueSkill and RPI. Suppose that Syracuse, with an RPI of 0.6823, plays Missouri, with an RPI of 0.6234, and the same night UCF, with an RPI of 0.5723 plays Oregon State with an RPI of 0.516.  Would we expect the same outcome in those games?  In both cases, the better team is about 0.06 better in RPI.  But Syracuse is about 10% better than Missouri, while UCF is about 12% better than OSU.  If it's the relative strength that matters, we would expect UCF to win (on average) by more than Syracuse.

To test this out, I generated the relative strengths for measures like TrueSkill and ran them through my testing setup.  In every case, the relative strengths had no predictive value above and beyond the value of the absolute strengths.  And when the relative strengths alone were used for prediction, they underperformed the absolutes used alone.

I then did the same thing for the statistical attributes like offensive rebounding and got the same result.  The relative strengths of the two teams provided no additional predictive accuracy.

I find this result fairly intriguing.  My strong intuition was that at least a portion of the game outcome would be better explained by the relative strengths of the two teams. It's hard to believe that Syracuse should win its game against Missouri by more points simply because they're both stronger teams than UCF and OSU.  But (as has often proven to be the case!) my intuition was just wrong, and relative strength is much less important than I would guess.

Friday, February 3, 2012

The Continued (Slow) Pursuit of Statistical Prediction (Part II)

Continuing on from last time, I had set up the infrastructure to allow me to easily test the value of derived variables in statistical prediction.  Before testing any of these derived variables, we need a baseline.  In this case, the baseline is the performance of a linear regression using all the base variables.  I don't know that I've ever documented the base variables, but they are basically all that can be created from the full game statistics available at Yahoo! Sports.  These are averaged by game, so for example one of the base statistics is "Average free throw attempts per game."  I also have the capability to average statistics by possession (e.g., "Average free throw attempts per possession" but unlike some other researchers, I've never found per possession averages to be any more useful than per game averages, so I generally don't produce them.

For most statistics, I also produce the average for the team's opponents.  So to continue the example above, I produce "Average free throws per game for this team's opponents."  I also produce a small number of simple derived statistics, such as "Average Margin of Victory (MOV)", and winning percentages at home and on the road.

When we get to predicting game outcomes, of course we have all of these statistics for both the home and the away team.  (And that home/road distinction is important, obviously.)  If we use all these base statistics to create a linear regression, we get the following performance:

  Predictor    % Correct    MOV Error  
Base Statistical Predictor72.3%11.10

This is the same performance I have reported earlier, and tracks fairly well with the best performance from the predictors based upon strength ratings.

Now we want to augment that predictor with derived statistics to see if they offer any performance improvement.  As mentioned last time, we have 1200 derived statistics, so we have to do some feature selection to thin that crop for testing. 

One possibility (as discussed here) is to build a decision tree, and use the features identified in the tree.  If we do that (and force the tree to be small), we identify these derived features as important:

  1. The home team's average margin of victory per possession over the overall winning percentage
  2. The away team's average number of field goals made by opponents over average score
  3. The home team's average assists by opponents over the field goals made
  4. The home teams average MOV per game over the home winning percentage
That is, you'd have to admit, quite a goulash of statistics.  I can probably come up with some rationale about some of those, but I won't bother.  All I really care about is whether they will improve my predictive accuracy.

To test that, I add those statistics to my base statistics and re-run the linear regression.  In this case, what I find is that while some of the derived statistics are identified as having high value by the linear regression, the overall performance does not improve.

There are other methods for feature selection, of course.  RapidMiner has an extension focused solely on feature extension.  This offers a variety of approaches, including selecting based on Maximum Relevance, Correlation-Based Feature Selection, and Recursive Conditional Correlation Weighting.  All of these methods identified "important" derived statistics, but none produced a set of features that out-performed the base set.

A final approach is a brute force approach called forward search.  In this approach, we start with the base set of statistics, add each of the derived statistics in turn, and test each combination.  If any of those combinations improve on the base set, we pick the best combination and repeat the process.  We continue this way until we can find no further improvement.

There are a couple of advantages to this approach.  First, there's no guessing about what features will be useful -- instead we're actually running a full test every time and determining whether a feature is useful or not.  Second, we're testing all combinations in our search space, so we know we'll find the best combination.  The caveat here is that we assume that improvement is monotonic with regards to adding features.  If the best feature set is "A, B, C" then we're assuming we can find that by adding A first (because it offers the most improvement at the first step), then B to that, and so on.  That isn't always true, but in this case it seems a reasonable assumption.

The big drawback of this approach is that it is very expensive.  We have to try lots of combinations of features, and we have to run a full test for each combination.  In this case, the forward search took about 54 hours to complete -- and since I had to run it several times because of errors or tweaks to the process in ended up taking about a solid week of computer time.

In the end, the forward search identified ten derived features, with this performance:

  Predictor    % Correct    MOV Error  
Base Statistical Predictor72.3%11.10
w/ Forward Search Features74.0%10.73

This is a fairly significant improvement.  The most important derived features in the resulting model were:
  1. The away team's opponent scoring average over the away team's winning percentage.
  2. The away team's offensive rebounding average over the away team's # of field goals attempted
  3. The away team's scoring average over the away team's winning percentage
  4. The away team's opponent treys attempted over the away team's rebounds
The ten statistics were actually evenly divided between home team statistics and away team statistics, but it turned out that the most significant five were all the away team statistics.

I'll leave it to the reader to contemplate the meaning of these statistics, but there are some interesting suggestions here.  The first and third statistics seem to be saying something about whether the away team is winning games through defense or offense.  The second and fourth statistics seem to be saying something about rebounding efficiency, and perhaps about whether the team is good at getting "long" rebounds.  (The statistics for the home team are completely different, by the way.)

Next time I'll begin looking at a different set of derived statistics.

Friday, January 20, 2012

The Continued (Slow) Pursuit of Statistical Prediction

When we last met on this topic, I was inspired by the Four Factors to look at derived statistics created from the ratio of two existing statistics, e.g.,


Offensive Balance = (# 3-Pt Attempts) / (# FG Attempts)

My previous work in this area has convinced me of the value of looking at all possibilities, no matter how non-intuitive, my approach was to look exhaustively at all the possible ratios between the ~35 base statistics.  That leads to some crazy statistics such as:

(Average # of fouls per possession by the home team's previous opponents) / 
    (Average offensive rebounds per game by the home team in previous games)

There turn out to be a number of difficulties with this approach.  (Perhaps not unsurprisingly, although crazy nonsensical statistics are not one of them.)

First, it's a lot of work just to generate the 1060 derived statistics.  (Only 1060 because I avoided inverse ratios, and avoided "cross-ratios" between the two teams.)  Initially I was generating a subset of these from within the Lisp code that pre-processes the game data.  That was painful to set up and slow to execute.  Eventually I discovered a way to generate the derived statistics within RapidMiner. I was able to drive this from a data file, so I wrote a small Lisp program to generate the data file that RapidMiner could use to construct all 1060 derived statistics.

Second, this amount of data tended to overwhelm my tools.  With the derived attributes, each game has about 1200 attributes total.  My training corpus has about 12K games.  The combination tended to break most of the data modeling features of RapidMiner, usually by overwhelming the memory capacity of Java.  Even when the software was capable of handling the data volume, operations like a linear regression might take hours (or days!) to complete, so testing and experimenting was laborious at best.

One way to reduce this problem is to thin the dataset, by testing on (say) a tenth of the full corpus.  But that introduced a new problem: overfitting.  If I used (say) a tenth of the data, I had about 1200 games in my test set -- just about the same number of test games as attributes.  The result of that is almost invariably a very specific model that does extremely well on the test data and very poorly on any other data.

Another approach is to thin the attributes.  This is the feature selection problem.  The idea is select the best (or at least some reasonably good) set of features for a model.  The stupid (but foolproof) way to do this is to try every possible combination of features, and select the best combination.  But of course that's infeasible in many cases (such as mine), so a variety of alternative approaches have been created.  RapidMiner has some built-in capabilities for feature selection, and there's a nice extension to add more feature selection capabilities here.

I experimented with a variety of different feature selection approaches.  I was hopeful that different approaches would show overlap and help identify derived attributes that were important, but for the most part that did not happen.  However, taking all the attributes recommended by any of the feature selection approaches did give me a more reasonable sized population of derived statistics to test.

More on this topic next time.

Wednesday, October 12, 2011

More on Statistical Prediction

I am continuing to explore statistical prediction.  In particular, after implementing the Four Factors as described here, I became interested in examining other statistics generated from the base set of statistics.  A subset of these generated statistics are ratios of the base statistics, like the "Offensive Balance" statistic I defined in my earlier post:
Offensive Balance = (# 3 Pt Attempts) / (# FG Attempts)
You can probably come up with a few sensible statistics like these off the top of your head.  But since I've seen time and again the value of exploring all options -- even the ones that make no "sense" -- I decided to calculate and test all of these sorts of ratios to see which of them (if any) have predictive value.

That's a more difficult job than you might imagine.  In my data sets there are 13 base statistics per team per game (FG Made, FG Attempted, 3PT Made, 3PT Attempted, FT Made, FT Attempted, Offensive Rebounds, Total Rebounds, Assists, Turnovers, Steals, Fouls, Score, and MOV).  For predictive purposes, we want to use the average of these over a team's previous games [1] and we can average by either game or possession - so that's 26 base statistics per team.  There are 26*25 = 650 possible ratios of those statistics.  But we also want to consider ratios not only of a team with itself but also of the team with its opponent, e.g., the ratio of the team's average number of 3 PT attempts in past games to it's opponents average number of 3 PT attempts in past games.  That adds another 676 possible ratios.  Finally, we also want to consider the statistics for a team's past opponents, e.g., the average number of 3 PT attempts in past games of a team's opponents in those games.  Adding those in creates a lot more ratios.  Multiply all that by the 12K games in my training data, and it's a lot of data.

My approach is to generate a subset of the possible ratios and test them for predictive value.  For various reasons I settled on generating all the ratios with a particular numerator, e.g.,
(FG Made) / (# Fouls)
(FG Made) / (Opponent's # Fouls)
(FG Made) / (# Fouls by Opponents in Past Games)
etc.
This ends up adding about 96 new statistics to every game in the database.  I can then take this expanded data and pump it through the usual linear regressions, etc., to find the statistics that have predictive value.  But this is a slow process -- for each numerator, it takes hours to generate all the statistics and run them through iterations of the predictive model.  (This has the disadvantage that I may miss some combination of generated statistics with different numerators that are only valuable in combination.)

So far, I haven't identified any ratios that result in significantly better predictions.  But I have been surprised that (at least so far) the models have selected a number of unexpected ratios as being of value.  For example:
(Away team's Average FG Made) / (Away team's Average 3PTs Attempted)
(Away team's Average FG Made) / (Away team's Average 3PTs Made)
These ratios seem to be capturing something about the Away team's offensive balance between inside and outside play.  Interestingly, both the ratio with 3 PTs Attempted and 3 PTs Made are significant -- it may be that the first captures the "offensive strategy" (whether a team plays outside first or inside first) and the second captures something about how effective they are at executing that strategy.  It's also interesting that these ratios are only significant for the Away team -- apparently the home team's performance doesn't depend strongly on what sort of offensive strategy it uses.

Another interesting statistic:
(Home team's Average FG Made) / (Home team's Past Opponents' Average Offensive Rebounds)
It takes a moment's thought to grasp this statistic.  It compares the average number of FGs made by a team to the offensive rebounding of the opponents the team faced.  If we take Offensive Rebounds as an indicator of how strongly teams are contesting inside play, then this ratio would seem to say something about how effective the home team's inside play has been relative to its opponents.

Hopefully working through all the ratio statistics will turn up a set of statistics that provide significantly better predictive value.

[1] Averaging isn't the only option here, and there are other possibilities for generated statistics that might be useful, but I feel that ratios are a reasonably fertile area for exploration.

Wednesday, September 21, 2011

Statistical Prediction: Pace-Adjusted Statistics & The Four Factors

There is much talk in sports statistics circles about pace-adjusted statistics.  As Wikipedia puts it:
A key tenet for many modern basketball analysts is that basketball is best evaluated at the level of possessions.
The notion here is that because teams play at different paces, game-level statistics can be misleading.  A team that averages 95 points per game is not necessarily better than one that averages 78 points per game.  The higher-scoring team may simply be playing at a much faster pace.  We can account for this by measuring statistics per possession rather than per game.

While this makes a lot of intuitive sense, I always like to test my intuitions.  So I took the same set of statistics used in this posting and re-calculated them as per-possession statistics.  (See here for how to estimate the number of possessions in a game.)  Then I ran the prediction model using the per-possession statistics.  (Obviously some statistics, like "Field Goal Shooting Percentage" are not calculated on a per-game basis, so those don't get pace-adjusted.)  Here is the performance comparison:

  Predictor    % Correct    MOV Error  
Govan + Averaging73.5%10.80
Statistical prediction (per-game stats)72.2%11.09
Statistical prediction (per-possession stats) 72.2%11.10

As you can see, the two approaches were indistinguishable.  Not only was performance nearly identical, but they both selected the same statistics for the prediction model.  So at least for this case, it doesn't appear that adjusting for pace improves performance.

My guess is that the relative unimportance of pace is due to the shot clock and the copycat nature of coaching.  There probably isn't enough pace variation across teams to make it a significant factor.

If you search around for "pace-adjusted statistics" you'll eventually stumble across Ken Pomeroy's Four Factors page.  The four factors are derived statistics that are intended to give additional insight into how teams play.  The factors are:
  • Effective field goal percentage
  • Turnover percentage
  • Offensive rebounding percentage
  • Free throw rate   
(Definitions can be found on Ken Pomeroy's page.)

"Effective FG%" is not of interest to me because the linear regression can adjust the relative importance of field goals versus three-point attempts.  "Turnover %" is turnovers per possession; that's one of the statistics I calculated as part of the per-possession statistics experiment above.  (It had no value in the predictor, fwiw.)  "Offensive rebounding %" is a more interesting statistics, and since offensive rebounds are used by the statistical prediction model, this seems like a worthwhile statistics to investigate.  "Free throw rate" seems to capture some notion about how often a team draws a foul.  I think that's already captured, but it isn't difficult to generate this statistic.

If I generate these two new statistics and run the prediction model, I find that performance remains the same, but the "Offensive rebounding %" statistics replace the per-game or per-possession offensive rebounding statistics.  ("Free throw rate" has no predictive value and is eliminated in the linear regression.)

Since three point shooting percentages are used in the predictor, I decided to define a new statistic to capture how much a team relies on the three-point shot (and impacts its opponents use of the three-point shot).  I defined this as:
Offensive Balance = (# 3 Pt Attempts) / (# FG Attempts)
and re-ran the predictor.  The new statistic has no predictive value.  An alternative formulation is to look at the made 3 pointers versus the made field goals:

Offensive Balance = 3*(# 3 Pt Made) / 2*(# FG Attempts)
but again, this statistic has no predictive value.

I'm open to suggestions if anyone out there has any thoughts on similar "derived statistics" that might be of value in prediction.

Friday, September 16, 2011

Statistical Prediction

With this post, I'm going to start taking a look at predicting game outcomes based upon team-level statistical measures other than won-loss or MOV, i.e., measures like "team scoring average," "average number of offensive rebounds per game," etc.

There are a number of ways to slice & dice these statistics, but the most straightforward approach is to use season-to-date averages.  So, when I'm trying to predict the Illinois-Purdue game on 2/15, I'll be looking at the statistics for those two teams averaged over all the games for that season before 2/15.  And I also want to include average statistics for a team's opponents.  So I want to know both Purdue's scoring average for all of its previous games, and also the scoring average of its opponents in those games.  For every game, I'll typically have four values for a statistic: the home team's average, the home team's opponents' average, the away team's average, and the away team's opponents' average.

To begin with, let's look at how well we can predict games using the most obvious statistic: the scoring average.  Using just the (four) scoring average statistics, and the usual methodology, here's our performance:

  Predictor    % Correct    MOV Error  
Govan + Averaging73.5%10.80
Scoring averages72.1%11.18

That's pretty encouraging.   Just using the scoring averages delivers performance comparable with some of our better W-L and MOV-based predictors.  The bad news is that this is still highly correlated with our best other predictors (around 96%), meaning that it probably can't be used in an ensemble to improve our overall predictive performance.

If we look at adding other statistics we find (as would be expected from the literature) that they offer little improvement.  The best combination I could find (in order of importance) was (1) scoring, (2) 3 pt percentage, and (3) opponent's average offensive rebounding:

  Predictor    % Correct    MOV Error  
Govan + Averaging73.5%10.80
Scoring averages72.1%11.18
Scoring + 3 pt % + Opponent's off rebounding 72.2%11.09

As you can see, the improvement was not huge.  The inclusion of "average number of offensive rebounds by opponents" is interesting because it is not scoring-related.  That statistic would seem to capture some aspect of a team's defensive performance -- a team that gives up a lot of offensive rebounds to its opponents is probably doing something wrong at the defensive end of the court.  That suggests that we might want to think about a better measure of defensive performance -- for example, we might want to look at offensive rebounding percentage rather than just the raw total.

Monday, August 15, 2011

More on Possessions

In a previous post, I mentioned that I was working on calculating the number of possessions in a game.  There's no direct stat for this, so it has to be approximated by formula from the existing stats.  The number of possessions in a college game averages about 67.  (In the previous post I reported the average as 68, but that number was skewed by overtime games.  I've since fixed that problem by normalizing to 40 minutes rather than per game.)

The next step is to try to predict the number of possessions when two teams play each other.  Why would we want to do this?  Consider the situation where Maryland beats Duke by 5 points.  How strong is that evidence that Maryland is better than Duke?  Well, if the final score was 45-40 we might consider that stronger evidence than if the score was 125-120.  Number of possessions is also used to calculate "tempo-free statistics" which allow us to better make apples-to-apples comparisons between games that are played at different paces.

So how do we predict the number of possessions?  The model I'm using right now supposes that each team has a preferred pace -- i.e., an ideal number of possessions.  Some teams would like to play games with lots of running up and down the court and many possessions; others would like to play a very slow, controlled game.  When two teams meet up, they each try to play the game at their preferred pace, and as a result the game is played somewhere in-between:
Posspred = (Preferred PossHome + Preferred PossAway)/2

Of course, we don't know the "preferred pace" for a team, so we have to try to discover that from the game data.  One way to do that is gradient descent, as was used for Danny Tarlow's PMM.  If we do that, and then test the predictor in the same way we've tested the others, we get this performance:

  Predictor    Error  
Possessions5.20

Is that good performance?  If we look at the distribution of possessions/game:


I'm inclined to say the predictor is "meh" -- not particularly good, but probably good enough to be useful.