Showing posts with label model performance. Show all posts
Showing posts with label model performance. Show all posts

Thursday, February 23, 2012

3PT Attempt Percentage

Ken Pomeroy recently made a couple of blog postings concerning defense, and specifically a statistic he calls the "3 Point Attempt Percentage" (3PA%).  He defines this statistic as the "percentage of field-goal attempts that are from three-point range."  Ken Pomeroy thinks this is a better measure of defense than 3PT%.  His reasoning is that most teams only take 3 point shots when they are relatively unguarded; the effect of defense is not to make these shots harder, but to cut down on the number of opportunities.  Hence the claim that it's really how many 3 pointers your opponent takes that reveals the quality of your 3PT defense.  Near the end of the second posting he says:
People that are unaware of 3PA% (which is to say nearly everyone) are missing a very telling statistic that explains a lot of how defense works.
This is a strong statement and worthy of a little research to see whether it is true (at least so far as predicting outcomes is concerned).

3PA% is similar to Effective Field Goal Percentage, one of Dean Oliver's Four Factors.  I have previously considered the Four Factors and concluded that they didn't add any predictive value to my models, but 3PA% captures a slightly different slice of information.

When I recently looked at derived statistics, one of the derived statistics was pretty close to 3PA%:

(Ave. number of 3PT attempts by the opposing team)
------------------------------------------------------
(Ave. number of FG attempts by the opposing team)
This isn't quite the same statistic, because it is using game averages rather than cumulatives, but it is close.  This statistic turned out to have no predictive value, but a couple of statistics based upon 3PT attempts did have value:

(Ave. number of 3PT attempts by the opposing team)
-----------------------------------------------------
               (Ave. number of turnovers)

(Ave. number of 3PT attempts by the opposing team)
-----------------------------------------------------
               (Ave. number of rebounds)
Note that these statistics are relating the number of 3PT attempts by the opponent to a statistic for the defending team.  I'm not entirely sure what these statistics are capturing, but I don't think it is 3PT defense.  (The latter might be indirectly saying something about how a team defends against the three pointer, from how it is positioned to rebound effectively or not after a taken three pointer.)

That aside, I modified my models to generate four new statistics:  the 3PA% for the home team in previous games, the 3PA% for the away team in previous games, the 3PA% for the home team's opponents in previous games, and the 3PA% for the away team's opponents in previous games.  I then tested the model both with and without these statistics:

  Model   Error      %Correct  
Base Statistical model  11.0672.7%
Base Statistical model + 3PA% statistics 11.0572.7%

There's a very small improvement in RMSE with the added 3PA% statistics.  So at least for my model, the 3PA% statistics don't seem to add any significant new information.

Friday, February 17, 2012

Performance Versus "The Line"

As I've mentioned earlier, I use my models to bet (in some theoretical sense) against the "the line".  Typically I bet the games where my model differs significantly from the line (e.g., >4 points or so).  As I've documented here, I have a number of different models, all of which have around the same performance (~11 points RMSE).

In the past I've usually averaged the predictions of these models for betting purposes, but for some time I've wondered whether they all perform equally well against the line.  Although they all have similar errors, it's possible that some of the models error more consistently to the winning side of the line.  To test this, I gathered three seasons worth of Vegas closing line data (about 7700 games) and tested each model for how often its predictions were correct versus the line.  (The predictor is "correct" if it would make a winning bet given the line.)  I also looked at each predictor's error versus the line (i.e., how accurately it predicted the line).

  ModelPerformance
vs. Line
Error vs. Line 
TrueSkill49.89%3.75
Govan49.28%3.49
BGD49.58%3.51
Base Statistical50.12%4.34
Statistical w/ Derived 50.15%4.34
All 52.00%3.49
All (Difference > 2)53.15%

The "All" model here is a linear predictor using all the inputs to TrueSkill, Govan, BGD and Statistics w/ Derived.  (I also tested some voting models, but they all under-perform the Statistical/All models.)

There are a couple of interesting results.

Most noticeably, the "All" predictor is at break-even versus the line.  (Due to "house cut" on sports bets, you need to win about 52% of your bets to break even.)  If we restrict ourselves to bets where the predictor differs from the line by at least two points, performance moves into (barely) positive territory.  This is very good performance; the best predictors tracked at The Prediction Tracker do not even break 50%.  (Furthermore, I am using the "closing" line, which is a tougher measure [by about one point] than the opening line used at the Prediction Tracker.)

It's also intriguing that TrueSkill/Govan/BGD all underperform the line but track it noticeably better than the statistical predictor.  This suggests to me that the line is set not by wily veteran gamblers in the smoky back rooms, but by a computer program using some kind of team strength measure.

A (possibly interesting) side-note:  All models that under-perform the line are going to fall into the seemingly miniscule range of 48-52%.  (If a model performs worse than 48% against the line, we would simply bet against the model.)  Pick any crazy model you like -- "Always bet the home team," "Always bet on the team whose trainer's name is first alphabetically," etc. -- and the performance is almost certainly going to fall in that 48-52% range against the line.  (If it doesn't, you've found the key to beating Vegas!)

Tuesday, January 17, 2012

Basketball Season Underway

I spent the last few days scraping game data, dusting off code and generally getting the basketball predictor back online.  The current version of the predictor uses an average of 4 linear regressions.  These models are based upon: (1) the Govan rating, (2) the TrueSkill rating, (3) a Batch Gradient Descent (BGD) rating, and (4) a rating based on a wide variety of statistical measures (such as "offensive rebounds per possession").   Individually, each of these models has a RMSE of less than 11 on my test corpus.   Unfortunately, they're all highly correlated, so the combined model doesn't do any better than the best of the underlying models.  Currently it has an RMSE of 10.79 on my test corpus.

During the season I compare the model predictions against the line and "bet" games where the prediction differs significantly from the line.  "Significantly" is a relative term.  When I first started doing this, my model often differed from the line by 10 points or more.  As the model has improved, those differences have narrowed considerably.  (As would be expected.  The line is usually the best predictor.)  In my testing so far this year, I've only seen a difference of more than 5 points once.  There is some good mathematical work on sizing wagers based upon bankroll, perceived advantage, etc., but I've gone to a simple approach of betting $10 with an advantage of < 5 points and $20 with an advantage of >5 points.  (Adopted after the 1/14 games shown below.)

Here are the games the model has "bet" so far (no real money was harmed):

Date Home            Score Away                   Score MOV Line Pred Adv Risk Win Result Won v.Line
1/14 Tennessee St. 52 SIU Edwardsville 49 3 16 8.8 -7.2 20 17.39 17.39 1 1
1/14 LA Lafayette 87 Florida Intl. 81 6 10 5.1 -4.9 20 19.05 19.05 1 1
1/14 Murray St. 81 Tennessee Tech 73 8 12 16.5 4.5 20 18.18 -20 1 0
1/14 Houston 55 Memphis 89 -34 -8.5 -4.2 4.3 20 17.39 -20 1 0
1/15 Ohio St. 80 Indiana 63 17 13.5 9.1 -4.4 10 9.09 -10 1 0
1/15 Bradley 78 Northern Iowa 67 11 -10 -7.2 2.8 10 8.70 8.70 0 1
1/15 USC 47 UCLA 66 -19 2 1.5 -0.5 10 9.09 9.09 0 1
1/16 Syracuse 71 Pittsburgh 63 8 13.5 17.3 3.8 10 9.09 -10 1 0

So far this season the model is 50% against the line (and subsequently down about $5) and 75% picking the correct outcome.  The (evolving) model picked 38 games last year, and over the two seasons so far is at a 63% win percentage and 60% versus the line (+$133).  Both are probably short-term aberrations -- the model has a 74% win percentage when tested against my corpus of 12K games.

I won't generally be posting predictions, but I will try to summarize the model's performance a few times during the season, as I'm sure it makes for interesting reading :-).