Showing posts with label mov-based. Show all posts
Showing posts with label mov-based. Show all posts

Tuesday, November 8, 2011

One Bad (Good) Game

As mentioned in my previous posting, I recently looked at the effect of dropping a football team's best game (highest Margin of Victory) and their worst game (lowest MOV).  The intuitive notion is that everybody has bad days, where everything goes wrong, and good days, where everything goes right, and maybe those days don't tell us anything useful about the real strength of a team.  If that's so, then dropping those games might give us ratings that are more accurate.

To test this hypothesis I implemented this "drop the worst score" grading system for a couple of the rating systems I use for football and measured performance in the usual way.  Here are the results for one of the rating systems:

  Predictor    % Correct    MOV Error  
BGD Baseline73.7%16.52
BGD w/o blowouts or lowouts 72.6%16.77
BGD w/o lowouts72.9%16.69
BGD w/o blowouts73.6%16.62

Here I'm using the whimsical "lowout" to indicate the worst loss for a team.

As this shows, eliminating the blowouts/lowouts hurts predictive performance.  For what it's worth, the losses seem to be more important than the wins.  (I saw the same effect in basketball when I looked at this last year.)

Friday, November 4, 2011

The Impact of MOV Cutoffs in Football Ratings

I was prompted to start my football predictions by a discussion on an email list of the value of MOV cutoffs in rating systems.  Roger Dendy believed that capping the MOV in blowout victories improved his rating system.  My testing of MOV cutoffs in basketball has shown just the opposite -- that no matter how big the blowout, there's always information in the margin of victory.  Capping MOV at any level (in both blowouts and nailbiters) always reduces the prediction value of a rating.

Of course, just because that's true in basketball doesn't mean it's true in football.  I was pretty sure it was true, but I believe in "trust but verify."  So I put together the football predictor and tested a couple of different rating systems both with and without MOV caps.

I have many rating systems that use MOV, so I picked one and measured it's performance with a 100-fold X-validation  across my archive of college football scores from 2005 to date.  It had a RMS of 16.78 and predicted 71% of the games correctly.

Then I experimented with adding a cutoff to the MOV.  I set the cutoff to 32 points, so that all the games where the MOV exceeded 32, it would be treated as 32.  I just picked 32 arbitrarily as a good figure for a blowout win.  The performance degraded to RMS=17.19 and 69%.  I then bumped up the cutoff to 48 points, and the performance was RMS=17.01 and 70%.

The other rating system showed a similar pattern of performance.

What this shows -- at least for the two rating systems I tested and these performance metrics -- is that even huge margins of victory have value in assessing future performance.  People argue intuitively that there's "no difference between winning by 48 and winning by 52" but that appears not to be true.

Recently I got to wondering if it might not make more sense to drop a blowout victory entirely.  This would be like "drop your lowest score" grading in high school.  The intuitive notion here is that sometimes teams just have a bad day -- a few unlucky bounces and worse goes to worse.  Or lucky bounces and better goes to better, from the other side of the coin.  More on that notion next time.

Wednesday, August 10, 2011

More PMM

In my previous post, I replaced the prediction model from Danny Tarlow's PMM:
Predi = Offensei * Defensej
with this model:
Predi = Offensei + Defensej
with predictably terrible results.   There are other models we could try that would likely be more reasonable, but I want to detour a bit into a two-stage model.

The basic idea is that we predict the game outcome using the original Offense*Defense model, and minimize the error in the prediction across all the games using gradient descent.  However, we then add a second stage, where we attempt to predict the residual error between our best Offense*Defense model and the actual scores.  The value we're going to try to predict is:
Residuali = Scorei  - (Offensei * Defensej)
Now there wouldn't be much sense in trying to predict this Residual by the same sort of Offense*Defense model -- if that would work, it would presumably be captured in our original model.  So we need to pick some different sort of model for the Residual, and in this case we'll use the additive model we used before, except applied this time to predict the Residual:
Pred Residuali = Ri + Sj
and we'll determine R and S by the same sort of gradient descent we use for Offense and Defense.  Our final predicted score will be:

Predi = Offensei * Defensej + Ri + Sj

Here's how that performs:

  Predictor    % Correct    MOV Error  
PMM71.7%11.23
PMM (w Residual Prediction) 72.0%11.19

The improvement isn't huge, but it does show some promise.  Intuitively, if we think of the performance of a team as a sum of a number of different factors plus some noise, then different models may be capable of accurately modeling different factors.  Some factors may be well-modeled by "Offense*Defense" while others are better modeled by "Offense+Defense".

There are several avenues to explore from here.  One is to look at other alternate models, for both the primary model and the residual model.  Another is to look at combining the two models, so we optimize both at once -- this would have some advantages if the two models are interdependent.  Another interesting notion is to use an entirely different approach -- say, TrueSkill -- for either the primary or the residual model.

Tuesday, August 9, 2011

An Alternate PMM Model

As mentioned in the last post, I'm looking at the code for implementing Danny Tarlow's PMM.  In Danny's (1 dimensional) model, each team has an Offense rating and a Defense rating and when Teami plays Teamj, the predicted score is:
Predi = Offensei * Defensej
In other words, we have a model where teams have a certain scoring potential (Offense) and the effect of defense is to reduce that potential proportionately.  So if a team has a Defense of 0.85, and plays a team with an Offense of 100, we expect the opponent to score 85 points -- 15 points below its potential.  On the other hand, if we play a team with an Offense of 60, we expect them to score 51 points -- 9 points below its potential.

That seems reasonable, but it's not the only possible model for the relationship between Offense and Defense.  For example, we might model defense as reducing the opposing team's offense by a fixed amount, e.g., Maryland holds opposing teams to 4 points less than they would score against an "average" defender.  With that model, our predicted score would look like this:
Predi = Offensei + Defensej
Which model is more reasonable?  My intuition says the first model, and it's certainly widely used, but I don't know of any work that has tried to determine the best model of the interaction between offense and defense in basketball.  (Although looking around did lead me to this interesting blog.)  Please correct my ignorance if you're aware of something relevant.  Easy enough, though, to test this model:

  Predictor    % Correct    MOV Error  
PMM71.7%11.23
PMM (alternate model) 64.5%13.65

So, not very good.  No real surprise, but it does raise the question of whether some different model for the interaction between Offense and Defense could out-perform the Offense*Defense model.

Monday, August 8, 2011

Regularization in PMM

I'm back from vacation and slowly finding some time to work on prediction.  One of the things I'm doing is revisiting the code for "Probabilistic Matrix Model" (PMM).  This model is based upon the code Danny Tarlow released for his tournament predictor, which he discusses here. At the heart of this code is an algorithm to minimize the error between the predicted scores and the actual scores using batch gradient descent.  This is something I'll want to do frequently in the next stage of development (e.g., to predict the number of possessions in a game), so I'm looking at adapting Danny's code.  (Or rather, adapting my adaptation of Danny's code :-)

Danny's code differs in a couple of ways from a straightforward batch gradient descent.  One difference is that Danny has added in a regularization step.  Danny made this comment about regularization:
In addition, I regularize the latent vectors by adding independent zero-mean Gaussian priors (or equivalently, a linear penalty on the squared L2 norm of the latent vectors). This is known to improve these matrix-factorization-like models by encouraging them to be simpler, and less willing to pick up on spurious characteristics of the data.
I theorize that with a large, diverse training set such as I'm using, regularization is unnecessary.  To test that, I re-ran the PMM without any regularization:

  Predictor    % Correct    MOV Error  
PMM71.7%11.23
PMM (w/o regularization) 71.8%11.20

Performance is almost identical, so indeed there doesn't seem to be any value in regularization.

Friday, May 27, 2011

The Offense-Defense Model & Probabilistic Matrix Model (PMM)

The first two MOV-based models will look at are similar:  they both calculate an "Offense" and a "Defense" for each team.  The "Offense" number represents a teams offensive capability and the "Defense" the defensive capability.  When Duke plays UNC, Duke's predicted score is Duke's "Offense" times UNC's "Defense."  Generally speaking, these numbers are calculated by initializing all teams to some baseline numbers (e.g., so that the expected score OxD = 65) and then iteratively adjusting the values so that they more closely match actual game outcomes.  If Arizona State University consistently scores few points, it's "Offense" value will drop (and at the same time it's opponents' "Defense" value will also drop).  After some number of iterations adjusting the numbers, the total error (across all games) will be minimized.

The Offense-Defense model is a version of this I first implemented for the 2010 March Madness Predictive Analytics Challenge.  It's a fairly simple model.  For each team, for each game, it predicts a score based upon the current Offense and Defense ratings.  It then determines the error between the prediction and the actual game result, and adjusts the appropriate Offense and Defense ratings to remove 75% of the error.  It then iterates this across all the games for a fixed number of iterations.  (This algorithm isn't guaranteed to converge, although in practice it usually does.)

Testing this algorithm with our usual methodology gives these results (for comparison, I show the best non-MOV predictor as well):

  Predictor    % Correct    MOV Error  
TrueSkill + iRPI72.9%11.01
Offense-Defense69.6%11.84

This performance (with some adjustments) was enough to win the 2010 Challenge, but seems disappointing in comparison to the TrueSkill + iRPI performance.  In particular, we might expect ratings based upon MOV to have lower MOV error rates, but that is not the case here.  I also implemented a version of the Dick Vitale methodology, where I calculated separate home and away ratings for all teams.  In this case, our predicted score is the home team's home Offense times the away team's away Defense (and vice versa).  Here's how that performs:

  Predictor    % Correct    MOV Error  
TrueSkill + iRPI72.9%11.01
Offense-Defense (home & away)68.2%12.26

Surprisingly (at least to me) this is significantly worse than the undifferentiated ratings.  Perhaps this is additional evidence that teams don't play differently at home than away; the home court advantage would then be due primarily to the referees -- a conclusion shared by Sports Illustrated.

The second model I tested is the "Probabilistic Matrix Model" (PMM).  This model is based upon the code Danny Tarlow released for his tournament predictor, which he discusses here.  This is similar in spirit to the Offense-Defense model, if much more sophisticated mathematically.  (You can tell this because the code has variables like s_hat_i in it.)  Testing PMM gives these results:

  Predictor    % Correct    MOV Error  
TrueSkill + iRPI72.9%11.01
Offense-Defense69.6%11.84
PMM71.7%11.23

The PMM does better than my naive Offense-Defense model (apparently there's something to all that math stuff) but still does not approach the performance of TrueSkill + iRPI.  I did not implement separate home & away ratings, but there's no reason to think they would provide improved performance.