Showing posts with label limits of prediction. Show all posts
Showing posts with label limits of prediction. Show all posts

Tuesday, January 22, 2013

Team to Team Variation in Predictions

The current version of the Prediction Machine averages about 11 points of error in predicting games, across all teams and all seasons.  I've speculated that there are some subsets of games where the error is significantly less -- for example, it might be the case that we can predict much more accurately when a good rebounding team plays a poor rebounding team.  However, my efforts to identify those subsets have been largely futile and there's some circumstantial evidence to suggest that no subsets exist -- primarily that a Support Vector Machine does no better than a Linear Regression at prediction.  (We would expect a SVM to do better in a data set with significant subsets.)

Last week I thought it would be interesting to look at what teams the PM has done the best at predicting this year and which ones the worst.  (For some reason, it's never occurred to me previously to look at this.)  So I gathered up all the predictions and results for this season and segmented them out by teams.  (Note: I'm only looking at the games after the first 1000 games of the season and not the last 100 in this sample.)
The overall best team is Idaho, which the PM has predicted with about 4.6 points of error.  The PM has gotten 5 of Idaho's games within 2 points.  It missed one game by 10 points, but that was by far the worst.
The overall worst team is Mississippi St with 18 points of error.  The PM missed games by 35, 29, 17, and 15 points.  So the overall range of predictions runs from less than 1/2 the average error to almost 2x the average.
 
I also took a look at the error for home games and away games separately.

Just looking at home games, the most predictable is TX Pan American (2.3 points error) and the worst Youngstown State (28.7 points error).  For away games only, the most predictable is Portland St (0.86 points error!) and the worst is Maryland (20.5 points error).   For Portland State's four away games, the PM was off by 0.7, 1.3, 0.4 and 0.7 points (!).  Maryland is a bit deceptive -- they only have two away games in the sample, and one was the Northwestern game which they were expected to lose by 9 and won by 20.

Some of this is no doubt just random variation.  Just by chance the PM will get some team's games close and some team's far off.  That effect should diminish the more games we sample, so I took a look at the entire 2012 season.

The overall best team to predict in 2012 was Dartmouth, with 6.3 points of error on average.  The worst was New Orleans, with 19.3 points of average error.  Once again we see a range of roughly 1/2 to 2x the average error.  The best home team to predict was Indiana St, at 4.7 points of error, and the worst Longwood at 17.15 points of error.  The best and worst away teams were Gardner-Webb (4.14) and New Orleans (22).   So again we see the overall range of predictions runs about 1/2 the average error to about 2x the average error.

Another test we can do is to look at how many teams that were predictable in the first half of the season are also predictable in the second half of the season.  If the effect is random, we'd expect to see a random level of overlap.  For the 2012 season, if we look at the most predictable half of the teams in both the first part of the season and the second part of the season, there's almost exactly 50% overlap -- a strong indication that the effect is just random variation.

The conclusion is that the error range on the PM's predictions for particular teams runs from about 1/2 the overall average to about 2 times the average, but that this variation is probably random.

Tuesday, April 19, 2011

New Data on Limits & Other Interesting Links

In a previous posting, I looked at the limits of prediction and concluded that the best performance we could hope for from a predictor would be in the 70-80% range for correct predictions.  Today I happened to run across "The Prediction Tracker" which (amongst other things) tracks the performance of various college basketball rating systems as predictors of game performance.  For the season that just ended, the best computer predictor belonged to Jon Doktor, and managed a "% Correct" measure of 73%.  All of the predictors tracked by that site cluster in the lower end of the 70-80% range.  Given that they include early season games, that's fairly solid performance.  It's also interesting to note that (1) none of the predictors managed even a 1% advantage betting against the spread, and (2) all of them had MOV errors in the 9-10 point range.  (Most of these predictors use margin of victory, so we would expect them to perform better on MOV than systems like RPI which use only win-loss.)

I got to the Prediction Tracker via the TeamRankings.com blog, which has a 4 part series discussing their rating systems starting here.  The discussion lacks any concrete details on the algorithms but covers some interesting ground and is worth a look.

Wednesday, April 13, 2011

As Good As We Can Do

The 1-Bit Predictor gives us a useful lower bound for prediction performance.  Let's turn now to the other end: What's the theoretical "best performance" we can hope to achieve?  Comparing RPI, Massey, Sagarin and LMRC predictions over six seasons of tournament games, [Sokol 2006] found performances in the 70-75% range.  How much can we hope to improve that number?

We can think of a college basketball game as having both a deterministic and a random component.  If the random component is zero, then there would be no variability in outcome -- every time two teams matched up (all other things being equal) the same result would occur.  If the deterministic component is zero, then results would be completely random.  Reality obviously lies somewhere in-between those two extremes.

By definition, there's no way to predict the random component of the outcome.  If we assume that we can predict the deterministic component perfectly, our performance is then limited by the magnitude of the random component.  So what is the magnitude of the random component?

There are a couple of different ways to explore this question.  One thought experiment is to imagine a game in which there is no random component except in the the last possession of the game, which is completely random.  Intuitively, that seems much less random than reality, so it provides a lower bound on estimating the magnitude of the random component.  So how would that affect the final outcome?

On the last possession, the team with the ball can score 2 or 3 points (or even 4 points), or might turn the ball over leading to the other team scoring -- a potential swing of 6 or more points.  So in this case, if we could predict the deterministic component of the game perfectly, we'd still have an average error of 3+ points.

Of course, in reality the last possession isn't entirely random.  But more importantly, the first 120 possessions aren't entirely deterministic!  This suggests that the best performance we can hope for is going to be significantly worse than +/- 3 points.

Another method to gain insight into this question is to look at repeat matchups of teams. Home and home conference matchups along with conference tournament matchups provide a lot of data that can be used to estimate the variability in college basketball games.  For example, in the course of a month in 2011, Duke and UNC played home-and-home and an ACC tournament game with the following results (margin from Duke's perspective):

                Result
           @Duke        +6
           @UNC     -14
           @Neutral     +17

This shows an enormous amount of variability.  Of course, there's a systemic bias in these numbers -- the home court advantage.  Sagarin estimates that at about 4 points for 2011.  If we factor that out, the results are:

                 Result
           @Duke         +2
           @UNC      -10
           @Neutral      +17

which still suggests double-digit variability in game outcomes. Looking at all the home-and-home matchups for a season, [Sokol 2006] found that a team had to win by 21 points at home to have an even chance to win on the road.  Part of that margin is due to home court advantage, but since most estimates of HCA are in the 4 point range, the rest of the margin is probably required to "overcome" significant variability.

These sorts of analysis suggest that the random component in game outcomes is at least +/- 8 points.  So what does that say about trying to predict the outcomes of college basketball games?

Looking at 185 tournament games from 2009 & 2010 (both NCAA and NIT), the average margin of victory was about 11 points.  40% of the games were decided by 8 points or less.  If we look at just our first performance metric (picking the correct winnger), we need only get the outcome correct (not the final margin).  A predictor that accounts perfectly for everything except (say) 8 points of variability would get 60% of the predictions correct along with some portion (say 65%) of the remaining 40% -- for a final performance of ~85%.

In reality, of course, our predictor won't be perfect on the deterministic component of games, either.  Taken all together, this suggests that a realistic upper limit for picking the correct winner of a game is in the 70-80% range.  Since [Sokol 2006] showed that RPI and other schemes are already predicting in the lower part of this range, our progress is likely to be very incremental.  Improvements of 1% will be significant progress!

On our other performance metric (predicting the MOV) the story may be better, but that's an analysis for another day.

Tuesday, April 12, 2011

The 1-Bit Predictor

In the previous posting, we established a bottom threshold for predictor performance: 50% correct and about 14.5 points error in the MOV. Of course, that "predictor" doesn't use any information at all about the game, so it's mostly a curiousity. It's more interesting to ask how well we can do with a very simple predictor that actually uses information about the game. It turns out that with the smallest amount of extra information we can do much better.

Information Theory wonks define the smallest unit of information as a bit -- essentially the answer to one yes-no question. (Disclaimer: yes, I know the real definition is more complex :-) Let us suppose, then, that we can get the answer to one yes-no question about a basketball game. What question should we ask, and how much can we improve our prediction based upon that information?

Basketball aficionados will know that "home court advantage" (HCA) is a major factor in college basketball. We will examine HCA in more detail soon, but for now let us use our one bit of information to determine the home team. How much does knowing the answer to that question improve our prediction?

It turns out that in college basketball, the home team wins an astonishing 66% of the games, and outscores the visiting team by an average of 4.5 points. We can use this information to create our "1-Bit Predictor":

The home team will win by 4.5 points.

This predictor gets 66% of its games correct with an error of about 13.5 points over all games in the 2009-2011 seasons. So with that one bit of information we've improved our predictor by 32% on one performance metric! (But only by about 7% on the other metric, which will prove a tougher nut to crack.)

  Predictor    % Correct    MOV Error  
Naive50%14.5
1-Bit62.6%14.17

(For comparison purposes to later predictors, the performance I show in this table corresponds to a standard testing methodology, to be explained shortly.)

So the "1-Bit Predictor" provides a reasonable lower bound on prediction performance. This may seem like a trivial result, but consider that [Sokol 2006] compared four rating systems and the Las Vegas betting line over six seasons of tournament games and found that they picked the correct winner in 70-75% of the games. The 1-Bit Predictor is already within a few percentage points of these much more sophisticated systems. If nothing else, this suggests that improving prediction performance is going to be a difficult task.

Having established a reasonable lower bound, the obvious next question is "What is a reasonable upper bound for prediction performance?" That is the topic of the next posting.

Monday, April 11, 2011

You can't improve what you don't measure

My day job involves a lot of process improvement work, and one of our catch-phrases is "You can't improve what you don't measure." The oft-unstated corollary is that what you choose to measure determines what you'll improve. Measure the wrong thing and you'll find yourself optimizing the wrong thing. In our quest to create a good basketball predictor, what should we take as our measure of performance?

The Predictive Analytics Challenge (as well as countless office pools) takes as a measure accuracy in predicting the NCAA tournament. That makes for an interesting challenge (and interesting office discussions) but has a few problems as a metric. First, the sample size for testing is rather small -- only 63 (or so) games a year. Second, picking all the games before any have been played and scoring different rounds with different values introduces a host of strategic complications. Finally, unlike 95% of the college basketball games, the tournament is played on a neutral court.

For these reasons, I prefer to measure the predictor against individual regular season games. Obviously I'll also use it to try to predict the Tournament -- I just won't measure its performance against Tournament games.

So how should we measure the performance of our predictor? The obvious (and simplest) measure is whether it predicts the correct outcome. That's a good metric, but it does have some flaws. For one thing, predicting the correct outcome of many games is trivial. When Duke plays Wake Forest, it isn't too difficult to predict with some confidence that Duke will win. Secondly, it's really only useful for entering Tournament contests.

A second measure we can use is to try to predict the Margin of Victory (MOV) and measure how close we got. This measure makes predicting the Duke-Wake Forest matchup more interesting -- Duke is very likely to win, but by how much? It's also useful if we want to match our predictor against the Las Vegas bookmakers, who release a "line" on every game that represents their best prediction for Margin of Victory. Given the strong financial motivation the bookmakers have to be good predictors, they should be a good test of our predictor.

(Strictly speaking, the bookmakers may not set the line to their best prediction of MOV. They may set or move the line to equalize betting on both sides of the game to minimize their financial risk.)

I will use both measures to assess the performance of the predictor. I've assessed a number of prediction models with both metrics, and it's almost always the case that optimizing one measure tends to optimize the other. In some cases that may not be true, and I'll rather arbitrarily weigh the trade-off and pick one over the other.

Now that we've established our metrics of performance, let's think about how good our predictor can be. Actually, let's start off by thinking about how bad our predictor can be.

If we know absolutely nothing about a game and randomly choose one of the two teams to win, we will predict the correct team 50% of the time. And, as it happens, the average MOV is about 15 points. So that sets a lower bound on prediction:

  Predictor    % Correct    MOV Error  
Naive50%14.5

(If you have a predictor that does worse than that, take the opposite of it's predictions and you'll have a better predictor :-)

Interestingly, with a tiny bit more information (and I mean that literally), we can do much better. That's a topic for the next posting.