It seems like he would have been elected based on the voters preferences in recent years (I have been analyzing voting patterns and what I write below will be based on that-scroll down to see these studies). But first I want briefly to discuss the sabermetric case for or against.
Garvey had 279 career "Win Shares" (WS), the Bill James stat which incorporates all phases of the game. That tied him for 222nd place all-time among all players and pitchers through 2001. Not bad, since about 200 guys are in the Hall. But this is marginal.
His career TPR or "total player rating," from Pete Palmer, editor of the Baseball Encyclopedia was actually -6.1. That means that if an average first baseman had played instead of Garvey, those teams would have won 6.1 more games during his career. Most of his seasons were negative and his best was only +1.2.
But the baseball writers who vote don't necessarily take sabermetric stats into account. The analysis I have posted recently used more conventional stats. In one model I used logit analysis to predict the probability of any player getting elected. That model had career AVG, seasons with 100+ RBIs, ALLSTAR games, career plate appearances (PAs), MVP awards, a variable for world series performance, being in the 3000 hit club and a positional adjustment for being a catcher. That model gave Garvey a 94.7% probability of being elected to the Hall of Fame. The model itself was 98.9% accurate
Another logit model (also 98.9% accurate) had the following variables:
Career HRs
2B
SS
3B
CF
C
WSIMP
ALLSTAR
MVPSH
500SB
Career NON-HRs
3000 HIT
The variables after Career HRs are positional adjustments. The WSIMP is for world series play. This model had Garvey's probability at 64.8%. Tony Perez has 52.6% and Jim Rice has 10.6% and both are in the Hall.
Another model simply predicted the % of votes received in the first year of eligibility. This model took into account MVP awards, a variable for world series performance, being in the 3000 hit club, ALLSTAR games, being in the 500 HR club, being in the 500 SB club, Gold Glove awards and career PAs. It predicted that Garvey would be named on 48.9% of the ballots in his first year while he actually got 41.6%. His predicted 48.9% is more than what was predicted for the following players who did eventually make it:
Ryne Sandberg -0.460
Kirby Puckett -0.418
Gary Carter-0.380
Carlton Fisk-0.375
Tony Perez-0.299
Jim Rice-0.219
And Garvey's actual first year % of 41.6 is higher than that of Rice (29.8%) and very close to Carter's 42.3%.
So from 3 different regressions, it looks like Garvey had the stats or qualifications to make it in, based on what the voters seem to like.
It is also very easy to find some impressive achievements that would go on Garvey's plaque, if he ever made it. They include:
-5 100 RBI seasons
-.294 career AVG
-batted over .300 7 times
-had 200 or more hits in a season 6 times
-1974 NL MVP
-batted .319 in 5 World Series
-batted .356 in 5 league championship series
-batted .393 in 10 all-star games
-won 4 Gold Glove awards
-2 time MVP of the all-star game
-2 time MVP of the league championship series
-finished in the top 5 in total bases 7 times
-set a NL record by playing 193 straight games without committing an error
-set a ML record with his .996 fielding percentage at first base.
-played in 1,207 consecutive games, an NL record and 4th longest overall
He also has 2.46 Career MVP Shares which is the 55th best total. An MVP share is what % of the total possible points a player got in the voting in a given year. A first place vote is 14 points, 2nd place 9, 3rd, 8points, etc. A guy might come in 5th but if he had 40 points out of a maximum of, say, 400, he gets a .10. Garvey's high rank here means the voters liked him when he played, alot more than they liked other players. He finished in the top 6 in MVP voting 5 times.
Baeball Reference lists the Hall of Fame Monitor for which they say:
"This is another Jamesian creation. It attempts to assess how likely (not how deserving) an active player is to make the Hall of Fame. It's rough scale is 100 means a good possibility and 130 is a virtual cinch. It isn't hard and fast, but it does a pretty good job. Here are the batting rules."
Garvey gets a 130 which is 104th best among position players. Now this is a very complicated point system with so many points for this or that. But this shows that Garvey fits the statistical profile of the kind of player the voters very much like to put in the Hall.
So why isn't he in? I found some theories.
The Baseball Page said, among other things, the following:
"In the 1980s it became clear that "Mr. Dodger" was far from wholesome. Several paternity suits and a tell all book from his ex-wife tarnished his image irreparably. Where he once was considered a candidate for state or even national office, Garvey became a leper, destined to host game shows and infomercials (really).
He had the reputation as a selfish, egotistical player. The media didn't like him as much as it seemed. His "Mr. Dodger" persona was created by Dodger PR and a few well-placed friends in the press. More than a few teammates quickly tired of Garvey's habit of staying in front of the camera or microphone.
In August of 1978, Garvey took offense to a comment made by teammate Don Sutton and the two men ended up wrestling their way across the visitors' clubhouse in Shea Stadium. The fight cemented a bitter feud between the two men and it damaged Garvey's reputation in the league.
He aged quickly. By the time he was 31-32, his skills were rapidly diminishing. He would have benefited from a day off here and there, but he didn't do it."
Chris Jaffe over at the Harball Times had an interesting article called Hitler. Stalin. Garvey. Here is an exerpt:
"There was always a sense he was a fake. With the Dodgers, he got in a big fistfight in the clubhouse with teammate Don Sutton. He had a nasty divorce in the early 1980s. When he started to get hit with paternity suits, though, his reputation was shattered.
In some ways, though, it's even deeper than that. Our society can forgive—or at least cease baiting—a hypocrite, provided he asks for some degree of atonement. Jim Bakker wrote his book, I Was Wrong, for instance.
Garvey hasn't done that."
Jeff Sackmann also has an interesting article called Steve Garvey Gets No Respect
Update Dec. 3, 2010: I did a follow up post in July, 2010. Click here to read it.
Monday, May 25, 2009
Friday, May 15, 2009
The Marginal Impact Statistics Have on Hall of Fame Voting
A couple of weeks ago I presented a binary (logit) model of Hall of Fame voting. I looked at all players whose first year of eligibility was from 1990-2009 (except for Pete Rose). The model's equation estimates the probability that a player would be elected to the Hall of Fame (though not necessarily in his first year of eligibility). Of course, there are coefficient estimates and what I do here is give the change in the probability of being elected due to a change in a hypothetical player's stats.
I made up a player who had a 0.290 career AVG, had 6 100 RBI seasons, won 1 MVP award, played in 8 all-star games and had 8,894 career plate appearances. This player was not a catcher and did not achieve 3,000 hits. He had a world series impact of 18 (that is, his world series PAs times his world series OPS). It is like he had a .750 OPS in 24 world series PAs.
The first line in the table below shows that his probability of being elected was .50 or 50%. The rest of the table shows what his probability would be if his stats changed. For instance, if his career average had been .300 instead of .290 (with nothing else changing), his probability (or PR) rises to .663. Another 10 point gain in average brings it to .795. But if he had only hit .280, the PR is just .337. Of course, other things might have changed if his averaged had changed. Maybe 100 RBI seasons or all-star games played would have been different. But I will have to assume that they did not. The different cases for AVG are in red. More discussion follows the table. You can see a larger version of the table if you click on it.

Adding 1 100 RBI seasons raises PR to .638. Taking 1 away lowers it to .362. So the 100 RBI seasons has a powerful impact, like AVG. Adding an MVP award raises PR to .59.But all-star games played in (AS) might have the strongest effect. Going from 8 to 9 all-star games raises the PR to .80. If all-star games falls to 7, PR is just .20.
Having 3000 hits, not surprisingly, raises the PR to 1.00 or 100% (actually its a little less but it is above 99.99%). If the player had been a catcher (see the 1 in the C column), it jumps to .973. The WS or world series impact is very slight. Adding or subtracting 500 career PAs matters alot. Adding 500 PAs bumps PR up to .657 and taking 500 away drops it to .343.
I also did the same test for an actual player, Jim Rice. His predicted PR was the closest of any of the players in the study to .50 at .595. The next table has the marginal changes like the one above. Maybe the biggest impact for Rice is all-star games. If he had just one less, his PR falls to .269. In general, his table looks similar to that of the hypothetical player.
I made up a player who had a 0.290 career AVG, had 6 100 RBI seasons, won 1 MVP award, played in 8 all-star games and had 8,894 career plate appearances. This player was not a catcher and did not achieve 3,000 hits. He had a world series impact of 18 (that is, his world series PAs times his world series OPS). It is like he had a .750 OPS in 24 world series PAs.
The first line in the table below shows that his probability of being elected was .50 or 50%. The rest of the table shows what his probability would be if his stats changed. For instance, if his career average had been .300 instead of .290 (with nothing else changing), his probability (or PR) rises to .663. Another 10 point gain in average brings it to .795. But if he had only hit .280, the PR is just .337. Of course, other things might have changed if his averaged had changed. Maybe 100 RBI seasons or all-star games played would have been different. But I will have to assume that they did not. The different cases for AVG are in red. More discussion follows the table. You can see a larger version of the table if you click on it.

Adding 1 100 RBI seasons raises PR to .638. Taking 1 away lowers it to .362. So the 100 RBI seasons has a powerful impact, like AVG. Adding an MVP award raises PR to .59.But all-star games played in (AS) might have the strongest effect. Going from 8 to 9 all-star games raises the PR to .80. If all-star games falls to 7, PR is just .20.
Having 3000 hits, not surprisingly, raises the PR to 1.00 or 100% (actually its a little less but it is above 99.99%). If the player had been a catcher (see the 1 in the C column), it jumps to .973. The WS or world series impact is very slight. Adding or subtracting 500 career PAs matters alot. Adding 500 PAs bumps PR up to .657 and taking 500 away drops it to .343.
I also did the same test for an actual player, Jim Rice. His predicted PR was the closest of any of the players in the study to .50 at .595. The next table has the marginal changes like the one above. Maybe the biggest impact for Rice is all-star games. If he had just one less, his PR falls to .269. In general, his table looks similar to that of the hypothetical player.
Friday, May 8, 2009
Does The "High" April Slugging Percentage Mean Anything?
(note: I will have another post on Hall of Fame voting next weekend)
It sure seemed like there was lots of slugging going on in April. The SLG for MLB was .420 for the month, much higher than the .402 and .401 for the two previous seasons. The table below shows the SLG for April and post-April for each of the last 10 years (5 years had some March games, so I included that data in April for those years).

The average April SLG over the last 10 years is just about .420. So it may not be that high, although in beats 5 of the last 7 seasons.
The graph below shows the relationship between April SLG (horizontal axis) and non-April SLG (vertical axis).

The average non-April SLG is about .426. I also ran a regression with non-April SLG as the dependent variable and April SLG as the independent variable. Here is the equation:
non-April SLG = 0.335*(April SLG) + 0.2856
The r-squared was .68 and the standard error was .0037, which seems pretty low. The equation predicts all non-April SLGs with +/- .006. So I guess we can expect SLG the rest of the year to be between .414 and .426 and it may be a good bet for it to be between .416 and .424. But so far in May it is only .410, so who knows. Maybe the weather matters and maybe this April was unusually warm.
It sure seemed like there was lots of slugging going on in April. The SLG for MLB was .420 for the month, much higher than the .402 and .401 for the two previous seasons. The table below shows the SLG for April and post-April for each of the last 10 years (5 years had some March games, so I included that data in April for those years).

The average April SLG over the last 10 years is just about .420. So it may not be that high, although in beats 5 of the last 7 seasons.
The graph below shows the relationship between April SLG (horizontal axis) and non-April SLG (vertical axis).

The average non-April SLG is about .426. I also ran a regression with non-April SLG as the dependent variable and April SLG as the independent variable. Here is the equation:
non-April SLG = 0.335*(April SLG) + 0.2856
The r-squared was .68 and the standard error was .0037, which seems pretty low. The equation predicts all non-April SLGs with +/- .006. So I guess we can expect SLG the rest of the year to be between .414 and .426 and it may be a good bet for it to be between .416 and .424. But so far in May it is only .410, so who knows. Maybe the weather matters and maybe this April was unusually warm.
Sunday, May 3, 2009
Predicting Who Makes The Hall Of Fame Using A Logit Model
The last two weeks I presented regression results on first year Hall of Fame vote percentage. I used a linear regression. This week I use a logit model where 1 means the player has made it in and 0 means not. The probability that a player was voted in is
(1) P = 1/(1 + exp(-Z))
where exp is approximately 2.78 or Euhler's number. The -Z is the following equation times -1 for each player. This is the estimatated regression equation:
(2) -46.1955 + 67.81*CAVG + .5658*100RBI* + 1.386*ALLSTAR + .0013*PA + .3645*MVP + .0001*WSIMP + 14.93*3000HIT + 3.586*C
CAVG is a player's career batting average, 100RBI is the number of seasons with 100+ RBIs, ALLSTAR is number of all-star games played in, PA is career plate appearances, MVP is number of MVP awards won, WSIMP is world series PAs times world series OPS (this is world series impact, a combination of quantity of quality), 3000HIT is a dummy variable (1 or 0) for reaching that milestone and C is the same if the player was a catcher. So all the data gets plugged in for each player and "Z" is calculated. Then the negative of that is plugged in to equation (1) to get the probability of each player being elected into the Hall.
The statistical results can be viewed at logit results. The data for each player and their calculated probability can be seen at logit probabilities.
The results page shows something called a "classification table." It says that if a probability is .5 or greater, the player will be in the Hall. My data includes all players whose first year of eligibility was 1990 or later (except for Pete Rose). The classification table says that all 20 players who actually have made it in should be (all 20 have a probability of .5 or higher). Of the other 161 players it predicts that 159 would not make it. The two who are predicted to make it are Steve Garvey and Andre Dawson. Garvey's 15 years of eligibility to be voted in by the writers has ended. But Dawson could still make it. Garvey had a probability of 95.7% according to the model. Dawson had 60.8%. The model is 98.9% correct.
I tried lots of variables and this model works the best in terms of getting all the actual Hall of Famers right and the overall correct%. Also, some models might have done a bit better but a variable would be negative (like career HRs) that should not be. Some models had many more variables. My Acastat programs sometimes could not complete a regression no matter how long I waited. That might have been if I had more than 12 variables. So another model might be a bit better. I just don't know.
Now it is hard to see what exactly the impact of each variable is since this is a non-linear model. So I will just mention a few players and how their probability (P)would change if their data changed.
Yount-If he does not have 3000 hits, his P falls to about 1% from 99%! Molitor would fall from 99% to 45%.
Ozzie Smith-If he falls from 14 all-star games to 10, his P falls from 99% to 36%
Fisk-If he is not a catcher, his P falls from 96% to 46%. Gary Carter falls from 65% to 36%.
Carew-If he does not have 3000 hits, his P falls only about 1% from 100% to 99%! (Boggs, Gwynn and Brett have the same thing)
Eddie Murray-If he does not have 3000 hits, his P falls to about 83% from 99%.
Puckett-If his CAVG falls from .318 to .301 his P falls from 75% to 49% (a NO). If you don't change his CAVG but take his All-Star games from 10 to 9, his P falls to 43%. Incidentally, he finished his career with 280 Win Shares and if he could have played 5 more years he could have easily gotten to 360 Win Shares, what Bill James says is almost a lock for the Hall (but Tim Raines has 392 and might not make it-if he went from 7 to 10 all-star games his P would be about 75%).
Tony Perez-If you take his All-Star games from 7 to 6, his P falls to 29% from 62%.
McGwire-Maybe his not getting in is why 500 HRs was not working very well in the model. The only others were Murray, Jackson and Schmidt. But if he had 9200 PAs, his P value jumps to .5. Of course, he has the steroid scandal. It would be nice to quantify scandals, but that may be impossible and I have already take Rose out of the model. Back to McGwire, if he had 11 all-star games instead of 9, his P would go up to 69%.
Ryne Sandberg-Take away his MVP award and his P falls from 63% to 54% (do that and take away 1 all-star game, he falls to 23%). For some other guys, it makes almost no difference. Take all 3 of Schmidt's awards away and he still has a 99% P. Take away 1 of Morgan's awards and he falls from 66% to 58%. Take away the other and he falls to 48%.
Al Oliver would go from a P of 10% to over 50% if he had 6 100 RBI seasons instead of 2. Same for Ted Simmons if he jumped from 3 to 7 100 RBI seasons. Same for Harold Baines.
So all-star games and 3000 hits matter alot (and 100 RBI seasons, too). If Dave Parker had 8 all-star game instead of 6, he goes from a P of 8% to over 60%. Harold Baines would have a P of 99% if he had 3000 hits. The 3 strike shortened seasons of 1981, 1994 and 1995 might have cost him 3000 hits. Will Clark and Keith Hernandez would have P's of 99% if they made 3000 hits.
I can email my spread sheet to anyone if you want to play around with these kinds of possibilities.
(1) P = 1/(1 + exp(-Z))
where exp is approximately 2.78 or Euhler's number. The -Z is the following equation times -1 for each player. This is the estimatated regression equation:
(2) -46.1955 + 67.81*CAVG + .5658*100RBI* + 1.386*ALLSTAR + .0013*PA + .3645*MVP + .0001*WSIMP + 14.93*3000HIT + 3.586*C
CAVG is a player's career batting average, 100RBI is the number of seasons with 100+ RBIs, ALLSTAR is number of all-star games played in, PA is career plate appearances, MVP is number of MVP awards won, WSIMP is world series PAs times world series OPS (this is world series impact, a combination of quantity of quality), 3000HIT is a dummy variable (1 or 0) for reaching that milestone and C is the same if the player was a catcher. So all the data gets plugged in for each player and "Z" is calculated. Then the negative of that is plugged in to equation (1) to get the probability of each player being elected into the Hall.
The statistical results can be viewed at logit results. The data for each player and their calculated probability can be seen at logit probabilities.
The results page shows something called a "classification table." It says that if a probability is .5 or greater, the player will be in the Hall. My data includes all players whose first year of eligibility was 1990 or later (except for Pete Rose). The classification table says that all 20 players who actually have made it in should be (all 20 have a probability of .5 or higher). Of the other 161 players it predicts that 159 would not make it. The two who are predicted to make it are Steve Garvey and Andre Dawson. Garvey's 15 years of eligibility to be voted in by the writers has ended. But Dawson could still make it. Garvey had a probability of 95.7% according to the model. Dawson had 60.8%. The model is 98.9% correct.
I tried lots of variables and this model works the best in terms of getting all the actual Hall of Famers right and the overall correct%. Also, some models might have done a bit better but a variable would be negative (like career HRs) that should not be. Some models had many more variables. My Acastat programs sometimes could not complete a regression no matter how long I waited. That might have been if I had more than 12 variables. So another model might be a bit better. I just don't know.
Now it is hard to see what exactly the impact of each variable is since this is a non-linear model. So I will just mention a few players and how their probability (P)would change if their data changed.
Yount-If he does not have 3000 hits, his P falls to about 1% from 99%! Molitor would fall from 99% to 45%.
Ozzie Smith-If he falls from 14 all-star games to 10, his P falls from 99% to 36%
Fisk-If he is not a catcher, his P falls from 96% to 46%. Gary Carter falls from 65% to 36%.
Carew-If he does not have 3000 hits, his P falls only about 1% from 100% to 99%! (Boggs, Gwynn and Brett have the same thing)
Eddie Murray-If he does not have 3000 hits, his P falls to about 83% from 99%.
Puckett-If his CAVG falls from .318 to .301 his P falls from 75% to 49% (a NO). If you don't change his CAVG but take his All-Star games from 10 to 9, his P falls to 43%. Incidentally, he finished his career with 280 Win Shares and if he could have played 5 more years he could have easily gotten to 360 Win Shares, what Bill James says is almost a lock for the Hall (but Tim Raines has 392 and might not make it-if he went from 7 to 10 all-star games his P would be about 75%).
Tony Perez-If you take his All-Star games from 7 to 6, his P falls to 29% from 62%.
McGwire-Maybe his not getting in is why 500 HRs was not working very well in the model. The only others were Murray, Jackson and Schmidt. But if he had 9200 PAs, his P value jumps to .5. Of course, he has the steroid scandal. It would be nice to quantify scandals, but that may be impossible and I have already take Rose out of the model. Back to McGwire, if he had 11 all-star games instead of 9, his P would go up to 69%.
Ryne Sandberg-Take away his MVP award and his P falls from 63% to 54% (do that and take away 1 all-star game, he falls to 23%). For some other guys, it makes almost no difference. Take all 3 of Schmidt's awards away and he still has a 99% P. Take away 1 of Morgan's awards and he falls from 66% to 58%. Take away the other and he falls to 48%.
Al Oliver would go from a P of 10% to over 50% if he had 6 100 RBI seasons instead of 2. Same for Ted Simmons if he jumped from 3 to 7 100 RBI seasons. Same for Harold Baines.
So all-star games and 3000 hits matter alot (and 100 RBI seasons, too). If Dave Parker had 8 all-star game instead of 6, he goes from a P of 8% to over 60%. Harold Baines would have a P of 99% if he had 3000 hits. The 3 strike shortened seasons of 1981, 1994 and 1995 might have cost him 3000 hits. Will Clark and Keith Hernandez would have P's of 99% if they made 3000 hits.
I can email my spread sheet to anyone if you want to play around with these kinds of possibilities.
Sunday, April 26, 2009
What Determines Vote Percentage In The First Year Of Hall Of Fame Eligibility? (Part 2)
I have tried to improve the model I used last week. It was an OLS (linear) regression model (next week I will present a logit or binary model which is non-liner simply looks to see if a player has made the Hall or not-it seems very accurate). This week I have converted some of last week's variables into a non-linear variables and I have added two new variables. At the end of this post I have links to some other research and discussion on this issue.
Here is the regression equation:
PCT = -.041 + .054*MVP + .432*3000H + .172*500HR + .004*ASSQ10 + .001*GGSQ7 + .074*500SB + .00001*WSIMPSQ50 + .102*10000PA
I will explain what the variables mean below. The adjusted r-squared was .901 (last week it was .839). So 90.1% of the difference across players is explained by the equation. The standard error was .081, down from .102. There were 181 players, all of those who came up for the first time from 1990-2009, except for Pete Rose(see last week). MVP is number of MVP awards won, 3000H is a dummy variable (1 if a player reached it, 0 otherwise). The 500HR is also a dummy variable as it is for 500SB and 10000PA (if you made it to 10,000 career plate appearances, you get a 1, 0 otherwise). I used all the voting data from 1990-2009.
What is ASSQ10? It is the square of the number of All-star games played in squared. But AS games played is maxed out at 10. The assumption here is that being an all-star has a positive exponential effect but only up to a point where no more games helps (I have a graph below to help explain this). The GGSQ7 is the same thing for Gold Gloves.
WSIMPSQ50 involves World Series play. First, WSIMP is World Series PAs times OPS. The idea here that the more you play in the World Series the more votes you would get, but by multiplying it by OPS, it also includes how well you played (or just hit). This gets maxed out at 50 and is squared, for the same reason as all-star games (yes, Reggie Jackson is first here and way ahead of everyone else at 141, with Dave Justice and Lonnie Smith tied for 2nd at 101).
All of the variables were significant at the 10% level except for GGSQ7, which came close with a p-value of .13. The other variables all had p-values of under .02 with 6 under .01.
This has been the best linear model I can come up with so far. Many different variables have been tried. 154 of the 181 players were predicted to within 10 percentage points. 126 or 69.6% were within 5 points.
The graph below illustrates what I mentioned above about squaring and capping variables. Notice that the line is increasing exponentially then flatlines.

Below are the players who had the biggest negative prediction differentials. Lynn, for example, was predicted to get 35.8% (.358) of the vote but got only 5.5% (.055).
0.055 *** -0.303 Lynn, Fred
0.235 *** -0.230 McGwire, Mark
0.017 *** -0.180 Bell, Buddy
0.083 *** -0.154 Nettles, Graig
0.053 *** -0.152 Baines, Harold
0.017 *** -0.151 Parrish, Lance
0.005 *** -0.144 Lopes, Davey
0.068 *** -0.140 Concepcion, Dave
0.019 *** -0.131 Cey, Ron
0.051 *** -0.130 Hernandez, Keith
Now the players who had the biggest positive prediction differentials.
0.298 *** 0.080 Rice, Jim
0.852 *** 0.095 Molitor, Paul
0.157 *** 0.109 Trammell, Alan
0.965 *** 0.111 Schmidt, Mike
0.775 *** 0.126 Yount, Robin
0.818 *** 0.164 Morgan, Joe
0.500 *** 0.189 Perez, Tony
0.664 *** 0.296 Fisk, Carlton
0.917 *** 0.319 Smith, Ozzie
0.821 *** 0.394 Puckett, Kirby
I also ran a regression which had the following variables. None of them were squared or made non-linear. 2B, SS and C are positional dummies. All of these variables were significant at the 10% level. 9 were significant at the 5% level and 7 the 1% level. But the adjusted r-squared was .763 and he standard error was .126. So it does not work nearly as well as the model mentioned above.
AVG
MVP
3000 HIT
2B
SS
C
SB
WSIMP
10000PA
HR
Now some links to other research
Baseball Hall of Fame voting: a test of the customer discrimination by Arna Desser, James Monks and Michael Robinson
Social Science Quarterly Sept 1999 v80 i3 p591(13)
Discussion of above article
A neural-net Hall of Fame prediction method
Teaching Statistical Thinking Using the Baseball Hall of Fame by Steven Wang
Who's missing from the hall of fame by JC Bradbury
Modeling Election to the Major League Baseball Hall of Fame through the use of Genetic Algorithms By David Cohen
Here is the regression equation:
PCT = -.041 + .054*MVP + .432*3000H + .172*500HR + .004*ASSQ10 + .001*GGSQ7 + .074*500SB + .00001*WSIMPSQ50 + .102*10000PA
I will explain what the variables mean below. The adjusted r-squared was .901 (last week it was .839). So 90.1% of the difference across players is explained by the equation. The standard error was .081, down from .102. There were 181 players, all of those who came up for the first time from 1990-2009, except for Pete Rose(see last week). MVP is number of MVP awards won, 3000H is a dummy variable (1 if a player reached it, 0 otherwise). The 500HR is also a dummy variable as it is for 500SB and 10000PA (if you made it to 10,000 career plate appearances, you get a 1, 0 otherwise). I used all the voting data from 1990-2009.
What is ASSQ10? It is the square of the number of All-star games played in squared. But AS games played is maxed out at 10. The assumption here is that being an all-star has a positive exponential effect but only up to a point where no more games helps (I have a graph below to help explain this). The GGSQ7 is the same thing for Gold Gloves.
WSIMPSQ50 involves World Series play. First, WSIMP is World Series PAs times OPS. The idea here that the more you play in the World Series the more votes you would get, but by multiplying it by OPS, it also includes how well you played (or just hit). This gets maxed out at 50 and is squared, for the same reason as all-star games (yes, Reggie Jackson is first here and way ahead of everyone else at 141, with Dave Justice and Lonnie Smith tied for 2nd at 101).
All of the variables were significant at the 10% level except for GGSQ7, which came close with a p-value of .13. The other variables all had p-values of under .02 with 6 under .01.
This has been the best linear model I can come up with so far. Many different variables have been tried. 154 of the 181 players were predicted to within 10 percentage points. 126 or 69.6% were within 5 points.
The graph below illustrates what I mentioned above about squaring and capping variables. Notice that the line is increasing exponentially then flatlines.

Below are the players who had the biggest negative prediction differentials. Lynn, for example, was predicted to get 35.8% (.358) of the vote but got only 5.5% (.055).
0.055 *** -0.303 Lynn, Fred
0.235 *** -0.230 McGwire, Mark
0.017 *** -0.180 Bell, Buddy
0.083 *** -0.154 Nettles, Graig
0.053 *** -0.152 Baines, Harold
0.017 *** -0.151 Parrish, Lance
0.005 *** -0.144 Lopes, Davey
0.068 *** -0.140 Concepcion, Dave
0.019 *** -0.131 Cey, Ron
0.051 *** -0.130 Hernandez, Keith
Now the players who had the biggest positive prediction differentials.
0.298 *** 0.080 Rice, Jim
0.852 *** 0.095 Molitor, Paul
0.157 *** 0.109 Trammell, Alan
0.965 *** 0.111 Schmidt, Mike
0.775 *** 0.126 Yount, Robin
0.818 *** 0.164 Morgan, Joe
0.500 *** 0.189 Perez, Tony
0.664 *** 0.296 Fisk, Carlton
0.917 *** 0.319 Smith, Ozzie
0.821 *** 0.394 Puckett, Kirby
I also ran a regression which had the following variables. None of them were squared or made non-linear. 2B, SS and C are positional dummies. All of these variables were significant at the 10% level. 9 were significant at the 5% level and 7 the 1% level. But the adjusted r-squared was .763 and he standard error was .126. So it does not work nearly as well as the model mentioned above.
AVG
MVP
3000 HIT
2B
SS
C
SB
WSIMP
10000PA
HR
Now some links to other research
Baseball Hall of Fame voting: a test of the customer discrimination by Arna Desser, James Monks and Michael Robinson
Social Science Quarterly Sept 1999 v80 i3 p591(13)
Discussion of above article
A neural-net Hall of Fame prediction method
Teaching Statistical Thinking Using the Baseball Hall of Fame by Steven Wang
Who's missing from the hall of fame by JC Bradbury
Modeling Election to the Major League Baseball Hall of Fame through the use of Genetic Algorithms By David Cohen
Monday, April 20, 2009
What Determines Vote Percentage In The First Year Of Hall Of Fame Eligibility?
I tried several models and maybe I will update this in the next several days, explaining what some of them were and why I am posting this one first. But using OLS regression, here is the equation I came up with:
Pct = -.08 + .00011*SB + .00741*GG + .071*MVP + .032*AS + .512*3000H + .29*500HR
These are all career totals. GG is number of Gold Gloves won, MVP is number of MVP awards won, 3000H is a dummy variable (1 if a player reached it, 0 otherwise). The 500HR is also a dummy variable. I used all the voting data from 1990-2009. Any player who had received votes before 1990 was not counted. Pete Rose was not included since leaving him out improved the results and he got nowhere near what the model predicts (usually about 70-80 points lower-in 1992 he got just 9.5% of the vote). So the scandals and controversies probably played a role. Mark McGwire got a lower % than predicted, but it was not anywhere near as bad as it was for Rose.
There were 182 players. The adjusted r-squared was .839 and the standard error was .104 (that is the lowest I got and I looked at many models with lots of different combinations of many variables). Perhaps this kind of fishing expedition, just looking for the most accurate regression, is not legitimate. But maybe it is the only way to figure out what the voters care about. All of the variables were significant at the 5% level (the highest p-value was about .04 and the next highest was .012)
McGwire did have the biggest negative difference between his predicted % and what he actually got. The equation predicts about 50.8% while he only got 23.5%, for a differential of -.273. But Fred Lynn came in .262 below his predicted value of .317. Here are the bottom ten in differential
-0.27344 McGwire, Mark
-0.26287 Lynn, Fred
-0.19284 Hernandez, Keith
-0.1864 Ripken, Cal
-0.15364 Parrish, Lance
-0.14969 Concepcion, Dave
-0.14829 Murphy, Dale
-0.14131 Cedeno, Cesar
-0.13065 Fernandez, Tony
-0.13031 McGee, Willie
Now the top ten
0.12798 Brett, George
0.12806 Schmidt, Mike
0.15458 Carter, Gary
0.17142 Molitor, Paul
0.24404 Jackson, Reggie
0.34928 Perez, Tony
0.35425 Morgan, Joe
0.38621 Smith, Ozzie
0.40061 Fisk, Carlton
0.5199 Puckett, Kirby
So the question is why did these guys get so many more votes than predicted (and why did those other guys get so many less)? I will have to think about that some more and maybe I can improve the model if I come up with something.
Brett and Schmidt would both still have made it in on the first ballot. But Puckett was only predicted to get .301. Many of the players here in the top ten also got alot more than predicted in at least one other model. That model had the following variables
HR
AVG
SB
GG
MVP
3000 HIT
2B
SS
C
The last 3 are positional dummies. But the standard error on this model was .132, much higher than the other model.
Pct = -.08 + .00011*SB + .00741*GG + .071*MVP + .032*AS + .512*3000H + .29*500HR
These are all career totals. GG is number of Gold Gloves won, MVP is number of MVP awards won, 3000H is a dummy variable (1 if a player reached it, 0 otherwise). The 500HR is also a dummy variable. I used all the voting data from 1990-2009. Any player who had received votes before 1990 was not counted. Pete Rose was not included since leaving him out improved the results and he got nowhere near what the model predicts (usually about 70-80 points lower-in 1992 he got just 9.5% of the vote). So the scandals and controversies probably played a role. Mark McGwire got a lower % than predicted, but it was not anywhere near as bad as it was for Rose.
There were 182 players. The adjusted r-squared was .839 and the standard error was .104 (that is the lowest I got and I looked at many models with lots of different combinations of many variables). Perhaps this kind of fishing expedition, just looking for the most accurate regression, is not legitimate. But maybe it is the only way to figure out what the voters care about. All of the variables were significant at the 5% level (the highest p-value was about .04 and the next highest was .012)
McGwire did have the biggest negative difference between his predicted % and what he actually got. The equation predicts about 50.8% while he only got 23.5%, for a differential of -.273. But Fred Lynn came in .262 below his predicted value of .317. Here are the bottom ten in differential
-0.27344 McGwire, Mark
-0.26287 Lynn, Fred
-0.19284 Hernandez, Keith
-0.1864 Ripken, Cal
-0.15364 Parrish, Lance
-0.14969 Concepcion, Dave
-0.14829 Murphy, Dale
-0.14131 Cedeno, Cesar
-0.13065 Fernandez, Tony
-0.13031 McGee, Willie
Now the top ten
0.12798 Brett, George
0.12806 Schmidt, Mike
0.15458 Carter, Gary
0.17142 Molitor, Paul
0.24404 Jackson, Reggie
0.34928 Perez, Tony
0.35425 Morgan, Joe
0.38621 Smith, Ozzie
0.40061 Fisk, Carlton
0.5199 Puckett, Kirby
So the question is why did these guys get so many more votes than predicted (and why did those other guys get so many less)? I will have to think about that some more and maybe I can improve the model if I come up with something.
Brett and Schmidt would both still have made it in on the first ballot. But Puckett was only predicted to get .301. Many of the players here in the top ten also got alot more than predicted in at least one other model. That model had the following variables
HR
AVG
SB
GG
MVP
3000 HIT
2B
SS
C
The last 3 are positional dummies. But the standard error on this model was .132, much higher than the other model.
Sunday, April 12, 2009
The Incredible Dominance of The 1936-39 Yankees
They won 4 straight world series. So you might be wondering why I don't write about the 1949-53 Yankees, who won 5 in a row. Those latter Yankees seem pretty dominiating. But I don't think in any 4 year period any team has done what those earlier Yankees did. The 1936-39 team had winning percentages of .667, .662, .651 and .702. Their only losing month in the 4 years was Sept. 1938 when they went 13-14.
First, they lead their league in ERA, HRs, Runs, fewest runs allowed and SLG in all 4 years. And, if I recall correctly, their run differential of 411 in 1939 is the highest of all time (967-556). The following table only amplifies their regular season dominance (the lines are in order of the years starting with 1936).

They spent 530 days in first place. That was about 80% of the time. The Indians were a half game ahead of the Yankees on July 12, 1938. But neither team had played 77 games yet. So no team besides the Yankees was in first place during the 2nd half of any of these 4 seasons. Where I have "Games left to play on clinch date," I simply added wins and losses and then subtracted from 154. I did not try to take ties or actual games left into account. So it is an approximation. But the lowest total was 12. That was the closest "race."
UPDATE: (April 16) I also figured out what the closest any team got to the Yankees was in each year between Sept. 1 and the date they clinched. Here are those games behind numbers in order of the years: 16-9-13-11.5. So between Sept. 1, 1936 and the date they clinched, the fewest games behind for the 2nd place team was 16. And in all 4 years, the closest any team came to them between Sept. 1 and the clinch date was 9 games, in 1937.
Then notice that no team ever finished closer than 9.5 games with the average games ahead being 14.75. Also notice that going into Sept., the closest anyone got was 11 games. In 1938, they went from being a half game behind on July 12 to being up 14 games up by Aug 31. In about 7 weeks, they gained 14 games.
You might be wondering if they kept up this dominance in the World Series. After all, they were playing the best the National League could throw at them. The table below summarizes what happened.

Notice that their edges in SLG and OBP are much larger than their edge in AVG (maybe they had a sabermetrician working for them back then). OBP just used walks, hits and atbats. They had 23 more walks while having huge leads in HRs and TB. Their winning pct against the NL champs was .842 while never losing more than 2 games in any one series. Their Pythagorean winning pct was .825 (runs scored squared divided by the sum of runs scored squared + runs allowed squared).
From other research, I have found that winning pct is about 1.21*(OPS differential) + .500. OPS is OBP + SLG. The Yankees had a .760 OPS while their opponents had .590. That would give the Yankees a pct of .706, not as dominat as what actually happened but still pretty impressive considering the competition.
First, they lead their league in ERA, HRs, Runs, fewest runs allowed and SLG in all 4 years. And, if I recall correctly, their run differential of 411 in 1939 is the highest of all time (967-556). The following table only amplifies their regular season dominance (the lines are in order of the years starting with 1936).

They spent 530 days in first place. That was about 80% of the time. The Indians were a half game ahead of the Yankees on July 12, 1938. But neither team had played 77 games yet. So no team besides the Yankees was in first place during the 2nd half of any of these 4 seasons. Where I have "Games left to play on clinch date," I simply added wins and losses and then subtracted from 154. I did not try to take ties or actual games left into account. So it is an approximation. But the lowest total was 12. That was the closest "race."
UPDATE: (April 16) I also figured out what the closest any team got to the Yankees was in each year between Sept. 1 and the date they clinched. Here are those games behind numbers in order of the years: 16-9-13-11.5. So between Sept. 1, 1936 and the date they clinched, the fewest games behind for the 2nd place team was 16. And in all 4 years, the closest any team came to them between Sept. 1 and the clinch date was 9 games, in 1937.
Then notice that no team ever finished closer than 9.5 games with the average games ahead being 14.75. Also notice that going into Sept., the closest anyone got was 11 games. In 1938, they went from being a half game behind on July 12 to being up 14 games up by Aug 31. In about 7 weeks, they gained 14 games.
You might be wondering if they kept up this dominance in the World Series. After all, they were playing the best the National League could throw at them. The table below summarizes what happened.

Notice that their edges in SLG and OBP are much larger than their edge in AVG (maybe they had a sabermetrician working for them back then). OBP just used walks, hits and atbats. They had 23 more walks while having huge leads in HRs and TB. Their winning pct against the NL champs was .842 while never losing more than 2 games in any one series. Their Pythagorean winning pct was .825 (runs scored squared divided by the sum of runs scored squared + runs allowed squared).
From other research, I have found that winning pct is about 1.21*(OPS differential) + .500. OPS is OBP + SLG. The Yankees had a .760 OPS while their opponents had .590. That would give the Yankees a pct of .706, not as dominat as what actually happened but still pretty impressive considering the competition.
Monday, April 6, 2009
How Many Home Runs Would Ruth Have Hit If Baseball Had Been Integrated In His Era?
(Note: This is a slightly revised version of an article that was published in 2007 in the now defunt print periodical called "The Chicago Sports Weekly." I also had posted something like this at "Beyond the Boxscore" called How Would Integration Have Affected Ruth and Cobb?)
Maybe you have seen the images on TV of fans around the country holding up asterisk signs when Barry Bonds comes to the plate, hinting that his HR record is tainted, due to his alleged steroid use. But others counter that Babe Ruth might deserve an asterisk since he never faced blacks or dark-skinned Hispanics (there were a few players with Hispanic names before 1947 whose skin was generally pretty light).
But this raises the question of how many HRs would Ruth have hit had there not been a color barrier? I know the answer because Clio, the Greek muse of history, whispered it in my ear. You see, my Ph. D. thesis was in the field of economic history and its application of statistics is called “cliometrics.” What I am about to attempt here is something dangerous called a “counterfactual” in this field. So don’t try it at home. Leave it to the trained professionals.
Robert Fogel, economic historian at the University of Chicago, won a Noble Prize, partly for using counterfactuals. He supposed what if railroads had not been built. What other kind of transportation system (like canals) would have emerged? How would this have affected economic growth? He concluded that GDP in 1890 would have been about 5% lower than it actually was.
Not everyone was thrilled with this approach. The historian Fritz Redlich referred to counterfactuals as figments, probably of an imagination gone wild. So maybe you will think this analysis is a figment of my imagination. So. Maybe you’re a figment of my imagination. In any case, here it is.
First, we need an estimate of how many non-white pitchers there might have been. Since 1947, about 15% of all the IP by pitchers with 1,000+ IP in their careers have been by non-whites. All of the 1,000+ IP pitchers made up about 58% of all the IP since 1947, so it is a good sample. Therefore, I assume that in Ruth’s day 15% of the IP were by non-whites.
How good would those pitchers have been? Good enough to replace some white guys, who would be the worst pitchers in the league. You don’t add Satchel Paige to your team and then get rid of Lefty Grove. You dump Grover Lowdermilk (who really was not a bad pitcher but his name sounds funny, unlike mine). The non-whites with 1,000+ IP since 1947 actually had a collective ERA just about the same as the whites. So pre-1947, you dump the worst 15% of the pitchers by ERA and re-calculate the league HR rate using the remaining pitchers or the top 85%
After getting rid of the bottom 15% of the IP in each season from 1920 to 1934 (when Ruth played with the Yankees and had all of his great seasons) in the AL, I recalculated the HRs allowed per IP and found how much lower than the league average the new figures were. The average fall in HRs per IP for the years 1920-1934 was about 5%. That is, the best 85% of the pitchers had a HR per IP rate that was 5% lower than the league average (which includes all pitchers). So if you improve the pitching quality in a way that is consistent with integration, Ruth would hit 5% fewer HRs or hit about 678. Even if we cut him 10%, he still hits 643.
Some things I have not considered: when Aaron and Mays were hitting HRs in the 1950s, there still were not that many non-whites pitching. So their totals might need to be reduced. We also don’t know if all batters would be affected in the same way. The best HR hitters might have had their totals reduced more than the average hitter. Also, we don’t know what percentage of pitchers would have been non-white. Probably it is more than 15% today. Suppose it is 25%. I looked at the 1927 AL and if you only count the best 75% of the pitchers, the HR rate falls about 9%.
Suppose we only looked at the best 50% of the pitchers from 1927, HRs would fall about 18.3%. If that happened to Ruth over his whole career, he still hits 583 HRs. It is about 17% for 1921. For 1934, it would be 20%. Given that I am only counting the best 50% of the pitchers, we can safely say that integration would have reduced his HR’s by no more than 20% (the top 50% of pitchers in 2008 gave up about 20% fewer HRs than average as well). So he ends up with 571 HRs. That would have stood as a record for quite awhile. And remember that we would have to reduce Aaron and Mays since they played a good part of their careers when there were not as many non-white pitchers as today.
Here is the link that shows the white and non-white pitchers since 1947
http://cyrilmorong.com/RuthAsterisk/Pitchers.htm
Maybe you have seen the images on TV of fans around the country holding up asterisk signs when Barry Bonds comes to the plate, hinting that his HR record is tainted, due to his alleged steroid use. But others counter that Babe Ruth might deserve an asterisk since he never faced blacks or dark-skinned Hispanics (there were a few players with Hispanic names before 1947 whose skin was generally pretty light).
But this raises the question of how many HRs would Ruth have hit had there not been a color barrier? I know the answer because Clio, the Greek muse of history, whispered it in my ear. You see, my Ph. D. thesis was in the field of economic history and its application of statistics is called “cliometrics.” What I am about to attempt here is something dangerous called a “counterfactual” in this field. So don’t try it at home. Leave it to the trained professionals.
Robert Fogel, economic historian at the University of Chicago, won a Noble Prize, partly for using counterfactuals. He supposed what if railroads had not been built. What other kind of transportation system (like canals) would have emerged? How would this have affected economic growth? He concluded that GDP in 1890 would have been about 5% lower than it actually was.
Not everyone was thrilled with this approach. The historian Fritz Redlich referred to counterfactuals as figments, probably of an imagination gone wild. So maybe you will think this analysis is a figment of my imagination. So. Maybe you’re a figment of my imagination. In any case, here it is.
First, we need an estimate of how many non-white pitchers there might have been. Since 1947, about 15% of all the IP by pitchers with 1,000+ IP in their careers have been by non-whites. All of the 1,000+ IP pitchers made up about 58% of all the IP since 1947, so it is a good sample. Therefore, I assume that in Ruth’s day 15% of the IP were by non-whites.
How good would those pitchers have been? Good enough to replace some white guys, who would be the worst pitchers in the league. You don’t add Satchel Paige to your team and then get rid of Lefty Grove. You dump Grover Lowdermilk (who really was not a bad pitcher but his name sounds funny, unlike mine). The non-whites with 1,000+ IP since 1947 actually had a collective ERA just about the same as the whites. So pre-1947, you dump the worst 15% of the pitchers by ERA and re-calculate the league HR rate using the remaining pitchers or the top 85%
After getting rid of the bottom 15% of the IP in each season from 1920 to 1934 (when Ruth played with the Yankees and had all of his great seasons) in the AL, I recalculated the HRs allowed per IP and found how much lower than the league average the new figures were. The average fall in HRs per IP for the years 1920-1934 was about 5%. That is, the best 85% of the pitchers had a HR per IP rate that was 5% lower than the league average (which includes all pitchers). So if you improve the pitching quality in a way that is consistent with integration, Ruth would hit 5% fewer HRs or hit about 678. Even if we cut him 10%, he still hits 643.
Some things I have not considered: when Aaron and Mays were hitting HRs in the 1950s, there still were not that many non-whites pitching. So their totals might need to be reduced. We also don’t know if all batters would be affected in the same way. The best HR hitters might have had their totals reduced more than the average hitter. Also, we don’t know what percentage of pitchers would have been non-white. Probably it is more than 15% today. Suppose it is 25%. I looked at the 1927 AL and if you only count the best 75% of the pitchers, the HR rate falls about 9%.
Suppose we only looked at the best 50% of the pitchers from 1927, HRs would fall about 18.3%. If that happened to Ruth over his whole career, he still hits 583 HRs. It is about 17% for 1921. For 1934, it would be 20%. Given that I am only counting the best 50% of the pitchers, we can safely say that integration would have reduced his HR’s by no more than 20% (the top 50% of pitchers in 2008 gave up about 20% fewer HRs than average as well). So he ends up with 571 HRs. That would have stood as a record for quite awhile. And remember that we would have to reduce Aaron and Mays since they played a good part of their careers when there were not as many non-white pitchers as today.
Here is the link that shows the white and non-white pitchers since 1947
http://cyrilmorong.com/RuthAsterisk/Pitchers.htm
Sunday, March 29, 2009
Should Curt Schilling Get Into The Hall Of Fame?
This has been discussed since he retired this week and he may have good case. I will borrow some of what I put in my entry a few weeks ago about Bunning called Does Jim Bunning Belong In The Hall Of Fame?
Click on this next ink to see where Schilling ranks in strikeout-to-walk ratio (relative to the league average) for all pitchers with 3,000+ IP. He is 2nd to Mathewson.
K/BB Ratio
I found the best fielding independent ERAs since 1920 (an imputed ERA based on walks, strikeouts and HRs with HRs being adjusted for park effects). Schilling ranked 7th among pitchers with 1500+ IP. The list went up through 2005. He is now over 3000 and my guess is that he has not slid too much since then (he was not on the 3000 list yet after 2005).
I found that he was 44th in Park-Adjusted Pitching Wins Above Replacement Level. I think that was also through 2005. He has probably moved up several slots since then.
So he has some impressive ranks. I would like to have seen a consecutive 3-year period in there that was Cy Young worthy. 2001-02 are in there and 2004 is but he missed alot of 2003, which was a great year for about 160 IP. He still finished 7th in runs saved over average with park adjustments. In 01, 02 and 04 he had 20+ wins, lead league in K/BB ratio, winning pct over .750, 2nd place in runs saved each year. Had a 1st, a 2nd and a 3rd in IP. Not saying he should not be there. But if he had pitched a full season in 2003, then he is a no-brainer. I guess I am 90% for him (or so)
Click on this next ink to see where Schilling ranks in strikeout-to-walk ratio (relative to the league average) for all pitchers with 3,000+ IP. He is 2nd to Mathewson.
K/BB Ratio
I found the best fielding independent ERAs since 1920 (an imputed ERA based on walks, strikeouts and HRs with HRs being adjusted for park effects). Schilling ranked 7th among pitchers with 1500+ IP. The list went up through 2005. He is now over 3000 and my guess is that he has not slid too much since then (he was not on the 3000 list yet after 2005).
I found that he was 44th in Park-Adjusted Pitching Wins Above Replacement Level. I think that was also through 2005. He has probably moved up several slots since then.
So he has some impressive ranks. I would like to have seen a consecutive 3-year period in there that was Cy Young worthy. 2001-02 are in there and 2004 is but he missed alot of 2003, which was a great year for about 160 IP. He still finished 7th in runs saved over average with park adjustments. In 01, 02 and 04 he had 20+ wins, lead league in K/BB ratio, winning pct over .750, 2nd place in runs saved each year. Had a 1st, a 2nd and a 3rd in IP. Not saying he should not be there. But if he had pitched a full season in 2003, then he is a no-brainer. I guess I am 90% for him (or so)
Sunday, March 22, 2009
An All-Time Ranking of Players By Wins Above Replacement Level
I have attemtped to rank players by their value above the replacement level player using Pete Palmer's "Total Player Rating" or TPR (it is now actually called BFW for batting wins + fielding wins). This comes from Palmer's "linear weights" method. For example, here are the run values of various events:
1B: .47
2B: .78
3B: 1.09
HR: 1.4
BB: .33
Outs have a negative value, usually around -.25. For each player he calculates how many batting runs they had. This has a win value (usually around 10 runs per win). Something similar is done with fielding. He figures out how man runs a player saved compared to the league average as a fielder and this is converted to wins. Both batting runs and fielding runs can be negative. If a player had 20 batting runs that would be 2 batting wins. If he also had -10 fielding runs, that would be about -1 fielding wins. So his TPR or BFW would be 1 for that season. A player can have a negative value for a season. So he would be below average.
But that does not always mean he had no value. He might have still been better than the next best player available (the replacement) who might have been more negative. There is no clear consensus on exactly what the replacement level in TPR is.
So I did two lists. One with a TPR or a BFW of -2 per 700 PAs as replacement level and one with -3. I divided each guy's career PAs by 700. Then I multiplied that times 2 or 3. That result got added to his career TPR to get career value over replacement. For example, in a season with 700 PAs and a TPR of 0, the player is still either 2 or 3 wins better than replacement level. A player with a TPR of 6, which puts you in the top 200 seasons all-time, would be 8 or 9 wins better than replacement.
Suppose a player had 7,000 career PAs. So that is 10 full seasons. If he had a career TPR of 20, his wins above replacement would be 34 or 41. Click here to see the all time rankings through 2004 using -2 BFW as the replacement level. I used everyone that had 2,000 PAs through 2004 (with PAs including just ABs and walks). Click here to see the all time rankings through 2004 using -3 BFW as the replacement level. The players were compiled using the annual listings in BFW from Retrosheet.
Most of the top ranked players will not be a big surprise. Bonds is first followed by Ruth. Two guys in the top 40 using either a -2 or -3 TPR for replacement level that are not in the Hall of Fame are Ron Santo and Bobby Grich. Bill Dahlen, too.
1B: .47
2B: .78
3B: 1.09
HR: 1.4
BB: .33
Outs have a negative value, usually around -.25. For each player he calculates how many batting runs they had. This has a win value (usually around 10 runs per win). Something similar is done with fielding. He figures out how man runs a player saved compared to the league average as a fielder and this is converted to wins. Both batting runs and fielding runs can be negative. If a player had 20 batting runs that would be 2 batting wins. If he also had -10 fielding runs, that would be about -1 fielding wins. So his TPR or BFW would be 1 for that season. A player can have a negative value for a season. So he would be below average.
But that does not always mean he had no value. He might have still been better than the next best player available (the replacement) who might have been more negative. There is no clear consensus on exactly what the replacement level in TPR is.
So I did two lists. One with a TPR or a BFW of -2 per 700 PAs as replacement level and one with -3. I divided each guy's career PAs by 700. Then I multiplied that times 2 or 3. That result got added to his career TPR to get career value over replacement. For example, in a season with 700 PAs and a TPR of 0, the player is still either 2 or 3 wins better than replacement level. A player with a TPR of 6, which puts you in the top 200 seasons all-time, would be 8 or 9 wins better than replacement.
Suppose a player had 7,000 career PAs. So that is 10 full seasons. If he had a career TPR of 20, his wins above replacement would be 34 or 41. Click here to see the all time rankings through 2004 using -2 BFW as the replacement level. I used everyone that had 2,000 PAs through 2004 (with PAs including just ABs and walks). Click here to see the all time rankings through 2004 using -3 BFW as the replacement level. The players were compiled using the annual listings in BFW from Retrosheet.
Most of the top ranked players will not be a big surprise. Bonds is first followed by Ruth. Two guys in the top 40 using either a -2 or -3 TPR for replacement level that are not in the Hall of Fame are Ron Santo and Bobby Grich. Bill Dahlen, too.
Friday, March 13, 2009
Is Ryan Howard The New Mickey Vernon? (Or Is His Career Really In Decline?)
Howard has had big drops in his OWP the last 2 years. OWP or offensive winning percentage is a Bill James stat that says what a team's winning percentage would be if it had a lineup of 9 identical players who all hit alike and they gave up an average number of runs. Since I got the data from the Lee Sinins Complete Baseball Encyclopedia, it is park adjusted. Here are Howard's OWP for each of the last 3 seasons with his age in parantheses:
.777 (26)
.675 (27)
.582 (28)
So the declines are .102 and .093. For players who had long careers, such big back to back drops in OWP are somewhat rare, especially for someone under 30. Some stories I read using google news search indicate he is in better shape this spring and is hitting better than usual this pre-season. So maybe he has taken the necessary steps to stop the decline.
To see how unusual his declines are, I used a list of players that I have compiled before. This list includes all players who had 15+ seasons with 400+ plate appearances from age 20-40. Then I found all cases of players having back to back seasons of a drop in OWP of .075 or more. The tables below show all of these cases (more discussion of the tables below).
There are 22 such cases. But in only 3 of them, did the fall in OWP start before the age of 30. Those belong to Robin Yount, Jake Beckley and Mickey Vernon. Beckley's started at age 23 and neither of my biographical encyclopedia's mention anything. Same for Vernon's decline. Yount had shoulder problems during the 1984 and 1985 seasons and actually had surgery twice. But both Beckley and Yount are in the Hall of Fame. So if Howard can end up with 15+ seasons with 400+ plate appearances he has a 2 out of 3 chance of making the Hall of Fame.
If I had limited the study to declines of .093 or more, there were only 8 guys. The only one whose decline started before age 30 was Vernon. So who knew that he had something in common with Howard?
In the tables below, the numbers in red are the decline years. The year before the decline is there for each player for reference. Two guys had 3 straight years that fit the criteria. They were Jimmy Dykes and Willie Keeler. Yount and Honus Wagner had two such streaks. The average age at which the decline started was 33.95. 17 of the 22 cases started at age 33 or older.
There were 91 players with 15+ seasons with 400+ plate appearances and 22 back to back seasons of a drop in OWP of .075 or more. So nearly 25% of the 91 players had these big back to back declines. That makes it look like what has happened to Howard is not that rare. But what is rare is the age at which it has happened to him.
What happened to these guys in the third year? Did they finally rebound? Well, we know that Keeler and Dykes each had declines that fit the criteria in the third year. 5 players did not get 400+ PAs the next year. The average change in the third year, including Keeler and Dykes (who were the only declines), was a positive .078. 8 of the 17 had changes of +.100 or more. So Howard has a good chance to bounce back but also has a chance not to make it to 400 PAs. Mickey Vernon improved .295 in the third year.



.777 (26)
.675 (27)
.582 (28)
So the declines are .102 and .093. For players who had long careers, such big back to back drops in OWP are somewhat rare, especially for someone under 30. Some stories I read using google news search indicate he is in better shape this spring and is hitting better than usual this pre-season. So maybe he has taken the necessary steps to stop the decline.
To see how unusual his declines are, I used a list of players that I have compiled before. This list includes all players who had 15+ seasons with 400+ plate appearances from age 20-40. Then I found all cases of players having back to back seasons of a drop in OWP of .075 or more. The tables below show all of these cases (more discussion of the tables below).
There are 22 such cases. But in only 3 of them, did the fall in OWP start before the age of 30. Those belong to Robin Yount, Jake Beckley and Mickey Vernon. Beckley's started at age 23 and neither of my biographical encyclopedia's mention anything. Same for Vernon's decline. Yount had shoulder problems during the 1984 and 1985 seasons and actually had surgery twice. But both Beckley and Yount are in the Hall of Fame. So if Howard can end up with 15+ seasons with 400+ plate appearances he has a 2 out of 3 chance of making the Hall of Fame.
If I had limited the study to declines of .093 or more, there were only 8 guys. The only one whose decline started before age 30 was Vernon. So who knew that he had something in common with Howard?
In the tables below, the numbers in red are the decline years. The year before the decline is there for each player for reference. Two guys had 3 straight years that fit the criteria. They were Jimmy Dykes and Willie Keeler. Yount and Honus Wagner had two such streaks. The average age at which the decline started was 33.95. 17 of the 22 cases started at age 33 or older.
There were 91 players with 15+ seasons with 400+ plate appearances and 22 back to back seasons of a drop in OWP of .075 or more. So nearly 25% of the 91 players had these big back to back declines. That makes it look like what has happened to Howard is not that rare. But what is rare is the age at which it has happened to him.
What happened to these guys in the third year? Did they finally rebound? Well, we know that Keeler and Dykes each had declines that fit the criteria in the third year. 5 players did not get 400+ PAs the next year. The average change in the third year, including Keeler and Dykes (who were the only declines), was a positive .078. 8 of the 17 had changes of +.100 or more. So Howard has a good chance to bounce back but also has a chance not to make it to 400 PAs. Mickey Vernon improved .295 in the third year.



Sunday, March 8, 2009
Does Jim Bunning Belong In The Hall Of Fame?
This issue came up recently on the SABR list. One of the issues was why were pitchers from his era who seem to have been about as good he was not in. Then someone else mentioned that maybe the Veterans committee put him because he is a Senator. I posted some evidence on this. Basically it was references to research I had done in the past and seeing where Bunning ranked. I will put that post below, but first something new, although it is a simple, rough estimate of his value (I used Fielding Independent Pitching ERA or FIP ERA to find an imputed winning percentage for Bunning which is fairly high-it's all based on how good he was at strikeouts, walks and HRs). I find that there is some evidence for him being in the Hall, but I don't think it is all on his side. He had a great strikeout-to-walk ratio, which is one indicator of how good a pitcher is.
Below are the top 25 pitchers with 3000+ IP in strikeout-to-walk ratio relative to the league average. Bunning is 19th, which is very good. Mathewson is 1st. He had a strikeout-to-walk ratio of 2.96 while the league average was 1.29. Since 2.96/1.29 = 2.30, Mathewson gets a 230. Data came from the Lee Sinins Complete Baseball Encyclopedia.

I calculated his Fielding Independent Pitching ERA or FIP ERA. The idea is that a pitcher controls HRs, BBs and Ks and hits on balls in play not so much (if you have not heard of this, google Voros McCracken).
Here are the key calculations. HRs, BBs and Ks are per 9 IP.
(1) FIP ERA = Constant + 1.44*HR + .33*BB - .22*K
(2) The constant = League ERA - (1.44*HR + .33*BB - .22*K)
I used Bunning's stats and adjusted them to the AL stats of 2008. Bunning's Ks per 9 IP was 6.83 or about 25% above average. In the 2008 AL K/9IP = 6.36. Raising that 25%leaves 7.97. He walked 2.39 batters per 9 IP or about 26% fewer than average. In the 2008 AL BB/9IP = 3.32. Lowering that 26% leaves 2.47. So those numbers will get plugged into equation (1). The league ERA in the AL in 2008 was 4.35 and the constant for equation (1) works out to 3.21.
We still need to calculate his HRs per 9 IP. He actually gave up 372 HRs while the average was 346. So it looks like Bunning did poorly here. But he pitched in Tiger Stadium for part of his career where an above average number of HRs were hit. So I adjusted his HRs allowed in each season based on the HR park factors from the STATS, INC. All-Time Baseball Sourcebook. For example, if Tiger stadium gave up 20% more HRs than average in a season, I reduced his HRs for that year by 10% (only half of the 20% since he only pitched half his games there). In some of his years with the Phillies, the park factor was below average. After doing this for each of his seasons, his HR total came out to 349, or almost exactly average.
In the AL in 2008, there was just about 1 HR per 9 IP. So I used that for equation (1). With the HR, BB and K data done, I found a FIP ERA of 3.73 for Bunning (adjusted for the 2008 AL).
I then calculated what Bill James calls the Pythagorean winning percentage for Bunning if he pitched on an average team. It is
(runs scored squared)/(((runs scored squared) + (runs allowed squared))
For Bunning, adjusted to the 2008 AL, we get
4.35*4.35/(4.35*4.35 + 3.73*3.73) = .577
So Bunning pitching for an average team would have a .577 winning pct. For pitchers with 3000+ IP, he would be tied for 42nd (with Jack Morris). But I did not calculate the FIP ERA or Pythagorean winning percentage for anyone else. I am just assuming if I did it for everyone, just as many guys would move ahead of Bunning as would fall behind. 42nd is pretty good and seems high enough for a starter to make the Hall.
Now for the post to the SABR list.
I found the best fielding independent ERAs since 1920 (an imputed ERA based on walks, strikeouts and HRs with HRs being adjusted for park effects). Bunning ranked 13th among pitchers with 3000+ IP.
I found that he was 51st in Park-Adjusted Pitching Wins Above Replacement Level.
It also looks like he out pitched Koufax in neutral parks while they were both in the NL.
But he only had 257 Win Shares through 2001, tied for 291st. Not sure where that ranked among pitchers.
He ranks 67th in adjusted pitching wins in Pete Palmer's baseball encyclopedia.
Below are the top 25 pitchers with 3000+ IP in strikeout-to-walk ratio relative to the league average. Bunning is 19th, which is very good. Mathewson is 1st. He had a strikeout-to-walk ratio of 2.96 while the league average was 1.29. Since 2.96/1.29 = 2.30, Mathewson gets a 230. Data came from the Lee Sinins Complete Baseball Encyclopedia.

I calculated his Fielding Independent Pitching ERA or FIP ERA. The idea is that a pitcher controls HRs, BBs and Ks and hits on balls in play not so much (if you have not heard of this, google Voros McCracken).
Here are the key calculations. HRs, BBs and Ks are per 9 IP.
(1) FIP ERA = Constant + 1.44*HR + .33*BB - .22*K
(2) The constant = League ERA - (1.44*HR + .33*BB - .22*K)
I used Bunning's stats and adjusted them to the AL stats of 2008. Bunning's Ks per 9 IP was 6.83 or about 25% above average. In the 2008 AL K/9IP = 6.36. Raising that 25%leaves 7.97. He walked 2.39 batters per 9 IP or about 26% fewer than average. In the 2008 AL BB/9IP = 3.32. Lowering that 26% leaves 2.47. So those numbers will get plugged into equation (1). The league ERA in the AL in 2008 was 4.35 and the constant for equation (1) works out to 3.21.
We still need to calculate his HRs per 9 IP. He actually gave up 372 HRs while the average was 346. So it looks like Bunning did poorly here. But he pitched in Tiger Stadium for part of his career where an above average number of HRs were hit. So I adjusted his HRs allowed in each season based on the HR park factors from the STATS, INC. All-Time Baseball Sourcebook. For example, if Tiger stadium gave up 20% more HRs than average in a season, I reduced his HRs for that year by 10% (only half of the 20% since he only pitched half his games there). In some of his years with the Phillies, the park factor was below average. After doing this for each of his seasons, his HR total came out to 349, or almost exactly average.
In the AL in 2008, there was just about 1 HR per 9 IP. So I used that for equation (1). With the HR, BB and K data done, I found a FIP ERA of 3.73 for Bunning (adjusted for the 2008 AL).
I then calculated what Bill James calls the Pythagorean winning percentage for Bunning if he pitched on an average team. It is
(runs scored squared)/(((runs scored squared) + (runs allowed squared))
For Bunning, adjusted to the 2008 AL, we get
4.35*4.35/(4.35*4.35 + 3.73*3.73) = .577
So Bunning pitching for an average team would have a .577 winning pct. For pitchers with 3000+ IP, he would be tied for 42nd (with Jack Morris). But I did not calculate the FIP ERA or Pythagorean winning percentage for anyone else. I am just assuming if I did it for everyone, just as many guys would move ahead of Bunning as would fall behind. 42nd is pretty good and seems high enough for a starter to make the Hall.
Now for the post to the SABR list.
I found the best fielding independent ERAs since 1920 (an imputed ERA based on walks, strikeouts and HRs with HRs being adjusted for park effects). Bunning ranked 13th among pitchers with 3000+ IP.
I found that he was 51st in Park-Adjusted Pitching Wins Above Replacement Level.
It also looks like he out pitched Koufax in neutral parks while they were both in the NL.
But he only had 257 Win Shares through 2001, tied for 291st. Not sure where that ranked among pitchers.
He ranks 67th in adjusted pitching wins in Pete Palmer's baseball encyclopedia.
Saturday, February 28, 2009
Arbitration Wrap-up – 2009 (A Guest Post By Bill Gilbert)
Bill Gilbert has been involved in arbitration hearings and is the president of the South Texas chapter of SABR (aka the Rogers Hornsby chapter)
In 2009, 111 players filed for salary arbitration. Before players and clubs exchanged figures on January 20, sixty five of these players had agreed to contracts with their clubs. Of the remaining 46 players only three players actually went to an arbitration hearing, tying the low point set in 2005.

(editors note: the third column shows what the player wanted in $1,000s and the next column shows what the club offered)
It was the first time that the majority of the decisions went in favor of the players since 1996. Since the first hearings were held in 1974, the clubs have won on 280 occasions and the players have prevailed 207 times.
By my count, here is the breakdown of the 111 cases.
93 players signed one-year contracts
15 players signed multi-year contracts.
3 players had their salary determined at an arbitration hearing.
There are two situations where arbitration can come into play. By far the most common is the one involving players, under control of their clubs, with 3 to 6 years of major league service (MLS), plus the 17% most senior MLS-2 players, referred to as “super twos”. Of the 111 players who filed this year, 109 were in this category.
The other situation involves free agents. When a player with 6 or more years of major league service files for free agency, his club has the option of offering arbitration. A club must offer arbitration in order to get compensation in the form of draft picks if the player signs with another club. If the player accepts arbitration, he is no longer considered a free agent and he becomes bound to that club. If a player refuses arbitration, as most players do, he is a free agent who can sign with any club including the one he played for last year. Of the 24 free agents who were offered arbitration this year, the only two that accepted were Darren Oliver of the Los Angeles Angels and David Weathers of Cincinnati.
With the economic uncertainties this year, the market was more difficult to read. A number of free agents, such as Jason Varitek of Boston, Orlando Hudson of Arizona, Orlando Cabrera of the Chicago White Sox and Jon Garland of the Los Angeles Angels would have fared much better if they had accepted arbitration. Clubs also had difficult decisions to make and recognized that in a declining market, some players would likely be paid much more in the arbitration process than their market value. This led to non-tendering arbitration eligible players like Ty Wigginton of Houston, Willy Taveras of Colorado, Takashi Saito of the Los Angeles Dodgers and Tim Redding of Washington.
The arbitration process is designed to promote a settlement at a salary in line with that of other players with comparable performance and service time. Players eligible for arbitration for the first time receive a large increase in salary since they have no leverage in their pre-arbitration years when their salaries are under control of the clubs. Players who have been through the process before also generally receive salary increases depending on their performance in the preceding year.
Players that settle prior to hearings are frequently able to include performance bonuses, based on playing time, and awards bonuses in their contracts.
The big winners in the arbitration process this year were Nick Markakis of Baltimore and Ryan Howard of Philadelphia. Markakis, in his first year of arbitration eligibility, signed a six year contract for $66.1 million, which carries him through 3 years of arbitration eligibility and 3 years of free agency. Howard, who won at a hearing in 2008, received a three year contract for $54 million which takes him through his arbitration years.
The relatively quiet arbitration season this year suggest that the system is working as designed in achieving benefits for both sides. Players with 3 to 6 years of major league service receive salaries that are influenced by their market value and Clubs are able to retain the rights to these players through 6 years of major league service before they become eligible for free agency.
In 2009, 111 players filed for salary arbitration. Before players and clubs exchanged figures on January 20, sixty five of these players had agreed to contracts with their clubs. Of the remaining 46 players only three players actually went to an arbitration hearing, tying the low point set in 2005.

(editors note: the third column shows what the player wanted in $1,000s and the next column shows what the club offered)
It was the first time that the majority of the decisions went in favor of the players since 1996. Since the first hearings were held in 1974, the clubs have won on 280 occasions and the players have prevailed 207 times.
By my count, here is the breakdown of the 111 cases.
93 players signed one-year contracts
15 players signed multi-year contracts.
3 players had their salary determined at an arbitration hearing.
There are two situations where arbitration can come into play. By far the most common is the one involving players, under control of their clubs, with 3 to 6 years of major league service (MLS), plus the 17% most senior MLS-2 players, referred to as “super twos”. Of the 111 players who filed this year, 109 were in this category.
The other situation involves free agents. When a player with 6 or more years of major league service files for free agency, his club has the option of offering arbitration. A club must offer arbitration in order to get compensation in the form of draft picks if the player signs with another club. If the player accepts arbitration, he is no longer considered a free agent and he becomes bound to that club. If a player refuses arbitration, as most players do, he is a free agent who can sign with any club including the one he played for last year. Of the 24 free agents who were offered arbitration this year, the only two that accepted were Darren Oliver of the Los Angeles Angels and David Weathers of Cincinnati.
With the economic uncertainties this year, the market was more difficult to read. A number of free agents, such as Jason Varitek of Boston, Orlando Hudson of Arizona, Orlando Cabrera of the Chicago White Sox and Jon Garland of the Los Angeles Angels would have fared much better if they had accepted arbitration. Clubs also had difficult decisions to make and recognized that in a declining market, some players would likely be paid much more in the arbitration process than their market value. This led to non-tendering arbitration eligible players like Ty Wigginton of Houston, Willy Taveras of Colorado, Takashi Saito of the Los Angeles Dodgers and Tim Redding of Washington.
The arbitration process is designed to promote a settlement at a salary in line with that of other players with comparable performance and service time. Players eligible for arbitration for the first time receive a large increase in salary since they have no leverage in their pre-arbitration years when their salaries are under control of the clubs. Players who have been through the process before also generally receive salary increases depending on their performance in the preceding year.
Players that settle prior to hearings are frequently able to include performance bonuses, based on playing time, and awards bonuses in their contracts.
The big winners in the arbitration process this year were Nick Markakis of Baltimore and Ryan Howard of Philadelphia. Markakis, in his first year of arbitration eligibility, signed a six year contract for $66.1 million, which carries him through 3 years of arbitration eligibility and 3 years of free agency. Howard, who won at a hearing in 2008, received a three year contract for $54 million which takes him through his arbitration years.
The relatively quiet arbitration season this year suggest that the system is working as designed in achieving benefits for both sides. Players with 3 to 6 years of major league service receive salaries that are influenced by their market value and Clubs are able to retain the rights to these players through 6 years of major league service before they become eligible for free agency.
Sunday, February 22, 2009
How Good Has Albert Pujols Been And How Good Will He Be?
This issue came up recently on the listserv of the Hornsby (or south Texas) chapter of SABR. Bill Gilbert wrote a report on Pujols, pointing out that he is now 4th all-time in SLG and 5th in OPS. But as he gets older, those ranks might slip. This got me thinking about how much a player might slip in percentage rankings since they might not be as good as they age.
So I looked at where some players (probably not a very scientifically selected group) ranked at early and late stages of their careers. I tried to find periods that parallel Pujols so far. But I was not always able to. In the table below, I either divided a players career roughly in half, or did the first 8 years and next 8 years 9since Pujols has 8 years so far). Then I also simply used when a guy really started to dropoff as a dividing line. The last 10 lines are all power hitting 1B men. Those guys I found by getting the top 10 all-time in SLG relative to the league average with 5000+ PAs using the Lee Sinins Complete Baseball Encyclopedia. That should give a group that is like Pujols (Greenberg and Mize don't appear due to WW II gaps).
Then I found each guys offensive winning percentage (OWP) for the given period. OWP is a Bill James stat that says what a team's winning percentage would be if it had a lineup of 9 identical players who all hit alike and they gave up an average number of runs. Since I got the data from the Lee Sinins Complete Baseball Encyclopedia, it is park adjusted. I also found where he ranked all-time for the stated years or up to a certain year. The normal PA minimum was 5000. But if a guy had, say, 4900 PAs in his first 8 years (or whatever the time periods was), I used that for both periods shown.

There is quite a variety of outcomes. Some guys fall quite a bit in the rankings due to a much lower performance in their 2nd half. Some actually did better and rose in the ranks. So based on this, where Pujols ends up is not clear.
I also took all the players who had 5000+ PAs before the age of 28 who also had at least 2500 PAs from ages 29-36. Of the 75 players in the first group, 49 made it into the2nd group. Only 16 of the 49 had a higher OWP from 29-36 than they did up to age 28. Roberto Clemente was the one real big gainer, .162 (from .531 to .723). The average change was a loss of .030. Only 8 of the 75 guys from the first group had 5000+ PAs from age 29-36. It seems like the chances Pujols will even get 5000+ PAs over the next 8 years is low. But there might be something I am missing here.
The other thing I did was to look at the normal performance trajectory as players age. I found all the players who had 15+ seasons with 400+ PAs up through 2005. Then I found the average OWP for each age from 20-40. The graph of that is below.

It looks like the ages 21-28 are symetric with the next 8 years. So the overall OWP is the same in each period. If that happens for Pujols, then he will not change much. Without getting into the details, I calculated his OWP will be .768 over the next 8 years based on what he has done and what the historical trends are. I project that if he plays until 40, he will end up with about a .755 career OWP, staying 10th (Musial is 11th at .752). For what is probably a more scientific treatment of aging and performance in baseball, see PEAK ATHLETIC PERFORMANCE AND AGEING: EVIDENCE FROM BASEBALL by J.C. BRADBURY.
Below is the current top 10 in career OWP with 5000+ PAs
1 Babe Ruth .852
2 Ted Williams .832
3 Barry Bonds .810
4 Mickey Mantle .801
5 Lou Gehrig .797
6 Rogers Hornsby .787
7 Ty Cobb .781
8 Joe Jackson .780
9 Dan Brouthers .770
10 Albert Pujols .769
One last thing about Pujols. His career really does not have the kind of rising arc that the historical trend shows. So we can't be sure how any of this applies to him. Here is the chart of his OWP by age:

If we take his .827 at age 28 and then project each year forward using the changes in the typical trend, he would get .830 at age 29, then starting at age 30 and going on through age 40, he would get
0.820
0.820
0.824
0.791
0.791
0.787
0.783
0.755
0.749
0.748
0.712
Roughly he will have an OWP of .784 over the the rest of his career.
So I looked at where some players (probably not a very scientifically selected group) ranked at early and late stages of their careers. I tried to find periods that parallel Pujols so far. But I was not always able to. In the table below, I either divided a players career roughly in half, or did the first 8 years and next 8 years 9since Pujols has 8 years so far). Then I also simply used when a guy really started to dropoff as a dividing line. The last 10 lines are all power hitting 1B men. Those guys I found by getting the top 10 all-time in SLG relative to the league average with 5000+ PAs using the Lee Sinins Complete Baseball Encyclopedia. That should give a group that is like Pujols (Greenberg and Mize don't appear due to WW II gaps).
Then I found each guys offensive winning percentage (OWP) for the given period. OWP is a Bill James stat that says what a team's winning percentage would be if it had a lineup of 9 identical players who all hit alike and they gave up an average number of runs. Since I got the data from the Lee Sinins Complete Baseball Encyclopedia, it is park adjusted. I also found where he ranked all-time for the stated years or up to a certain year. The normal PA minimum was 5000. But if a guy had, say, 4900 PAs in his first 8 years (or whatever the time periods was), I used that for both periods shown.

There is quite a variety of outcomes. Some guys fall quite a bit in the rankings due to a much lower performance in their 2nd half. Some actually did better and rose in the ranks. So based on this, where Pujols ends up is not clear.
I also took all the players who had 5000+ PAs before the age of 28 who also had at least 2500 PAs from ages 29-36. Of the 75 players in the first group, 49 made it into the2nd group. Only 16 of the 49 had a higher OWP from 29-36 than they did up to age 28. Roberto Clemente was the one real big gainer, .162 (from .531 to .723). The average change was a loss of .030. Only 8 of the 75 guys from the first group had 5000+ PAs from age 29-36. It seems like the chances Pujols will even get 5000+ PAs over the next 8 years is low. But there might be something I am missing here.
The other thing I did was to look at the normal performance trajectory as players age. I found all the players who had 15+ seasons with 400+ PAs up through 2005. Then I found the average OWP for each age from 20-40. The graph of that is below.

It looks like the ages 21-28 are symetric with the next 8 years. So the overall OWP is the same in each period. If that happens for Pujols, then he will not change much. Without getting into the details, I calculated his OWP will be .768 over the next 8 years based on what he has done and what the historical trends are. I project that if he plays until 40, he will end up with about a .755 career OWP, staying 10th (Musial is 11th at .752). For what is probably a more scientific treatment of aging and performance in baseball, see PEAK ATHLETIC PERFORMANCE AND AGEING: EVIDENCE FROM BASEBALL by J.C. BRADBURY.
Below is the current top 10 in career OWP with 5000+ PAs
1 Babe Ruth .852
2 Ted Williams .832
3 Barry Bonds .810
4 Mickey Mantle .801
5 Lou Gehrig .797
6 Rogers Hornsby .787
7 Ty Cobb .781
8 Joe Jackson .780
9 Dan Brouthers .770
10 Albert Pujols .769
One last thing about Pujols. His career really does not have the kind of rising arc that the historical trend shows. So we can't be sure how any of this applies to him. Here is the chart of his OWP by age:

If we take his .827 at age 28 and then project each year forward using the changes in the typical trend, he would get .830 at age 29, then starting at age 30 and going on through age 40, he would get
0.820
0.820
0.824
0.791
0.791
0.787
0.783
0.755
0.749
0.748
0.712
Roughly he will have an OWP of .784 over the the rest of his career.
Sunday, February 15, 2009
Which Pitchers Improved The Most In 2008?
I used two different stats to determine this. The first was an imputed value for ERA based on HRs allowed, walks and strikeouts (sort of a poor man's DIPS ERA). The other was RSAA from the Lee Sinins Complete Baseball Encyclopedia. "RSAA--Runs saved against average. It's the amount of runs that a pitcher saved vs. what an average pitcher would have allowed." It is park adjusted.
For the first measure of imputed ERA, I ran a regression using all pitchers who had 100+ IP in either of the last two seasons. There were 284 cases. The resulting regression equation was
ERA = 3.16 + 1.37*HR + .365*BB - .226*SO
Those are all per 9 IP. Walks include HBP. Then I took only pitchers who had 100+ IP in both seasons, found an imputed ERA for each of them in each season, and then found their change. They were then ranked from lowest to highest. Lowest would be most negative, so those are the ones that improved the most. Here are the ten best:

Now the ten who declined the most.

Now the ten best using RSAA. It is on a per 9 IP basis.

So Sanatana allowed 1.08 runs more than average per 9 IP in 2007 and allowed .95 less than average in 2008. So that is a swing of 2.03, which was the best improvement.
Now for the ten who declined the most.

Gorzelany really had a miserable year, being worst in both methods. The first method only takes into account what the pitchers did (but it is not park adjusted). The second method is park adjusted but is not solely determined by the pitcher. Anyone who made both lists either really did alot better or alot worse than the year before. Interesting that Mussina made the best by the RSAA method. He was also 14th by the imputed ERA method. There were 93 pitchers in all. The correlation between the change calculated by the two methods is -.67 (that negative makes sense since a positive RSAA is good). The more runs you save, the more your ERA will be below average.
For the first measure of imputed ERA, I ran a regression using all pitchers who had 100+ IP in either of the last two seasons. There were 284 cases. The resulting regression equation was
ERA = 3.16 + 1.37*HR + .365*BB - .226*SO
Those are all per 9 IP. Walks include HBP. Then I took only pitchers who had 100+ IP in both seasons, found an imputed ERA for each of them in each season, and then found their change. They were then ranked from lowest to highest. Lowest would be most negative, so those are the ones that improved the most. Here are the ten best:

Now the ten who declined the most.

Now the ten best using RSAA. It is on a per 9 IP basis.

So Sanatana allowed 1.08 runs more than average per 9 IP in 2007 and allowed .95 less than average in 2008. So that is a swing of 2.03, which was the best improvement.
Now for the ten who declined the most.

Gorzelany really had a miserable year, being worst in both methods. The first method only takes into account what the pitchers did (but it is not park adjusted). The second method is park adjusted but is not solely determined by the pitcher. Anyone who made both lists either really did alot better or alot worse than the year before. Interesting that Mussina made the best by the RSAA method. He was also 14th by the imputed ERA method. There were 93 pitchers in all. The correlation between the change calculated by the two methods is -.67 (that negative makes sense since a positive RSAA is good). The more runs you save, the more your ERA will be below average.
Sunday, February 8, 2009
Positional Hitting Over Time (Part 2)
Part 1 was a few weeks ago (you can scroll down to see it). I looked at the slugging percentage (SLG) divided by the league average for all 8 every day fielding positions. Here I look at how many players in each decade were among the top 100 or 200 seasons at each position in offensive winning percentage (OWP). OWP is a Bill James stat that says what a team's winning percentage would be if it had a lineup of 9 identical players who all hit alike and they gave up an average number of runs. Since I got the data from the Lee Sinins Complete Baseball Encyclopedia, it is park adjusted. The PA minimum was 400.
The first table is the top 100. The second table has the top 200. After the tables is a little discussion then there are two more tables. In those tables I adjust the figures to account for the different number of teams in baseball at different times.


One general comment is that using the Sinins database, a player is listed at the position he played the most. Jimmy Dykes is a SS in a year he only plaed 60 games there. It might be better to only use seasons with 100+ games at a position. Maybe if I get time someday I will do that. Then guys like Dykes could be put into the utiltiy category. Musial had 3 at 1B, 1 in LF, 1 in CF and 4 in RF. He gets lost in the shuffle and deserve to be remembered because he is Polish.
1B-In the teens, the only one in the top 100 was Jack Fournier in 1915. As you might guess, Gehrig(7) and Foxx (6) dominate in the 1930s. Greenberg and Mize had 2 each. In the 1950s, the only one is Musial (he also appears as CFer in 1952). Thomas (6), McGwire (5) and Bagwell (3) are the big names that caused the surge in the 1990s.
2B-The first three decades are dominated by guys like Lajoie (9), Collins (10) and Hornsby (9). Hornsby had another in 1931 (and 2 at 3B in the teens). Then there is a big drought in the 1940s through the 1960s. Joe Morgan (6) and Rod Carew (3) are the big names in the 1970s. But Mike Andrews has one, too! In the 1990s, Alomar and Biggio each had 4.
SS-Honus Wagner has 9 in the first decade (and 2 in the teens plus 1 in RF in decade 1). Wagner has 9 of the top ten all-time. The only 2 in the 1920s were Dykes and Joe Sewell. Arky Vaughn has 6 in the 1930s (plus 2 in the 1940s). Boudreau leads with 4 in the 1940s. In the 1980s it was Trammell (4), Ripken (3) and Yount (3). Larkin had 5 in the 1990s. AROD has 2 in the 1990s and 6 in the 2000s (plus 3 at 3B). Jeter has 2 in each decade.
3B-In the teens, Baker has 4. But Hornsby has 2 more. The drought from the 1920s-1940 is incredible (maybe defense was considered more important). One of the few is actually Mel Ott in 1938 (he played 113 games at 3B). Mathews has 6 in the 1950s (and 2 in the 1960s), with Rosen getting 3. Minnine Minoso got 1! (68 games, more than either LF (44) or RF (42)). Dick Allen leads the 1960s with 4, Santo had 3. Schmidt had 3 in the 1970s and 5 in the 1980s. Brett was 2 & 3. Boggs had 5 in the 1980s (and 1 in the 1990s). Randy Ready had 1 in the 1980s. Chipper Jones had 5 in the 2000s and 2 in the 1990s. AROD has 3 in the 2000s.
LF-Not counting Bonds, Ruth and Ted Williams, Rickey Henderson has the highest ever for a LFer (.859, 1990). The only two from the teens are Sherry Magee and Ruth, who has 3 in the 1920s. Williams has 7 in the 1940s (he missed 3 years in the military, remember!) and 7 in the 1950s (missing 2 years to the military and in 1959 he had 331 PAs (although OWP = .555)). Keller had 4. 2 of the 5 in the 1960s are Carl Yastrzemski. Boog Powell is one! But intersting how the 1960-80s are light. Bonds has 9 in the 1990s and 6 in the 2000s.
CF-Besides Mantle, Cobb, Speaker and DiMaggio, the best ever was by Cy Seymour (.825, 1905). Cobb has 10 in the teens (and 3 in RF in decade 1) and Speaker has 7. DiMaggio has 4 in the 1940s and 2 in the 1930s. Maybe it is no surprise how incredible the 1950s are. Mantle had 8 and Mays had 5 (missing 2 years). Snider and Doby each had 3. Then there is the Musial season and one for Tito Francona. Mays had 6 in the 1960s while Mantle had 4. Aaron and Kaline each had 1. Then we have quite another drought. Griffey had 4 of those all in the 1990s. I can't imagine any reason for this. Has defense become more important for CF?
RF-In the first decade, Cobb and Flick each had 3. Honus Wagner had 1, too. Joe Jackson has 3 in the teens. Ruth has 6 in the 1920s (and 4 in the 1930s). Heilman had 5. Ott has 5 in the 1930s (and 1 in the 1920s and 2 in the 1940s). Musial has 4 in the 1940s. Aaron had 1 in the 1950s and 2 in the 1960s. Maybe he does not have more since he was so consistent. Frank Robinson had 4 in the 1960s. Only 2 guys since 1970 have as many as 3. Reggie Jackson and Shefield, 3 each.
C-Not many before 1920. Bresnahan had all 3 in the first decade. Dickey had 5 in the 1930s and Cochrane had 4. Hartnett had 2 in the 1920s and 2 in the 1930s. The only 1 from the 1940s was Lombardi and that was in a war year, 1945. Berra had 5 in the 1950s and Campanella had 3. Bench, Simmons and Tenace all had 3 each in the 1970s. Fisk had 2. Simmons has 1 more in the 1980s. Piazza had 6 in the 1990s plus 2 more in the 2000s.
For the tables below, I divided the absolute total for each position in each decade by the number of teams in MLB that decade. That got multiplied by 30. The first one has the top 100. The second has the top 200.

The first table is the top 100. The second table has the top 200. After the tables is a little discussion then there are two more tables. In those tables I adjust the figures to account for the different number of teams in baseball at different times.


One general comment is that using the Sinins database, a player is listed at the position he played the most. Jimmy Dykes is a SS in a year he only plaed 60 games there. It might be better to only use seasons with 100+ games at a position. Maybe if I get time someday I will do that. Then guys like Dykes could be put into the utiltiy category. Musial had 3 at 1B, 1 in LF, 1 in CF and 4 in RF. He gets lost in the shuffle and deserve to be remembered because he is Polish.
1B-In the teens, the only one in the top 100 was Jack Fournier in 1915. As you might guess, Gehrig(7) and Foxx (6) dominate in the 1930s. Greenberg and Mize had 2 each. In the 1950s, the only one is Musial (he also appears as CFer in 1952). Thomas (6), McGwire (5) and Bagwell (3) are the big names that caused the surge in the 1990s.
2B-The first three decades are dominated by guys like Lajoie (9), Collins (10) and Hornsby (9). Hornsby had another in 1931 (and 2 at 3B in the teens). Then there is a big drought in the 1940s through the 1960s. Joe Morgan (6) and Rod Carew (3) are the big names in the 1970s. But Mike Andrews has one, too! In the 1990s, Alomar and Biggio each had 4.
SS-Honus Wagner has 9 in the first decade (and 2 in the teens plus 1 in RF in decade 1). Wagner has 9 of the top ten all-time. The only 2 in the 1920s were Dykes and Joe Sewell. Arky Vaughn has 6 in the 1930s (plus 2 in the 1940s). Boudreau leads with 4 in the 1940s. In the 1980s it was Trammell (4), Ripken (3) and Yount (3). Larkin had 5 in the 1990s. AROD has 2 in the 1990s and 6 in the 2000s (plus 3 at 3B). Jeter has 2 in each decade.
3B-In the teens, Baker has 4. But Hornsby has 2 more. The drought from the 1920s-1940 is incredible (maybe defense was considered more important). One of the few is actually Mel Ott in 1938 (he played 113 games at 3B). Mathews has 6 in the 1950s (and 2 in the 1960s), with Rosen getting 3. Minnine Minoso got 1! (68 games, more than either LF (44) or RF (42)). Dick Allen leads the 1960s with 4, Santo had 3. Schmidt had 3 in the 1970s and 5 in the 1980s. Brett was 2 & 3. Boggs had 5 in the 1980s (and 1 in the 1990s). Randy Ready had 1 in the 1980s. Chipper Jones had 5 in the 2000s and 2 in the 1990s. AROD has 3 in the 2000s.
LF-Not counting Bonds, Ruth and Ted Williams, Rickey Henderson has the highest ever for a LFer (.859, 1990). The only two from the teens are Sherry Magee and Ruth, who has 3 in the 1920s. Williams has 7 in the 1940s (he missed 3 years in the military, remember!) and 7 in the 1950s (missing 2 years to the military and in 1959 he had 331 PAs (although OWP = .555)). Keller had 4. 2 of the 5 in the 1960s are Carl Yastrzemski. Boog Powell is one! But intersting how the 1960-80s are light. Bonds has 9 in the 1990s and 6 in the 2000s.
CF-Besides Mantle, Cobb, Speaker and DiMaggio, the best ever was by Cy Seymour (.825, 1905). Cobb has 10 in the teens (and 3 in RF in decade 1) and Speaker has 7. DiMaggio has 4 in the 1940s and 2 in the 1930s. Maybe it is no surprise how incredible the 1950s are. Mantle had 8 and Mays had 5 (missing 2 years). Snider and Doby each had 3. Then there is the Musial season and one for Tito Francona. Mays had 6 in the 1960s while Mantle had 4. Aaron and Kaline each had 1. Then we have quite another drought. Griffey had 4 of those all in the 1990s. I can't imagine any reason for this. Has defense become more important for CF?
RF-In the first decade, Cobb and Flick each had 3. Honus Wagner had 1, too. Joe Jackson has 3 in the teens. Ruth has 6 in the 1920s (and 4 in the 1930s). Heilman had 5. Ott has 5 in the 1930s (and 1 in the 1920s and 2 in the 1940s). Musial has 4 in the 1940s. Aaron had 1 in the 1950s and 2 in the 1960s. Maybe he does not have more since he was so consistent. Frank Robinson had 4 in the 1960s. Only 2 guys since 1970 have as many as 3. Reggie Jackson and Shefield, 3 each.
C-Not many before 1920. Bresnahan had all 3 in the first decade. Dickey had 5 in the 1930s and Cochrane had 4. Hartnett had 2 in the 1920s and 2 in the 1930s. The only 1 from the 1940s was Lombardi and that was in a war year, 1945. Berra had 5 in the 1950s and Campanella had 3. Bench, Simmons and Tenace all had 3 each in the 1970s. Fisk had 2. Simmons has 1 more in the 1980s. Piazza had 6 in the 1990s plus 2 more in the 2000s.
For the tables below, I divided the absolute total for each position in each decade by the number of teams in MLB that decade. That got multiplied by 30. The first one has the top 100. The second has the top 200.

Sunday, February 1, 2009
MVP Awards And Award Shares By Position
I used the baseball writers award given since 1931. I added up all the MVP awards and shares by regular postions. An award share is figured by dividing the points he got by the maximum possible points. The first I ever saw of this was by Bill James back in the 1980s. If you came in 2nd, but your points added up to 25% of the max (if you got all first place votes), you get a .25 share. Right now, the system is 14 for a first place vote, 9 for second and so on. It may have been different in earlier years. The last paragraph has more technical notes. Anyway, here are the awards by position since 1931
1B 26.1
2B 10
3B 14.35
SS 15
LF 21.67
CF 14.21
RF 20.09
C 13.74
DH 1.83
Now for the shares. BR lists the top 200 in MVP vote shares. But they include the different awards from before 1931. I removed any shares from those cases. Here how the positions ranked
1B 78.89
RF 68.30
LF 56.65
3B 34.82
SS 34.71
CF 33.18
C 23.86
2B 22.72
DH 9.30
Now some of the guys actually had quite a bit of their shares from before 1931 and only a little after. So I removed anyone who had any shares from before 1931 (so even now their post 1931 data does not count).
1B 67.68
RF 61.66
LF 52.88
3B 34.77
SS 33.72
CF 32.92
2B 21.48
C 20.05
DH 9.30
Which ever way I do it, it seems that the writers like to reward 1B men, RFers and LFers and don't like to reward 2B men and catchers. Maybe things would look better for 2B men if I had included pre 1931 info. There were 2B men like Hornsby, Lajoie, Frisch and Collins. But I was mainly interested in looking at who the writers like.
If a player split time between two or more positions, I divided up the award or share proportionately. If he played 25% of the time at one position and 75% at another, he got .25 for one and .75 for the other. I did not count time at any position that was less than 10% of the total. I used Baseball Reference for the data. I used innings played where possible and games other wise. If games added up to more than 154 or 162, I just had to suppose that the percentages still held. Players do switch between positions during the game. When innings are not known, it can add up to more than 154. For DH cases, I also used games. DH is never listed by innings played, just games.
1B 26.1
2B 10
3B 14.35
SS 15
LF 21.67
CF 14.21
RF 20.09
C 13.74
DH 1.83
Now for the shares. BR lists the top 200 in MVP vote shares. But they include the different awards from before 1931. I removed any shares from those cases. Here how the positions ranked
1B 78.89
RF 68.30
LF 56.65
3B 34.82
SS 34.71
CF 33.18
C 23.86
2B 22.72
DH 9.30
Now some of the guys actually had quite a bit of their shares from before 1931 and only a little after. So I removed anyone who had any shares from before 1931 (so even now their post 1931 data does not count).
1B 67.68
RF 61.66
LF 52.88
3B 34.77
SS 33.72
CF 32.92
2B 21.48
C 20.05
DH 9.30
Which ever way I do it, it seems that the writers like to reward 1B men, RFers and LFers and don't like to reward 2B men and catchers. Maybe things would look better for 2B men if I had included pre 1931 info. There were 2B men like Hornsby, Lajoie, Frisch and Collins. But I was mainly interested in looking at who the writers like.
If a player split time between two or more positions, I divided up the award or share proportionately. If he played 25% of the time at one position and 75% at another, he got .25 for one and .75 for the other. I did not count time at any position that was less than 10% of the total. I used Baseball Reference for the data. I used innings played where possible and games other wise. If games added up to more than 154 or 162, I just had to suppose that the percentages still held. Players do switch between positions during the game. When innings are not known, it can add up to more than 154. For DH cases, I also used games. DH is never listed by innings played, just games.
Sunday, January 25, 2009
Which Hitters Improved The Most In 2008?
The stat I used for this was "offensive winning percentage" or OWP. It is a Bill James stat that says what a team's winning percentage would be if it had a lineup of 9 identical players who all hit alike and they gave up an average number of runs. Since I got the data from the Lee Sinins Complete Baseball Encyclopedia, it is park adjusted. I included all players who had at least 300 plate appearances in both 2007 and 2008. The top 25 are below:

I was surprised to see so many players aged 30 or more (11) plus 5 more aged 29. I thought that it would be younger players who improved. Maybe the older guys fluctuate alot more so bigger improvements are possible. But the leader and the #6 guy were both 36. One guy was even 38. The number of players aged 33 or more equalled the number aged 24 or less. The next table shows how much these guys improved in more conventional stats.

I was surprised to see so many players aged 30 or more (11) plus 5 more aged 29. I thought that it would be younger players who improved. Maybe the older guys fluctuate alot more so bigger improvements are possible. But the leader and the #6 guy were both 36. One guy was even 38. The number of players aged 33 or more equalled the number aged 24 or less. The next table shows how much these guys improved in more conventional stats.
Sunday, January 18, 2009
Jim Rice vs. Jose Cruz
I thought this might make an interesting comparison since I belong to a chapter of SABR in Texas (the Austin one or Hornsby chapter). The point is not that Rice does or does not belong in the Hall of Fame, just that Cruz compares so well. Cruz got 2 votes in 1994. Those are the only votes he has ever gotten.
Career PA
Rice-9058
Cruz-8931
Career Offensive Winning Percentage
Rice-.593
Cruz-.611
Highest 3 year OWP
Rice-.698 (1977-79)
Cruz-.687(1983-85)
Full seasons with .700 OWP or better
Rice-2
Cruz-3
Full season means 400+ PAs.
Full seasons with .600 OWP or better
Rice-5
Cruz-8
Cruz had an additional season with 346 PA
Career Win Shares per 648 PA
Rice-20.17
Cruz-22.71
Since it takes about 3 WS to make 1 win in Bill James' system, Cruz was worth .85 more wins per season. See
http://us.share.geocities.com/cyrilmorong@sbcglobal.net/WSperPA.htm
Career Win Shares
Rice-282
Cruz-313
Seasons with 20+ WS (all-star type seasons)
Rice-7
Cruz-8
Seasons with 30+ WS (MVP type seasons)
Rice-1
Cruz-1
Best 3 Consecutive years in WS
Rice-90 (1977-79, 26-36-28)
Cruz-80 (1983-85, 30-29-21)
I have also attemtped to rank players by their value above replacement. I did two lists. One with a TPR or a BFW (from Pete Palmer) of -2 per 700 PAs as replacement level and one with -3. I divided each guy's career PAs by 700. Then I multiplied that times 2 or 3. That result got added to his career TPR to get career value over replacement. Here is the all-time ranking through 2004.
http://www.geocities.com/cyrilmorong@sbcglobal.net/REP.htm
For VAR using -2 TPR per season
Rice-44.8
Cruz-46.72
For VAR using -3 TPR per season
Rice-57.42
Cruz-59.48
Best 3 Consecutive years in TPR or BFW
Rice-10.3 (1977-79, 3.0-4.2-3.1)
Cruz-7.7 (1983-85, 2.8-3.6-1.3)
MVP award shares
Rice-3.15 (tied for 29th, 6 top 5 finishes)
Cruz-.96 (248th, Al Oliver is higher with 1.25, only 1 top 5 finish)
Career PA
Rice-9058
Cruz-8931
Career Offensive Winning Percentage
Rice-.593
Cruz-.611
Highest 3 year OWP
Rice-.698 (1977-79)
Cruz-.687(1983-85)
Full seasons with .700 OWP or better
Rice-2
Cruz-3
Full season means 400+ PAs.
Full seasons with .600 OWP or better
Rice-5
Cruz-8
Cruz had an additional season with 346 PA
Career Win Shares per 648 PA
Rice-20.17
Cruz-22.71
Since it takes about 3 WS to make 1 win in Bill James' system, Cruz was worth .85 more wins per season. See
http://us.share.geocities.com/cyrilmorong@sbcglobal.net/WSperPA.htm
Career Win Shares
Rice-282
Cruz-313
Seasons with 20+ WS (all-star type seasons)
Rice-7
Cruz-8
Seasons with 30+ WS (MVP type seasons)
Rice-1
Cruz-1
Best 3 Consecutive years in WS
Rice-90 (1977-79, 26-36-28)
Cruz-80 (1983-85, 30-29-21)
I have also attemtped to rank players by their value above replacement. I did two lists. One with a TPR or a BFW (from Pete Palmer) of -2 per 700 PAs as replacement level and one with -3. I divided each guy's career PAs by 700. Then I multiplied that times 2 or 3. That result got added to his career TPR to get career value over replacement. Here is the all-time ranking through 2004.
http://www.geocities.com/cyrilmorong@sbcglobal.net/REP.htm
For VAR using -2 TPR per season
Rice-44.8
Cruz-46.72
For VAR using -3 TPR per season
Rice-57.42
Cruz-59.48
Best 3 Consecutive years in TPR or BFW
Rice-10.3 (1977-79, 3.0-4.2-3.1)
Cruz-7.7 (1983-85, 2.8-3.6-1.3)
MVP award shares
Rice-3.15 (tied for 29th, 6 top 5 finishes)
Cruz-.96 (248th, Al Oliver is higher with 1.25, only 1 top 5 finish)
Monday, January 12, 2009
Positional Hitting Over Time
There are four graphs below. Each one shows the slugging percentage (SLG) divided by the league average for all 8 every day fielding positions. Each data point is a five year average. The first two graphs are the AL and the next two are the NL.
Some interesting trends:
-Shortstops have been rising quite a bit in both leagues since the 1970s.
-2B men started declining in the AL in the 1940s, then starting rising again in the 1960s. But even now they have not reached their earlier peak. They started declining in the 1920s in the NL but then started back up in the late 1950s.
-3B men started to rise in the AL in the 1920s and in the 1930s in the NL. But they have tailed off in the AL since 1980.
-1B men had a big spike in the AL in the 1920s and 1930s. In both leagues, in general, they have been high but have fluctuated.
-CFers seem to have been in decline since the 1970s in both leagues.
-LFers seem to have been in decline in the AL for some time but it does not seem that way in the NL.



Some interesting trends:
-Shortstops have been rising quite a bit in both leagues since the 1970s.
-2B men started declining in the AL in the 1940s, then starting rising again in the 1960s. But even now they have not reached their earlier peak. They started declining in the 1920s in the NL but then started back up in the late 1950s.
-3B men started to rise in the AL in the 1920s and in the 1930s in the NL. But they have tailed off in the AL since 1980.
-1B men had a big spike in the AL in the 1920s and 1930s. In both leagues, in general, they have been high but have fluctuated.
-CFers seem to have been in decline since the 1970s in both leagues.
-LFers seem to have been in decline in the AL for some time but it does not seem that way in the NL.



Wednesday, January 7, 2009
Which Players Had The Most Surprising Walk Rates
A few years ago I noticed that Miller Huggins walked quite a bit yet did not seem to be much of a hitter. From 1904-1916 he lead the league 4 times in walks and was in the top ten 7 other times. So what kind of fearsome hitter was he that he got walked so much? His career batting average (AVG) was .265, not bad in the dead ball era. The league average was .260 during his career. His career slugging percentage (SLG) was .314 while the league average was .343.
But a better measure of power is isolated power (ISO) or SLG - AVG. It tells us extra bases per AB (after all, if a guy can get a single every time up, his SLG would be 1.000 yet he has now power). Huggins' ISO was .049 while the league average was .083. So I was very impressed that he knew the strike zone well enough and had so much discipline that he could walk so frequently yet not have much hitting ability in general.
I thought it would be interesting to come up with some kind of measure of this ability. At first I used walk rate divided by ISO. But his turned out to be unfair to many sluggers since ISO can go very high (theoretically as high as 3.000). Their walk rate would end up being divided by a very large number so their Walk rate/ISO would be low.
So I ran a regression. A player's walk rate (relative to the league average) was the dependent variable and his ISO (relative to the league average) was the independent variable. The idea is that power hitters would get walked more than other hitters. I used all players with 5000+ career plate appearances (885 players). The data comes from the Lee Sinins Complete Baseball Encyclopedia. Here is the regression equation:
Walk Rate = 61.5 + .428*ISO
(In the Sinins encyclopedia, a walk rate of 150, for example, means that the player walked 50% more than average). The r-squared was .159, meaning that only 15.9% of the variation in hitter's walk rates is explained by variation in ISO. But the t-value for the coefficient on ISO was over 12, so it was statistically very significant.
Then each player's walk rate was predicted using the equation and the difference between their actual rate and the predicted rate was found. Then all the players were ranked from highest to lowest by this difference. The table below shows the top 25.

So Roy Thomas did the best. His actual walk rate was 2.49 times the average but his ISO was only .55 or 55% of the average. The equation predicts that he would have a walk rate of 85.06 (or 85.06% of the league average). Since 249 - 85.06 = 163.94, his walk rate was that many points above expected and he had the highest difference. Miller Huggins does very well, coming in at number 7.
The players who walked the least (based on the equation) are below:

I also did something similar using SLG. Here are the best players followed by the worst.

But a better measure of power is isolated power (ISO) or SLG - AVG. It tells us extra bases per AB (after all, if a guy can get a single every time up, his SLG would be 1.000 yet he has now power). Huggins' ISO was .049 while the league average was .083. So I was very impressed that he knew the strike zone well enough and had so much discipline that he could walk so frequently yet not have much hitting ability in general.
I thought it would be interesting to come up with some kind of measure of this ability. At first I used walk rate divided by ISO. But his turned out to be unfair to many sluggers since ISO can go very high (theoretically as high as 3.000). Their walk rate would end up being divided by a very large number so their Walk rate/ISO would be low.
So I ran a regression. A player's walk rate (relative to the league average) was the dependent variable and his ISO (relative to the league average) was the independent variable. The idea is that power hitters would get walked more than other hitters. I used all players with 5000+ career plate appearances (885 players). The data comes from the Lee Sinins Complete Baseball Encyclopedia. Here is the regression equation:
Walk Rate = 61.5 + .428*ISO
(In the Sinins encyclopedia, a walk rate of 150, for example, means that the player walked 50% more than average). The r-squared was .159, meaning that only 15.9% of the variation in hitter's walk rates is explained by variation in ISO. But the t-value for the coefficient on ISO was over 12, so it was statistically very significant.
Then each player's walk rate was predicted using the equation and the difference between their actual rate and the predicted rate was found. Then all the players were ranked from highest to lowest by this difference. The table below shows the top 25.

So Roy Thomas did the best. His actual walk rate was 2.49 times the average but his ISO was only .55 or 55% of the average. The equation predicts that he would have a walk rate of 85.06 (or 85.06% of the league average). Since 249 - 85.06 = 163.94, his walk rate was that many points above expected and he had the highest difference. Miller Huggins does very well, coming in at number 7.
The players who walked the least (based on the equation) are below:

I also did something similar using SLG. Here are the best players followed by the worst.

Saturday, December 20, 2008
Was Jim Rice A Feared Hitter?
This issue came up on the SABR list this week. Someone suggested that batters in the lineup slot ahead of him were helped by his presence. That is, since pitchers knew Rice was up next, they gave good pitches to the current batter. Did batting in front of Rice actually help anyone? I address this below but first I discuss Rice and intentional walks.
My recollection is that Rice was very feared and he was very imposing. So many HRs (and so many long ones) were probably the reason. But he only finished in the top 10 in IBBs 3 times in his career (thanks to Lee Sinins Complete Baseball Encyclopedia). A 5th a tied for 10th and a tied for 9th. Also, he was only tied for 12th in the AL in IBBs from 1975-89. Here are the leaders:
1 George Brett 187
2 Eddie Murray 131
3 Rod Carew 111
4 Ben Oglivie 95
5 Harold Baines 89
6 Wade Boggs 87
T7 Reggie Jackson 85
T7 Ken Singleton 85
T9 Don Baylor 82
T9 Don Mattingly 82
11 Carlton Fisk 78
T12 Kent Hrbek 77
T12 Jim Rice 77
I would expect a feared hitter to rank higher. There are lots of factors that go into IBBs. Maybe he always had someone good behind him (but these other guys might have, too). The guys ahead of him tend to be lefties or switch hitters. Maybe that is the reason (I think there is another interesting issue here about IBBs that I address below after I discuss if batting in front of Rice actually help anyone).
As for how batters in front of him did, I looked at 4 seasons, using Retrosheet, 1977-79 and 1983, arguably his 4 best years. I threw out 1979 since Rice batted 4th all year and Lynn pretty much was the only 3rd place hitter and Lynn did not bat anywhere else.
Let's start with 1977. Rice pretty much batted third. Below are the players who had a significant number of ABs batting both 2nd and in other slots. I show there ABs, AVG, SLG. First I show there stats batting 2nd (in front of Rice) and then the others (combining all ABs not in front of Rice)
Doyle (137-.219-.285) (318-0.248-.318)
Lynn (364-.253-.453) (133-0.278-.428)
Now 1978 (Rice was pretty much 3rd)
Burleson (76-.197-.263) (550-.255-.349)
Lynn (80-.275-.463) (461-.302-.497)
Remy (418-.280-.349) (165-.273-.352)
Now 1983 (Rice was pretty much 3rd)
Boggs (315-.352-.470) (267-.371-.506)
Evans (258-.225-0.419) (212-.255-.458)
Stapleton (54-.259-.352) (488-.246-.365)
It does not look like hitters did alot better in front of Rice than they did elsewhere.
I mentioned the leaders in the AL in IBBs from 1975-89 in my last post. I also just checked the NL. Below are the top 20 in each league. It looks like the AL only had 5 righties while the NL had 10. Also, the top 2 in the NL were righties while in the AL the highest ranked righty was tied for 9th. Seems like a big difference between the two leagues. Also looks like all 10 righties in the AL had more IBBs than the highest ranked AL righty (Baylor). Maybe it is jut a fluke. My apologies if I miss labeled anyone below. I put in R for the righties and nothing for lefites and switch hitters.
AL
1 George Brett 187
2 Eddie Murray 131
3 Rod Carew 111
4 Ben Oglivie 95
5 Harold Baines 89
6 Wade Boggs 87
T7 Reggie Jackson 85
T7 Ken Singleton 85
T9 Don Baylor-R 82
T9 Don Mattingly 82
11 Carlton Fisk-R 78
T12 Kent Hrbek 77
T12 Jim Rice-R 77
T14 Cecil Cooper 73
T14 Fred Lynn 73
16 Mike Hargrove 68
T17 Alvin Davis 67
T17 Robin Yount-R 67
19 Bruce Bochte 65
20 Buddy Bell-R 62
NL
1 Mike Schmidt-R 184
2 Dale Murphy-R 141
3 Dave Parker 139
4 Garry Templeton 134
5 Keith Hernandez 127
6 Ted Simmons 124
7 Jose Cruz 123
8 Bill Madlock-R 112
9 Tim Raines 110
10 Jack Clark-R 104
11 Andre Dawson-R 103
12 Steve Garvey-R 100
13 Gary Carter-R 98
T14 Ron Cey-R 96
T14 Pedro Guerrero-R 96
T14 Leon Durham 96
T14 George Foster-R 96
18 Darryl Strawberry 93
19 Ron Oester 92
20 Dan Driessen 91
My recollection is that Rice was very feared and he was very imposing. So many HRs (and so many long ones) were probably the reason. But he only finished in the top 10 in IBBs 3 times in his career (thanks to Lee Sinins Complete Baseball Encyclopedia). A 5th a tied for 10th and a tied for 9th. Also, he was only tied for 12th in the AL in IBBs from 1975-89. Here are the leaders:
1 George Brett 187
2 Eddie Murray 131
3 Rod Carew 111
4 Ben Oglivie 95
5 Harold Baines 89
6 Wade Boggs 87
T7 Reggie Jackson 85
T7 Ken Singleton 85
T9 Don Baylor 82
T9 Don Mattingly 82
11 Carlton Fisk 78
T12 Kent Hrbek 77
T12 Jim Rice 77
I would expect a feared hitter to rank higher. There are lots of factors that go into IBBs. Maybe he always had someone good behind him (but these other guys might have, too). The guys ahead of him tend to be lefties or switch hitters. Maybe that is the reason (I think there is another interesting issue here about IBBs that I address below after I discuss if batting in front of Rice actually help anyone).
As for how batters in front of him did, I looked at 4 seasons, using Retrosheet, 1977-79 and 1983, arguably his 4 best years. I threw out 1979 since Rice batted 4th all year and Lynn pretty much was the only 3rd place hitter and Lynn did not bat anywhere else.
Let's start with 1977. Rice pretty much batted third. Below are the players who had a significant number of ABs batting both 2nd and in other slots. I show there ABs, AVG, SLG. First I show there stats batting 2nd (in front of Rice) and then the others (combining all ABs not in front of Rice)
Doyle (137-.219-.285) (318-0.248-.318)
Lynn (364-.253-.453) (133-0.278-.428)
Now 1978 (Rice was pretty much 3rd)
Burleson (76-.197-.263) (550-.255-.349)
Lynn (80-.275-.463) (461-.302-.497)
Remy (418-.280-.349) (165-.273-.352)
Now 1983 (Rice was pretty much 3rd)
Boggs (315-.352-.470) (267-.371-.506)
Evans (258-.225-0.419) (212-.255-.458)
Stapleton (54-.259-.352) (488-.246-.365)
It does not look like hitters did alot better in front of Rice than they did elsewhere.
I mentioned the leaders in the AL in IBBs from 1975-89 in my last post. I also just checked the NL. Below are the top 20 in each league. It looks like the AL only had 5 righties while the NL had 10. Also, the top 2 in the NL were righties while in the AL the highest ranked righty was tied for 9th. Seems like a big difference between the two leagues. Also looks like all 10 righties in the AL had more IBBs than the highest ranked AL righty (Baylor). Maybe it is jut a fluke. My apologies if I miss labeled anyone below. I put in R for the righties and nothing for lefites and switch hitters.
AL
1 George Brett 187
2 Eddie Murray 131
3 Rod Carew 111
4 Ben Oglivie 95
5 Harold Baines 89
6 Wade Boggs 87
T7 Reggie Jackson 85
T7 Ken Singleton 85
T9 Don Baylor-R 82
T9 Don Mattingly 82
11 Carlton Fisk-R 78
T12 Kent Hrbek 77
T12 Jim Rice-R 77
T14 Cecil Cooper 73
T14 Fred Lynn 73
16 Mike Hargrove 68
T17 Alvin Davis 67
T17 Robin Yount-R 67
19 Bruce Bochte 65
20 Buddy Bell-R 62
NL
1 Mike Schmidt-R 184
2 Dale Murphy-R 141
3 Dave Parker 139
4 Garry Templeton 134
5 Keith Hernandez 127
6 Ted Simmons 124
7 Jose Cruz 123
8 Bill Madlock-R 112
9 Tim Raines 110
10 Jack Clark-R 104
11 Andre Dawson-R 103
12 Steve Garvey-R 100
13 Gary Carter-R 98
T14 Ron Cey-R 96
T14 Pedro Guerrero-R 96
T14 Leon Durham 96
T14 George Foster-R 96
18 Darryl Strawberry 93
19 Ron Oester 92
20 Dan Driessen 91
Monday, December 15, 2008
Maybe Joe Gordon Does Belong In The Hall Of Fame
Gordon had 242 career win shares. Through 2001, that was tied for 334th among all players including pitchers. But he did miss two seasons due to the war. He missed 1944 and 1945. The two previous seasons he had 28 and 31 (although the competition in 1943 was not so good). In 1946 he only had 9, must have been hurt. In the next two years he had 25 and 24. Suppose we give him 50 for the two years missed. That brings him up to 292. That would be tied for 187 through 2001. Not too bad of a ranking. Good enough for the Hall? I don't know.
But in general, 2B men have an average wins shares per PA that is lower than other positions. Win Shares is supposed to allow us to compare players across positions. Gordon might deserve even more win shares. Maybe he deserves another 10-20. To see the data on win shares per PA for different positions, go to
http://us.share.geocities.com/cyrilmorong@sbcglobal.net/WSperPA.htm (Update Jan. 10, 2016: Here is the new, correct link http://cyrilmorong.com/WSperPA.htm)
If we do give him all these extra win shares he gets close to the top 150 through 2001. I really don't know what type of adjustments to make for him, but I guess a good case could be made for him.
One other thing I thought of is that he was a right handed batter in Yankee Stadium. Win Shares uses runs created get offensive value. Runs created are adjusted for park effects but to the extent that I understand them, no adjustemt is made if a park favors lefties over righties. Gordon hit 69 HRs in home games at Yankee stadium and 84 in road games. You would expect more at home. In his Cleveland years, he had 50 both home and away. Perhaps, on balance, over his career, he was hurt by his parks.
I was a little surprised by his, and his only, selection. But it may be okay. Joe McCarthy said Gordon was the best all around player he ever saw.
But in general, 2B men have an average wins shares per PA that is lower than other positions. Win Shares is supposed to allow us to compare players across positions. Gordon might deserve even more win shares. Maybe he deserves another 10-20. To see the data on win shares per PA for different positions, go to
http://us.share.geocities.com/cyrilmorong@sbcglobal.net/WSperPA.htm (Update Jan. 10, 2016: Here is the new, correct link http://cyrilmorong.com/WSperPA.htm)
If we do give him all these extra win shares he gets close to the top 150 through 2001. I really don't know what type of adjustments to make for him, but I guess a good case could be made for him.
One other thing I thought of is that he was a right handed batter in Yankee Stadium. Win Shares uses runs created get offensive value. Runs created are adjusted for park effects but to the extent that I understand them, no adjustemt is made if a park favors lefties over righties. Gordon hit 69 HRs in home games at Yankee stadium and 84 in road games. You would expect more at home. In his Cleveland years, he had 50 both home and away. Perhaps, on balance, over his career, he was hurt by his parks.
I was a little surprised by his, and his only, selection. But it may be okay. Joe McCarthy said Gordon was the best all around player he ever saw.
Sunday, December 7, 2008
Two Follow Ups: Underpaid Second Basemen And What Happens When Players Cut Down On Strikeouts
A recent report called Increase in MLB salary slowed in 2008 shows that only relief pitchers get paid less than second basemen. Here is the key exerpt:
"Among regulars at positions, designated hitters had the highest average at $7.5 million, followed by first basemen ($7.1 million), third basemen ($6.6 million), shortstops ($5 million), outfielders ($4.8 million), catchers ($3.7 million), second basemen ($3.5 million) and relief pitchers ($1.9 million)."
So second basemen are only half as valuable as designated hitters? Hard to believe. A few months ago I posted a study called Have Second Basemen Been Underpaid?. I found in regressions that, holding hitting performance constant and accounting for free agent/arbitration status, that being a second baseman had a negative effect on salaries.
For the other issue, two weeks ago, I posted Should Ryan Howard Try To Strikeout Less?. The basic idea was that from year to year, there was a positive correlation between player's change in strikeout frequency and change in contact average.
But a commentor named Vince at the The Sabernomics blog said:
"Could this just be a selection effect? If your strikeout rate rises and your contact rate falls, then you might get benched and not show up in the sample."
My response was:
"There were 267 players in 2005 who had 300+ ABs. 200 of them also had 300+ ABs in 2006. So it is possible that those 67 who did not make it to 300 in 2006 were benched for poor performance (which would include a low contact average).
But I took those 67 guys and found the ones who had atleast 100 ABs in 2006 (I think anything less is a small sample size). That left 41 guys. The correlation between their change in strikeout frequency and change in contact average was .037. So it was still positive for the ones who were “selected out” but not as strong an effect."
"Among regulars at positions, designated hitters had the highest average at $7.5 million, followed by first basemen ($7.1 million), third basemen ($6.6 million), shortstops ($5 million), outfielders ($4.8 million), catchers ($3.7 million), second basemen ($3.5 million) and relief pitchers ($1.9 million)."
So second basemen are only half as valuable as designated hitters? Hard to believe. A few months ago I posted a study called Have Second Basemen Been Underpaid?. I found in regressions that, holding hitting performance constant and accounting for free agent/arbitration status, that being a second baseman had a negative effect on salaries.
For the other issue, two weeks ago, I posted Should Ryan Howard Try To Strikeout Less?. The basic idea was that from year to year, there was a positive correlation between player's change in strikeout frequency and change in contact average.
But a commentor named Vince at the The Sabernomics blog said:
"Could this just be a selection effect? If your strikeout rate rises and your contact rate falls, then you might get benched and not show up in the sample."
My response was:
"There were 267 players in 2005 who had 300+ ABs. 200 of them also had 300+ ABs in 2006. So it is possible that those 67 who did not make it to 300 in 2006 were benched for poor performance (which would include a low contact average).
But I took those 67 guys and found the ones who had atleast 100 ABs in 2006 (I think anything less is a small sample size). That left 41 guys. The correlation between their change in strikeout frequency and change in contact average was .037. So it was still positive for the ones who were “selected out” but not as strong an effect."
Subscribe to:
Posts (Atom)