There is a book about them. It is 1922 St. Louis Browns: Best of the American League's Worst by by Roger A. Godin. But I have not read it and I don't know if it attempted to rank them among all teams in baseball history (I will post something on this tomorrow and it will show that they rank very high).
I started thinking about this when I read about a simulation called the Seamheads Near Miss League at wezen-ball. The Browns did extremely well. Actually, they were an outstanding team statistically.
The Brown's team strikeout-to-walk ratio was about 34% above the league average. That is the 110th best ever in AL & NL history. With well over 2000 team-seasons, that puts them in the top 5%. Their team ERA was about 19% below the league average, good for 195th, still in the top 10%. But their park factor was 108, meaning that they pitched in a somewhat high run environment. They did give up a few more HRs than the average AL team in 1922. They gave up 71 HRs while the average for the other 7 teams was about 65. But their park gave up 91% more HRs than average.
On the hitting side, their team offensive winning percentage of .577, which ranks 191st. Again, that is in the top 10%. OWP is the Bill James stat that tells us if a team had 9 identical hitters and gave up an average number of runs, what would their winning percentage be? Since I used the Lee Sinins Complete Baseball Encyclopedia, it is park adjusted. So pretty impressive, being in the top 10% in both hitting and pitching.
What happened to them? Why didn't they win the pennant? Using Retrosheet, here are some interesting facts about the 1922 season:
The Browns finished 1 game out, behind the Yankees. But the Browns had a run differential of 224 to the Yankees 140 (of course Ruth missed 44 games, most probably due to his suspension. He also got off to a slow start in May, batting just .190 with 2 HRs in 42 ABs).
The Browns had 256 more hits than their opponents that year, 50 more walks and 27 more HRs. For the Yankees, it was 101-69-25. So the Browns stats look much better.
The Yankees beat the Browns 14 times out of 22 even though the Browns outscored the Yankees 105-100 in those games. It looks like the Yankees won all 8 of the 1-run games between the two teams that year, with 4 in extra innings.
In mid-Sept, the Yankees came to Stl. for the last series between the two teams that year. The Yanks were a half game ahead when the series started and won 2 out of 3. In the last game, the Browns led 2-0 after 7, but NY won it 3-2 with 2 in the 8th and 1 in the 9th. 2 of the 3 runs were unearned as the Browns made 3 errors.
Later, with 2 games for each team left, the Browns were 2 back. But both of them won game 153 and so it was over. But if the Browns had just one more win against NY, they would have been tied with 2 games left.
Udate at 7:43 am central time, 8-11-2009: Chirs Jaffe had a good discusssion of this team last year at October country’s refugees (part 2 of 2)
Monday, August 10, 2009
Saturday, August 8, 2009
Rangers Power Update
(or maybe we should call them the "Power Rangers") With 2/3 of the season having been played, the Rangers have 168 HRs. If we simply increse that by 50%, they would finish with 252, 4th best all time. I first wrote about this issue last May with Texas Rangers On A Pace To Set Power Hitting Records. Obviously the Rangers have tailed off in their power hitting since then. But their isolated power is still .201, which would be tied for 3rd all-time. Their HR% is 4.56. If they finish with that, they would be 2nd.
But they are only scoring an average number of runs per game (4.84, the league average is 4.83). The reason they are only averge in runs per game even though they are doing all this power hitting is that their OBP is only .317 while the league average is .335.
But they are only scoring an average number of runs per game (4.84, the league average is 4.83). The reason they are only averge in runs per game even though they are doing all this power hitting is that their OBP is only .317 while the league average is .335.
Monday, August 3, 2009
Jim Rice and the Hall of Fame (Revisited)
Bob Ryan of boston.com recently wrote an article called The big picture is that Rice earned his plaque. He takes a swipe at "SABR people." But we're not monolithic and I don't know if all or even most SABR members agree with my views on Rice. But anyway, here is part of what Ryan said and I follow by summarizing some of my past posts on Rice that counter what Ryan says.
One of my posts was Was Jim Rice A Feared Hitter?. I showed that he did not draw very many intentional walks compared to other top hitters and that players who batted in front of him were not especially helped.
With Jim Rice and the Hall of Fame I showed that his clutch hitting stats, although better than his overall stats, were very close. I wrote "According to retrosheet, with ROB [runners on base], his AVG-SLG were .305 & .509. With RISP [runners in scoring position] he had .308 & .501. These are very close to his overall stats of .298 & .502." In fact, he was more likely to come up with runners on base in Fenway, a good hitters park. So naturally he would hit better in those situations. He most likely had a disproportionate number of ABs with RISP & ROB in Fenway. His close and late AVG-SLG were .274-.453. So that does not look very clutch.
He was helped by Fenway. His AVG-SLG in home games was .320 & .546 while on the road they were only .277 & .459. I also showed that his RBI-to-GDP ratio was very poor, even below average.
One commentor at Ryan's article mentioned that Rice had alot of clutch hits in Septmber 1986 when the Red Sox were in a tight divisional race. But his AVG was .310 in Sept while it was .324 for the whole season. He also grouned into 6 double plays that month. He had 19 for the whole season, so he had close to 1/3 in Sept. His SLG was .560 in Sept. while it was .490 for the whole season. So he did slug better even if he got fewer hits.
I have also found that Jose Cruz of the Astros may be just as Hall worthy as Rice. Go to Jim Rice vs. Jose Cruz.
"The SABR people are resolutely anti-Rice. They’ve got numbers parsed by the truckload to downplay his impact, and to this I say, “Phooey,’’ or maybe even something stronger. For SABR people refuse to acknowledge the concept of anecdotal evidence when evaluating a ballplayer (no, not you, Bill James). So when I speak of the time Milwaukee manager Alex Grammas confirmed for me that, yes, indeed, he had ordered a sizzling Jim Rice pitched around (like, four straight unhittable balls out of the strike zone) in a sixth-inning, bases-loaded situation, or when fellow inductee Rickey Henderson says, as he did yesterday, that when the A’s had pitchers meetings prior to Red Sox series in the Rice era guys “trembled,’’ they say that’s nice, but irrelevant.
Sorry, it matters.
There was a three-year period from 1977-79 when Rice was The Man in the American League, averaging 41 homers, 127 RBIs, and 206 hits a year. And did you know he had back-to-back seasons (’77-78) of 15 triples? He was a feared - yeah, SABR people, feared - hitter, because he was very content to get a base hit in a key situation. He was, after all, just trying to win the game."
One of my posts was Was Jim Rice A Feared Hitter?. I showed that he did not draw very many intentional walks compared to other top hitters and that players who batted in front of him were not especially helped.
With Jim Rice and the Hall of Fame I showed that his clutch hitting stats, although better than his overall stats, were very close. I wrote "According to retrosheet, with ROB [runners on base], his AVG-SLG were .305 & .509. With RISP [runners in scoring position] he had .308 & .501. These are very close to his overall stats of .298 & .502." In fact, he was more likely to come up with runners on base in Fenway, a good hitters park. So naturally he would hit better in those situations. He most likely had a disproportionate number of ABs with RISP & ROB in Fenway. His close and late AVG-SLG were .274-.453. So that does not look very clutch.
He was helped by Fenway. His AVG-SLG in home games was .320 & .546 while on the road they were only .277 & .459. I also showed that his RBI-to-GDP ratio was very poor, even below average.
One commentor at Ryan's article mentioned that Rice had alot of clutch hits in Septmber 1986 when the Red Sox were in a tight divisional race. But his AVG was .310 in Sept while it was .324 for the whole season. He also grouned into 6 double plays that month. He had 19 for the whole season, so he had close to 1/3 in Sept. His SLG was .560 in Sept. while it was .490 for the whole season. So he did slug better even if he got fewer hits.
I have also found that Jose Cruz of the Astros may be just as Hall worthy as Rice. Go to Jim Rice vs. Jose Cruz.
Tuesday, July 28, 2009
Buehrle Sets New Record For Consecutive Batters Retired
He retired the first 17 tonight. Plus he retired the last batter he faced in the game before the perfect game. That makes 45 straight batters retired. That breaks the record of 41 held by Jim Barr and Bobby Jenks. But it was reported in several places just about as soon as it happened. I was not first.
Two Grand Slams in One Game vs. Perfect Games
Yesterday, Josh Willingham of the Washington Nationals hit two grandslams in one game, becoming the 13th player to ever do this. It was just 4 days after Mark Buehrle of the Chicago White Sox pitched a perfect game. So I wondered which event is rarer. To see all of the occurrences of these events, you can to to the following Baseball Almanac sites:
Two Grand Slams in One Game
Perfect Games
The first player to hit 2 grand slams in one game was Tony Lazzeri, in 1936. Maybe it is not surprising no one did it before 1920, in the dead ball era, when HRs were rare. So I use 1920 as the starting point for the comparison between these two rare events. I only include regular season games, so Don Larsen's perfect game in the World Series will not count. And I also include all games where a pitcher had 9 perfect innings from the start of the game, regardless of what happened after the 9th inning. So I include Harvey Haddix's 1959 game and Pedro Martinez's 1995 game. After all, it was not their fault their teammates could not score just one run for them. They did match what these other pitchers did.
Using the data from the Baseball Almanac site, that leaves 15 perfect games since 1920. That is two more than games when a player hit 2 grandslams. There were about 149,000 major league games played from 1920-2008. So prior to this year, a perfect game happened once every 10,642 games. Two grandslams in one game happened once every 12,416 games.
Looking at the Baseball Almanac sites, you can see that in 1968, 1995, 1998 and 1999 both events occurred. So 2009 is the 5th year that both happened. But this time it was only 4 days apart. In those other years the two events were always atleast a month apart.
Two Grand Slams in One Game
Perfect Games
The first player to hit 2 grand slams in one game was Tony Lazzeri, in 1936. Maybe it is not surprising no one did it before 1920, in the dead ball era, when HRs were rare. So I use 1920 as the starting point for the comparison between these two rare events. I only include regular season games, so Don Larsen's perfect game in the World Series will not count. And I also include all games where a pitcher had 9 perfect innings from the start of the game, regardless of what happened after the 9th inning. So I include Harvey Haddix's 1959 game and Pedro Martinez's 1995 game. After all, it was not their fault their teammates could not score just one run for them. They did match what these other pitchers did.
Using the data from the Baseball Almanac site, that leaves 15 perfect games since 1920. That is two more than games when a player hit 2 grandslams. There were about 149,000 major league games played from 1920-2008. So prior to this year, a perfect game happened once every 10,642 games. Two grandslams in one game happened once every 12,416 games.
Looking at the Baseball Almanac sites, you can see that in 1968, 1995, 1998 and 1999 both events occurred. So 2009 is the 5th year that both happened. But this time it was only 4 days apart. In those other years the two events were always atleast a month apart.
Saturday, July 11, 2009
How Important is Home Field Advantage in the World Series?
That is the name of an article I wrote a few years ago for Beyond the Boxscore. It is at How Important is Home Field Advantage in the World Series?. I figured out the probability of the team with HFA winning in 4, 5, 6 and 7 games and then added those together to get 51.52%. Not everyone agreed with my methods and you can read about that in the comments. I got reminded of this because Sky Andrecheck recently wrote an interesting article at Baseball Analysts called Is The All-Star Game The Biggest Remaining Game for Dodgers?. He came up with 51.26%. We both used the historical average of home teams winning 54% of the time in regular season games.
Monday, July 6, 2009
My Sabermetrics Page Has Moved
Click on this link to go to the new address Cyril Morong's Sabermetric Research. This link has alot of my research on it. If some of those links don't work, please be patient. I am working on getting it all straightened out. Geocities is going away and I just switched to the new Yahoo service. So the new address is
http://cyrilmorong.com
http://cyrilmorong.com
Sunday, July 5, 2009
Do Pitchers Differ In Their Ability To Prevent HRs? (and does it persist over time?)
There was a very intriguing post a few weeks ago at the Hardball Times by Derek Carty called Using FIP to evaluate pitchers? I wouldn’t. The FIP refers to "fielding independent ERA," the idea that pitchers should be evaluated only on outcomes that don't involve the fielders. That would include HRs allowed. But, the article said:
"Here's how things work: a pitcher can influence the rate of fly balls he gives up. By this logic, the more fly balls allowed, the more total balls will clear the fences for home runs (all else being equal). However, while a starting pitcher can control the rate of fly balls allowed, he cannot do a very good job of controlling the rate at which those fly balls become home runs (with very few exceptions).
To put it more simply, starting pitchers don't have any underlying ability to prevent home runs—the best they can do is prevent fly balls. If those fly balls are clearing the fence at too high a rate (or too low), we say that the pitcher has been unlucky (or lucky)."
I am not sure I completely agree with this. It could be that there is a difference in flyballs allowed that accounts for the HR rates allowed across pitchers. But whatever the reason, the year-to-year correlation of HR rates allowed by pitchers, although not as high as they are for their walk rates and strikeout rates, they are not small.
The data I looked at involves year-to-year correlations of various years for pitchers who faced at least 500 batters in both of two consecutive seasons. The table below summarizes the results. Starting with the 1955 season, I eliminated IBBs from the calculations. HBP were counted as walks in all years. The columns show the correlation between the rates allowed for each stat year-to-year. The last line is the simple average of all the correlations.
Overall, the correlations are much higher for strikeout rates and walk rates (the denominator I used in all cases was batters faced). But the correlations do seem to be getting higher for the HR rates. It was very surprising to see how low they were in some of the earlier years.
One more thing that I tried (and this really makes me think that we should keep looking at HR rates) is that I found a high correlation in HR rates from one period to the next using more years. For that, I found all the pitchers that had 1000+ batters faced in both the 2003-05 period and the 2006-08 period. The correlations for walk rates and strikeout rates from period 1 to period 2 were 0.736and 0.767, respectively. But for HR rates it was 0.505. This seems high enough to say that, yes, pitchers do differ in the HR rates they allow, even if the reason is their flyball rates.
"Here's how things work: a pitcher can influence the rate of fly balls he gives up. By this logic, the more fly balls allowed, the more total balls will clear the fences for home runs (all else being equal). However, while a starting pitcher can control the rate of fly balls allowed, he cannot do a very good job of controlling the rate at which those fly balls become home runs (with very few exceptions).
To put it more simply, starting pitchers don't have any underlying ability to prevent home runs—the best they can do is prevent fly balls. If those fly balls are clearing the fence at too high a rate (or too low), we say that the pitcher has been unlucky (or lucky)."
I am not sure I completely agree with this. It could be that there is a difference in flyballs allowed that accounts for the HR rates allowed across pitchers. But whatever the reason, the year-to-year correlation of HR rates allowed by pitchers, although not as high as they are for their walk rates and strikeout rates, they are not small.
The data I looked at involves year-to-year correlations of various years for pitchers who faced at least 500 batters in both of two consecutive seasons. The table below summarizes the results. Starting with the 1955 season, I eliminated IBBs from the calculations. HBP were counted as walks in all years. The columns show the correlation between the rates allowed for each stat year-to-year. The last line is the simple average of all the correlations.
Overall, the correlations are much higher for strikeout rates and walk rates (the denominator I used in all cases was batters faced). But the correlations do seem to be getting higher for the HR rates. It was very surprising to see how low they were in some of the earlier years.
One more thing that I tried (and this really makes me think that we should keep looking at HR rates) is that I found a high correlation in HR rates from one period to the next using more years. For that, I found all the pitchers that had 1000+ batters faced in both the 2003-05 period and the 2006-08 period. The correlations for walk rates and strikeout rates from period 1 to period 2 were 0.736and 0.767, respectively. But for HR rates it was 0.505. This seems high enough to say that, yes, pitchers do differ in the HR rates they allow, even if the reason is their flyball rates.
Saturday, July 4, 2009
Albert Pujols Has A Good Chance To Win The Triple Crown
He leads the NL in both HRs and RBIs by 7. He is 2nd in AVG with .336 while Hanley Ramirez is hitting .344. But I compared Pujols to Ramirez and the rest of the NL top ten in AVG, and based on previous performances, Pujols has done much better in AVG. The table below shows the current averages of the NL top 10. It also shows what they hit in 2008, their current career average, and their highest average before 2008.
In terms of what he hit last year, his career average and his high average, Pujols is well ahead of the other guys in the top ten. Sandoval only had 145 ABs in 2008 and only has 422 so far in his career. Pujols also has a history of hitting well after the All-Star break. Last year it was .366 and for his career it is .344. Looks like he has a good chance to lead in AVG once the season is over.
In terms of what he hit last year, his career average and his high average, Pujols is well ahead of the other guys in the top ten. Sandoval only had 145 ABs in 2008 and only has 422 so far in his career. Pujols also has a history of hitting well after the All-Star break. Last year it was .366 and for his career it is .344. Looks like he has a good chance to lead in AVG once the season is over.
Monday, June 29, 2009
Which Players Had The Best HR-To-Strikeout Ratios?
I looked at every player with 5000+ PAs since 1920. I found their relative HRs and their relative strikeouts. Then found the ratio of the two. Ken Williams, for example, hit 3.70 times as many HRs as the average player of his time and league while striking out only 75% as often as the average player. Since his ratio of ratios (3.7/.75 = 4.93) is the highest of anyone in the study, he is ranked first. The data comes from the Lee Sinins Complete Baseball Encyclopedia. The table below shows the top 25:

DiMaggio hit only 41% of his HRs at home in his career while Williams hit 72%. So it is likely the case that DiMaggio would rank first, and probably by a wide margin, if HRs were park adjusted. Ted Williams hit less than 50% of his HRs at home.
The next table shows which players had the lowest relative strikeout rates among guys who hit 40+ HRs. Again, no pikers here. In 2004, Bonds had only 41 strikeouts while the average player would have had 100. I am so proud to see the demonstration of Polish power with 3 for Ted Kluszewski and 1 for Carl Yastrzemski (whose 1970 season ranks 27th). Don't forget Stan Musial is 13th on the above list.

DiMaggio hit only 41% of his HRs at home in his career while Williams hit 72%. So it is likely the case that DiMaggio would rank first, and probably by a wide margin, if HRs were park adjusted. Ted Williams hit less than 50% of his HRs at home.
The next table shows which players had the lowest relative strikeout rates among guys who hit 40+ HRs. Again, no pikers here. In 2004, Bonds had only 41 strikeouts while the average player would have had 100. I am so proud to see the demonstration of Polish power with 3 for Ted Kluszewski and 1 for Carl Yastrzemski (whose 1970 season ranks 27th). Don't forget Stan Musial is 13th on the above list.
Monday, June 22, 2009
Harold Reynolds And Using Context To Evaluate Hitters
ESPN analyst and former major league player wrote a blog entry called Enjoy it for what it's worth. Sky Kalkman at "Beyond the Boxscore" wrote a response called Defending Harold Reynolds. Reynolds criticizes some of the "newer" stats like OPS:
"Not all statistics work. Some do, some don't. And one of the stats that has become real popular is OPS. On-base plus slugging. All of a sudden, it's this stat that defines whether a guy is a good ball player or not. And the fact of the matter is, if you're a power hitter then the situation will dictate what a pitcher does with you - either walk you or pitch you real careful. So more than likely you're going to end up on base and therefore your on-base percentage goes up. This in my mind has become the stat the everyone thinks is the be all and end all. It is not. If you have a ball club that's a great offensive team then that changes everything. But if you have a guy like Adrian Gonzalez, for example, his OPS is going to be high - he's got a lot of home runs and walks a lot...because you're not going to pitch to him. Power guys like Giambi and Dunn have always had high OPS because no one wants to pitch to them. But it takes two hits to score them from first."
Reynolds began by saying that context and situation matters and it probably does. But this raises the question of how much? Some of my past research touches on these issues and I will discuss that below. But first, even if you don't like OPS, or OBP + SLG, it is still better than the traditional stats (for example, he mentions that Ichiro Suzuki gets 200+ hits every year). The 1998 STATS, INC. Baseball Scoreboard book had a nice little study that showed that the team with the highest OPS in a game has a winning percentage of .852 while it was .804 for batting average (they looked at several other stats and OPS had the highest winning percentage).
But let's look at some of what Reynolds said specifically. He seems to be saying that when a slugger walks on a weak hitting team, it is not so valuable. But I had done some analysis on this. It was called The Value of OBP and SLG by Lineup Position for High-Scoring and Low-Scoring Teams. If you go to this link, you will see that the marginal run value for the cleanup hitter's OBP is actually higher on the low scoring team.
Now how much might context matter or change our evaluation of hitters if we are using OBP and SLG? My analysis on this is called Evaluating Hitters Based on Their Lineup Slot. The most anyone was adjusted was a +6.2 runs per season, for Luis Castillo. So if I took into account that he was a leadoff hitter instead of a generic hitter, his value to his team would be about 6 more runs a year. This seems pretty small. So context does not change our valuation much.
Then there is the issue of situational hitting. My analysis on this is called The Problem With “Total Clutch” Hitting Statistics. What I found was that OPS was highly correlated with how much impact a hitter had on winning and losing depending upon the situation. The stat I used was Ed Oswalt’s measure “player’s win value” (or PWV). It makes a HR in a close and late game more valuable than one in a blowout. It calculates how much each hitter's result changed his team's chances of winning. The correlation between PWV/PA and OPS was .948 (a perfect correlation is 1.00). The relationship was even stronger when I broke down OPS into its separate components of OBP and SLG. So the bottom line is that we really don't need to know the situations a player faced to evaluate him. His regular stats tell us that.
"Not all statistics work. Some do, some don't. And one of the stats that has become real popular is OPS. On-base plus slugging. All of a sudden, it's this stat that defines whether a guy is a good ball player or not. And the fact of the matter is, if you're a power hitter then the situation will dictate what a pitcher does with you - either walk you or pitch you real careful. So more than likely you're going to end up on base and therefore your on-base percentage goes up. This in my mind has become the stat the everyone thinks is the be all and end all. It is not. If you have a ball club that's a great offensive team then that changes everything. But if you have a guy like Adrian Gonzalez, for example, his OPS is going to be high - he's got a lot of home runs and walks a lot...because you're not going to pitch to him. Power guys like Giambi and Dunn have always had high OPS because no one wants to pitch to them. But it takes two hits to score them from first."
Reynolds began by saying that context and situation matters and it probably does. But this raises the question of how much? Some of my past research touches on these issues and I will discuss that below. But first, even if you don't like OPS, or OBP + SLG, it is still better than the traditional stats (for example, he mentions that Ichiro Suzuki gets 200+ hits every year). The 1998 STATS, INC. Baseball Scoreboard book had a nice little study that showed that the team with the highest OPS in a game has a winning percentage of .852 while it was .804 for batting average (they looked at several other stats and OPS had the highest winning percentage).
But let's look at some of what Reynolds said specifically. He seems to be saying that when a slugger walks on a weak hitting team, it is not so valuable. But I had done some analysis on this. It was called The Value of OBP and SLG by Lineup Position for High-Scoring and Low-Scoring Teams. If you go to this link, you will see that the marginal run value for the cleanup hitter's OBP is actually higher on the low scoring team.
Now how much might context matter or change our evaluation of hitters if we are using OBP and SLG? My analysis on this is called Evaluating Hitters Based on Their Lineup Slot. The most anyone was adjusted was a +6.2 runs per season, for Luis Castillo. So if I took into account that he was a leadoff hitter instead of a generic hitter, his value to his team would be about 6 more runs a year. This seems pretty small. So context does not change our valuation much.
Then there is the issue of situational hitting. My analysis on this is called The Problem With “Total Clutch” Hitting Statistics. What I found was that OPS was highly correlated with how much impact a hitter had on winning and losing depending upon the situation. The stat I used was Ed Oswalt’s measure “player’s win value” (or PWV). It makes a HR in a close and late game more valuable than one in a blowout. It calculates how much each hitter's result changed his team's chances of winning. The correlation between PWV/PA and OPS was .948 (a perfect correlation is 1.00). The relationship was even stronger when I broke down OPS into its separate components of OBP and SLG. So the bottom line is that we really don't need to know the situations a player faced to evaluate him. His regular stats tell us that.
Monday, June 15, 2009
Which Players Had The Most Surprising Walk Rates? (Part 2)
Click on Part 1 to see what I did last January. Then I looked at walk rates relative to the league average as a function of isolated power, relative to the league average with the idea being that it is harder to walk alot if you are not a power hitter.
What inspired me to go back and do more on this was a discussion of walks between Bill James and Joe Posnanski at Talkin' about the underappreciated base on balls, with Bill James. Another interesting take on walks appeared in Baseball Magazine in 1917. The article was by FC Lane and seems ahead of its time. It was called The Base on Balls: Why Should the Records Ignore This Powerful Factor in Brainy Baseball?
This time I also included a variable for height and one for stealing. Height was in inches and stealing was stolen bases divided by singles + walks + HBP. Sort of a frequency. That was also relative to the league average. The idea is that shorter guys have an easier time walking and guys who steal alot won't get walked too much if the pitcher can help it. Here is the regression equation. Everything is relative to the league average except height. My data sourse in the Lee Sinins Complete Baseball Encyclopedia.
Walks = 195.58 - 1.25*SB - 1.8*HT + .369*ISO
The stats are all converted to a number relative to 100. If you were average at something, then you get a 100 (except for SB where 1.00 was average). Height and isolated power were significant but stealing was not.
The graph below shows the players with the most surprising walk rates. That is, their walks relative to the league average were the most above league average compared to what the equation predicted.

So Thomas walked 2.19 times as often as the average hitter. His isolated power was only 57% of the league average, he was 71 inches tall and his stolen base rate was only 68% of the league average. Now the guys who walked the least compared to expectations.

I will try to give more details later. But time to give a test.
I am back. The r-squared was .148 and the standard error was about 30. I also tried taking logs of all the variables but the results were no better. For the linear regression there was no correlation between the prediction error and any of the independent variables. I also wonder if height should be relative to the league average. But it raises the question if a 6'0" tall pitcher has a harder time throwing strikes to a 5'6" batter than a 5'6" pitcher. I don't but I assumed the height of the pitcher did not matter.
What inspired me to go back and do more on this was a discussion of walks between Bill James and Joe Posnanski at Talkin' about the underappreciated base on balls, with Bill James. Another interesting take on walks appeared in Baseball Magazine in 1917. The article was by FC Lane and seems ahead of its time. It was called The Base on Balls: Why Should the Records Ignore This Powerful Factor in Brainy Baseball?
This time I also included a variable for height and one for stealing. Height was in inches and stealing was stolen bases divided by singles + walks + HBP. Sort of a frequency. That was also relative to the league average. The idea is that shorter guys have an easier time walking and guys who steal alot won't get walked too much if the pitcher can help it. Here is the regression equation. Everything is relative to the league average except height. My data sourse in the Lee Sinins Complete Baseball Encyclopedia.
Walks = 195.58 - 1.25*SB - 1.8*HT + .369*ISO
The stats are all converted to a number relative to 100. If you were average at something, then you get a 100 (except for SB where 1.00 was average). Height and isolated power were significant but stealing was not.
The graph below shows the players with the most surprising walk rates. That is, their walks relative to the league average were the most above league average compared to what the equation predicted.
So Thomas walked 2.19 times as often as the average hitter. His isolated power was only 57% of the league average, he was 71 inches tall and his stolen base rate was only 68% of the league average. Now the guys who walked the least compared to expectations.
I will try to give more details later. But time to give a test.
I am back. The r-squared was .148 and the standard error was about 30. I also tried taking logs of all the variables but the results were no better. For the linear regression there was no correlation between the prediction error and any of the independent variables. I also wonder if height should be relative to the league average. But it raises the question if a 6'0" tall pitcher has a harder time throwing strikes to a 5'6" batter than a 5'6" pitcher. I don't but I assumed the height of the pitcher did not matter.
Saturday, June 13, 2009
What Luis Castillo teaches us about the internet
I did the following google search (searching blogs only)
"Luis Castillo" yankees
I restricted it at first to the the following dates: June 10-11. 25 hits came up. Then I had the search cover the past 12 hours and there were 1100 hits. The last day gave over 1500.
In case you don't know, last night Castillo dropped a pop up that allowed 2 runs to score in the bottom of the 9th inning, giving the Yankees a 1-run win over the Mets. If he had caught it, the game would have ended.
"Luis Castillo" yankees
I restricted it at first to the the following dates: June 10-11. 25 hits came up. Then I had the search cover the past 12 hours and there were 1100 hits. The last day gave over 1500.
In case you don't know, last night Castillo dropped a pop up that allowed 2 runs to score in the bottom of the 9th inning, giving the Yankees a 1-run win over the Mets. If he had caught it, the game would have ended.
Wednesday, June 10, 2009
My Interview With Vince Gennaro
Last year I interviewed him for the now defunct paper, the Chicago Sports Weekly. Here is a link. It is a PDF file.
Vince Gennaro interview
Vince is the author of Diamond Dollars: The Economics of Winning in Baseball. He also just got elected as secretary of SABR. So it is good for us to have someone with his training and experience on board.
He was also interviewed at Hardball Times
Vince Gennaro interview
Vince is the author of Diamond Dollars: The Economics of Winning in Baseball. He also just got elected as secretary of SABR. So it is good for us to have someone with his training and experience on board.
He was also interviewed at Hardball Times
Monday, June 8, 2009
Should We Keep An Eye On The Rockies?
They just swept the Cardinals in a 4 game series in St. Louis. Prior to the series, the Rockies were 21-32 and the Cards were 31-23. And last year the Cards won 12 more games (84 vs. 76).
The Rockies also won their last game in Houston just before going on to St. Louis, beating the Astros 10-3. The scores against the Cards were 11-4, 10-1, 7-2 and 5-2. So no game was close. Also, before that win against the Astros, the Rockies had been averaging only 3.83 runs per game on the road. Then they score 43 in 5 games. The starting pitchers they faced were not a bad lot. Below are their names followed first by their ERAs going into the game with the Rockies and their ERAs from last year.
Wandy RodrÃguez-2.26, 3.54
Adam Wainwright-3.38, 3.20
Todd Wellemeyer-5.05, 3.71
Joel Piñeiro-3.86, 5.15
Brad Thompson-4.12, 5.15
No all-stars, but they are capable major league hurlers. Combined they have to be at least average. Fellow SABR member Dean Hendrickson suggested that firing Clint Hurdle is what changed things. I sure don't know.
The Rockies also won their last game in Houston just before going on to St. Louis, beating the Astros 10-3. The scores against the Cards were 11-4, 10-1, 7-2 and 5-2. So no game was close. Also, before that win against the Astros, the Rockies had been averaging only 3.83 runs per game on the road. Then they score 43 in 5 games. The starting pitchers they faced were not a bad lot. Below are their names followed first by their ERAs going into the game with the Rockies and their ERAs from last year.
Wandy RodrÃguez-2.26, 3.54
Adam Wainwright-3.38, 3.20
Todd Wellemeyer-5.05, 3.71
Joel Piñeiro-3.86, 5.15
Brad Thompson-4.12, 5.15
No all-stars, but they are capable major league hurlers. Combined they have to be at least average. Fellow SABR member Dean Hendrickson suggested that firing Clint Hurdle is what changed things. I sure don't know.
Saturday, June 6, 2009
Rick Reuschel for the Hall of Fame
(this was posted on the Chicago Sports Review site a few years ago but as far as I can tell, alot of my articles there have been taken down. Someone at Joe Posnanski's site mentioned that Reuschel was very under rated. So I have been wanting to post some of my old CSR articles, so this seems like a good time) First, some highlights:
-His strike-out-to-walk ratio was 31% better than the league average
-He gave up 21.6% fewer HRs than average
Yes, I’m crazy. But not because I think Reuschel deserves to be in the Hall of Fame. What makes me crazy is that I was doing a lot ridiculously time-consuming analysis that opened my eyes. As of now, I don’t know how he has done in the voting or if he is still eligible for induction based on the writer’s vote or if he has to wait for the Veterans Committee.
Before I get into explaining the actual analysis, lets review a few things. A pitcher needs to prevent the other team from scoring and, in this endeavor, he can be aided or hindered by his fielders. So one thing to look at are his defense-independent pitching stats. Many of you probably know that I am taking a page out of Voros McCracken’s book on this. Readers who have not heard of him, Google his name. He came up with the idea a few years ago that on balls in play (not HRs, walks or strikeouts), the batting average that most pitchers allow is not too much different from other pitchers and that it appears to be mainly influenced by the fielders and hitters. Analysts Tom Tippett and Mike Emeigh, to name just two, have challenged McCracken’s thesis. But it is still pretty good.
So we should look at the things a pitcher controls, like HRs, walks and strikeouts (and even if the pitcher has some influence on the batting average on balls in play, the fielders and hitters still play a role, so we are still isolating just what the pitcher does and so it is a legitimate analysis). But how a pitcher does in HRs, walks or strikeouts must be put into context. They should be adjusted for the league average and for park effects.
So using the linear regression technique, I came up with a formula for estimating a pitcher’s ERA. I looked at all pitching seasons from 1920-2000 with 150+ IP. Here is the formula that I got
ERA = .44*HR + .4*BB - .3*K
The intercept or constant term was less than .0001. The r-squared was .422 and the standard error was .75. But for each pitcher I used not his ERA, but how much he differed from the league average in the given year (actually how many standard deviations from the mean he was). Same thing for the HRs, walks and strikeouts. Using standard deviations, since they measure dispersion, does a better job of placing a pitcher in the context of his season than what percent above or below the mean they are in some stat (or the absolute difference) since we see where a pitcher fits in the statistical distribution of the given year.
Once I had this equation, I plugged in each pitcher's data on HRs, BBs and Ks and ranked them all (2000+ IP minimum during the 1920-2000 period) in terms of how many standard deviations below the mean they were in their careers (using only seasons when they pitched 150+ IP-data came from the Sean Lahman data base). Each pitcher’s HRs were adjusted for park effects before their HRs were plugged into the formula. I used park effect data from fellow economist and SABR member Ron Selter. So if a pitcher was 1 standard deviation below the mean, he got a -.44. If he were, say, half a standard deviation better on walks and strikeouts, he gets a -.2 and a -.15. So he would come out -.79. How I used the park data is explained below in technical notes (I don’t have any data on how parks affect strikeouts and walks so those were not adjusted).
Rick Reuschel was 14th! Yes. That seems to be a high enough ranking in an 80-year period to merit the Hall of Fame. Here are the top 20 in terms of how many standard deviations below the average ERA they were for their careers:
Dazzy Vance –1.25
Lefty Grove –1.21
Roger Clemens –1.18
Greg Maddux –1.14
Carl Hubbell -.94
Randy Johnson -.92
Kevin Brown -.91
Dwight Gooden -.89
Mike Garcia -.84
Hal Newhouser -.82
Sandy Koufax -.81
Bert Blyleven -.75
Ron Guidry -.73
Rick Reuschel -.71
G. Alexander -.70
Gaylord Perry -.68
Urban Schocker -.65
John Smoltz -.65
Lefty Gomez -.65
Bob Gibson -.62
(I also looked at ERA relative to the league average in addition to this standard deviation technique and he would be 20th, still a very high ranking). Reuschel is in obviously very good company. This means that, when only looking at pitcher controlled factors, he was outstanding at preventing runs in the context of his era and parks. This was done while pitching 3500 innings (38th since 1920) and winning over 200 games.
Some conventional stats back up my claim. From 1972 –1984, the years Reuschel was on the Cubs, he was 23rd among all major league pitchers with 1000+ IP in HRs allowed relative to the league average (thanks to the Lee Sinins Sabermetric Encyclopedia). He allowed 25% HRs fewer than the average pitcher would have, pitching in Wrigley Field! Wrigley was a great HR park during this period, compared to other NL parks, allowing 42% more HRs than average. Yet Reuschel was one of the best in baseball at preventing HRs during this period!
For his entire career, Reuschel gave up 21.6% fewer HRs than average. This is 41st for all pitchers from 1920-2004 with 2000+ IP. Pitchers he is ahead of include
Randy Johnson
Dwight Gooden
Bob Feller
Bob Lemon
Dazzy Vance
Allie Reynolds
Lefty Gomez
Warren Spahn
Herb Pennock
Waite Hoyt
Jim Palmer
He is 50th in strike-out-to-walk ratio relative to the league average in all of baseball history for pitchers with 3000+IP. His strike-out-to-walk ratio was 2.16, 31% better than the league average of 1.65. Pitchers he is ahead of include
Catfish Hunter
Warren Spahn
Bob Gibson
Whitey Ford
Nolan Ryan
Waite Hoyt
Jack Morris
Orel Hershiser
Phil Niekro
Jim Palmer
Ted Lyons
Jim Palmer, for example, had the luxury of great fielders behind him, like Brooks Robinson, Mark Belanger, Dave Johnson, Bobby Grich and Paul Blair. Did the Cubs have anyone that good from 1972-1984?
So, we can see that Reuschel was very, very good in things that the pitcher mostly controls: HRs, BBs and Ks. His high rankings in these stats warrant his induction into the Hall of Fame, especially when we see some of the pitchers he is ahead of.
Technical note: In using the park effects to adjust for HRs, I take the factor, say 120, and find the number that is half-way between it and 100 since a pitcher only pitches half his innings in his own park. Then I would use 110. If a pitcher allowed 1 HR per 9 IP, I divide 1 by 1.10 and get .91. Then I see how far from the league average that is. That difference gets divided by the standard deviation. Then that gets plugged into the formula.
-His strike-out-to-walk ratio was 31% better than the league average
-He gave up 21.6% fewer HRs than average
Yes, I’m crazy. But not because I think Reuschel deserves to be in the Hall of Fame. What makes me crazy is that I was doing a lot ridiculously time-consuming analysis that opened my eyes. As of now, I don’t know how he has done in the voting or if he is still eligible for induction based on the writer’s vote or if he has to wait for the Veterans Committee.
Before I get into explaining the actual analysis, lets review a few things. A pitcher needs to prevent the other team from scoring and, in this endeavor, he can be aided or hindered by his fielders. So one thing to look at are his defense-independent pitching stats. Many of you probably know that I am taking a page out of Voros McCracken’s book on this. Readers who have not heard of him, Google his name. He came up with the idea a few years ago that on balls in play (not HRs, walks or strikeouts), the batting average that most pitchers allow is not too much different from other pitchers and that it appears to be mainly influenced by the fielders and hitters. Analysts Tom Tippett and Mike Emeigh, to name just two, have challenged McCracken’s thesis. But it is still pretty good.
So we should look at the things a pitcher controls, like HRs, walks and strikeouts (and even if the pitcher has some influence on the batting average on balls in play, the fielders and hitters still play a role, so we are still isolating just what the pitcher does and so it is a legitimate analysis). But how a pitcher does in HRs, walks or strikeouts must be put into context. They should be adjusted for the league average and for park effects.
So using the linear regression technique, I came up with a formula for estimating a pitcher’s ERA. I looked at all pitching seasons from 1920-2000 with 150+ IP. Here is the formula that I got
ERA = .44*HR + .4*BB - .3*K
The intercept or constant term was less than .0001. The r-squared was .422 and the standard error was .75. But for each pitcher I used not his ERA, but how much he differed from the league average in the given year (actually how many standard deviations from the mean he was). Same thing for the HRs, walks and strikeouts. Using standard deviations, since they measure dispersion, does a better job of placing a pitcher in the context of his season than what percent above or below the mean they are in some stat (or the absolute difference) since we see where a pitcher fits in the statistical distribution of the given year.
Once I had this equation, I plugged in each pitcher's data on HRs, BBs and Ks and ranked them all (2000+ IP minimum during the 1920-2000 period) in terms of how many standard deviations below the mean they were in their careers (using only seasons when they pitched 150+ IP-data came from the Sean Lahman data base). Each pitcher’s HRs were adjusted for park effects before their HRs were plugged into the formula. I used park effect data from fellow economist and SABR member Ron Selter. So if a pitcher was 1 standard deviation below the mean, he got a -.44. If he were, say, half a standard deviation better on walks and strikeouts, he gets a -.2 and a -.15. So he would come out -.79. How I used the park data is explained below in technical notes (I don’t have any data on how parks affect strikeouts and walks so those were not adjusted).
Rick Reuschel was 14th! Yes. That seems to be a high enough ranking in an 80-year period to merit the Hall of Fame. Here are the top 20 in terms of how many standard deviations below the average ERA they were for their careers:
Dazzy Vance –1.25
Lefty Grove –1.21
Roger Clemens –1.18
Greg Maddux –1.14
Carl Hubbell -.94
Randy Johnson -.92
Kevin Brown -.91
Dwight Gooden -.89
Mike Garcia -.84
Hal Newhouser -.82
Sandy Koufax -.81
Bert Blyleven -.75
Ron Guidry -.73
Rick Reuschel -.71
G. Alexander -.70
Gaylord Perry -.68
Urban Schocker -.65
John Smoltz -.65
Lefty Gomez -.65
Bob Gibson -.62
(I also looked at ERA relative to the league average in addition to this standard deviation technique and he would be 20th, still a very high ranking). Reuschel is in obviously very good company. This means that, when only looking at pitcher controlled factors, he was outstanding at preventing runs in the context of his era and parks. This was done while pitching 3500 innings (38th since 1920) and winning over 200 games.
Some conventional stats back up my claim. From 1972 –1984, the years Reuschel was on the Cubs, he was 23rd among all major league pitchers with 1000+ IP in HRs allowed relative to the league average (thanks to the Lee Sinins Sabermetric Encyclopedia). He allowed 25% HRs fewer than the average pitcher would have, pitching in Wrigley Field! Wrigley was a great HR park during this period, compared to other NL parks, allowing 42% more HRs than average. Yet Reuschel was one of the best in baseball at preventing HRs during this period!
For his entire career, Reuschel gave up 21.6% fewer HRs than average. This is 41st for all pitchers from 1920-2004 with 2000+ IP. Pitchers he is ahead of include
Randy Johnson
Dwight Gooden
Bob Feller
Bob Lemon
Dazzy Vance
Allie Reynolds
Lefty Gomez
Warren Spahn
Herb Pennock
Waite Hoyt
Jim Palmer
He is 50th in strike-out-to-walk ratio relative to the league average in all of baseball history for pitchers with 3000+IP. His strike-out-to-walk ratio was 2.16, 31% better than the league average of 1.65. Pitchers he is ahead of include
Catfish Hunter
Warren Spahn
Bob Gibson
Whitey Ford
Nolan Ryan
Waite Hoyt
Jack Morris
Orel Hershiser
Phil Niekro
Jim Palmer
Ted Lyons
Jim Palmer, for example, had the luxury of great fielders behind him, like Brooks Robinson, Mark Belanger, Dave Johnson, Bobby Grich and Paul Blair. Did the Cubs have anyone that good from 1972-1984?
So, we can see that Reuschel was very, very good in things that the pitcher mostly controls: HRs, BBs and Ks. His high rankings in these stats warrant his induction into the Hall of Fame, especially when we see some of the pitchers he is ahead of.
Technical note: In using the park effects to adjust for HRs, I take the factor, say 120, and find the number that is half-way between it and 100 since a pitcher only pitches half his innings in his own park. Then I would use 110. If a pitcher allowed 1 HR per 9 IP, I divide 1 by 1.10 and get .91. Then I see how far from the league average that is. That difference gets divided by the standard deviation. Then that gets plugged into the formula.
Friday, June 5, 2009
Can Joe Mauer Bat .400 This Year?
His average right now is .436 and people are already talking about it. You can read a couple of articles about this here and here and here. But he only has 110 at-bats so far. He will need to hit about .390 the rest of the way to finish at .400 (assuming about 3.5 at-bats per game and 107 more games).
Just about a year ago I had a post on Chipper Jones batting .400. You can read that here. He was at .421 on June 6, in over 200 ABs but he finished at .364. In that post, I looked at the previous .400 hitters. What I found was that on average, since 1900, the mean of their previous career average was .342, they hit .378 on average the year before, their average age was 27 and the league average in the year they hit .400 was, on average, .282.
How does Mauer compare to this profile? He came in to 2009 with a career average of .317 and hit .328 last year. So those numbers are not very close to the .400 hitters. His age is about right. But the league average in the AL this year is just .267, well below the norm of .282. When Ted Williams hit .406 in 1941 (the last time anyone hit .400), the league average was .266. But that included the pitchers. Taking them out, the average was .276. So except for age, Mauer does not fit the profile of a .400 hitter.
One more thing. The highest average ever recorded by a catcher who qualified for the batting title was .362 by Mike Piazza in 1997. That means that he (Mauer) would be 38 points ahead of the next best. Here are the top 2 in season average for the other 7 positions. No one else has a 38 point lead
1B
George Sisler 1922 .420
George Sisler 1920 .407
Bill Terry 1930 .401
2B
Nap Lajoie 1901 .426
Rogers Hornsby 1924 .424
SS
Hughie Jennings 1896 .401
Luke Appling 1936 .388
3B
John McGraw 1899 .391
George Brett 1980 .390
LF
Tip O'Neill 1887 .435
Ed Delahanty 1899 .410
Jesse Burkett 1896 .410
CF
Hugh Duffy 1894 .440
Ty Cobb 1911 .420
RF
Willie Keeler 1897 .424
Joe Jackson 1911 .408
The only one with a lead even close to 38 points is O'Neill. So a .400 average from Mauer would be simply unprecedented by this other standard
Just about a year ago I had a post on Chipper Jones batting .400. You can read that here. He was at .421 on June 6, in over 200 ABs but he finished at .364. In that post, I looked at the previous .400 hitters. What I found was that on average, since 1900, the mean of their previous career average was .342, they hit .378 on average the year before, their average age was 27 and the league average in the year they hit .400 was, on average, .282.
How does Mauer compare to this profile? He came in to 2009 with a career average of .317 and hit .328 last year. So those numbers are not very close to the .400 hitters. His age is about right. But the league average in the AL this year is just .267, well below the norm of .282. When Ted Williams hit .406 in 1941 (the last time anyone hit .400), the league average was .266. But that included the pitchers. Taking them out, the average was .276. So except for age, Mauer does not fit the profile of a .400 hitter.
One more thing. The highest average ever recorded by a catcher who qualified for the batting title was .362 by Mike Piazza in 1997. That means that he (Mauer) would be 38 points ahead of the next best. Here are the top 2 in season average for the other 7 positions. No one else has a 38 point lead
1B
George Sisler 1922 .420
George Sisler 1920 .407
Bill Terry 1930 .401
2B
Nap Lajoie 1901 .426
Rogers Hornsby 1924 .424
SS
Hughie Jennings 1896 .401
Luke Appling 1936 .388
3B
John McGraw 1899 .391
George Brett 1980 .390
LF
Tip O'Neill 1887 .435
Ed Delahanty 1899 .410
Jesse Burkett 1896 .410
CF
Hugh Duffy 1894 .440
Ty Cobb 1911 .420
RF
Willie Keeler 1897 .424
Joe Jackson 1911 .408
The only one with a lead even close to 38 points is O'Neill. So a .400 average from Mauer would be simply unprecedented by this other standard
Thursday, June 4, 2009
Is Lefty Grove The Most Underrated Player In History?
Joe Posnanski thinks so. I concur (and not just because Joe has a Polish sounding last name, although that is often reason enough). See his post Lefty. Grove did very poorly in a recent opinion poll on who was the greatest lefthanded pitcher ever.
I have written several articles that show Grove may have been the best ever. I took park effects into account, normalized to the league average, used fielding independent ERA, and calculated wins above replacement level. I also looked at the best 5-year performances and Grove almost always comes out on top. In fact, he often had 2 distinct 5-year periods among the leaders. Here are the articles:
The Best Five-Year Pitching Performances
The Best Five-Year Pitching Performances Since 1920 Based on Fielding Independent ERA
The Best Pitchers Since 1920
The All-Time Leaders in Park-Adjusted Pitching Wins Above Replacement Level
Grove appears to be the top lefty on this last one, although if 2006 through today were counted, Randy Johnson might be better. But through last year, Randy Johnson was still 133 runs saved behind Grove in about 100 more IP (adjusting for park effects and league average, from the Lee Sinins Complete Baseball Encyclopedia).
I have written several articles that show Grove may have been the best ever. I took park effects into account, normalized to the league average, used fielding independent ERA, and calculated wins above replacement level. I also looked at the best 5-year performances and Grove almost always comes out on top. In fact, he often had 2 distinct 5-year periods among the leaders. Here are the articles:
The Best Five-Year Pitching Performances
The Best Five-Year Pitching Performances Since 1920 Based on Fielding Independent ERA
The Best Pitchers Since 1920
The All-Time Leaders in Park-Adjusted Pitching Wins Above Replacement Level
Grove appears to be the top lefty on this last one, although if 2006 through today were counted, Randy Johnson might be better. But through last year, Randy Johnson was still 133 runs saved behind Grove in about 100 more IP (adjusting for park effects and league average, from the Lee Sinins Complete Baseball Encyclopedia).
Tuesday, June 2, 2009
Homeruns And The New Yankee Stadium (they are about 55% more likely than elsewhere)
Just in case no one else has posted this, Yankee batters have a HR% at home of 5.6% while it is 3.5% on the road. So the frequency is 60% higher at home (5.6/3.5 = 1.6). Yankee pitchers allow a HR% of 4.9% at home and 3.24% on the road. So that frequency is 51% higher at the new Yankee stadium (4.9/3.24 = 1.51).
That 5.6% the Yankee batters have at home would mean that each batter in your lineup would hit about 35 HRs for the whole season, assuming each batter in the lineup got 620 ABs (last year the average AL team had 5,580 ABs and that divided by 9 is 620). Then .056*620 = 34.72. Of course, you would have to do that both home and away.
That 5.6% the Yankee batters have at home would mean that each batter in your lineup would hit about 35 HRs for the whole season, assuming each batter in the lineup got 620 ABs (last year the average AL team had 5,580 ABs and that divided by 9 is 620). Then .056*620 = 34.72. Of course, you would have to do that both home and away.
Monday, June 1, 2009
Are Mauer And Morneau The New M & M Boys?
Back in the 1960s, slugging Yankee teammates Roger Maris and Mickey Mantle were sometimes called the "M & M boys." I don't recall if Willie Mays and Willie McCovey got that nickname (although a Topps baseball card called them "fence busters" even though they did not study to be come cops like Shaquille O'Neal). Now, the Twins have Justin Morneau and Joe Mauer. They both just had great months in May. Morneau hit .361 with a .459 OBP and a .713 SLG (for an OPS of 1.172). Mauer did even better, with the line of .414-.500-.838, for an OPS of 1.338. So I started to wonder how that stacked up against the best months from the earlier versions of the M & M boys.
I found all the months when the three pairs of teammates both had an OPS of 1.000 or higher (minimum 20 games played). Then I added them and also multiplied them (multiplying might give a better idea of a great 1-2 punch since both players have to do well). The results are ranked by the product of their monthly OPS in the table below:

Mantle and Maris were sensational in July 1961. Mantle's line was .375-.508-.854. He hit 14 HRs and had 28 RBIs in 29 games. Maris had .330-.403-.755 with 13 HRs and 31 RBIs in 28 games. The Yankees went 20-9 while scoring 162 runs. The Twins this May did not fare so well, going only 14-16, although they did score 168 runs. Morneau had 9 HRs and 29 RBIs while Mauer had 11 & 32.
The performance of Mays and McCovey in September 1968 is amazing because it was the year of the pitcher. But the Giants entered September 12 games behind the Cardinals and still finished 9 back in 2nd place. They were eliminated on September 15 and only went 15-12.
I did try to adjust each player's OPS for the league average. In doing so, I took off .024 for Mauer & Morneau in June 2006 since in 2006 that was the difference between the 2006 NL OPS with and without pitchers included. For this past May, I used .021, the difference from the 2007 NL. So each player had his OPS divided by the relevant league average and normalized it by multiplying it by an OPS of .725. Then I summed them and multiplied them as before. The results, ranked by product, are in the table below:

Sources included Retrosheet, The Lee Sinins Complete Baseball Encyclopedia, ESPN site, and Yahoo site
I found all the months when the three pairs of teammates both had an OPS of 1.000 or higher (minimum 20 games played). Then I added them and also multiplied them (multiplying might give a better idea of a great 1-2 punch since both players have to do well). The results are ranked by the product of their monthly OPS in the table below:

Mantle and Maris were sensational in July 1961. Mantle's line was .375-.508-.854. He hit 14 HRs and had 28 RBIs in 29 games. Maris had .330-.403-.755 with 13 HRs and 31 RBIs in 28 games. The Yankees went 20-9 while scoring 162 runs. The Twins this May did not fare so well, going only 14-16, although they did score 168 runs. Morneau had 9 HRs and 29 RBIs while Mauer had 11 & 32.
The performance of Mays and McCovey in September 1968 is amazing because it was the year of the pitcher. But the Giants entered September 12 games behind the Cardinals and still finished 9 back in 2nd place. They were eliminated on September 15 and only went 15-12.
I did try to adjust each player's OPS for the league average. In doing so, I took off .024 for Mauer & Morneau in June 2006 since in 2006 that was the difference between the 2006 NL OPS with and without pitchers included. For this past May, I used .021, the difference from the 2007 NL. So each player had his OPS divided by the relevant league average and normalized it by multiplying it by an OPS of .725. Then I summed them and multiplied them as before. The results, ranked by product, are in the table below:

Sources included Retrosheet, The Lee Sinins Complete Baseball Encyclopedia, ESPN site, and Yahoo site
Saturday, May 30, 2009
Texas Rangers On A Pace To Set Power Hitting Records
I did not check to see if this was mentioned anywhere, but the Rangers have a team isolated power (ISO) of .217 through 48 games (a .485 SLG and .268 AVG). The all-time single season record for a team is .204 by the 1997 Mariners. The Yankees have .205 so far this year. The Phillies have .199, equal to the NL record set by the 2000 Astros. The Rangers have 80 HRs in 48 games which is a pace to hit 270, more that the record of 264 by 1997 Mariners. The Rangers are on a pace to get 651 extrabase hits. The 2003 Red Sox hold the record of 649.
The Rangers have 80 HRs in 1,671 ABs for a HR% of 4.79%. That is higher than the 4.70% record of the 1997 Mariners.
The Rangers only have a team OBP of .327 (the AL avg is .338). The power is probably why they are above average in runs per game (5.27 vs. 4.88). Just imagine if Josh Hamilton were hitting.
The Rangers have 80 HRs in 1,671 ABs for a HR% of 4.79%. That is higher than the 4.70% record of the 1997 Mariners.
The Rangers only have a team OBP of .327 (the AL avg is .338). The power is probably why they are above average in runs per game (5.27 vs. 4.88). Just imagine if Josh Hamilton were hitting.
Wednesday, May 27, 2009
Some Baseball Pioneers Are "Heroes Of Capitalism"
There is a great blog called Heroes of Capitalism. What is the blog about? The authors say "What is a hero of capitalism? Someone who used private property* to produce wealth."
*Private property includes the tangible (like land) and intangible (like ideas).
Here are links to their posts on baseball people they have honored:
John A. "Bud" Hillerich
Bill Veeck
Branch Rickey
Bill Beane
Hal Richman (inventor of the table top baseball game Strat-o-matic, which simulates the performance of actual major league players)
Daniel Okrent, Robert Sklar, Steve Wulf and Glen Waggoner (creators of Rotisserie League baseball, the forerunner of fantasy baseball)
Okrent has also written baseball books, including Nine Innings: The Anatomy of a Baseball Game in which you can learn alot about baseball and its history through his analysis of a single game, bewteen the Brewers and Orioles on June 10, 1982.
*Private property includes the tangible (like land) and intangible (like ideas).
Here are links to their posts on baseball people they have honored:
John A. "Bud" Hillerich
Bill Veeck
Branch Rickey
Bill Beane
Hal Richman (inventor of the table top baseball game Strat-o-matic, which simulates the performance of actual major league players)
Daniel Okrent, Robert Sklar, Steve Wulf and Glen Waggoner (creators of Rotisserie League baseball, the forerunner of fantasy baseball)
Okrent has also written baseball books, including Nine Innings: The Anatomy of a Baseball Game in which you can learn alot about baseball and its history through his analysis of a single game, bewteen the Brewers and Orioles on June 10, 1982.
Monday, May 25, 2009
Why Isn't Steve Garvey In The Hall Of Fame?
It seems like he would have been elected based on the voters preferences in recent years (I have been analyzing voting patterns and what I write below will be based on that-scroll down to see these studies). But first I want briefly to discuss the sabermetric case for or against.
Garvey had 279 career "Win Shares" (WS), the Bill James stat which incorporates all phases of the game. That tied him for 222nd place all-time among all players and pitchers through 2001. Not bad, since about 200 guys are in the Hall. But this is marginal.
His career TPR or "total player rating," from Pete Palmer, editor of the Baseball Encyclopedia was actually -6.1. That means that if an average first baseman had played instead of Garvey, those teams would have won 6.1 more games during his career. Most of his seasons were negative and his best was only +1.2.
But the baseball writers who vote don't necessarily take sabermetric stats into account. The analysis I have posted recently used more conventional stats. In one model I used logit analysis to predict the probability of any player getting elected. That model had career AVG, seasons with 100+ RBIs, ALLSTAR games, career plate appearances (PAs), MVP awards, a variable for world series performance, being in the 3000 hit club and a positional adjustment for being a catcher. That model gave Garvey a 94.7% probability of being elected to the Hall of Fame. The model itself was 98.9% accurate
Another logit model (also 98.9% accurate) had the following variables:
Career HRs
2B
SS
3B
CF
C
WSIMP
ALLSTAR
MVPSH
500SB
Career NON-HRs
3000 HIT
The variables after Career HRs are positional adjustments. The WSIMP is for world series play. This model had Garvey's probability at 64.8%. Tony Perez has 52.6% and Jim Rice has 10.6% and both are in the Hall.
Another model simply predicted the % of votes received in the first year of eligibility. This model took into account MVP awards, a variable for world series performance, being in the 3000 hit club, ALLSTAR games, being in the 500 HR club, being in the 500 SB club, Gold Glove awards and career PAs. It predicted that Garvey would be named on 48.9% of the ballots in his first year while he actually got 41.6%. His predicted 48.9% is more than what was predicted for the following players who did eventually make it:
Ryne Sandberg -0.460
Kirby Puckett -0.418
Gary Carter-0.380
Carlton Fisk-0.375
Tony Perez-0.299
Jim Rice-0.219
And Garvey's actual first year % of 41.6 is higher than that of Rice (29.8%) and very close to Carter's 42.3%.
So from 3 different regressions, it looks like Garvey had the stats or qualifications to make it in, based on what the voters seem to like.
It is also very easy to find some impressive achievements that would go on Garvey's plaque, if he ever made it. They include:
-5 100 RBI seasons
-.294 career AVG
-batted over .300 7 times
-had 200 or more hits in a season 6 times
-1974 NL MVP
-batted .319 in 5 World Series
-batted .356 in 5 league championship series
-batted .393 in 10 all-star games
-won 4 Gold Glove awards
-2 time MVP of the all-star game
-2 time MVP of the league championship series
-finished in the top 5 in total bases 7 times
-set a NL record by playing 193 straight games without committing an error
-set a ML record with his .996 fielding percentage at first base.
-played in 1,207 consecutive games, an NL record and 4th longest overall
He also has 2.46 Career MVP Shares which is the 55th best total. An MVP share is what % of the total possible points a player got in the voting in a given year. A first place vote is 14 points, 2nd place 9, 3rd, 8points, etc. A guy might come in 5th but if he had 40 points out of a maximum of, say, 400, he gets a .10. Garvey's high rank here means the voters liked him when he played, alot more than they liked other players. He finished in the top 6 in MVP voting 5 times.
Baeball Reference lists the Hall of Fame Monitor for which they say:
"This is another Jamesian creation. It attempts to assess how likely (not how deserving) an active player is to make the Hall of Fame. It's rough scale is 100 means a good possibility and 130 is a virtual cinch. It isn't hard and fast, but it does a pretty good job. Here are the batting rules."
Garvey gets a 130 which is 104th best among position players. Now this is a very complicated point system with so many points for this or that. But this shows that Garvey fits the statistical profile of the kind of player the voters very much like to put in the Hall.
So why isn't he in? I found some theories.
The Baseball Page said, among other things, the following:
"In the 1980s it became clear that "Mr. Dodger" was far from wholesome. Several paternity suits and a tell all book from his ex-wife tarnished his image irreparably. Where he once was considered a candidate for state or even national office, Garvey became a leper, destined to host game shows and infomercials (really).
He had the reputation as a selfish, egotistical player. The media didn't like him as much as it seemed. His "Mr. Dodger" persona was created by Dodger PR and a few well-placed friends in the press. More than a few teammates quickly tired of Garvey's habit of staying in front of the camera or microphone.
In August of 1978, Garvey took offense to a comment made by teammate Don Sutton and the two men ended up wrestling their way across the visitors' clubhouse in Shea Stadium. The fight cemented a bitter feud between the two men and it damaged Garvey's reputation in the league.
He aged quickly. By the time he was 31-32, his skills were rapidly diminishing. He would have benefited from a day off here and there, but he didn't do it."
Chris Jaffe over at the Harball Times had an interesting article called Hitler. Stalin. Garvey. Here is an exerpt:
"There was always a sense he was a fake. With the Dodgers, he got in a big fistfight in the clubhouse with teammate Don Sutton. He had a nasty divorce in the early 1980s. When he started to get hit with paternity suits, though, his reputation was shattered.
In some ways, though, it's even deeper than that. Our society can forgive—or at least cease baiting—a hypocrite, provided he asks for some degree of atonement. Jim Bakker wrote his book, I Was Wrong, for instance.
Garvey hasn't done that."
Jeff Sackmann also has an interesting article called Steve Garvey Gets No Respect
Update Dec. 3, 2010: I did a follow up post in July, 2010. Click here to read it.
Garvey had 279 career "Win Shares" (WS), the Bill James stat which incorporates all phases of the game. That tied him for 222nd place all-time among all players and pitchers through 2001. Not bad, since about 200 guys are in the Hall. But this is marginal.
His career TPR or "total player rating," from Pete Palmer, editor of the Baseball Encyclopedia was actually -6.1. That means that if an average first baseman had played instead of Garvey, those teams would have won 6.1 more games during his career. Most of his seasons were negative and his best was only +1.2.
But the baseball writers who vote don't necessarily take sabermetric stats into account. The analysis I have posted recently used more conventional stats. In one model I used logit analysis to predict the probability of any player getting elected. That model had career AVG, seasons with 100+ RBIs, ALLSTAR games, career plate appearances (PAs), MVP awards, a variable for world series performance, being in the 3000 hit club and a positional adjustment for being a catcher. That model gave Garvey a 94.7% probability of being elected to the Hall of Fame. The model itself was 98.9% accurate
Another logit model (also 98.9% accurate) had the following variables:
Career HRs
2B
SS
3B
CF
C
WSIMP
ALLSTAR
MVPSH
500SB
Career NON-HRs
3000 HIT
The variables after Career HRs are positional adjustments. The WSIMP is for world series play. This model had Garvey's probability at 64.8%. Tony Perez has 52.6% and Jim Rice has 10.6% and both are in the Hall.
Another model simply predicted the % of votes received in the first year of eligibility. This model took into account MVP awards, a variable for world series performance, being in the 3000 hit club, ALLSTAR games, being in the 500 HR club, being in the 500 SB club, Gold Glove awards and career PAs. It predicted that Garvey would be named on 48.9% of the ballots in his first year while he actually got 41.6%. His predicted 48.9% is more than what was predicted for the following players who did eventually make it:
Ryne Sandberg -0.460
Kirby Puckett -0.418
Gary Carter-0.380
Carlton Fisk-0.375
Tony Perez-0.299
Jim Rice-0.219
And Garvey's actual first year % of 41.6 is higher than that of Rice (29.8%) and very close to Carter's 42.3%.
So from 3 different regressions, it looks like Garvey had the stats or qualifications to make it in, based on what the voters seem to like.
It is also very easy to find some impressive achievements that would go on Garvey's plaque, if he ever made it. They include:
-5 100 RBI seasons
-.294 career AVG
-batted over .300 7 times
-had 200 or more hits in a season 6 times
-1974 NL MVP
-batted .319 in 5 World Series
-batted .356 in 5 league championship series
-batted .393 in 10 all-star games
-won 4 Gold Glove awards
-2 time MVP of the all-star game
-2 time MVP of the league championship series
-finished in the top 5 in total bases 7 times
-set a NL record by playing 193 straight games without committing an error
-set a ML record with his .996 fielding percentage at first base.
-played in 1,207 consecutive games, an NL record and 4th longest overall
He also has 2.46 Career MVP Shares which is the 55th best total. An MVP share is what % of the total possible points a player got in the voting in a given year. A first place vote is 14 points, 2nd place 9, 3rd, 8points, etc. A guy might come in 5th but if he had 40 points out of a maximum of, say, 400, he gets a .10. Garvey's high rank here means the voters liked him when he played, alot more than they liked other players. He finished in the top 6 in MVP voting 5 times.
Baeball Reference lists the Hall of Fame Monitor for which they say:
"This is another Jamesian creation. It attempts to assess how likely (not how deserving) an active player is to make the Hall of Fame. It's rough scale is 100 means a good possibility and 130 is a virtual cinch. It isn't hard and fast, but it does a pretty good job. Here are the batting rules."
Garvey gets a 130 which is 104th best among position players. Now this is a very complicated point system with so many points for this or that. But this shows that Garvey fits the statistical profile of the kind of player the voters very much like to put in the Hall.
So why isn't he in? I found some theories.
The Baseball Page said, among other things, the following:
"In the 1980s it became clear that "Mr. Dodger" was far from wholesome. Several paternity suits and a tell all book from his ex-wife tarnished his image irreparably. Where he once was considered a candidate for state or even national office, Garvey became a leper, destined to host game shows and infomercials (really).
He had the reputation as a selfish, egotistical player. The media didn't like him as much as it seemed. His "Mr. Dodger" persona was created by Dodger PR and a few well-placed friends in the press. More than a few teammates quickly tired of Garvey's habit of staying in front of the camera or microphone.
In August of 1978, Garvey took offense to a comment made by teammate Don Sutton and the two men ended up wrestling their way across the visitors' clubhouse in Shea Stadium. The fight cemented a bitter feud between the two men and it damaged Garvey's reputation in the league.
He aged quickly. By the time he was 31-32, his skills were rapidly diminishing. He would have benefited from a day off here and there, but he didn't do it."
Chris Jaffe over at the Harball Times had an interesting article called Hitler. Stalin. Garvey. Here is an exerpt:
"There was always a sense he was a fake. With the Dodgers, he got in a big fistfight in the clubhouse with teammate Don Sutton. He had a nasty divorce in the early 1980s. When he started to get hit with paternity suits, though, his reputation was shattered.
In some ways, though, it's even deeper than that. Our society can forgive—or at least cease baiting—a hypocrite, provided he asks for some degree of atonement. Jim Bakker wrote his book, I Was Wrong, for instance.
Garvey hasn't done that."
Jeff Sackmann also has an interesting article called Steve Garvey Gets No Respect
Update Dec. 3, 2010: I did a follow up post in July, 2010. Click here to read it.
Friday, May 15, 2009
The Marginal Impact Statistics Have on Hall of Fame Voting
A couple of weeks ago I presented a binary (logit) model of Hall of Fame voting. I looked at all players whose first year of eligibility was from 1990-2009 (except for Pete Rose). The model's equation estimates the probability that a player would be elected to the Hall of Fame (though not necessarily in his first year of eligibility). Of course, there are coefficient estimates and what I do here is give the change in the probability of being elected due to a change in a hypothetical player's stats.
I made up a player who had a 0.290 career AVG, had 6 100 RBI seasons, won 1 MVP award, played in 8 all-star games and had 8,894 career plate appearances. This player was not a catcher and did not achieve 3,000 hits. He had a world series impact of 18 (that is, his world series PAs times his world series OPS). It is like he had a .750 OPS in 24 world series PAs.
The first line in the table below shows that his probability of being elected was .50 or 50%. The rest of the table shows what his probability would be if his stats changed. For instance, if his career average had been .300 instead of .290 (with nothing else changing), his probability (or PR) rises to .663. Another 10 point gain in average brings it to .795. But if he had only hit .280, the PR is just .337. Of course, other things might have changed if his averaged had changed. Maybe 100 RBI seasons or all-star games played would have been different. But I will have to assume that they did not. The different cases for AVG are in red. More discussion follows the table. You can see a larger version of the table if you click on it.

Adding 1 100 RBI seasons raises PR to .638. Taking 1 away lowers it to .362. So the 100 RBI seasons has a powerful impact, like AVG. Adding an MVP award raises PR to .59.But all-star games played in (AS) might have the strongest effect. Going from 8 to 9 all-star games raises the PR to .80. If all-star games falls to 7, PR is just .20.
Having 3000 hits, not surprisingly, raises the PR to 1.00 or 100% (actually its a little less but it is above 99.99%). If the player had been a catcher (see the 1 in the C column), it jumps to .973. The WS or world series impact is very slight. Adding or subtracting 500 career PAs matters alot. Adding 500 PAs bumps PR up to .657 and taking 500 away drops it to .343.
I also did the same test for an actual player, Jim Rice. His predicted PR was the closest of any of the players in the study to .50 at .595. The next table has the marginal changes like the one above. Maybe the biggest impact for Rice is all-star games. If he had just one less, his PR falls to .269. In general, his table looks similar to that of the hypothetical player.
I made up a player who had a 0.290 career AVG, had 6 100 RBI seasons, won 1 MVP award, played in 8 all-star games and had 8,894 career plate appearances. This player was not a catcher and did not achieve 3,000 hits. He had a world series impact of 18 (that is, his world series PAs times his world series OPS). It is like he had a .750 OPS in 24 world series PAs.
The first line in the table below shows that his probability of being elected was .50 or 50%. The rest of the table shows what his probability would be if his stats changed. For instance, if his career average had been .300 instead of .290 (with nothing else changing), his probability (or PR) rises to .663. Another 10 point gain in average brings it to .795. But if he had only hit .280, the PR is just .337. Of course, other things might have changed if his averaged had changed. Maybe 100 RBI seasons or all-star games played would have been different. But I will have to assume that they did not. The different cases for AVG are in red. More discussion follows the table. You can see a larger version of the table if you click on it.

Adding 1 100 RBI seasons raises PR to .638. Taking 1 away lowers it to .362. So the 100 RBI seasons has a powerful impact, like AVG. Adding an MVP award raises PR to .59.But all-star games played in (AS) might have the strongest effect. Going from 8 to 9 all-star games raises the PR to .80. If all-star games falls to 7, PR is just .20.
Having 3000 hits, not surprisingly, raises the PR to 1.00 or 100% (actually its a little less but it is above 99.99%). If the player had been a catcher (see the 1 in the C column), it jumps to .973. The WS or world series impact is very slight. Adding or subtracting 500 career PAs matters alot. Adding 500 PAs bumps PR up to .657 and taking 500 away drops it to .343.
I also did the same test for an actual player, Jim Rice. His predicted PR was the closest of any of the players in the study to .50 at .595. The next table has the marginal changes like the one above. Maybe the biggest impact for Rice is all-star games. If he had just one less, his PR falls to .269. In general, his table looks similar to that of the hypothetical player.
Friday, May 8, 2009
Does The "High" April Slugging Percentage Mean Anything?
(note: I will have another post on Hall of Fame voting next weekend)
It sure seemed like there was lots of slugging going on in April. The SLG for MLB was .420 for the month, much higher than the .402 and .401 for the two previous seasons. The table below shows the SLG for April and post-April for each of the last 10 years (5 years had some March games, so I included that data in April for those years).

The average April SLG over the last 10 years is just about .420. So it may not be that high, although in beats 5 of the last 7 seasons.
The graph below shows the relationship between April SLG (horizontal axis) and non-April SLG (vertical axis).

The average non-April SLG is about .426. I also ran a regression with non-April SLG as the dependent variable and April SLG as the independent variable. Here is the equation:
non-April SLG = 0.335*(April SLG) + 0.2856
The r-squared was .68 and the standard error was .0037, which seems pretty low. The equation predicts all non-April SLGs with +/- .006. So I guess we can expect SLG the rest of the year to be between .414 and .426 and it may be a good bet for it to be between .416 and .424. But so far in May it is only .410, so who knows. Maybe the weather matters and maybe this April was unusually warm.
It sure seemed like there was lots of slugging going on in April. The SLG for MLB was .420 for the month, much higher than the .402 and .401 for the two previous seasons. The table below shows the SLG for April and post-April for each of the last 10 years (5 years had some March games, so I included that data in April for those years).

The average April SLG over the last 10 years is just about .420. So it may not be that high, although in beats 5 of the last 7 seasons.
The graph below shows the relationship between April SLG (horizontal axis) and non-April SLG (vertical axis).

The average non-April SLG is about .426. I also ran a regression with non-April SLG as the dependent variable and April SLG as the independent variable. Here is the equation:
non-April SLG = 0.335*(April SLG) + 0.2856
The r-squared was .68 and the standard error was .0037, which seems pretty low. The equation predicts all non-April SLGs with +/- .006. So I guess we can expect SLG the rest of the year to be between .414 and .426 and it may be a good bet for it to be between .416 and .424. But so far in May it is only .410, so who knows. Maybe the weather matters and maybe this April was unusually warm.
Sunday, May 3, 2009
Predicting Who Makes The Hall Of Fame Using A Logit Model
The last two weeks I presented regression results on first year Hall of Fame vote percentage. I used a linear regression. This week I use a logit model where 1 means the player has made it in and 0 means not. The probability that a player was voted in is
(1) P = 1/(1 + exp(-Z))
where exp is approximately 2.78 or Euhler's number. The -Z is the following equation times -1 for each player. This is the estimatated regression equation:
(2) -46.1955 + 67.81*CAVG + .5658*100RBI* + 1.386*ALLSTAR + .0013*PA + .3645*MVP + .0001*WSIMP + 14.93*3000HIT + 3.586*C
CAVG is a player's career batting average, 100RBI is the number of seasons with 100+ RBIs, ALLSTAR is number of all-star games played in, PA is career plate appearances, MVP is number of MVP awards won, WSIMP is world series PAs times world series OPS (this is world series impact, a combination of quantity of quality), 3000HIT is a dummy variable (1 or 0) for reaching that milestone and C is the same if the player was a catcher. So all the data gets plugged in for each player and "Z" is calculated. Then the negative of that is plugged in to equation (1) to get the probability of each player being elected into the Hall.
The statistical results can be viewed at logit results. The data for each player and their calculated probability can be seen at logit probabilities.
The results page shows something called a "classification table." It says that if a probability is .5 or greater, the player will be in the Hall. My data includes all players whose first year of eligibility was 1990 or later (except for Pete Rose). The classification table says that all 20 players who actually have made it in should be (all 20 have a probability of .5 or higher). Of the other 161 players it predicts that 159 would not make it. The two who are predicted to make it are Steve Garvey and Andre Dawson. Garvey's 15 years of eligibility to be voted in by the writers has ended. But Dawson could still make it. Garvey had a probability of 95.7% according to the model. Dawson had 60.8%. The model is 98.9% correct.
I tried lots of variables and this model works the best in terms of getting all the actual Hall of Famers right and the overall correct%. Also, some models might have done a bit better but a variable would be negative (like career HRs) that should not be. Some models had many more variables. My Acastat programs sometimes could not complete a regression no matter how long I waited. That might have been if I had more than 12 variables. So another model might be a bit better. I just don't know.
Now it is hard to see what exactly the impact of each variable is since this is a non-linear model. So I will just mention a few players and how their probability (P)would change if their data changed.
Yount-If he does not have 3000 hits, his P falls to about 1% from 99%! Molitor would fall from 99% to 45%.
Ozzie Smith-If he falls from 14 all-star games to 10, his P falls from 99% to 36%
Fisk-If he is not a catcher, his P falls from 96% to 46%. Gary Carter falls from 65% to 36%.
Carew-If he does not have 3000 hits, his P falls only about 1% from 100% to 99%! (Boggs, Gwynn and Brett have the same thing)
Eddie Murray-If he does not have 3000 hits, his P falls to about 83% from 99%.
Puckett-If his CAVG falls from .318 to .301 his P falls from 75% to 49% (a NO). If you don't change his CAVG but take his All-Star games from 10 to 9, his P falls to 43%. Incidentally, he finished his career with 280 Win Shares and if he could have played 5 more years he could have easily gotten to 360 Win Shares, what Bill James says is almost a lock for the Hall (but Tim Raines has 392 and might not make it-if he went from 7 to 10 all-star games his P would be about 75%).
Tony Perez-If you take his All-Star games from 7 to 6, his P falls to 29% from 62%.
McGwire-Maybe his not getting in is why 500 HRs was not working very well in the model. The only others were Murray, Jackson and Schmidt. But if he had 9200 PAs, his P value jumps to .5. Of course, he has the steroid scandal. It would be nice to quantify scandals, but that may be impossible and I have already take Rose out of the model. Back to McGwire, if he had 11 all-star games instead of 9, his P would go up to 69%.
Ryne Sandberg-Take away his MVP award and his P falls from 63% to 54% (do that and take away 1 all-star game, he falls to 23%). For some other guys, it makes almost no difference. Take all 3 of Schmidt's awards away and he still has a 99% P. Take away 1 of Morgan's awards and he falls from 66% to 58%. Take away the other and he falls to 48%.
Al Oliver would go from a P of 10% to over 50% if he had 6 100 RBI seasons instead of 2. Same for Ted Simmons if he jumped from 3 to 7 100 RBI seasons. Same for Harold Baines.
So all-star games and 3000 hits matter alot (and 100 RBI seasons, too). If Dave Parker had 8 all-star game instead of 6, he goes from a P of 8% to over 60%. Harold Baines would have a P of 99% if he had 3000 hits. The 3 strike shortened seasons of 1981, 1994 and 1995 might have cost him 3000 hits. Will Clark and Keith Hernandez would have P's of 99% if they made 3000 hits.
I can email my spread sheet to anyone if you want to play around with these kinds of possibilities.
(1) P = 1/(1 + exp(-Z))
where exp is approximately 2.78 or Euhler's number. The -Z is the following equation times -1 for each player. This is the estimatated regression equation:
(2) -46.1955 + 67.81*CAVG + .5658*100RBI* + 1.386*ALLSTAR + .0013*PA + .3645*MVP + .0001*WSIMP + 14.93*3000HIT + 3.586*C
CAVG is a player's career batting average, 100RBI is the number of seasons with 100+ RBIs, ALLSTAR is number of all-star games played in, PA is career plate appearances, MVP is number of MVP awards won, WSIMP is world series PAs times world series OPS (this is world series impact, a combination of quantity of quality), 3000HIT is a dummy variable (1 or 0) for reaching that milestone and C is the same if the player was a catcher. So all the data gets plugged in for each player and "Z" is calculated. Then the negative of that is plugged in to equation (1) to get the probability of each player being elected into the Hall.
The statistical results can be viewed at logit results. The data for each player and their calculated probability can be seen at logit probabilities.
The results page shows something called a "classification table." It says that if a probability is .5 or greater, the player will be in the Hall. My data includes all players whose first year of eligibility was 1990 or later (except for Pete Rose). The classification table says that all 20 players who actually have made it in should be (all 20 have a probability of .5 or higher). Of the other 161 players it predicts that 159 would not make it. The two who are predicted to make it are Steve Garvey and Andre Dawson. Garvey's 15 years of eligibility to be voted in by the writers has ended. But Dawson could still make it. Garvey had a probability of 95.7% according to the model. Dawson had 60.8%. The model is 98.9% correct.
I tried lots of variables and this model works the best in terms of getting all the actual Hall of Famers right and the overall correct%. Also, some models might have done a bit better but a variable would be negative (like career HRs) that should not be. Some models had many more variables. My Acastat programs sometimes could not complete a regression no matter how long I waited. That might have been if I had more than 12 variables. So another model might be a bit better. I just don't know.
Now it is hard to see what exactly the impact of each variable is since this is a non-linear model. So I will just mention a few players and how their probability (P)would change if their data changed.
Yount-If he does not have 3000 hits, his P falls to about 1% from 99%! Molitor would fall from 99% to 45%.
Ozzie Smith-If he falls from 14 all-star games to 10, his P falls from 99% to 36%
Fisk-If he is not a catcher, his P falls from 96% to 46%. Gary Carter falls from 65% to 36%.
Carew-If he does not have 3000 hits, his P falls only about 1% from 100% to 99%! (Boggs, Gwynn and Brett have the same thing)
Eddie Murray-If he does not have 3000 hits, his P falls to about 83% from 99%.
Puckett-If his CAVG falls from .318 to .301 his P falls from 75% to 49% (a NO). If you don't change his CAVG but take his All-Star games from 10 to 9, his P falls to 43%. Incidentally, he finished his career with 280 Win Shares and if he could have played 5 more years he could have easily gotten to 360 Win Shares, what Bill James says is almost a lock for the Hall (but Tim Raines has 392 and might not make it-if he went from 7 to 10 all-star games his P would be about 75%).
Tony Perez-If you take his All-Star games from 7 to 6, his P falls to 29% from 62%.
McGwire-Maybe his not getting in is why 500 HRs was not working very well in the model. The only others were Murray, Jackson and Schmidt. But if he had 9200 PAs, his P value jumps to .5. Of course, he has the steroid scandal. It would be nice to quantify scandals, but that may be impossible and I have already take Rose out of the model. Back to McGwire, if he had 11 all-star games instead of 9, his P would go up to 69%.
Ryne Sandberg-Take away his MVP award and his P falls from 63% to 54% (do that and take away 1 all-star game, he falls to 23%). For some other guys, it makes almost no difference. Take all 3 of Schmidt's awards away and he still has a 99% P. Take away 1 of Morgan's awards and he falls from 66% to 58%. Take away the other and he falls to 48%.
Al Oliver would go from a P of 10% to over 50% if he had 6 100 RBI seasons instead of 2. Same for Ted Simmons if he jumped from 3 to 7 100 RBI seasons. Same for Harold Baines.
So all-star games and 3000 hits matter alot (and 100 RBI seasons, too). If Dave Parker had 8 all-star game instead of 6, he goes from a P of 8% to over 60%. Harold Baines would have a P of 99% if he had 3000 hits. The 3 strike shortened seasons of 1981, 1994 and 1995 might have cost him 3000 hits. Will Clark and Keith Hernandez would have P's of 99% if they made 3000 hits.
I can email my spread sheet to anyone if you want to play around with these kinds of possibilities.
Sunday, April 26, 2009
What Determines Vote Percentage In The First Year Of Hall Of Fame Eligibility? (Part 2)
I have tried to improve the model I used last week. It was an OLS (linear) regression model (next week I will present a logit or binary model which is non-liner simply looks to see if a player has made the Hall or not-it seems very accurate). This week I have converted some of last week's variables into a non-linear variables and I have added two new variables. At the end of this post I have links to some other research and discussion on this issue.
Here is the regression equation:
PCT = -.041 + .054*MVP + .432*3000H + .172*500HR + .004*ASSQ10 + .001*GGSQ7 + .074*500SB + .00001*WSIMPSQ50 + .102*10000PA
I will explain what the variables mean below. The adjusted r-squared was .901 (last week it was .839). So 90.1% of the difference across players is explained by the equation. The standard error was .081, down from .102. There were 181 players, all of those who came up for the first time from 1990-2009, except for Pete Rose(see last week). MVP is number of MVP awards won, 3000H is a dummy variable (1 if a player reached it, 0 otherwise). The 500HR is also a dummy variable as it is for 500SB and 10000PA (if you made it to 10,000 career plate appearances, you get a 1, 0 otherwise). I used all the voting data from 1990-2009.
What is ASSQ10? It is the square of the number of All-star games played in squared. But AS games played is maxed out at 10. The assumption here is that being an all-star has a positive exponential effect but only up to a point where no more games helps (I have a graph below to help explain this). The GGSQ7 is the same thing for Gold Gloves.
WSIMPSQ50 involves World Series play. First, WSIMP is World Series PAs times OPS. The idea here that the more you play in the World Series the more votes you would get, but by multiplying it by OPS, it also includes how well you played (or just hit). This gets maxed out at 50 and is squared, for the same reason as all-star games (yes, Reggie Jackson is first here and way ahead of everyone else at 141, with Dave Justice and Lonnie Smith tied for 2nd at 101).
All of the variables were significant at the 10% level except for GGSQ7, which came close with a p-value of .13. The other variables all had p-values of under .02 with 6 under .01.
This has been the best linear model I can come up with so far. Many different variables have been tried. 154 of the 181 players were predicted to within 10 percentage points. 126 or 69.6% were within 5 points.
The graph below illustrates what I mentioned above about squaring and capping variables. Notice that the line is increasing exponentially then flatlines.

Below are the players who had the biggest negative prediction differentials. Lynn, for example, was predicted to get 35.8% (.358) of the vote but got only 5.5% (.055).
0.055 *** -0.303 Lynn, Fred
0.235 *** -0.230 McGwire, Mark
0.017 *** -0.180 Bell, Buddy
0.083 *** -0.154 Nettles, Graig
0.053 *** -0.152 Baines, Harold
0.017 *** -0.151 Parrish, Lance
0.005 *** -0.144 Lopes, Davey
0.068 *** -0.140 Concepcion, Dave
0.019 *** -0.131 Cey, Ron
0.051 *** -0.130 Hernandez, Keith
Now the players who had the biggest positive prediction differentials.
0.298 *** 0.080 Rice, Jim
0.852 *** 0.095 Molitor, Paul
0.157 *** 0.109 Trammell, Alan
0.965 *** 0.111 Schmidt, Mike
0.775 *** 0.126 Yount, Robin
0.818 *** 0.164 Morgan, Joe
0.500 *** 0.189 Perez, Tony
0.664 *** 0.296 Fisk, Carlton
0.917 *** 0.319 Smith, Ozzie
0.821 *** 0.394 Puckett, Kirby
I also ran a regression which had the following variables. None of them were squared or made non-linear. 2B, SS and C are positional dummies. All of these variables were significant at the 10% level. 9 were significant at the 5% level and 7 the 1% level. But the adjusted r-squared was .763 and he standard error was .126. So it does not work nearly as well as the model mentioned above.
AVG
MVP
3000 HIT
2B
SS
C
SB
WSIMP
10000PA
HR
Now some links to other research
Baseball Hall of Fame voting: a test of the customer discrimination by Arna Desser, James Monks and Michael Robinson
Social Science Quarterly Sept 1999 v80 i3 p591(13)
Discussion of above article
A neural-net Hall of Fame prediction method
Teaching Statistical Thinking Using the Baseball Hall of Fame by Steven Wang
Who's missing from the hall of fame by JC Bradbury
Modeling Election to the Major League Baseball Hall of Fame through the use of Genetic Algorithms By David Cohen
Here is the regression equation:
PCT = -.041 + .054*MVP + .432*3000H + .172*500HR + .004*ASSQ10 + .001*GGSQ7 + .074*500SB + .00001*WSIMPSQ50 + .102*10000PA
I will explain what the variables mean below. The adjusted r-squared was .901 (last week it was .839). So 90.1% of the difference across players is explained by the equation. The standard error was .081, down from .102. There were 181 players, all of those who came up for the first time from 1990-2009, except for Pete Rose(see last week). MVP is number of MVP awards won, 3000H is a dummy variable (1 if a player reached it, 0 otherwise). The 500HR is also a dummy variable as it is for 500SB and 10000PA (if you made it to 10,000 career plate appearances, you get a 1, 0 otherwise). I used all the voting data from 1990-2009.
What is ASSQ10? It is the square of the number of All-star games played in squared. But AS games played is maxed out at 10. The assumption here is that being an all-star has a positive exponential effect but only up to a point where no more games helps (I have a graph below to help explain this). The GGSQ7 is the same thing for Gold Gloves.
WSIMPSQ50 involves World Series play. First, WSIMP is World Series PAs times OPS. The idea here that the more you play in the World Series the more votes you would get, but by multiplying it by OPS, it also includes how well you played (or just hit). This gets maxed out at 50 and is squared, for the same reason as all-star games (yes, Reggie Jackson is first here and way ahead of everyone else at 141, with Dave Justice and Lonnie Smith tied for 2nd at 101).
All of the variables were significant at the 10% level except for GGSQ7, which came close with a p-value of .13. The other variables all had p-values of under .02 with 6 under .01.
This has been the best linear model I can come up with so far. Many different variables have been tried. 154 of the 181 players were predicted to within 10 percentage points. 126 or 69.6% were within 5 points.
The graph below illustrates what I mentioned above about squaring and capping variables. Notice that the line is increasing exponentially then flatlines.

Below are the players who had the biggest negative prediction differentials. Lynn, for example, was predicted to get 35.8% (.358) of the vote but got only 5.5% (.055).
0.055 *** -0.303 Lynn, Fred
0.235 *** -0.230 McGwire, Mark
0.017 *** -0.180 Bell, Buddy
0.083 *** -0.154 Nettles, Graig
0.053 *** -0.152 Baines, Harold
0.017 *** -0.151 Parrish, Lance
0.005 *** -0.144 Lopes, Davey
0.068 *** -0.140 Concepcion, Dave
0.019 *** -0.131 Cey, Ron
0.051 *** -0.130 Hernandez, Keith
Now the players who had the biggest positive prediction differentials.
0.298 *** 0.080 Rice, Jim
0.852 *** 0.095 Molitor, Paul
0.157 *** 0.109 Trammell, Alan
0.965 *** 0.111 Schmidt, Mike
0.775 *** 0.126 Yount, Robin
0.818 *** 0.164 Morgan, Joe
0.500 *** 0.189 Perez, Tony
0.664 *** 0.296 Fisk, Carlton
0.917 *** 0.319 Smith, Ozzie
0.821 *** 0.394 Puckett, Kirby
I also ran a regression which had the following variables. None of them were squared or made non-linear. 2B, SS and C are positional dummies. All of these variables were significant at the 10% level. 9 were significant at the 5% level and 7 the 1% level. But the adjusted r-squared was .763 and he standard error was .126. So it does not work nearly as well as the model mentioned above.
AVG
MVP
3000 HIT
2B
SS
C
SB
WSIMP
10000PA
HR
Now some links to other research
Baseball Hall of Fame voting: a test of the customer discrimination by Arna Desser, James Monks and Michael Robinson
Social Science Quarterly Sept 1999 v80 i3 p591(13)
Discussion of above article
A neural-net Hall of Fame prediction method
Teaching Statistical Thinking Using the Baseball Hall of Fame by Steven Wang
Who's missing from the hall of fame by JC Bradbury
Modeling Election to the Major League Baseball Hall of Fame through the use of Genetic Algorithms By David Cohen
Monday, April 20, 2009
What Determines Vote Percentage In The First Year Of Hall Of Fame Eligibility?
I tried several models and maybe I will update this in the next several days, explaining what some of them were and why I am posting this one first. But using OLS regression, here is the equation I came up with:
Pct = -.08 + .00011*SB + .00741*GG + .071*MVP + .032*AS + .512*3000H + .29*500HR
These are all career totals. GG is number of Gold Gloves won, MVP is number of MVP awards won, 3000H is a dummy variable (1 if a player reached it, 0 otherwise). The 500HR is also a dummy variable. I used all the voting data from 1990-2009. Any player who had received votes before 1990 was not counted. Pete Rose was not included since leaving him out improved the results and he got nowhere near what the model predicts (usually about 70-80 points lower-in 1992 he got just 9.5% of the vote). So the scandals and controversies probably played a role. Mark McGwire got a lower % than predicted, but it was not anywhere near as bad as it was for Rose.
There were 182 players. The adjusted r-squared was .839 and the standard error was .104 (that is the lowest I got and I looked at many models with lots of different combinations of many variables). Perhaps this kind of fishing expedition, just looking for the most accurate regression, is not legitimate. But maybe it is the only way to figure out what the voters care about. All of the variables were significant at the 5% level (the highest p-value was about .04 and the next highest was .012)
McGwire did have the biggest negative difference between his predicted % and what he actually got. The equation predicts about 50.8% while he only got 23.5%, for a differential of -.273. But Fred Lynn came in .262 below his predicted value of .317. Here are the bottom ten in differential
-0.27344 McGwire, Mark
-0.26287 Lynn, Fred
-0.19284 Hernandez, Keith
-0.1864 Ripken, Cal
-0.15364 Parrish, Lance
-0.14969 Concepcion, Dave
-0.14829 Murphy, Dale
-0.14131 Cedeno, Cesar
-0.13065 Fernandez, Tony
-0.13031 McGee, Willie
Now the top ten
0.12798 Brett, George
0.12806 Schmidt, Mike
0.15458 Carter, Gary
0.17142 Molitor, Paul
0.24404 Jackson, Reggie
0.34928 Perez, Tony
0.35425 Morgan, Joe
0.38621 Smith, Ozzie
0.40061 Fisk, Carlton
0.5199 Puckett, Kirby
So the question is why did these guys get so many more votes than predicted (and why did those other guys get so many less)? I will have to think about that some more and maybe I can improve the model if I come up with something.
Brett and Schmidt would both still have made it in on the first ballot. But Puckett was only predicted to get .301. Many of the players here in the top ten also got alot more than predicted in at least one other model. That model had the following variables
HR
AVG
SB
GG
MVP
3000 HIT
2B
SS
C
The last 3 are positional dummies. But the standard error on this model was .132, much higher than the other model.
Pct = -.08 + .00011*SB + .00741*GG + .071*MVP + .032*AS + .512*3000H + .29*500HR
These are all career totals. GG is number of Gold Gloves won, MVP is number of MVP awards won, 3000H is a dummy variable (1 if a player reached it, 0 otherwise). The 500HR is also a dummy variable. I used all the voting data from 1990-2009. Any player who had received votes before 1990 was not counted. Pete Rose was not included since leaving him out improved the results and he got nowhere near what the model predicts (usually about 70-80 points lower-in 1992 he got just 9.5% of the vote). So the scandals and controversies probably played a role. Mark McGwire got a lower % than predicted, but it was not anywhere near as bad as it was for Rose.
There were 182 players. The adjusted r-squared was .839 and the standard error was .104 (that is the lowest I got and I looked at many models with lots of different combinations of many variables). Perhaps this kind of fishing expedition, just looking for the most accurate regression, is not legitimate. But maybe it is the only way to figure out what the voters care about. All of the variables were significant at the 5% level (the highest p-value was about .04 and the next highest was .012)
McGwire did have the biggest negative difference between his predicted % and what he actually got. The equation predicts about 50.8% while he only got 23.5%, for a differential of -.273. But Fred Lynn came in .262 below his predicted value of .317. Here are the bottom ten in differential
-0.27344 McGwire, Mark
-0.26287 Lynn, Fred
-0.19284 Hernandez, Keith
-0.1864 Ripken, Cal
-0.15364 Parrish, Lance
-0.14969 Concepcion, Dave
-0.14829 Murphy, Dale
-0.14131 Cedeno, Cesar
-0.13065 Fernandez, Tony
-0.13031 McGee, Willie
Now the top ten
0.12798 Brett, George
0.12806 Schmidt, Mike
0.15458 Carter, Gary
0.17142 Molitor, Paul
0.24404 Jackson, Reggie
0.34928 Perez, Tony
0.35425 Morgan, Joe
0.38621 Smith, Ozzie
0.40061 Fisk, Carlton
0.5199 Puckett, Kirby
So the question is why did these guys get so many more votes than predicted (and why did those other guys get so many less)? I will have to think about that some more and maybe I can improve the model if I come up with something.
Brett and Schmidt would both still have made it in on the first ballot. But Puckett was only predicted to get .301. Many of the players here in the top ten also got alot more than predicted in at least one other model. That model had the following variables
HR
AVG
SB
GG
MVP
3000 HIT
2B
SS
C
The last 3 are positional dummies. But the standard error on this model was .132, much higher than the other model.
Sunday, April 12, 2009
The Incredible Dominance of The 1936-39 Yankees
They won 4 straight world series. So you might be wondering why I don't write about the 1949-53 Yankees, who won 5 in a row. Those latter Yankees seem pretty dominiating. But I don't think in any 4 year period any team has done what those earlier Yankees did. The 1936-39 team had winning percentages of .667, .662, .651 and .702. Their only losing month in the 4 years was Sept. 1938 when they went 13-14.
First, they lead their league in ERA, HRs, Runs, fewest runs allowed and SLG in all 4 years. And, if I recall correctly, their run differential of 411 in 1939 is the highest of all time (967-556). The following table only amplifies their regular season dominance (the lines are in order of the years starting with 1936).

They spent 530 days in first place. That was about 80% of the time. The Indians were a half game ahead of the Yankees on July 12, 1938. But neither team had played 77 games yet. So no team besides the Yankees was in first place during the 2nd half of any of these 4 seasons. Where I have "Games left to play on clinch date," I simply added wins and losses and then subtracted from 154. I did not try to take ties or actual games left into account. So it is an approximation. But the lowest total was 12. That was the closest "race."
UPDATE: (April 16) I also figured out what the closest any team got to the Yankees was in each year between Sept. 1 and the date they clinched. Here are those games behind numbers in order of the years: 16-9-13-11.5. So between Sept. 1, 1936 and the date they clinched, the fewest games behind for the 2nd place team was 16. And in all 4 years, the closest any team came to them between Sept. 1 and the clinch date was 9 games, in 1937.
Then notice that no team ever finished closer than 9.5 games with the average games ahead being 14.75. Also notice that going into Sept., the closest anyone got was 11 games. In 1938, they went from being a half game behind on July 12 to being up 14 games up by Aug 31. In about 7 weeks, they gained 14 games.
You might be wondering if they kept up this dominance in the World Series. After all, they were playing the best the National League could throw at them. The table below summarizes what happened.

Notice that their edges in SLG and OBP are much larger than their edge in AVG (maybe they had a sabermetrician working for them back then). OBP just used walks, hits and atbats. They had 23 more walks while having huge leads in HRs and TB. Their winning pct against the NL champs was .842 while never losing more than 2 games in any one series. Their Pythagorean winning pct was .825 (runs scored squared divided by the sum of runs scored squared + runs allowed squared).
From other research, I have found that winning pct is about 1.21*(OPS differential) + .500. OPS is OBP + SLG. The Yankees had a .760 OPS while their opponents had .590. That would give the Yankees a pct of .706, not as dominat as what actually happened but still pretty impressive considering the competition.
First, they lead their league in ERA, HRs, Runs, fewest runs allowed and SLG in all 4 years. And, if I recall correctly, their run differential of 411 in 1939 is the highest of all time (967-556). The following table only amplifies their regular season dominance (the lines are in order of the years starting with 1936).

They spent 530 days in first place. That was about 80% of the time. The Indians were a half game ahead of the Yankees on July 12, 1938. But neither team had played 77 games yet. So no team besides the Yankees was in first place during the 2nd half of any of these 4 seasons. Where I have "Games left to play on clinch date," I simply added wins and losses and then subtracted from 154. I did not try to take ties or actual games left into account. So it is an approximation. But the lowest total was 12. That was the closest "race."
UPDATE: (April 16) I also figured out what the closest any team got to the Yankees was in each year between Sept. 1 and the date they clinched. Here are those games behind numbers in order of the years: 16-9-13-11.5. So between Sept. 1, 1936 and the date they clinched, the fewest games behind for the 2nd place team was 16. And in all 4 years, the closest any team came to them between Sept. 1 and the clinch date was 9 games, in 1937.
Then notice that no team ever finished closer than 9.5 games with the average games ahead being 14.75. Also notice that going into Sept., the closest anyone got was 11 games. In 1938, they went from being a half game behind on July 12 to being up 14 games up by Aug 31. In about 7 weeks, they gained 14 games.
You might be wondering if they kept up this dominance in the World Series. After all, they were playing the best the National League could throw at them. The table below summarizes what happened.

Notice that their edges in SLG and OBP are much larger than their edge in AVG (maybe they had a sabermetrician working for them back then). OBP just used walks, hits and atbats. They had 23 more walks while having huge leads in HRs and TB. Their winning pct against the NL champs was .842 while never losing more than 2 games in any one series. Their Pythagorean winning pct was .825 (runs scored squared divided by the sum of runs scored squared + runs allowed squared).
From other research, I have found that winning pct is about 1.21*(OPS differential) + .500. OPS is OBP + SLG. The Yankees had a .760 OPS while their opponents had .590. That would give the Yankees a pct of .706, not as dominat as what actually happened but still pretty impressive considering the competition.
Monday, April 6, 2009
How Many Home Runs Would Ruth Have Hit If Baseball Had Been Integrated In His Era?
(Note: This is a slightly revised version of an article that was published in 2007 in the now defunt print periodical called "The Chicago Sports Weekly." I also had posted something like this at "Beyond the Boxscore" called How Would Integration Have Affected Ruth and Cobb?)
Maybe you have seen the images on TV of fans around the country holding up asterisk signs when Barry Bonds comes to the plate, hinting that his HR record is tainted, due to his alleged steroid use. But others counter that Babe Ruth might deserve an asterisk since he never faced blacks or dark-skinned Hispanics (there were a few players with Hispanic names before 1947 whose skin was generally pretty light).
But this raises the question of how many HRs would Ruth have hit had there not been a color barrier? I know the answer because Clio, the Greek muse of history, whispered it in my ear. You see, my Ph. D. thesis was in the field of economic history and its application of statistics is called “cliometrics.” What I am about to attempt here is something dangerous called a “counterfactual” in this field. So don’t try it at home. Leave it to the trained professionals.
Robert Fogel, economic historian at the University of Chicago, won a Noble Prize, partly for using counterfactuals. He supposed what if railroads had not been built. What other kind of transportation system (like canals) would have emerged? How would this have affected economic growth? He concluded that GDP in 1890 would have been about 5% lower than it actually was.
Not everyone was thrilled with this approach. The historian Fritz Redlich referred to counterfactuals as figments, probably of an imagination gone wild. So maybe you will think this analysis is a figment of my imagination. So. Maybe you’re a figment of my imagination. In any case, here it is.
First, we need an estimate of how many non-white pitchers there might have been. Since 1947, about 15% of all the IP by pitchers with 1,000+ IP in their careers have been by non-whites. All of the 1,000+ IP pitchers made up about 58% of all the IP since 1947, so it is a good sample. Therefore, I assume that in Ruth’s day 15% of the IP were by non-whites.
How good would those pitchers have been? Good enough to replace some white guys, who would be the worst pitchers in the league. You don’t add Satchel Paige to your team and then get rid of Lefty Grove. You dump Grover Lowdermilk (who really was not a bad pitcher but his name sounds funny, unlike mine). The non-whites with 1,000+ IP since 1947 actually had a collective ERA just about the same as the whites. So pre-1947, you dump the worst 15% of the pitchers by ERA and re-calculate the league HR rate using the remaining pitchers or the top 85%
After getting rid of the bottom 15% of the IP in each season from 1920 to 1934 (when Ruth played with the Yankees and had all of his great seasons) in the AL, I recalculated the HRs allowed per IP and found how much lower than the league average the new figures were. The average fall in HRs per IP for the years 1920-1934 was about 5%. That is, the best 85% of the pitchers had a HR per IP rate that was 5% lower than the league average (which includes all pitchers). So if you improve the pitching quality in a way that is consistent with integration, Ruth would hit 5% fewer HRs or hit about 678. Even if we cut him 10%, he still hits 643.
Some things I have not considered: when Aaron and Mays were hitting HRs in the 1950s, there still were not that many non-whites pitching. So their totals might need to be reduced. We also don’t know if all batters would be affected in the same way. The best HR hitters might have had their totals reduced more than the average hitter. Also, we don’t know what percentage of pitchers would have been non-white. Probably it is more than 15% today. Suppose it is 25%. I looked at the 1927 AL and if you only count the best 75% of the pitchers, the HR rate falls about 9%.
Suppose we only looked at the best 50% of the pitchers from 1927, HRs would fall about 18.3%. If that happened to Ruth over his whole career, he still hits 583 HRs. It is about 17% for 1921. For 1934, it would be 20%. Given that I am only counting the best 50% of the pitchers, we can safely say that integration would have reduced his HR’s by no more than 20% (the top 50% of pitchers in 2008 gave up about 20% fewer HRs than average as well). So he ends up with 571 HRs. That would have stood as a record for quite awhile. And remember that we would have to reduce Aaron and Mays since they played a good part of their careers when there were not as many non-white pitchers as today.
Here is the link that shows the white and non-white pitchers since 1947
http://cyrilmorong.com/RuthAsterisk/Pitchers.htm
Maybe you have seen the images on TV of fans around the country holding up asterisk signs when Barry Bonds comes to the plate, hinting that his HR record is tainted, due to his alleged steroid use. But others counter that Babe Ruth might deserve an asterisk since he never faced blacks or dark-skinned Hispanics (there were a few players with Hispanic names before 1947 whose skin was generally pretty light).
But this raises the question of how many HRs would Ruth have hit had there not been a color barrier? I know the answer because Clio, the Greek muse of history, whispered it in my ear. You see, my Ph. D. thesis was in the field of economic history and its application of statistics is called “cliometrics.” What I am about to attempt here is something dangerous called a “counterfactual” in this field. So don’t try it at home. Leave it to the trained professionals.
Robert Fogel, economic historian at the University of Chicago, won a Noble Prize, partly for using counterfactuals. He supposed what if railroads had not been built. What other kind of transportation system (like canals) would have emerged? How would this have affected economic growth? He concluded that GDP in 1890 would have been about 5% lower than it actually was.
Not everyone was thrilled with this approach. The historian Fritz Redlich referred to counterfactuals as figments, probably of an imagination gone wild. So maybe you will think this analysis is a figment of my imagination. So. Maybe you’re a figment of my imagination. In any case, here it is.
First, we need an estimate of how many non-white pitchers there might have been. Since 1947, about 15% of all the IP by pitchers with 1,000+ IP in their careers have been by non-whites. All of the 1,000+ IP pitchers made up about 58% of all the IP since 1947, so it is a good sample. Therefore, I assume that in Ruth’s day 15% of the IP were by non-whites.
How good would those pitchers have been? Good enough to replace some white guys, who would be the worst pitchers in the league. You don’t add Satchel Paige to your team and then get rid of Lefty Grove. You dump Grover Lowdermilk (who really was not a bad pitcher but his name sounds funny, unlike mine). The non-whites with 1,000+ IP since 1947 actually had a collective ERA just about the same as the whites. So pre-1947, you dump the worst 15% of the pitchers by ERA and re-calculate the league HR rate using the remaining pitchers or the top 85%
After getting rid of the bottom 15% of the IP in each season from 1920 to 1934 (when Ruth played with the Yankees and had all of his great seasons) in the AL, I recalculated the HRs allowed per IP and found how much lower than the league average the new figures were. The average fall in HRs per IP for the years 1920-1934 was about 5%. That is, the best 85% of the pitchers had a HR per IP rate that was 5% lower than the league average (which includes all pitchers). So if you improve the pitching quality in a way that is consistent with integration, Ruth would hit 5% fewer HRs or hit about 678. Even if we cut him 10%, he still hits 643.
Some things I have not considered: when Aaron and Mays were hitting HRs in the 1950s, there still were not that many non-whites pitching. So their totals might need to be reduced. We also don’t know if all batters would be affected in the same way. The best HR hitters might have had their totals reduced more than the average hitter. Also, we don’t know what percentage of pitchers would have been non-white. Probably it is more than 15% today. Suppose it is 25%. I looked at the 1927 AL and if you only count the best 75% of the pitchers, the HR rate falls about 9%.
Suppose we only looked at the best 50% of the pitchers from 1927, HRs would fall about 18.3%. If that happened to Ruth over his whole career, he still hits 583 HRs. It is about 17% for 1921. For 1934, it would be 20%. Given that I am only counting the best 50% of the pitchers, we can safely say that integration would have reduced his HR’s by no more than 20% (the top 50% of pitchers in 2008 gave up about 20% fewer HRs than average as well). So he ends up with 571 HRs. That would have stood as a record for quite awhile. And remember that we would have to reduce Aaron and Mays since they played a good part of their careers when there were not as many non-white pitchers as today.
Here is the link that shows the white and non-white pitchers since 1947
http://cyrilmorong.com/RuthAsterisk/Pitchers.htm
Sunday, March 29, 2009
Should Curt Schilling Get Into The Hall Of Fame?
This has been discussed since he retired this week and he may have good case. I will borrow some of what I put in my entry a few weeks ago about Bunning called Does Jim Bunning Belong In The Hall Of Fame?
Click on this next ink to see where Schilling ranks in strikeout-to-walk ratio (relative to the league average) for all pitchers with 3,000+ IP. He is 2nd to Mathewson.
K/BB Ratio
I found the best fielding independent ERAs since 1920 (an imputed ERA based on walks, strikeouts and HRs with HRs being adjusted for park effects). Schilling ranked 7th among pitchers with 1500+ IP. The list went up through 2005. He is now over 3000 and my guess is that he has not slid too much since then (he was not on the 3000 list yet after 2005).
I found that he was 44th in Park-Adjusted Pitching Wins Above Replacement Level. I think that was also through 2005. He has probably moved up several slots since then.
So he has some impressive ranks. I would like to have seen a consecutive 3-year period in there that was Cy Young worthy. 2001-02 are in there and 2004 is but he missed alot of 2003, which was a great year for about 160 IP. He still finished 7th in runs saved over average with park adjustments. In 01, 02 and 04 he had 20+ wins, lead league in K/BB ratio, winning pct over .750, 2nd place in runs saved each year. Had a 1st, a 2nd and a 3rd in IP. Not saying he should not be there. But if he had pitched a full season in 2003, then he is a no-brainer. I guess I am 90% for him (or so)
Click on this next ink to see where Schilling ranks in strikeout-to-walk ratio (relative to the league average) for all pitchers with 3,000+ IP. He is 2nd to Mathewson.
K/BB Ratio
I found the best fielding independent ERAs since 1920 (an imputed ERA based on walks, strikeouts and HRs with HRs being adjusted for park effects). Schilling ranked 7th among pitchers with 1500+ IP. The list went up through 2005. He is now over 3000 and my guess is that he has not slid too much since then (he was not on the 3000 list yet after 2005).
I found that he was 44th in Park-Adjusted Pitching Wins Above Replacement Level. I think that was also through 2005. He has probably moved up several slots since then.
So he has some impressive ranks. I would like to have seen a consecutive 3-year period in there that was Cy Young worthy. 2001-02 are in there and 2004 is but he missed alot of 2003, which was a great year for about 160 IP. He still finished 7th in runs saved over average with park adjustments. In 01, 02 and 04 he had 20+ wins, lead league in K/BB ratio, winning pct over .750, 2nd place in runs saved each year. Had a 1st, a 2nd and a 3rd in IP. Not saying he should not be there. But if he had pitched a full season in 2003, then he is a no-brainer. I guess I am 90% for him (or so)
Sunday, March 22, 2009
An All-Time Ranking of Players By Wins Above Replacement Level
I have attemtped to rank players by their value above the replacement level player using Pete Palmer's "Total Player Rating" or TPR (it is now actually called BFW for batting wins + fielding wins). This comes from Palmer's "linear weights" method. For example, here are the run values of various events:
1B: .47
2B: .78
3B: 1.09
HR: 1.4
BB: .33
Outs have a negative value, usually around -.25. For each player he calculates how many batting runs they had. This has a win value (usually around 10 runs per win). Something similar is done with fielding. He figures out how man runs a player saved compared to the league average as a fielder and this is converted to wins. Both batting runs and fielding runs can be negative. If a player had 20 batting runs that would be 2 batting wins. If he also had -10 fielding runs, that would be about -1 fielding wins. So his TPR or BFW would be 1 for that season. A player can have a negative value for a season. So he would be below average.
But that does not always mean he had no value. He might have still been better than the next best player available (the replacement) who might have been more negative. There is no clear consensus on exactly what the replacement level in TPR is.
So I did two lists. One with a TPR or a BFW of -2 per 700 PAs as replacement level and one with -3. I divided each guy's career PAs by 700. Then I multiplied that times 2 or 3. That result got added to his career TPR to get career value over replacement. For example, in a season with 700 PAs and a TPR of 0, the player is still either 2 or 3 wins better than replacement level. A player with a TPR of 6, which puts you in the top 200 seasons all-time, would be 8 or 9 wins better than replacement.
Suppose a player had 7,000 career PAs. So that is 10 full seasons. If he had a career TPR of 20, his wins above replacement would be 34 or 41. Click here to see the all time rankings through 2004 using -2 BFW as the replacement level. I used everyone that had 2,000 PAs through 2004 (with PAs including just ABs and walks). Click here to see the all time rankings through 2004 using -3 BFW as the replacement level. The players were compiled using the annual listings in BFW from Retrosheet.
Most of the top ranked players will not be a big surprise. Bonds is first followed by Ruth. Two guys in the top 40 using either a -2 or -3 TPR for replacement level that are not in the Hall of Fame are Ron Santo and Bobby Grich. Bill Dahlen, too.
1B: .47
2B: .78
3B: 1.09
HR: 1.4
BB: .33
Outs have a negative value, usually around -.25. For each player he calculates how many batting runs they had. This has a win value (usually around 10 runs per win). Something similar is done with fielding. He figures out how man runs a player saved compared to the league average as a fielder and this is converted to wins. Both batting runs and fielding runs can be negative. If a player had 20 batting runs that would be 2 batting wins. If he also had -10 fielding runs, that would be about -1 fielding wins. So his TPR or BFW would be 1 for that season. A player can have a negative value for a season. So he would be below average.
But that does not always mean he had no value. He might have still been better than the next best player available (the replacement) who might have been more negative. There is no clear consensus on exactly what the replacement level in TPR is.
So I did two lists. One with a TPR or a BFW of -2 per 700 PAs as replacement level and one with -3. I divided each guy's career PAs by 700. Then I multiplied that times 2 or 3. That result got added to his career TPR to get career value over replacement. For example, in a season with 700 PAs and a TPR of 0, the player is still either 2 or 3 wins better than replacement level. A player with a TPR of 6, which puts you in the top 200 seasons all-time, would be 8 or 9 wins better than replacement.
Suppose a player had 7,000 career PAs. So that is 10 full seasons. If he had a career TPR of 20, his wins above replacement would be 34 or 41. Click here to see the all time rankings through 2004 using -2 BFW as the replacement level. I used everyone that had 2,000 PAs through 2004 (with PAs including just ABs and walks). Click here to see the all time rankings through 2004 using -3 BFW as the replacement level. The players were compiled using the annual listings in BFW from Retrosheet.
Most of the top ranked players will not be a big surprise. Bonds is first followed by Ruth. Two guys in the top 40 using either a -2 or -3 TPR for replacement level that are not in the Hall of Fame are Ron Santo and Bobby Grich. Bill Dahlen, too.
Friday, March 13, 2009
Is Ryan Howard The New Mickey Vernon? (Or Is His Career Really In Decline?)
Howard has had big drops in his OWP the last 2 years. OWP or offensive winning percentage is a Bill James stat that says what a team's winning percentage would be if it had a lineup of 9 identical players who all hit alike and they gave up an average number of runs. Since I got the data from the Lee Sinins Complete Baseball Encyclopedia, it is park adjusted. Here are Howard's OWP for each of the last 3 seasons with his age in parantheses:
.777 (26)
.675 (27)
.582 (28)
So the declines are .102 and .093. For players who had long careers, such big back to back drops in OWP are somewhat rare, especially for someone under 30. Some stories I read using google news search indicate he is in better shape this spring and is hitting better than usual this pre-season. So maybe he has taken the necessary steps to stop the decline.
To see how unusual his declines are, I used a list of players that I have compiled before. This list includes all players who had 15+ seasons with 400+ plate appearances from age 20-40. Then I found all cases of players having back to back seasons of a drop in OWP of .075 or more. The tables below show all of these cases (more discussion of the tables below).
There are 22 such cases. But in only 3 of them, did the fall in OWP start before the age of 30. Those belong to Robin Yount, Jake Beckley and Mickey Vernon. Beckley's started at age 23 and neither of my biographical encyclopedia's mention anything. Same for Vernon's decline. Yount had shoulder problems during the 1984 and 1985 seasons and actually had surgery twice. But both Beckley and Yount are in the Hall of Fame. So if Howard can end up with 15+ seasons with 400+ plate appearances he has a 2 out of 3 chance of making the Hall of Fame.
If I had limited the study to declines of .093 or more, there were only 8 guys. The only one whose decline started before age 30 was Vernon. So who knew that he had something in common with Howard?
In the tables below, the numbers in red are the decline years. The year before the decline is there for each player for reference. Two guys had 3 straight years that fit the criteria. They were Jimmy Dykes and Willie Keeler. Yount and Honus Wagner had two such streaks. The average age at which the decline started was 33.95. 17 of the 22 cases started at age 33 or older.
There were 91 players with 15+ seasons with 400+ plate appearances and 22 back to back seasons of a drop in OWP of .075 or more. So nearly 25% of the 91 players had these big back to back declines. That makes it look like what has happened to Howard is not that rare. But what is rare is the age at which it has happened to him.
What happened to these guys in the third year? Did they finally rebound? Well, we know that Keeler and Dykes each had declines that fit the criteria in the third year. 5 players did not get 400+ PAs the next year. The average change in the third year, including Keeler and Dykes (who were the only declines), was a positive .078. 8 of the 17 had changes of +.100 or more. So Howard has a good chance to bounce back but also has a chance not to make it to 400 PAs. Mickey Vernon improved .295 in the third year.



.777 (26)
.675 (27)
.582 (28)
So the declines are .102 and .093. For players who had long careers, such big back to back drops in OWP are somewhat rare, especially for someone under 30. Some stories I read using google news search indicate he is in better shape this spring and is hitting better than usual this pre-season. So maybe he has taken the necessary steps to stop the decline.
To see how unusual his declines are, I used a list of players that I have compiled before. This list includes all players who had 15+ seasons with 400+ plate appearances from age 20-40. Then I found all cases of players having back to back seasons of a drop in OWP of .075 or more. The tables below show all of these cases (more discussion of the tables below).
There are 22 such cases. But in only 3 of them, did the fall in OWP start before the age of 30. Those belong to Robin Yount, Jake Beckley and Mickey Vernon. Beckley's started at age 23 and neither of my biographical encyclopedia's mention anything. Same for Vernon's decline. Yount had shoulder problems during the 1984 and 1985 seasons and actually had surgery twice. But both Beckley and Yount are in the Hall of Fame. So if Howard can end up with 15+ seasons with 400+ plate appearances he has a 2 out of 3 chance of making the Hall of Fame.
If I had limited the study to declines of .093 or more, there were only 8 guys. The only one whose decline started before age 30 was Vernon. So who knew that he had something in common with Howard?
In the tables below, the numbers in red are the decline years. The year before the decline is there for each player for reference. Two guys had 3 straight years that fit the criteria. They were Jimmy Dykes and Willie Keeler. Yount and Honus Wagner had two such streaks. The average age at which the decline started was 33.95. 17 of the 22 cases started at age 33 or older.
There were 91 players with 15+ seasons with 400+ plate appearances and 22 back to back seasons of a drop in OWP of .075 or more. So nearly 25% of the 91 players had these big back to back declines. That makes it look like what has happened to Howard is not that rare. But what is rare is the age at which it has happened to him.
What happened to these guys in the third year? Did they finally rebound? Well, we know that Keeler and Dykes each had declines that fit the criteria in the third year. 5 players did not get 400+ PAs the next year. The average change in the third year, including Keeler and Dykes (who were the only declines), was a positive .078. 8 of the 17 had changes of +.100 or more. So Howard has a good chance to bounce back but also has a chance not to make it to 400 PAs. Mickey Vernon improved .295 in the third year.



Sunday, March 8, 2009
Does Jim Bunning Belong In The Hall Of Fame?
This issue came up recently on the SABR list. One of the issues was why were pitchers from his era who seem to have been about as good he was not in. Then someone else mentioned that maybe the Veterans committee put him because he is a Senator. I posted some evidence on this. Basically it was references to research I had done in the past and seeing where Bunning ranked. I will put that post below, but first something new, although it is a simple, rough estimate of his value (I used Fielding Independent Pitching ERA or FIP ERA to find an imputed winning percentage for Bunning which is fairly high-it's all based on how good he was at strikeouts, walks and HRs). I find that there is some evidence for him being in the Hall, but I don't think it is all on his side. He had a great strikeout-to-walk ratio, which is one indicator of how good a pitcher is.
Below are the top 25 pitchers with 3000+ IP in strikeout-to-walk ratio relative to the league average. Bunning is 19th, which is very good. Mathewson is 1st. He had a strikeout-to-walk ratio of 2.96 while the league average was 1.29. Since 2.96/1.29 = 2.30, Mathewson gets a 230. Data came from the Lee Sinins Complete Baseball Encyclopedia.

I calculated his Fielding Independent Pitching ERA or FIP ERA. The idea is that a pitcher controls HRs, BBs and Ks and hits on balls in play not so much (if you have not heard of this, google Voros McCracken).
Here are the key calculations. HRs, BBs and Ks are per 9 IP.
(1) FIP ERA = Constant + 1.44*HR + .33*BB - .22*K
(2) The constant = League ERA - (1.44*HR + .33*BB - .22*K)
I used Bunning's stats and adjusted them to the AL stats of 2008. Bunning's Ks per 9 IP was 6.83 or about 25% above average. In the 2008 AL K/9IP = 6.36. Raising that 25%leaves 7.97. He walked 2.39 batters per 9 IP or about 26% fewer than average. In the 2008 AL BB/9IP = 3.32. Lowering that 26% leaves 2.47. So those numbers will get plugged into equation (1). The league ERA in the AL in 2008 was 4.35 and the constant for equation (1) works out to 3.21.
We still need to calculate his HRs per 9 IP. He actually gave up 372 HRs while the average was 346. So it looks like Bunning did poorly here. But he pitched in Tiger Stadium for part of his career where an above average number of HRs were hit. So I adjusted his HRs allowed in each season based on the HR park factors from the STATS, INC. All-Time Baseball Sourcebook. For example, if Tiger stadium gave up 20% more HRs than average in a season, I reduced his HRs for that year by 10% (only half of the 20% since he only pitched half his games there). In some of his years with the Phillies, the park factor was below average. After doing this for each of his seasons, his HR total came out to 349, or almost exactly average.
In the AL in 2008, there was just about 1 HR per 9 IP. So I used that for equation (1). With the HR, BB and K data done, I found a FIP ERA of 3.73 for Bunning (adjusted for the 2008 AL).
I then calculated what Bill James calls the Pythagorean winning percentage for Bunning if he pitched on an average team. It is
(runs scored squared)/(((runs scored squared) + (runs allowed squared))
For Bunning, adjusted to the 2008 AL, we get
4.35*4.35/(4.35*4.35 + 3.73*3.73) = .577
So Bunning pitching for an average team would have a .577 winning pct. For pitchers with 3000+ IP, he would be tied for 42nd (with Jack Morris). But I did not calculate the FIP ERA or Pythagorean winning percentage for anyone else. I am just assuming if I did it for everyone, just as many guys would move ahead of Bunning as would fall behind. 42nd is pretty good and seems high enough for a starter to make the Hall.
Now for the post to the SABR list.
I found the best fielding independent ERAs since 1920 (an imputed ERA based on walks, strikeouts and HRs with HRs being adjusted for park effects). Bunning ranked 13th among pitchers with 3000+ IP.
I found that he was 51st in Park-Adjusted Pitching Wins Above Replacement Level.
It also looks like he out pitched Koufax in neutral parks while they were both in the NL.
But he only had 257 Win Shares through 2001, tied for 291st. Not sure where that ranked among pitchers.
He ranks 67th in adjusted pitching wins in Pete Palmer's baseball encyclopedia.
Below are the top 25 pitchers with 3000+ IP in strikeout-to-walk ratio relative to the league average. Bunning is 19th, which is very good. Mathewson is 1st. He had a strikeout-to-walk ratio of 2.96 while the league average was 1.29. Since 2.96/1.29 = 2.30, Mathewson gets a 230. Data came from the Lee Sinins Complete Baseball Encyclopedia.

I calculated his Fielding Independent Pitching ERA or FIP ERA. The idea is that a pitcher controls HRs, BBs and Ks and hits on balls in play not so much (if you have not heard of this, google Voros McCracken).
Here are the key calculations. HRs, BBs and Ks are per 9 IP.
(1) FIP ERA = Constant + 1.44*HR + .33*BB - .22*K
(2) The constant = League ERA - (1.44*HR + .33*BB - .22*K)
I used Bunning's stats and adjusted them to the AL stats of 2008. Bunning's Ks per 9 IP was 6.83 or about 25% above average. In the 2008 AL K/9IP = 6.36. Raising that 25%leaves 7.97. He walked 2.39 batters per 9 IP or about 26% fewer than average. In the 2008 AL BB/9IP = 3.32. Lowering that 26% leaves 2.47. So those numbers will get plugged into equation (1). The league ERA in the AL in 2008 was 4.35 and the constant for equation (1) works out to 3.21.
We still need to calculate his HRs per 9 IP. He actually gave up 372 HRs while the average was 346. So it looks like Bunning did poorly here. But he pitched in Tiger Stadium for part of his career where an above average number of HRs were hit. So I adjusted his HRs allowed in each season based on the HR park factors from the STATS, INC. All-Time Baseball Sourcebook. For example, if Tiger stadium gave up 20% more HRs than average in a season, I reduced his HRs for that year by 10% (only half of the 20% since he only pitched half his games there). In some of his years with the Phillies, the park factor was below average. After doing this for each of his seasons, his HR total came out to 349, or almost exactly average.
In the AL in 2008, there was just about 1 HR per 9 IP. So I used that for equation (1). With the HR, BB and K data done, I found a FIP ERA of 3.73 for Bunning (adjusted for the 2008 AL).
I then calculated what Bill James calls the Pythagorean winning percentage for Bunning if he pitched on an average team. It is
(runs scored squared)/(((runs scored squared) + (runs allowed squared))
For Bunning, adjusted to the 2008 AL, we get
4.35*4.35/(4.35*4.35 + 3.73*3.73) = .577
So Bunning pitching for an average team would have a .577 winning pct. For pitchers with 3000+ IP, he would be tied for 42nd (with Jack Morris). But I did not calculate the FIP ERA or Pythagorean winning percentage for anyone else. I am just assuming if I did it for everyone, just as many guys would move ahead of Bunning as would fall behind. 42nd is pretty good and seems high enough for a starter to make the Hall.
Now for the post to the SABR list.
I found the best fielding independent ERAs since 1920 (an imputed ERA based on walks, strikeouts and HRs with HRs being adjusted for park effects). Bunning ranked 13th among pitchers with 3000+ IP.
I found that he was 51st in Park-Adjusted Pitching Wins Above Replacement Level.
It also looks like he out pitched Koufax in neutral parks while they were both in the NL.
But he only had 257 Win Shares through 2001, tied for 291st. Not sure where that ranked among pitchers.
He ranks 67th in adjusted pitching wins in Pete Palmer's baseball encyclopedia.
Subscribe to:
Posts (Atom)