Friday, March 18, 2011

How Good Are “Replacement” Pitchers? Evidence From Case Studies

When a star pitcher goes down with an injury, a team must get someone else to pitch. That new guy is called a replacement pitcher. To see how good those understudies might be, I looked at cases of when this happened from 1946-2005. I examined the ERAs of the new guys (only in their starts, data from Retrosheet) and then determined what their winning percentage would be based on that ERA. That winning percentage will be an estimate of how good a replacement pitcher is.

The cases come from any pitcher since WWII who was in the top 25 in career IP or in the top 25 in ERA relative to the league average (minimum of 2000 IP). Then I found seasons when they had a drop in IP of 25% or more from the previous season. This had to be followed by an increase back up of at least 25% (of the following year’s total). I did this to avoid cases of someone who had really declined or was near the end of his career. The idea is that their being out was somewhat temporary. If a pitcher had two seasons in a row of low IP, became a reliever, or was traded in the three-year period, I did not include it. I also only wanted to look at starters, so I did not check Hoyt Wilhelm, who ranked high in relative ERA. Also, IP totals were adjusted for strike seasons to be more in line with a normal full season.

The table below shows the results:


Jim Palmer, for example, pitched 140.333 fewer IP in 1979 than in 1978. The ERA of the worst combined 140.333 IP on the Orioles that year was 3.99 (after adjusting for park effects). With the league ERA of 4.23, that would give a Pythagorean winning pct of .529.

Since the Orioles got 140.333 fewer IP from Palmer that year than the previous year, they had to "replace" him. And we know he was not going way because he pitched over 200 IP the next year. Some of his IP literally had to be replaced. And the "worst" 140.333 IP on the Orioles that year were not really that bad. They had good replacements. But, as can be seen in the table, that was not always true. Jack Morris's replacements in 1989 combined to have a 7.94 ERA, good for a .194 pct. The overall winning pct of the replacements was for all 27 cases .348 (a weighted average based on the missing IP in each case).

There are a host of issues associated with this study. I could have used the increase in IP after the year in question for what the replacements would have (or an average of the two). ERA may not be the best way to evaluate the replacements. Sometimes many pitchers were combined to make the replacement and some of them had very low IP totals. Maybe they only pitched on the road or at home, so the park adjustment might not be a good idea. Still another issue is that teams don’t pitch the same number of innings each year. If IP go up, regardless of who is or is not hurt, someone has to pitch them. So maybe in some cases I set the replacement IP too low.

It might also be instructive to see how the worst pitchers on a given team did in the year before the guy in question got hurt. The next table shows the pythagorean pct for the year they were actually hurt and for the year before, covering the same number of IP (with adjustments for strike years being made). I also did not count the pitcher himself in determining the worst ERAs for starters on his team the year before he got hurt, no matter what his ERA was.


For Blyleven in 1981, for example, I looked at the worst 153.666 IP among starters (making an adjustment for the strike year). Once that was adjusted for park effects and the league average was used, their Pythagorean winning pct would have been .346. That is actually worse than in a comparable number of IP in the year when Blyeven was actually hut and had to be replaced. I would expect that the worst batch of starters' IP to have an even higher ERA the next year with Blyleven missing. But it did not work out that way.

Only 9 of the 27 cases actually saw the projected winning pct of the worst starters go down. One example is 1991 for the Angels, when Blyleven missed the whole season and they had to replace the 134 IP they got from him in 1990. The worst 134 starting IP for the Angels in 1991 would get a Pythagorean pct. of .263. The worst 134 IP in 1990 has .366. That is the kind of thing we would expect when a good starter needs to be replaced. Yet this was not the norm.

The overall composite or weighted average of the project percentages was .348 in the year these pitchers were hurt and it was .356 the year before. That seems like a very small difference. This tells us that when a star pitcher goes down, the missing IP are covered by a group of pitchers who perform about as well as the worst pitchers on that team the year before. It makes me wonder if every team is getting a fairly big chunk of its IP from replacement level pitchers each year.

The one other thing to look at would be who actually took over the missing starts and see how they did. In 1967, Nelson Briles came out of the bullpen to make 14 starts in the second half of the season when Cardinal ace Bob Gibson got hurt. Briles went 10-2 with a 1.89 ERA as a starter. He then became a regular in the Cards rotation the next two years and was pretty much a starter (and a good one) for the next 10 years.

Sunday, February 27, 2011

Bert Blyleven's peak vs. Sandy Koufax's Peak

I raise this issue because it came up the other day at "The Hardball Times." See Visual Baseball: Blyleven vs. (dare I say) Koufax? by Kevin Dame. Here is something I wrote in the comments:

In 2007 at “Beyond the Boxscore” I posted an article called

"Bert Blyleven: As Dominating as Sandy Koufax

Over a 5 year stretch, both of them were 22 runs better than the 2nd best pitcher."

That was using a stat called RSAA or Runs saved against average. "It's the amount of runs that a pitcher saved vs. what an average pitcher would have allowed," as explained in the Lee Sinins Complete Baseball Encyclopedia. It is also adjusted for park effects.

I then went on to show that Blyleven was comparable to Koufax in strikeout-to-walk ratio and HRs allowed when park effects were factored in. Blyleven may have been a dominant pitcher during that time and no one really noticed.

Sunday, February 20, 2011

Pitchers' Performance With Runners On And None On, 1974-2010

I used data from David Pinto's Day-by-Day Database. I found all the pitchers who had 2500+ PAs with runners on base (ROB) and no runners on (NONE). There were 274 pitchers.

The correlation between the batting average (BA) allowed with ROB and BA allowed with NONE was .723. For slugging percentage (Slug%) it was .740. Those are both fairly high.

But I did expect it to be even higher based on some recent posts where I compared BA differential and Slug% differential between the two cases for consecutive years were extremely low. That suggests very little ability to pitch to the situation.

The table below shows the 25 best pitchers in terms of lowering their BA with ROB compared to NONE.


Now for Slug%



The next two tables are the worst pitchers. That is, they gave up a higher BA and Slug% with ROB than with NONE.



Sunday, February 13, 2011

Bud Black Says Padres Are More Balanced Now And This Will Make Up For The Loss Of Adrian Gonzalez

See Black: ‘We’re more balanced’. (Hat Tip: Baseball Think Factory). Click here to go their discussion. Here is what Black said:

"There is definitely going to be a different look. The Padres made a major transition midway through the 2009 season. It has transitioned again. I like this team. The focal point of last season was always Adrian (Gonzalez). We’re more balanced. That balance has to make up for the loss of Adrian."

That will be tough to do since in the last two years Gonzalez finished 4th and then 2nd in WAR with 7.0 and 6.3 (from Baseball Reference). Black said they will be improved elsewhere. That might make up for losing Gonzalez, but it does not have anything to do with balance per se.

I have done some research on balance. See

The Impact of Lineup Balance on Scoring, 1920-89

and

Are "Balanced" Teams More Successful?

I found balance helps very little, if at all. I also have links to research by Keith Woolner and David Gassko at that first site. It looks like their research shows that at most better balance can add 1 per season. So if Adrian Gonzalez's replacement costs the Padres more than 1 win compared to him, the balance will not help, as Black says. That is assuming that they actually are more balanced.

Sunday, February 6, 2011

Does Clutch Pitching Persist Year-To-Year? (Part 2)

Part 1 was a couple of weeks ago, when I only looked at one pair of years. I have now done 5 pairs of years. Low correlations indicate that pitchers tend not to be clutch one year and again clutch the next year.

I looked at all pitchers that had at least 250 plate appearances against opposing hitters in both of five consecutive two-year periods both with and without runners on base. There were around 70 such pitchers in each of the five cases. Data from David Pinto's "Day by Day Database" which is based on Retrosheet.

I found the differential between the batting average they allowed with runners on base (ROB) and the batting average they allowed with no runners on (NONE). I did the same for Slug%. So if a pitcher allowed a .240 BA with ROB and a .260 BA with NONE, his differential was -.020. That means he did better in the clutch than otherwise.

Did pitchers maintain about the same clutch performance in each year? Probably not. The table below shows the correlation between the first year's differential and the second year's differential for both BA and Slug%. They all tend to be pretty low or negative. If a pitchers in general tended to have about the same differential in each year, the correlation would be much higher. That is what I would expect if they really had a constant ability (or inability) with runners on base. But what this means is that many pitchers have a good differential one year and a bad or mediocre one the next year.



In both cases, with runners on and with none on, there is a large number of PAs. So if there really is some clutch ability here, we should see it. Also, pitchers, unlike hitters, can actually bear down and throw a little harder with runners on (or try harder to break off good curve balls). A batter really can't swing harder, for example, depending on the situation. But pitchers could save a little extra for when they needed it with runners on. Batters probably don't save a little extra.

So the one group that could theoretically have some control over the clutch don't seem to have it even when looking at a fairly large number of observations.

Saturday, January 29, 2011

The Expanded Postseason and its Impact on True Champions

SABR member Steve Fall presented his research on this at the Hornsby Chapter annual meeting. Very interesting.

Click here to read more about this and there is a link to the power point presentation Steve did. There are 24 slides. Here is the conclusion:

"While other factors have had some impact on the results, the most successful regular season teams have a much tougher time prevailing as World Series Champions due to the extra round of playoffs they must survive."

Tuesday, January 25, 2011

An Attempt To Measure Park Effects From 1934

I heard about this from the South Texas Chapter of SABR (Hornsby Chapter). Jan Larson said that author and chapter member Norman Macht, author of many books on baseball and other topics including Connie Mack and the Early Years of Baseball, came across a Baseball Magazine article from 1934 that looked at park effects. Here is the link A New Way to Compile Batting Averages (1934). The article is below anyway. The big idea was to calculate the batting average for each park and see if a player did better or worse than that.

The article was by FLETCHER PRATT. Jim Baker pointed out that he was very interesting. He was a science fiction and history writer who created some type of popular naval war game. Go to

Fletcher Pratt Wikipedia article

The World's Most Complicated Game: Fletcher Pratt, a historian and naval expert, invented a complex pastime that used ballroom floors and up to 120 players (from Sports Illustrated)

Here is the article:

The Present System of Computing Batting Averages Has Long Been Criticized and it Has Many Defects. Many Schemes for Improving the Records Have Been Advanced From Time to Time. Here is One That May Have Some Merit.

Does it seem right that Chick Fullis, who didn't hit hard when he was with the Giants, should suddenly develop a batting average thirty points higher than Mel Ott's at Philadelphia? Does it seem fair that Chick Hafey, one of the best hitters in baseball, should be getting the Bronx cheer for weak hitting, when the only trouble is that at Redland Field he was in a park where he had to drive the ball a mile to get a single?

Averages to show how well a player should hit in a park would straighten out these in¬consistencies. They would also give a lot of other information. They might tell us, for example, whether the Braves' pitching staff is really good, or merely look good because they have a lot of room to work in, and they might give us a real line on how well Chuck Klein should hit when he gets into a uniform with "Chicago" written across the chest.

But how will you get the batting average for a ball park? Simply by taking all the at bats in that park during a season, no matter what teams were playing, and dividing by the hits. The result will be the park's batting average, and it will be very accurate, for it will be an average compiled from the work of all the players in the league. This composite percentage will show, with reasonable certainty, how well the average player, if there were any such thing as an average player, ought to hit in a given park.

Trying this process with the National League parks for 1933, what results do we get? Here's the list, with all the games played during the year included:


There's a. lot of information in this table. Note that there is more than 73 points’ differ¬ence between Baker Bowl and the Polo Grounds. When one analyzes these figures they become even more important. They mean, for instance, that any team will hit over .300 in Philadelphia, and that the poor Philly pitchers have to face solid line-ups of .307 hitters from the beginning to the end of the season. But, of course, the visiting pitchers are up against the problem of curbing the clouting teammates of Dick Bartell and Don Hurst.
On the other band, take a look at the batting averages compiled at the Polo Grounds and at Braves' Field. No wonder Terry and McKechnie have good pitching staffs; they couldn’t have anything else in parks where the batters can hit no better than .244 and .232. It will be noted that 47 more runs were made at the Polo Grounds than at Boston during the 1933 season. This may mean that part of the low batting averages at the New York park is due to the ability of the Giants' pitch¬ing staff rather than to the park.

But that brings up another point. What is a good hitting team? Not just a team that makes a big batting average. The Phillies have been up near the top of the batting averages in the National League for years, and for just as many years they have been down near the bottom of the standing of the clubs. The usual reason given for this is that the Philly pitching staff has been weak, which is just another way of saying that the Phillies, although they pounded the ball, didn't hit quite as well as the teams they stacked up against—not well enough to beat them, any¬way. The Yankees in the days of their glory were a really good hitting team; that is, they consistently made more hits than the oppo¬sition in the same parks and under the same conditions. A hitting team, then, is not a team with a good batting average, but a team that can hit better than the opposition; that is, better than the average.

Now if the batting average of all the teams in the league at a certain park is .260 and the team whose home park it is, smacks the ball for a team batting average of .280, that team will be consistently hitting better than the opposition. So that if we compare the team batting averages in the league with the bat-ting averages of the parks they play in, we will be in a fair way to finding out what teams hit better than the opposition they encountered.

First let's take the clouting Phillies, who had three of the National League's six leading hitters and ranked third in team batting, with a figure of .274. The average for their park was .307, as we just saw. That is, all the teams playing in the Phillies' park hit .307 in 1933. But the Phillies themselves hit only .274, so their average was 33 points under what it should have been if they had hit as well as the rest of the teams in the league.

This shows that in spite of that .274 average, the Phillies were a weak-hitting outfit, which may help to explain why they finished in seventh place.

Now compare the figures in the same way for the Giants. In 1933 they hit only .263 and ranked down near the bottom of the league in hitting, while people compared them to the old Hitless Wonders of Fielder Jones. But the park averages we just compiled show that all the teams in the league hit only .232 at the Polo Grounds. This means that the Giants were hitting 31 points better than the opposition they met all season. They look like light hitters because they were play¬ing in a light-hitting park, but when they got into a batter's park, you couldn't get them out, as the Senators discovered. It doesn't matter whether it was because the Giants' pitching staff always kept the enemy in check, or whether it was because Bill Terry and his merry men could always find the opposing pitchers in the pinch. The Giants looked like weak hitters in the averages, but on the ball field they were a better-hitting team than any they faced.

Suppose we go right through the list of the teams in the league then, and make new averages. The team batting average will be one element; that will show how each team actually hit. The park batting average will show how all the teams in the league hit in that team's park; and the difference will be the amount each team hit better or worse than the teams it played against. Here's what we get:



Note how this list brings the teams out almost exactly in the order they finished in the Standing of the Clubs. This shows why the rank of the teams in the team batting averages doesn’t mean much. It’s not how hard a team hits the ball that counts; it’s how much better it hits than the other teams. On this basis, it appears that the Braves, who seemed to be a punchless outfit, actually hit just as well as the Pirates, who were rated the leading hitters of the league. It also shows that the Cards didn’t get so much good out of their wonderful pitching staff because they themselves weren’t hitting. Sportmen’s Field shows up as one of the best batting parks in the league.

But if you can subtract the park averages from a team’s batting average and get a good line on the team’s hitting, why can’t you do the same thing with individual batting averages and get a very accurate judgment on how good they are as hitters? If all the players in the league compile a general average of .280 at Sportsmen’s Field and Pepper Martin hits .316 there, then Pepper Martin is hitting 36 points better than the average and deserves to rank high up among the batters of the league. And if a player hits .305 at the Phillies’ .307 park, he is a -2 hitter in spite of his good-looking average.

So when we apply the same system to the individual batting averages we begin to understand why Chick Fullis fattened his batting average by 50 points when he moved from New York to Philadelphia and why Chick Hafey’s average fell when he changed over from St. Louis to Cincinnati. Sportsmen’s Field is a .280 park, Redland Field a .258 park. This would normally take about 22 points off the Chick’s average. Now Hafey hit .303 in 1933. If you give him back the 22 points he lost by switching a Cardinals uniform for a Redlegs one, his average comes back to .325, which is just about right, allowing for the dead ball in the National League last year.

When we go right down the line and apply the same system to the ranking batters of the league, we get a lot of interesting results. The leaders come out as follows:


Look what this does to the individual batting averages. For one thing, the Phillies no longer have three of the six leading hitters. Klein and Virgil Davis, who would be hot hitters in any company, are still up near the top, but the others have melted away. And Wally Berger, a fine hitter, who breaks up many a ball game and doesn’t get enough credit for the good hitting he does, is placed up there where he belongs.

Incidentally, with Klein in the fold this means that the Cubs are going to have the hardest-hitting outfield in baseball, and it also gives us a line on the vexed question of how Klein and Virgil Davis will hit in their new monkey-suits. The averages show Klein hits about 61 points better than his park. In Chicago, he will be batting in a .257 park. Add 61 points to this, with about 20 more for the livelier ball he will be batting in 1934, and you have .338 for his average if he does as well as last year. Virgil Davis will be batting in a .280 park at St. Louis. Add the 42 points he hit over his park average at Philadelphia, and make the same allowance for the lively ball and you have .342 for his probable average in 1934. So that both Klein and Davis will be found somewhere around .340 when the 1934 World’s Series rolls around, and if they aren’t the writer of this article will eat an arithmetic book with cream and sugar.

Monday, January 24, 2011

Does Clutch Pitching Persist Year-To-Year?

Low correlations indicate that pitchers tend not to be clutch one year and again clutch the next year.

I looked at all pitchers that had at least 250 plate appearances against opposing hitters in both 2009 & 2010 both with and without runners on base. There were 76 such pitchers. Data from David Pinto's "Day by Day Database" which is based on Retrosheet.

I found the differential between the batting average they allowed with runners on base (ROB) and the batting average they allowed with no runners on (NONE). I did the same for SLG. So if a pitcher allowed a .240 AVG with ROB and a .260 AVG with NONE, his differential was -.020. That means he did better in the clutch than otherwise.

Did pitchers maintaing about the same clutch performance in each year? Probably not. The correlation between their differentials in each year was just -.064. A zero correlation means no relationship and this is very close to that.

It was higher for SLG, though. But it was -.196 . That means if a pitcher was good in the clutch one year, he tended not to be good the next year, although the effect is weak.

This is just one year. More years need to be looked at. I will try to do more when I can.

Tuesday, January 18, 2011

FC Lane on the Batting Order

The post below is the few pages from FC Lane's book called "Batting" that dealt with the batting order. Whether or not it matches up with some of the recent analysis on lineups I will leave up to readers. One expert mentioned that it was a good idea to bat Cy Williams 2nd. FC Lane was a great baseball writer and editor of Baseball Magazine in the early part of the 20th century.

How the Batting Order "Colors" Batting

JOHN McGRAW once said, "Every ball team is capable of being arranged in a way to produce the maximum batting punch. You need a fast man who is a good waiter for lead off, another fast man good at the bunt and the hit and run in second place, then a massing of your heavy artillery in the next three or four positions so as to deliver the hardest blow with the least possible slowing down in speed. A slow footed runner, for example, will often cripple an attack. He must hit uncommonly well to be placed high on the list. Naturally your pitchers come last, for even if they are good hitters, they change too frequently for a settled batting order."

Jack Coombs said, "Every manager models his batting order on a scientific basis. He has so many batters at his disposal. He wishes to align those men so that their combined efforts will appear to the best advantage. How can he do this? Few managers agree on the precise details, but all agree on certain essential points. For example, number one should be a good hitter, but above all a good waiter. If he is short of stature so much the better for he will be harder to pitch to. All managers agree that the second man on the list must be a foxy hitter and fast on his feet. They agree also that third, fourth and fifth positions should be filled by good hitters who are preferably sluggers. I believe the three most important positions on the line-up are first, fourth and seventh place. First is obviously important. He is the entering wedge of your attack. Fourth is the logical clean-up man, the fellow who drives home that entering wedge. Seventh is a kind of clean-up man, but I cannot afford to put too good a hitter there. If I do, the opposing pitcher will pass him to take a chance at the tail of the batting order. Rather I must station a hitter at seventh who is not easily excited but is cool and always likely to come through with a hit."

Miller Huggins said, "An attack which is distributed through six or seven men rather than centered in one or two is much more effective. The team with a bunch of good hitters is usually consistent in its stick work. It is the steadiness of the pace which counts. On some clubs the batting punch is supplied by a renowned hitter like Hans Wagner or Nap Lajoie. On other clubs there is no such individual star but a better balanced attack of several men who are all good hitters. I prefer such a batting attack for your one or two stars may have an off day. The average work of six or seven men doesn't vary so widely. Besides, it is difficult for the pitcher to side step such an attack. In a pinch he can pass one or two men, but he can't pass half a dozen in succession. Furthermore, the strain of pitching to a number of men who are always dangerous is cumulative. The pitcher gets no breathing space as he would when he had retired one or two formidable stars and then faced mediocre batters."

Not all experts agree on the relative importance of the various positions on the line-up. Most of them would rate the lead-off man as important, and the clean-up sluggers as even more so. Hugh Jennings, however, thought differently. He said, "The neck moves the head and what the lead-off man accomplishes depends pretty much on the follow-up assistance he gets from the second man in the line-up. I believe it is a bigger job to locate a man who can play second properly than it is to find a good lead-off man. The talents which the lead-off man must possess are well understood and everybody realizes that the clean-up man must be a slugger. But the second place man hasn't been studied so thoroughly. This batter must be a good bunter. Good bunters ought to be common, but they are really less numerous than good hitters. The second place man must have a good batting eye and be a good waiter. Above all, he must use his head. In general he should hit to right field for his main object is to advance the lead-off man who has presumably reached first either through hit, pass or error. By driving the ball to right field he can send the runner to third base. I f he hit to left field, that runner would be held at second. Above all the second place man must have the peculiar knack of knowing whether the second baseman or short stop is going to cover the bag. Then he must be able to hit in a manner to break up their defensive play. This is very important. In fact I consider it the prime qualification for the man playing second position on the line-up. Moreover, the second place hitter should be fast. Then he won't get snarled up ina double play. There are times when he will retire the base runner in spite of himself. Then his thoughts are bent on saving his own scalp. That's largely a matter of speed in getting to first base."

Arthur Fletcher once played Cy Williams, his heaviest slugger, in second place. He said, "Cy isn't much of a bunter, I will admit. But he has some qualifications that you can't overlook. First of all, he's a right field hitter. That's what you want, a man to advance the runner. Then Cy seldom strikes out. You can generally depend upon him to hit the ball and hit it hard. Thus he advances the runner even though he is thrown out himself. And that's as good as a sacrifice. Besides, Cy is always likely to come through with a hit which may be a homer. Placing him high in the batting order you get more of his work. He'll go to bat five times in many a game where he would appear but four times if he batted farther down the list."

Even the despised tail of the batting list may be a source of strength. Wilbur Cooper said, "I am convinced that a pitcher adds much to his effectivness by his own good hitting. I believe that my batting and fielding have won seven or eight games a season for me that would otherwise have been lost."

Bill McKechnie said, "In all my experience I have known just one batter who liked to play the lead-off position." Bill thought this an inexcusable attitude, but it's not difficult to fathom. Batters don't like the lead-off position because it interferes with their hitting. It cuts their batting average many points. For example, John Tobin said, "The man who bats number one on the list and hits for .280 is doing well. He must forget his own hase hits in an effort to get on and of course his average suffers. How much it suffers I couldn't say, but I believe it will drop twenty to thirty points. Of course, some one has to play that position, but I think the records ought to make some provision for lead-off man and not rate his batting on the same basis as that of the slugger who comes fourth or fifth on the list."

Max Carey said, "In fairness to myself, I shall claim special consideration for my batting. I would have done much better had I not been lead-off man for several seasons. It is well known that lead-off man can not expect to have as high an average as he could get lower down the list. There are two reasons for this. In the first place he often has to wait out the pitcher and try to work him for a pass. In the second place, the pitcher, when he faces the lead-off man, usually has no one on bases to bother him. He is able to take lis full wind up and concentrate on the batter."

Batting languishes at the tail of the list. The catcher is out of the game frequently while the pitcher appears only once in three or four days. As Babe Ruth says, "No man can get in the games twice a week and do himself justice at bat as he would do were he getting daily practice."

Some managers shift their batting lists infrequently, even though one or two positions are open to criticism. They prefer to suffer this disadvantage rather than the greater disadvantage of a general disorganization. Not a few managers, however, particularly on losing clubs, shake up their batting lists rather often in the effort to hit upon a better working combination. When they do, the batting of the various players on the list is apt to fluctuate widely, for there is a definite connection between a batting average and the particular position in the batting order which a player is called upon to fill. In general lead-off man is handicapped by orders to wait out the pitcher, second position is handicapped by orders to sacrifice. The tail of the batting order is handicapped by a variety of adverse conditions among which infrequency of batting practice ranks rather high. Only at the clean-up positions does batting flourish at its best, for those players are usually called upon to "hit it out." There is also a noticeable psychology in a batter's position on the list. Let the man, for example, who has hit seventh, be raised to fourth or fifth place and his new responsibilities often act as a tonic on his batting average.

To sum up, a batter's work is colored to a considerable degree by the particular position he is called upon to fill in the batting order.

Tuesday, January 11, 2011

Trevor Hoffman Retires With Highest Strikeout-to-Walk Ratio

The table below shows the top 15 pitchers in strikeout-to-walk ratio since 1955 with 1000+ IP. For walks I used BB + HBP - IBB (which is called BB*). Data from Baseball Reference and the Lee Sinins Complete Baseball Encyclopedia.


In 1089 IP, Hoffman only hit 9 batters. The average pitcher would have hit about 4 times as many. And about 19% of his walks were intentional. For the average pitcher it was about 9%. Now for the guys who had 2000+ IP.

Friday, January 7, 2011

Ted Williams, Pedro Ramos, Dizzy Trout And Autographs

This is a post about trying to get to the bottom of some mythic stories. Stories that sound good but don't seem to stand up to scrutiny. The following site

http://www.sheilaomalley.com/?p=3280

Has a passage from the book "The Teammates: A Portrait of a Friendship" by David Halberstam. Here it is:

"When [Ted] was generous there was no one more generous, and when he was petulant there was no one more petulant, and sometimes he was both within a few seconds. Once in the mid-1950s, Pedro Ramos, then a young pitcher with Washington, struck Ted out, which was a very big moment for Ramos. He rolled the ball into the dugout to save, and later, after the game, the Cuban right-hander ventured into the Boston dugout with the ball and asked Ted to sign it. Mel Parnell was watching and had expected an immediate explosion, Ted being asked to sign a ball he had struck out on, and he was not disappointed. Soon there was a rising bellow of blasphemy from Williams, and then he had looked over and seen Ramos, a kid of 20 or 21, terribly close to tears now. Suddenly Ted had softened and said, “Oh, all right, give me the goddamn ball,” and had signed it. Then about two weeks later he had come up against Ramos again and hit a tremendous home run, and as he rounded first he had slowed down just a bit and yelled to Ramos, “I’ll sign that son of a bitch too if you can ever find it.”"

Now a writer at Baseball Think Factory wrote a refutation of this. It is at

Tracer: The Ted Williams-Pedro Ramos Story

My problem is that "The Biographical Encyclopedia of Baseball" has almost the same story about Ted Williams and Dizzy Trout. Page 1145, the entry on Trout. It does not say which year. It is the one edited by Pietrusza, Silverman and Gershman. Does anyone know anything about these stories? Any of them true? Are they told about other players? When was the first one reported? I doubt they are all true!

In the Dizzy Trout version (the Biographical Encyclopedia quotes his son, Steve), he strikes out Williams to preserve a 2-1 victory. I could find no game even remotely close to something like this for Trout vs. Boston using the Baseball Reference and Retrosheet game logs.

Update January 9: Rob Neyer discussed both the Ramos story and the Trout story in his book on baseball legends (pages 127-130). The Ramos story, he says, comes from Hy Hurwitz of the Boston Globe. The Trout story, he says, comes from Bruce Nash and Allan Zullo.

There is something plausible about the Trout case. On August 29, 1946, the Tigers beat the Red Sox 9-8 in 14 innings. Trout pitched the last 6.2 innings to win the game (allowing no runs on 3 hits and 3 walks). He struck out 2 batters but neither was Williams since the Baseball Reference boxscore shows no K's for him in that game. It is possible that Trout retired Williams in a key situation with runners on base. But the boxscore does not show that and the only news stories I found did not describe much about the game.

The Red Sox sent 64 men to the plate in the game, meaning that the leadoff man made the last out of the game. So Trout certainly did not get Williams out in the last inning.

But on Sept. 11, 13 days later, Williams did hit a HR off of Trout. So the story is not that far off at least as far as the events on the field are concerned.

Update January 9, 2016: Trout walked Williams both times he faced him in that August 29th game. One was intentional. Williams did not bat in the last inning. The play by play is now at Retrosheet

http://www.retrosheet.org/boxesetc/1946/B08290BOS1946.htm

Tuesday, January 4, 2011

Rob Neyer Asks: Does Kevin Brown have Cooperstown case?

Click here to read what Rob has to say. He lays out a good case for Brown. Here is some info on Brown I posted recently at Baseball Think Factory. I was surprised by how good the case for Brown is.

I have not thought too much about Kevin Brown one way or another but looking at his Baseball Reference page seems to show he deserves it. 34th in career WAR among pitchers despite the lower usage of starters in recent times.

Here are his ranks in WAR in the NL from 1996-2000: 1, 3, 1, 3, 2 (also a 3 in 2003). The only pitcher with more WAR from 1996-2000 was Pedro Martinez, 36.6 vs. 34.6. He was in the top 7 in IP in all those years.

He was in the top 6 in ERA+ every year from 1995-2000 with one 1st place. 53rd in career ERA+.

He was in the top 10 in strikeout-to-walk ratio 7 straight years (94-2000) with 5 top 5 finishes. His ratio relative to the league average is only 69th all-time among pitchers with 2000+ IP (through 2009). But that includes 406 pithers, so he is in the top 17%.

Using that same group of pitchers he is 5th all-time in HRs allowed relative to the league average. I know he pitched 5 years in Dodger stadium (and maybe some other tough HR parks), but look at the top 15. Data from the Lee Sinins Complete Baseball Encyclopedia.

Jack Taylor 200
Eppa Rixey 200
Tim Keefe 180
Addie Joss 174
Kevin Brown 172
Ed Morris 170
Eddie Plank 166
Dean Chance 165
Ed Walsh 165
Pete Donohue 161
Eddie Cicotte 161
Cy Falkenberg 157
Harry Howell 156
Tim Hudson 154
Roger Clemens 154

The 172 for Brown comes from the fact he gave up 208 HRs while the average pitcher would have given up 358. 208/358 = .581. 1/.581 = 1.72. That gets multiplied by 100. Taylor, Keefe and Joss all pitched in the dead ball era or earlier. If we only look at 1920-2009, Brown is 2nd only to Rixey who pitched alot in Cincinnati and that park in those days was really hard to hit a HR in.

Here is the top 15 from 1920-2009

Eppa Rixey 202
Kevin Brown 172
Dean Chance 165
Pete Donohue 161
Tim Hudson 154
Roger Clemens 154
Mark Gubicza 153
Danny Jackson 152
Derek Lowe 150
Greg Maddux 149
Hoyt Wilhelm 147
Mike Garcia 146
Roy Halladay 145
Dizzy Trout 144
Andy Pettitte 144

Pretty impressive rank for Brown.

Brown's HR rate (HR/PA), home, road

H 1.57%
R 1.50%

So his HR prevention excellence is not due to pitching in pitcher friendly parks. Brown seems to have very high career value and very high peak value. He was also very good at preventing runs and homeruns. He was good at striking batters out and not walking them. All that covers quite a bit of what we expect pitchers to do.

To analyze how good Brown is just using walks, strikeouts and HRs, I ran a regression in which a pitcher's ERA relative to the league average was the dependent variable and walks, strikeouts and HRs (all relative to the league average) were the dependent variables (in this case being over 100 is better than average like with HRs allowed, as discussed above). I looked at all pitchers from 1920-2009 with 2000+ IP. Here is the regression equation:

ERA = 35.23 + .235*HR + .264*SO + .176*BB

Here are Brown's rates for each stat:

HR 172
SO 106
BB 140

Plugging those values into the equation gives him a relative ERA of 128.27. Here is the top 10

Dazzy Vance 140.23
Lefty Grove 134.04
Pedro Martinez 133.46
Roger Clemens 129.85
Eppa Rixey 128.72
Kevin Brown 128.27
Greg Maddux 127.80
Nolan Ryan 127.32
Bret Saberhagen 126.97
Roy Halladay 125.27

Brown is 6th. That is very, very good. Just based on HRs, BBs, and SOs, Brown was 28.27% better than the league average.

But none of this is park adjusted. I also found RSAA per 9 IP (that is Runs Saved Above Average and is park adjusted, from Lee Sinins). Here is the top 20

Pedro Martinez 1.579
Lefty Grove 1.526
Roger Clemens 1.340
Roy Halladay 1.152
Randy Johnson 1.147
Hoyt Wilhelm 1.126
Greg Maddux 0.992
Tim Hudson 0.961
Tommy Bridges 0.958
Curt Schilling 0.955
Hal Newhouser 0.929
Whitey Ford 0.911
Urban Shocker 0.901
Carl Hubbell 0.890
Lefty Gomez 0.866
Gro. Alexander 0.857
Sandy Koufax 0.852
Bret Saberhagen 0.846
Kevin Brown 0.840
Mark Buehrle 0.838

Brown is 19th. Still pretty good.

Thursday, December 30, 2010

How Much Of A Yankee Killer Was Frank Lary?

His career record against them was 28-13. From 1955-61, it was 27-10 with a 3.06 ERA while his ERA against everyone else was 3.42. But did he really pitch better or differently against the Yankees?

Let's start with strikeout-to-walk ratio. In those years, Lary's was 1.62 against non-Yankee teams (I included HBP and took out IBBs-all data from Retrosheet). Against NY, it was 1.71. That may seem consistent with the "Yankee Killer" nick name, but over those years the Yankees themselves had a 1.43 ratio while the rest of the league had 1.32. So the typical pitcher had a strikeout-to-walk ratio that was .11 higher against the Yanks than everyone else. Lary was .09 better. So he was doing just about what other pitchers did.

Now HRs or HR rate (I use HRs divided by PAs with IBBs taken out). Lary allowed the Yanks a 2.6988% while he allowed the rest of the league 1.75%. So the Yanks did about 0.948 percentage points better against Lary than the average team from the rest of the AL. But that is just about normal. Over these years, the Yankees had a rate of 3.0097% while the rest of the league had a rate of 2.162%. The Yankees were about 0.848 percentage points better than the league average. So again, Lary's relative performance vs. NY is about what it was for other pitchers.

What about other hits? Lary's non-HR hit% against NY was .199 while against other teams it was .217. So that is a fairly big improvement. Some how he was better at preventing hits against the Yankees than he was against other teams. The Yankees themselves had a .205 rate while the rest of the league had .204. So the typical pitcher allowed more hits (but not alot more) to the Yankees than they normally did.

So it seems like the one thing that Lary was good at when he faced the Yankees was in preventing them from getting singles, doubles and triples. But the difference was only .018. Over, say, 36 PAs per game, that is just .648 hits. The run value of those hits is about .55 (the weighted average of the linear weights values that Pete Palmer established). So that makes a run value of .36 (interesting that that is just about the difference between his ERA against other teams and the one he had against the Yankees, 3.42 vs. 3.06).

The Tigers did score 4.93 runs per game in his starts against the Yankees from 1955-61. They averaged 4.61 runs per game overall. So the hitters rose to the occassion to support him. And maybe the fielders played a role in lowering the rate of non-HR hits he allowed. So it is possible that Lary became the "Yankee Killer" due to the aid of his teammates.

Wednesday, December 22, 2010

Bert Blyleven vs. Jack Morris

It seems like people who favor Morris over Blyleven say Morris was better in the clutch or better in big games. So I try to look at those issues here.

The table below shows their stats in 3 situations: runners on base (ROB), runners in scoring position (RISP), and close and late (CL). Data from Retrosheet.


I did not try to adjust these numbers for the league average. Blyleven might get a slight edge since the early 70s were not a big hitting era. But much of their careers did overlap. The only place where either pitcher has a big edge is Morris's edge in AVG in CL situations. But that .021 does not add up to alot. Blyleven had 2,129 ABs faced in those cases. That amounts to about 44 hits or 2 per season. That seems pretty small.

The next table shows their post season stats. League Championship Series and World Series are combined.



Morris has just about twice the IP. So if you doubled Blyleven's stats, you can see that there is not much difference between the two. Blyleven would have 86 hits, just about what Morris has. Same for HRs. But he would have more strikeouts and fewer walks.

I also looked at how they did in September pennant races. If a team finished 10 or more games ahead or behind, it was not considered to be a pennant race. If a team finished less than 10 games ahead or behind and if they were 5 or fewer games ahead or behind at the end of play of Aug. 31, it was considered a pennant race. 1991 for the Twins was not considered a pennant race (Morris was on that team). They began Sept. 7 games ahead (GA). On Sept. 15 they were 7.5 GA and they finished the season 8 GA. 1981 was not included since it was a strike year with a split season. Many teams were within a few games in Sept. This is highly unusual and winning the 2nd half only gave you a chance to play for the divisional title.

So the years I have for Morris as Sept. pennant races are 83, 87, 88, 92, 93. For Blyleven they were 77-80, 87, 89. Each pitcher had a total of 231.66 IP (Oct. data was included). Some of this data might inlcude games pitched after the divisional title was decided. But I did not feel like spending the time to figure that out. The table below shows how each pitcher did in these cases.



Again, it does not look like there is much difference between the two. So given Blyleven's far superior career stats (and peak value as measured by stats like WAR), he still deserves to make the Hall of Fame ahead of Morris. Whatever edge in the clutch or big games Morris might have, it is definitely not enough to put him ahead of Blyleven.

Wednesday, December 15, 2010

A Crude Measure Of The Most "All-Around" Players Since 1957

I started thinking about this when Cooper Nielson in a Baseball Think Factory discussion said:

"I suppose the "best all-around player" argument could go like this (keep in mind this is not my argument and not one I even agree with, but one that could conceivably and logically put Walker #1 in his era): There are five traditional baseball tools: hitting (for average), hitting for power, running, playing defense, and throwing."

See Cooperstowners in Canada: Larry Walker should be the second Canadian player elected to Cooperstown.

So here is how the crude measure works:

Multiply Gold Glove awards times 30. The idea here was to scale a great player in this stat to a great player in HRs or SBs. Brooks Robinson had the most GGs among position players with 16 and 16*30 = 480, close to 500.

Divide non-HR hits by 5. If a player had 2500 non-HR hits, you get 500.

Multiply SB*HR*non-HR*GG (with the above mentioned adjustments being made for GG and non-HR). If player had no GGs, I stopped multiplying so they did not end up at zero.

For Willie Mays it was 42,129,996,480. That is way too high a number to work with. So I raised it to the .25 power. That gave him 453, a more familiar kind of number to baseball fans. But that was divided by PAs and then multiplied by 10 to get the final number. Mays then had .363 (a nice number, close to the highest all-time batting average of .366 belonging to Ty Cobb). Here is the top 25:

1 Willie Mays 0.363
2 Torii Hunter 0.362
3 Barry Bonds 0.357
4 Larry Walker 0.355
5 Ichiro Suzuki 0.352
6 Ryne Sandberg 0.349
7 Eric Davis 0.345
8 Cesar Cedeno 0.345
9 Roberto Alomar 0.337
10 Devon White 0.333
11 Andruw Jones 0.330
12 Andre Dawson 0.327
13 Garry Maddox 0.325
14 Bobby Bonds 0.316
15 Andy Van Slyke 0.313
16 Mike Schmidt 0.311
17 Ken Griffey Jr. 0.309
18 Carlos Beltran 0.302
19 Paul Blair 0.296
20 Joe Morgan 0.295
21 Marquis Grissom 0.293
22 Ivan Rodriguez 0.292
23 Dwayne Murphy 0.291
24 Bill White 0.285
25 Jimmy Rollins 0.284

If I started with his stats from 1957 on, when they started giving out Gold Gloves, Mays gets .378.

Sunday, December 12, 2010

What Might Explain Ron Santo's Low Hall Of Fame Voting Percentages?

It seems like it might be for the reasons I have have seen people give the last week or so: no post-season exposure, somewhat short career (he did not reach 10,000 PAs), lack of milestones like 3000 hits or 500 HRs and lack of MVP awards.

Last year and earlier this year I posted some regression generated equations that tried to explain the percentage of the Hall of Fame vote player got in their first year of eligibility (and also their highest percentage). The model I came up with was based on some trial and error. That seemed unavoidable, since it is hard to have priors on what exactly the voters are thinking. The model looked at all players that became eligible for the first time from 1980-2009.

The model uses the following data to explain vote percentage:

Reaching 10,000 PAs
500 HRs
3000 hits
500 SBs
Gold Gloves
All-Star games
World Series performance
MVP awards

Gold Gloves and All-Star games got capped at certain levels which were then squared. The idea was that those things have an exponential effect which tapers off. There were also interaction terms for World Series performance, Gold Gloves and All-Star games. The idea there was that getting lots of Gold Gloves and playing in lots of All-Star games has more than an additive effect (after I discuss what the model predicted for Santo, technical details like regression results and variable descriptons will be covered).

Santo's first year percentage was 3.9%. Normally, he would no longer be eligible in the writers' voting. But he and some other players were re-instated in 1985. He got 13.4%. The model predicted that he would get 17.65%. The standard error was .08. So even if we give him 8% more, that only jumps him up to 21.4%. Still a pretty low total for a first year (Billy Williams got 23.4% in his first year in 1982 and steadily increased until he got 85.7% in 1987).

Santo's highest percentage was 43%. The model predicted it would be 30%. So he actually did better than that. The standard error was .117. So he was predicted to be about 4 standard errors below what is needed for induction, 75%. And his actual highest percentage was still about 3 standard errors below 75%. Billy Williams highest predicted percentage was 29.6% while it was actually 85.7%. That differential of 56.1% is the highest positive differential. Why Williams is in and Santo isn't is an interesting question.

Here was the equation where the player's first year vote percentage was the dependent variable

PCT = -.010 + .00086(WSAS) + .048(GGAS) + .070(MVP) + .404(3000 HIT) + .280(500 HR) + .002(ASSQ10) - .00089(GGSQ7) + .071(500SB) - .006(WSIMPSQ50) + .100(10000PA)

The adjusted r-squared was .898 The standard error was .08.

Here was the equation where the player's highest vote percentage was the dependent variable

PCT = -.014 + .00037(WSAS/1000) + .025(GGAS/1000) + .067(MVP) + .257(3000 HIT) + .201(500 HR) + .0048(ASSQ10) - .0013(GGSQ7) + .071(500SB) - .00167(WSIMPSQ50/1000) + .137(10000PA)

The adjusted r-squared was .861 The standard error was .117.

MVP is number of MVP awards won, 3000H is a dummy variable (1 if a player reached it, 0 otherwise). The 500HR is also a dummy variable as it is for 500SB and 10000PA (if you made it to 10,000 career plate appearances, you get a 1, 0 otherwise). I used all the voting data from 1990-2009.

What is ASSQ10? It is the square of the number of All-star games played in squared. But AS games played is maxed out at 10. The assumption here is that being an all-star has a positive exponential effect but only up to a point where no more games helps (I have a graph below to help explain this). The GGSQ7 is the same thing for Gold Gloves.

WSIMPSQ50 involves World Series play. First, WSIMP is World Series PAs times OPS. The idea here that the more you play in the World Series the more votes you would get, but by multiplying it by OPS, it also includes how well you played (or just hit). This gets maxed out at 50 and is squared, for the same reason as all-star games (yes, Reggie Jackson is first here and way ahead of everyone else at 141, with Dave Justice and Lonnie Smith tied for 2nd at 101).

The last two variables are interaction variables. GGAS is the gold glove variable multiplied by the all-star variable and WSAS is the world series variable times the all-star game variable. It looks strange that the coefficient values on GGSQ7 and WSIMPSQ50 are negative. But you might notice that they are positive on the interactive variables. I think this is like when a regression uses both X and X-squared in a regression if the phenomena is non-linear (an inverted parabola, for example). The coefficient on X ends up being positive while the x-squared coefficient is negative. The reason I put in these interactive variables was to see if players who were strong in both got an extra boost, as if there was some synergy going on. It seems like they did get an extra boost.

Since the dependent variable can only go from 0 to 100, the coefficient would be very low. So I divided these three variables by 1000 (my stat package was showing coefficient values of .00000 before I did this).

Monday, December 6, 2010

Did Santo Play In An Era Of Poor Third Basemen?

Here are the offensive winning percentages for NL 3B men for different periods. Data from the Lee Sinins Complete baseball encyclopedia.

1941-50) .507
1951-60) .495
1961-70) .516
1971-80) .512
1981-90) .498

Santo had .618 from 1961-70. He was about 9% of the total, so without him it was probably about .506. Nothing unusual. The guys Santo got compared to were not sub par in hitting.

Santo lead the NL 4 straight years in Total Zone Runs (fielding) for 3rd basemen (from Baseball Reference). But his total over those 4 years, 39, is one of the lowest (BR starts this stat in the early 1950s-I calculated the cumulative total of the leaders over each 4 year period regardless of who it was). It is tied for the 7th lowest in the NL. The lowest is 35 and some of the periods that were lower include the 1981 strike year. The average cumulative 4 year total for the leaders was about 55 in the NL and 72 in the AL. Only two periods in the AL were below 40.

So it is possible that in some years Santo benefits by being compared to poor fielding 3rd basemen. But this is probably not alot of his overall value.

See Yearly League Leaders & Records for Total Zone Runs as 3B at BR. Santo's numbers from 1965-8, the years he lead, seem low compared to the AL in those years and the NL in the years both before and after.

Sunday, December 5, 2010

Santo Was Valuable Outside Of Wrigley Field

Santo did seem to benefit alot from Wrigley. But what if we tried to estimate only his value in road games? Doing a quick calculation to find his road OBP & SLG from 1960-73, I got .346 & .413. Does not sound that great. But in his time, it was pretty valuable. Here is the relationship from regression analysis between runs per game and OBP & SLG:

R/G =16.55*OBP + 10.56*SLG - 5.15

A team with an OBP of .342 and an SLG of .413 would score 4.93 R/G. The league average in those years was about 4.06. That would give us a Pythagorean pct of .596. Pretty darn good.

I also ran a regression with winning pct being the dependent variable and runs per game and opponents runs per game being the independent variables. Here is the equation

Pct = .515 + .111*RG - .114*ORG

If a team scored 4.93 runs per game and allowed 4.06 per game, they would have a .596 pct. That is how good Santo was just in road games.

Friday, November 26, 2010

Lefty Grove's Peak Vs. Sandy Koufax's Peak

I used a 5-year period for each guy. For Grove, it was 1928-32. For Koufax, it was 1962-66. The table below has some comparisons:


RSAA means "runs saved above average." It comes from Lee Sinins' Complete Baseball Encyclopedia. The numbers are park adjusted. So Grove has a big lead here, both in total and per 9 IP. I will come back to these numbers later when I plug them into the Pythagorean formula.

Grove was 60% better than average at preventing HRs (that is what the 160 means). He gave up 49 HRs while the average pitcher would have allowed 78 (100*(78/49) is about 160). This gives him a pretty big edge over Koufax. But they are not park adjusted. If they were, Grove would have an even bigger edge. Here are the HR park factors for the Philadelphia A's from 1928-32 from the STATS, Inc. All-Time Baseball Sourcebook: 126, 165, 153, 104, 199 (the 126 means that Shibe gave up 26% more HRs than the average park). Now Shibe Park may have had some asymmetries, so that lefties hit alot more HRs. With Grove more likely to face righties (being a lefty himself), it is possible the park did not hurt him as much as these factors suggest. But A's righties Foxx, Miller and Dykes generally had much higher slugging percentages at home than on the road (from Retrosheet). So my guess is that Grove certainly was not aided by his park in preventing HRs.

Koufax allowed 89 HRs while the league average was 124 and had the following HR park factors in his years: 50, 63, 62, 49, 70 (meaning Dodger Stadium allowed fewer HRs than average). So he was helped quite a bit yet Grove still has the big edge here. He allowed 89 HRs while the league average was 124.

Relative SO/BB is each pitcher's strikeout-to-walk ratio divided by the league average. Grove had a 2.67 strikeout-to-walk ratio while the league average was 0.95. The 2.67/0.95 is multiplied by 100 to get 281. That beats the 225 of Koufax or 100*(4.57/2.03).

The ERA+ comes from Baseball Reference. It is ERA relative to the league average but also adjusted for park effects. Grove only has a slight edge here.

WAR comes from Baseball Reference (and they get it from Sean Smith at Baseball Projections). It is "Wins Above Replacement for Pitchers. A single number that presents the number of wins the player added to the team above what a replacement player (think AAA or AAAA) would add. This value includes defensive support and includes additional value for high leverage situations."

It is not clear to me how Koufax beats Grove here. Grove has alot more RAR or "runs above replacement." It might have something to do with the leverage adjustments. None are made for Grove since the play-by-play data has not been posted at Retrosheet. The WAR and RAR numbers imply that for Grove's years, it took 11 extra runs to win a game (441/40.1 = 11) and only 8.26 for Koufax (347/42).

Baseball Projections says that it normally takes about 10 extra runs to get a win. I wonder if they are using the formula which says it takes 10 times the square root of the number of runs scored per inning by both teams. For Grove's years I calculated that to be 10.7 and for Koufax got 9.54. That would give Grove a WAR of 41.21 (441/10.7) and Koufax 36.37 (347/9.54).

Pitching Runs is "Adjusted Pitching Runs." It comes from Baseball Reference. It is "A set of formulas developed by Gary Gillette, Pete Palmer and others that estimates a pitcher’s total contributions to a team’s runs total via linear weights." Lee Sinins told me it might also be based on decisions, but I am not really sure. Anyway, Grove has a big lead here, too.

Now to come back to RSAA and try to calculate the Pythagorean pct for each guy using RSAA per 9 IP. The AL of 1928-32 averaged 5.12 runs per game (yearly averages weighted by Grove's IP) and 5.12 - 1.98 = 3.14. So if Grove allows 3.14 while his team scored 5.12, he would have a winning pct of .727. Koufax would allow 2.78 while his team would score 4.05 runs per game. That gives him a pct of .679.

One thing I have not mentioned yet or tried to take into account is integration. Last January, I compared Grove's career to Randy Johnson's. See How Might Integration Have Affected The Lefty Grove/Randy Johnson Debate? I tried to estimate how much better the hitters would have been during Grove's time if the percentage of players who were non-white was about the same as during Johnson's. I also tried to adjust for the number of non-white pitchers and non-white fielders. I came up with Grove's ERA going up about 10%. What if I did that here?

Then Grove would allow 3.45 runs per game and his pct would fall to .688. That is still higher than Koufax.

But if we use the adjusted pitching runs, Grove allows 3.32 runs per game (5.12 - 1.8). He would have a pct of .704. Koufax would allow 2.63 runs per game (4.05 - 1.42). He would have a pct of .703. That would make the two about even. Grove would get the edge due to more IP.

But if we raise Grove's runs per game by 10%, to 3.65, his pct would be only .663. That would put Koufax ahead.

Finally, if we knock down Grove's ERA+ from Baseball Reference of 172 by 10%, he would be at 155, below Koufax's 167. The 10% adjustment for integration is just an estimate. It is the same one I used when comparing Grove to Johnson. The % of players and pitchers who were non-whites during Koufax's time was probably lower than during Johnson's time. So adding 10% to Grove's ERA is probably too much. I don't think I know the right adjustment to make. But this gives us some idea of what the effect of integration might be.

If I lowered Grove's strikeouts per 9 IP by 10% from 5.91 to 5.32 and raised his walks per 9 IP from 2.21 to 2.43, his new strikeout-to-walk ratio would be 2.19. That divided by 0.95 would be 2.30. So his relative SO/BB would be 230, still higher than Koufax's 225.

If I raised Grove's HRs by 10%, he would have allowed 54 HRs. Then 78/54 = 1.45. That times 100 is 145. That is still higher than Koufax's relative HR rate of 139.

Sunday, November 21, 2010

Indispensable Seasons Go To WAR! (Or Did Willie Mays Have The Greatest Season Since 1950 in 1962?)

If you are still reading, thanks. I will try to explain.

Suppose a team comes in 1st place, finishing 1 game ahead of the 2nd place team. John Smith had a WAR (wins above a replacement player) of 6. Then is "INDWAR" would be 5 or 6 - 1. His team needed 5 of his WAR to get them into at least a tie for first.

My first post on this was The most indispensable seasons. In that case, instead of using WAR, I used what Pete Palmer calls "Total Player Rating" or TPR (more recently it has been called "Batting + Fielding Wins" or BFW). Here I used WAR from Baseball Reference to find the most indispensable seasons since 1900.

The table below shows the top 25.


When you see a "0" in the games ahead column, it means that player's team tied for first place with another team. Then they had a playoff. Their season's WAR included what they did in the playoff game(s). I tried to estimate their WAR from any playoff games. Probably the most anyone got was about .4 by Boudreau in the one game (he went 4-for-4 with 2 HRs). Some cases are teams that were wild cards, like the 2002 Giants. So they would have been so many games ahead of the next best team. Some 1st place teams in the wild card era were either compared to the 2nd place team in their division if that team was not the wild card or the team that finished 2nd in the wild card if their division's 2nd place team was the wild card.

It probably does not surprise anyone that Yaz is first. But Willie Mays 1962 is not far behind. Guidry 1978 is the highest pitcher. But he probably got a very small amount of WAR in the playoff game (he pitched well but not great). The 1980 Phillies needed great years from both Schmidt and Carlton just to eke out a 1 game victory.

Many of the players are Hall of Famers. Ruth also has the 37th best season in 1916, as a pitcher!

I wondered how well some of these guys did in the clutch that year. So I looked at the top 10 since 1950 when Retrosheet has stats like hitting with Runners in Scoring Position (RISP) and in Close and Late Situations (CL). The table below shows how well the top 10 hit in all situations.


Now with Runners on Base (ROB)


Now with RISP


Now Close and Late.


Now Sept/Oct



If you examine those numbers closely, you will see that the only player to have a higher AVG, OBP, and SLG in all the "clutch" cases than he did in Total was Mays. Click here to see all of these stats grouped by player. It might be easier to see that only Mays did better in all the clutch situations.

In fact, Willie Mays was the best of the ten in Tom Tango's clutch rating, which involves WPA or Win Probability Added. All plate appearances are rated for how much they affect the outcome based on score, inning, etc. Here is the definition from the Fangraphs cite:

"Clutch: A measurement of how much better or worse a player does in high leverage situations than he would have done in a context neutral environment."
Here is how well the top ten did:

Willie Mays 1.4
Jackie Robinson 0.5
Hank Aaron 0.3
Adrian Beltre 0.1
Alex Rodriguez -0.1
Mike Schmidt -0.3
Barry Bonds(98) -0.5
Robin Yount -0.8
Carl Yastrzemski -1.1
Barry Bonds(02) -1.3

This means that Mays' extra good hitting in high leverage situations added 1.4 wins. Seeing as how the Giants finished in a tie with the Dodgers in 1962, that is very important. Mays did well in the 3-game playoff series, too. In game 1, he went 3-for-3 with 2 HRs and a walk. One HR was off of Koufax, in the first inning with one on to get the scoring started. Giants won 8-0. In game 2, he was 1-for-5 and the Dodgers won 8-7. In game 3, he was 1-for-3 with 2 BBs. Giants won 6-4, getting 4 runs in the top of the 9th. Mays singled in a run in that rally and scored another.

When it was all said and done, a great player had to have one of his greatest seasons just to get his team into a playoff. Mays had to hit much better than usual in the clutch and come through in the playoff. What could be a more fantastic year than that?

Thursday, November 18, 2010

Rick Reuschel for the Hall of Fame (Revisited)

See my first post Rick Reuschel for the Hall of Fame .

Here is a brief summary (skipping the more advanced stats I used):

-His strike-out-to-walk ratio was 31% better than the league average

-He gave up 21.6% fewer HRs than average (pitching mainly in Wrigley Field!)

I am doing this again because when I looked at Halladay and the Cy Young award, I noticed that Reuschel is 30th in career Wins Above Replacement (WAR) among pitchers at Baseball Reference. His WAR is 66.3. That seems like a high enough rank in terms of career value. The Hall should have room for 30 pitchers. He had good longevity, pitching over 3500 innings in 19 seasons.

The only pitchers ahead of him in career WAR not in the Hall are: Clemens, Maddux, Randy Johnson, Bert Blyleven, Pedro Martinez, Mussina, Schilling and Glavine. Most, if not all, of them will make it. I counted about 26 Hall of Famers behind him, just in the top 100. He is ahead of Jim Palmer, Juan Marichal, Whitey Ford, Don Drysdale, Jim Bunning, just to name a few.

He had a pretty decent peak value, too. He was in the top 5 among NL pitchers in WAR each year from 1977-80 (1-5-3-4). He also had two other top 5 finishes in his career. He was the 2nd best pitcher in the NL over the 1977-80 period, according to WAR. Here is the top 10:

Phil Niekro 27.6
Rick Reuschel 24.8
Steve Carlton 20.6
Steve Rogers 19.3
J.R. Richard 18.1
Tom Seaver 17.4
Burt Hooton 16
Bruce Sutter 15.6
John Candelaria 15.3
Don Sutton 13.7

He beats Carlton, who had 2 Cy Young awards in those years.

Tuesday, November 16, 2010

Halladay And The Cy Young Award

It was unanimous. That seems a little surprising. No doubt among the voters. Here are the NL leaders in WAR for pitchers according to Baseball Reference:

1. Jimenez (COL) 7.1
2. Halladay (PHI) 6.9
3. Johnson (FLA) 6.4

Seems like these other two guys could have gotten some first place votes.

Halladay is the one of only 4 pitchers since 1980 to have 3 or more straight seasons with a strikeout-to-walk ratio greater than 5 while qualifying for the ERA title. The others are Maddux (4), Schilling (4) and Wells (3). Data from the Lee Sinins Complete Baseball Encyclopedia.

Halladay is one of only 6 pitchers since 1980 to have 3 or more straight seasons with an ERA less than 2.80 while qualifying for the ERA title. The others are Maddux (7), Rijo (4), Johnson (4), Clemens (3) and Brown (3).

Halladay has the most WAR over the last three years, 20.2. The last pitcher to have 20+ WAR over three years was Santana, 2004-6. Here is the top ten from 2008-10:

Roy Halladay 20.2
CC Sabathia 16.8
Tim Lincecum 16.7
Cliff Lee 16.6
Felix Hernandez 16.3
Jon Lester 16.2
John Danks 16.1
Zack Greinke 15.6
Ubaldo Jimenez 15.3
Johan Santana 14.4

Halladay now ranks 10th in Cy Young Award Voting Shares. Besides his two wins, he has four other top 5 finishes. He joins Gaylord Perry, Roger Clemens, Pedro Martinez and Randy Johnson as the only pitchers to win the award in both leagues. Here is the top 10 in award shares:

Randy Johnson (5 wins) 6.5
Greg Maddux (4 wins) 4.92
Steve Carlton* (4 wins) 4.29
Pedro Martinez (3 wins) 4.26
Tom Seaver* (3 wins) 3.85
Jim Palmer* (3 wins) 3.57
Tom Glavine (2 wins) 3.15
Sandy Koufax* (3 wins) 3.05
Roy Halladay (2 wins) 2.91

*Hall of Famer

Halladay has finished first in WAR 2 times and has 5 second place finishes. His career WAR is no 54.3. That is 62nd best ever. In a year or two he will be in the top 40. Maybe he will end his career in the top 25.

Friday, November 12, 2010

Players Who Won The Triple Crown Over A Two-Year Period Since 1920

I got started on this because I wanted to see if Albert Pujols did it for the last two years. He just missed. More on that later. I set the plate appearance (PA) minimum at 800. The tables below show the winners. Data came from the Baseball Reference Play Index.

What is interesting to me is that in some cases, the winners were far ahead of the other players in all three stats, that Al Rosen was the only guy to do it twice in a row, and that Albert Belle is the only guy to do it since 1954. Rosen was probably only able to do it because Ted Williams was in the Korean War. And Williams is the only other guy to do it twice.




Now back to Pujols. The table below shows the leaders over the 2009-2010 seasons with a 1,000 PA minimum. He's just a bit beind Votto and Ramirez in batting average (BA). I hate it when mere mortals get in the way. May the Gods show them no mercy.


Then I thought "what about a 3-year triple crown for Pujols?" No luck there either, since Ryan Howard beats him out in RBIs. This is the next table. 1500 PA minimum.


Not having the RBI lead is certainly not Pujols' fault. The next two tables show how both he and Howard hit with Men On and with Runners in Scoring Position (RISP). Pujols ends up walking alot more in those cases. In all cases, Pujols walks 13.5% of the time while it is 12.4% for Howard. But with Men On, Pujols' walk rate is about 21% while Howard has 12%. With RISP, those numbers are about 29% & 16%. So, even though Pujols hits alot better with Men On and with RISP, as you can see below, he ended up with fewer RBIs.



But Pujols gets the last laugh. He has won the 10-year triple crown, with a 2,000 PA minimum. This is the last table.

Friday, November 5, 2010

The Weather Was Nice For The World Series But How Was It In Some Other Major League Cities?

I have wondered what things would have been like if the Twins had made it to the World Series. Imagine night games, outside, in Minnesota, in the last week of October and the first week of November. So while I watched the series, I checked the temperatures in various cites using Accu Weather. The data I collected can be seen at World Series Weather 2010.

The first column gives the temperature and the second column says "wind ch" for wind chill. I think that is what Accu Weather means by real feel. I also recored the local time (they were all PM, inspite of my typos). In the last column I mention any description that Accu Weather gave, like showers if it was raining at that time or if showers were on the way. In some cases they said something about wind gusts. On Nov. 1 at 7:11 pm, it was 47 degrees in Minneapolis with wind gusts over 40 MPH. That would have made for a fun game as the night went on (I don't know why Accu Weather showed a "real feel" of 46 degrees with such strong winds-I also don't know why the real feel was sometimes higher than the stated temperature).

There were some low temps out there but probably not anything we have not seen in recent years. October 28 in Cleveland had a temp of 43, wind chil of 30 and showers. That would have been no fun to play in. So it seems that no city realized my worst fear of sub-freezing weather and snow when a game was supposed to be played.

Sunday, October 31, 2010

Have The Rangers And Giants Discovered A New (Old) Way To Win?

That is what a recent Wall Street Journal article says. See Hitting Baseballs, Just Not as Far: Giants and Rangers Win With Contact Hitting, Bunts and Baserunning; the 'Lost Arts'. But I don't think that they are doing anything so different from other teams that it helps them score extra runs. Here is an excerpt:

"San Francisco was 17th in runs scored and 13th in slugging percentage this season. But they ranked fifth in strikeouts and third in sacrifice bunts in the National League and fourth in all of baseball in sacrifice hits.

Texas was only ninth in slugging percentage, but the team had the most sacrifice bunts in the American League, the second-most sacrifice flies and the fourth fewest strikeouts. The Rangers were also seventh in the majors in stolen bases."

Both teams, however, are actually scoring just about the number of runs you would expect based on their OBP and SLG. From 2007-2009, the relationship between runs per game and those stats in MLB was:

R/G = 16.04*OBP + 11.595*SLG - 5.52

The Rangers had an OBP & SLG of .338 & .419. The equation predicts they would score 4.76 runs per game while it actually was 4.86. So just about what you would expect, meaning all those sacrifices and SBs are not making much difference.

The Giants had an OBP & SLG of .321 & .408, projecting to 4.36 runs per game while it was actually 4.3. Just like the Rangers, all these "small ball" strategies are not making much difference. (the equation comes from a linear regression analysis of all 90 teams from 2007-09-the r-squared was .904 and the standard error of the regression was .137 runs per game).

Another regression, based on all teams from 1989-2002, shows the relationship between team winning pct and OPS differential. Here it is:

Pct = .5 + 1.25*OPSDIFF

The Rangers hitters had an OPS (OBP + SLG) this year of .757 while they allowed an OPS of .709. The Giants had .729 & .683. So the two team's differentials, respectively, were .048 & .046. The numbers below show each team's predicted pct, and predicted wins, followed by their actual wins in parantheses:

Rangers) .560-90.72 (90)
Giants) .558-90.32 (92)

Each team won just about the number of games expected (each within two of the prediction). There are no extra wins due to using "lost arts." In fact, they have done well by some combination of hitting for power and getting on base and generally preventing their opponents from doing so. This is a time honored way of winning, as Branch Rickey explained back in 1954. I posted something about that earlier this year. See Scouts vs. Statheads: What Might Branch Rickey Say?.