Showing posts with label statistics. Show all posts
Showing posts with label statistics. Show all posts

Friday, May 7, 2010

Fun With Shutdowns and Meltdowns

Based on a conversation which originated at Beyond the Boxscore and continued over at The Book blog, FanGraphs created two new stats for relief pitchers based on WPA. Fortunately there are no intimidating acronyms. The positive one is "Shutdowns (SD)" and the negative one is "Meltdowns (MD)". I think everyone can understand that.

The Yankees are tied for the fewest meltdowns in the MLB with 8 but they've also had the second fewest appearances by relievers this year primarily due to the length their starting pitchers have provided for them. Yankee starters are averaging 6 1/3 IP per start and on nights when Javy Vazquez isn't pitching that number jumps over 6 2/3.

Looking at it another way, the Yankees' ratio of meltdowns to shutdowns is better than 1:2, which is good for 7th best in the league. Without getting my hands too dirty, that seems like a pretty good way to rank the performance of team's bullpens so far this season. When sorted by meltdowns/shutdowns, the teams fall out into three distinct tiers.

Tier 1: These teams all have a better than 2:1 shutdown to meltdown ratio. Most of them have an ERA better than league average with the exception of the Tigers (13th), Nationals (18th) and Pirates (dead last).

(The columns next to ERA and MD/SD are the
team's league-wide ranks in those categories)

What's apparently right away is that there is a decent general correlation to ERA but some extreme outliers (likely due to the fact that it's still early in the season).

The Tigers have a bullpen ERA that is slightly higher than league average but still have the best MD/SD in the MLB. Detroit has converted 80% of their saves (tied for 4th best in the league) which is a good indication that they have been getting the job done when it matters but giving up a lot of runs in low leverage situations. The Nationals are in the same boat with a bad ERA (19th), high SV% (11/13) and a solid MD/SD (5th). The Pirates are an interesting case, as their ERA is clearly inflated by having eight losses by seven runs or more, including two games they lost to the Brewers by a combined 33 runs. Those big losses have essentially no effect on WPA after a certain point but are given equal weight in ERA.

Tier 2: All of these teams have between a 1:2 and 3:4 MD/SD ratio.

The Reds have the fourth most relief appearances in the league so it makes sense that they have the most shutdowns with 33 and are in the top 5 for meltdowns with 18. They might have a good ratio but their ERA is indicative of the fact that their 'pen is over-taxed. Paging Aroldis Chapman. The Red Sox are in a similar situation with fewer shutdowns but a lower ERA. The Mariners are the only team in this cluster with an ERA in the top 1/3 of the league.

Tier 3: The final tier consists of teams who have more than 3 meltdowns for every four shutdowns. Four of them - the Cubs, Dodgers, Giants and Royals - have more MDs than SDs.

Incredibly, the Giants, who are leading the league in bullpen ERA, have only 10 shutdowns (last in the MLB) and are second to last in terms of their MD/SD ratio.

Clearly, over the course of the season, teams like the Giants and Tigers are going to see their ERA and MD/SD ranks converge on each other. The Giants will get some good performances from their relievers when it matters and the Tigers will see their guys give up some runs when it counts too.

Neither ERA nor MDs and SDs are perfect ways of measuring reliever's contributions but each brings their own unique perspective. And when combined with or contrasted against each other they have a way of teasing out information that we might not have known before. It should be interesting to keep on eye on these stats as the season goes forward and FanGraphs includes them on the player pages.

Thursday, April 29, 2010

Have The Yanks Over Or Underachieved So Far?

Good morning, Fackers. Through the first 1/8 of the season, the Yankees have won at a 65% clip, which, if continued throughout the entire season would net them 105 wins. Not too shabby, right?

Well, Darren Everson from the Wall Street Journal sees a bunch of guys who are well off their career numbers and determines that the Yanks haven't been that impressive:
The Yankees are hiding a dirty little secret: This team, despite its 13-7 record, actually hasn't played all that well.

We don't just mean over the past week, which saw them lose four of five before their 8-3 victory over the woeful Baltimore Orioles Wednesday night. We mean period.

...In total, there are probably 13 players on the 25-man roster who would be happy with their numbers if they maintained them all season. Yet not only are the Yankees playing .650 baseball, there's a feeling around the sport that this team is virtually certain to reach the postseason. Never mind that the Tampa Bay Rays, another team in the Yankees' division, have raced off to a 16-5 start.

This all means one of two things: Either the Yankees are going to be really scary once they get their individual acts together, or they're actually fortunate to be in the position they're in right now.
There are certainly guys on the Yankees who are struggling, but as a team they are balanced out by players who are performing better than expected. The Yanks' run differential aligns evenly with their actual record, so I disagree with Everson's implication that they "haven't played that well".

He names Mark Teixiera, Nick Johnson and Javier Vazquez are the prime Yankee underachievers and I think we can all agree that those guys are sure to improve as the season wears on. But the Yanks are also getting outstanding contributions from a lot of players that aren't likely to be maintain them once sample sizes start catching up to them.

Robinson Cano is hitting a scalding .390/.430/.701. He has 6 homers and 15 RBIs, paces that would add up to 48/120 for an entire season. Those aren't completely out of the question, but even the most ardent Cano supporters would settle for much less. Jorge Posada is at .316/.400/.649, a line that will be atrophied as the season wears on and he logs more innings behind the plate. Nick Swisher has a 140 OPS+ whereas the highest mark of his career is 129. Frankie Cervelli (185 OPS+) and Marcus Thames (240) have both contributed valiantly in their limited roles and will return to earth as they are given more and more chances.

On the pitching side of things, even with Vazquez's terrible start, the rotation has a 3.50 ERA. Andy Pettitte and Phil Hughes have been outstanding, with ERAs of 1.29 and 2.00, respectively. CC Sabathia is off to a great start. If A.J. Burnett could keep his mark around 3.20, it would be the lowest of his career. The bullpen has had some bad moments but overall, they've combined for a 4.30 ERA.

The only DL stint served by a Yankee player has been the one by Chan Ho Park and he's probably the 20th most important player on the team. It'd be nice if everyone was healthy all year long, but we know that's not going to happen.

I think the last paragraph of the quote from Everson's piece is a false dichotomy. Over any 20 game stretch of a season, individual players are going to be producing above, at or below what their numbers for the year turn out to be. However, when you combine everyone's production together, it typically evens out. When Teixeira, Johnson and Vazquez get their acts together, Posada, Cano and Pettitte will be returning to earth. The Yanks have played well so far, even if a few key players haven't.

Thursday, April 8, 2010

Why We Overvalue Relievers

Earlier this offseason, Joe DeLessio of New York Magazine did a countdown of the most important Yankee players. His number 1: Mariano Rivera.

Those of you who are sabermatrically-inclined probably just responded with a collective eye-roll. Wins Above Replacement ranked Rivera as only the 5th most valuable pitcher on the Yankees last season, behind CC Sabathia, A.J. Burnett, Andy Pettitte and Phil Hughes. When you include position players, Rivera drops to 15th, just behind Brett Gardner.

How is that possible? Dave Cameron explains:
While the quality of [relievers] work is very high, the quantity is low, which limits their total value. It’s nearly impossible to rack up huge win values while facing less than 300 batters per season. Yes, each of those batters faced are more critical to a win than a regular batter faced, but this is accounted for in WAR.
It's not to say that WAR is a perfect measure - I don't think anyone believes that Brett Gardner is more irreplaceable/valuable than Rivera - but if the numbers are even close, then it's clear that people (media, fans, etc.) tend to overvalue relievers. When you look at these numbers, it becomes clear that some front offices share this skewed view and are willing to overpay them as well.

Why is that? Perhaps a series like the one the Yanks just wrapped up can shed some light.

None of the six pitchers Joe Girardi and Terry Francona called upon to start the last three at Fenway games factored into a decision. Only Sabathia and Lackey were particularly close - both of them watching their lead evaporate under the watch of the man who relieved them. Therefore, each contest was decided by pitchers who entered the game via the bullpen.

Chan Ho Park was on both sides of that equation, taking the loss on Sunday night after allowing a two run homer to Dustin Pedroia and getting the win last night after throwing three scoreless innings while the game was knotted at 1. Jonathan Papelbon had a similar experience, converting a save on Friday and blowing one last night. Each played the role of goat and hero just a few nights apart.

While WAR can objectively weight the contributions of pitchers by leverage, we as fans can't hope to be nearly as unbiased. As a close game progresses, stress and anxiety in the attentive viewer build. Our joy and frustration are multiplied by those factors and relief pitchers are the one major variable in the equation. The lineups are essentially the same but as the stakes within the game increase, the faces on the mound change.

And that's why someone would try to make the case that Mariano Rivera is the most important player on the Yankees. He might not be the most important from a zero sum sabermetric perspective, but he is on a purely observational standpoint, if you have a rooting interest in the team, Mo is the man. CC Sabathia throws far more innings, but they don't seem to have as much on the line. Mark Teixiera plays in almost every game, but most of his contributions occur under ordinary circumstances.

When Rivera enters the game, as a fan, you can exhale. We've seen him do it so many times before, it's hard not to be confident. You trust that he's going to get the job does until he doesn't - and then you assume that he'll do it next time. Conversely, the three innings that Chan Ho park pitched felt significantly more tense and uncertain. The difference between them is in that respect more than commensurate with their respective abilities.

Of course, Mo did what he usually does during the past two nights. He gave up just one baserunner and the go-ahead run never came to the plate. A couple of late nights at the office and two saves in the book.

It might not show up in advanced stats or translate to as many wins above replacement as we would assume, but Rivera and other trustworthy relievers contribute greatly to the enjoyment of rooting for the team. If you spend enough time reading about and understanding the principles of sabermertics, you should realize that his importance is magnified in your mind. But when a save situation rolls around, he really does seems like the most important guy on the team.

Friday, February 5, 2010

Lackey And Vazquez

Even before John Lackey signed with the Red Sox for a deal nearly identical to the 5 year, $82.5M one that the Yankees gave A.J. Burnett a year prior, many saw the two as being very similar. Both pitchers are right handed, oft-injured about six and a half feet tall and around 30 years old with intimidating on-the-mound demeanors. Today, however, I wanted to compare Lackey to Javer Vazquez given their parallel entrance to the Yankees vs. Red Sox rivalry and their disparate reputations in regards to handling pressure.

Lackey has been in the league since 2002 and has averaged 188 innings per season since then. Over that same time period, Vazquez has averaged 215. Lackey's ERA is about a quarter of a run lower over stretch, but that's essentially erased by having to fill in those extra 27 innings a year with a replacement level pitcher.

In general, when John Lackey is healthy, he's a better pitcher than Javy Vazquez. But he's also the Red Sox highest paid player ($18M this year) and is expected to contribute to the top of their rotation. The Yankees are paying Vazquez only about 2/3 as much and hoping that he slots in as their number four.

But what about their reputations under pressure? The idea for that comparison between the two comes from fellow LoHud pinch hitter and editor at the Harvard Crimson, Yair Rosenberg. On Sunday, Yair dropped me an email with the following suggestion/request:
I think the more productive comparison for AL East purposes would not be to Pettitte or Glavine, but to Lackey, whose reputation is that of a big game pitcher, and who is essentially the corresponding addition to this year's Red Sox as Vasquez is to the Yankees. It would be really interesting to see if the stats bear out Lackey's clutch rep - and might go a long way towards predicting the key factors in the coming Yankees-Red Sox race. I'd love to see a post on that.
So here we go. Fighting in the red corner, we have John "Big Game" Lackey, the winning pitcher in Game 7 of the 2002 World Series and supposed consummate clutch performer as anointed by his former manager. In the blue corner is Javier "Can't Handle New York" Vazquez, the man responsible for one of the more infamous home runs in Yankee history, who was called out publicly by Ozzie Guillen for not stepping up when it counts.

There is no question as to who has the better postseason resume. Lackey was thrust into the spotlight at an early age, his team reaching the playoffs in his first season in the majors and asking him to start Game 7 of the World Series only four days after his 24th birthday. Since then, he's been back to the postseason 5 times and pitched a total of 78 innings to a 3.12 ERA.

Vazquez, on the other hand, was trapped on bad Expos teams (no offense, Jonah) for the first six years of his career, and didn't pitch during October until 2004. His performance in the postseason has been pretty dreadful (10.34 ERA), but he's only had a chance to throw 15 2/3 innings in the playoffs.

Do these reputations carry over into the regular season? Do their postseason resumes line up with how they handle pressure during games throughout the year? We know the answer to that question when it comes to Vazquez, as we have assessed his clutch reputation at length here and in other places.

That first and more in-depth inquiry into Vazquez's purported lack of clutchiferousness began with his FIP/ERA differential. Coincidentally, Lackey and Vazquez have identical 3.83 career FIPs. However, Lackey's career ERA is 3.81 while Vazquez's is 4.19. Leaving aside team defense - which would be very difficult to quantify over multiple years and teams - it's helpful to look at situational and leverage statistics when trying to explain FIP/ERA differentials.

There are some notable similarities between the Vazquez and Lackey in the chart to the right. Both pitch better with the bases empty than with runners on. They have similar tOPS+ distributions when the score of the game is within 4 runs.

Naturally, the biggest differences come in the smallest sample sizes. Lackey has done much better with the bases loaded than Vazquez and far worse when the game is out of hand.

Both Lackey's distributions are optimal and both are significant. If you could choose a situation to pitch your best in, it would be when the bases were loaded. If you had to give up runs, you would prefer to allow them when the margin of the game was greater than four runs. But there is a limit to how much these numbers can tell us. Lackey has only 141 plate appearances with the bases loaded while Vazquez has 163.

The sample sizes are larger for when the margin is greater than four (394 for Lackey, 811 for Vazquez) but those at bats are by definition less important. Lackey is obviously better in those situations, but not likely by as much as the numbers indicate.

What about the leverage index, though? While Lackey's numbers don't tell a coherent, progressive story like Vazquez's do, it's still clear that he pitches his worst in high leverage situations. Again, high leverage is based on the smallest sample size among the three levels, but each pitcher has over 1000 plate appearances to draw upon. So perhaps Lackey can't simply summon his best performances when the stakes increase.

If there was something about Lackey's internal constitution that gave him to ability to elevate his performance under pressure, wouldn't it show up in the leverage index? Shouldn't he be able to sense when the game is on the line and reach back for a little extra?

This contradiction begins to chip away not at Lackey's resume in particular but at the manufactured archetype of the "big game pitcher". It's one thing to have had good results in the postseason but it's another entirely to universally improve as the leverage increases. You can argue that the playoff results are more important, but Lackey has only faced 328 batters in postseason play. I think the regular season numbers tell us more.

While it may be convenient to label certain pitchers as big time performers and others as choke artists, they rarely fall neatly into one category or another. More correctly, there are players who have performed well in certain situations and others who have not.

As far as this season goes, it will be interesting to see who is better, Lackey or Vazquez. It's very likely that Lackey will have a lower ERA than Vazquez but based on their respective histories, Vazquez should be the better bet to throw more than 200 innings. However, perhaps this is the year that Vazquez's heroic workload over the past decade-plus catches up with him and it's also the first time in 3 years Lackey makes more than 30 starts.

Time will tell, but remember that the Yankees only need to get 2/3 the performance out of Vazquez to get as much value as the Sox do out of Lackey.

Monday, February 1, 2010

Baseball Braces For Bloomberg

For about four hours yesterday, a good portion of the baseball blogoshpere gathered at the Bloomberg headquarters in Midtown Manhattan. Many went through the painful process of putting on pants, emerging from their mother's basements and commuting to New York City in exchange for an up close and personal tour of Bloomberg's new baseball software, along with lunch, beverages, some especially delicious frosted peanut butter brownies and an awesome shirt.

Even those of us who don't work in finance were familiar with Bloomberg's capabilities in data gathering, organization and visualization and were anxious to see how they had applied it to baseball. The company is offering two separate products, one aimed towards fantasy baseball players and the other designed specifically for and in partnership with the teams in the MLB. The demonstration of the fantasy version came first.

Within the fantasy product, there are two main incarnations of the program, available for purchase separately. The first is the draft kit, in which you can enter the specifications of your league regardless of where it is hosted (CBS, ESPN, Yahoo, etc).

Bloomberg's system revolves around a proprietary B-Rank based on a 5x5 league, which the presenters disclosed next to nothing about. I understand reasons for the secrecy, but I think users might find this off-putting in contrast to the availability of data throughout the rest of the software. It's tough to get a someone with a sabermetric slat to put their faith behind a metric whose methodology is purposely concealed.

Beyond the B-Rank though, you can search for and prioritize players by any statistical category (and by multiple categories at once) to ensure you have a balanced roster. Conveiently, Bloomberg has already computed player's eligibility by position to save you that headache. With your subscription, you have access to a player's average draft position in relation to their projected production, providing a simple visual representation of their value among many, many, many other intuitive features and comparisons. If you take your fantasy draft preparation seriously, the biggest limitation of the utility of the software is the amount of time you want to put into it.

Similar to the draft tracker, the in-season version of the Bloomberg software offers an impressive depth of information and is very customizable. They don't have historical or minor league stats available yet, but those are included in the professional suite and could be added in the future. The in-season version will be the most useful to this site and you will likely see some of the Bloomberg visualizations in our posts once the software debuts on February 18th.

The draft kit is being offered for $19.95, the in-season tools for $24.95 and both of them together will set you back $31.95.

The professional level software was what was most tantalizing to most of us bloggers in attendance. With a less colorful and more data-heavy layout, I could imagine spending days on end poking through the spray charts, pitch predictors and countless other tools that were offered. The teams are being given a free trial which started at the Winter Meetings and extends through the All-Star break, during which time Bloomberg has been and will be making every effort to customize the software and give the teams exactly what they want.

One of the underlying themes Bloomberg bent over backwards to convey was that what we were looking at on the 7th floor of 731 Lexington Ave. yesterday was just the beginning. They welcomed and encouraged - and might have even demanded - our input and suggestions were that socially acceptable.

To a company like Bloomberg, the market for baseball information is pretty limited. The financial firm employs over 10,000 people and has 126 offices around the world, so their entrée into the world of sports statistics and information isn't likely to significantly affect their bottom line, even if their software becomes ubiquitous among teams, fantasy baseball players and bloggers alike. But they seem committed to constant improvement of the technology nonetheless.

What was most encouraging about yesterday's presentation was that Bloomberg seemed to embrace the sabermetric ideal, even though they had arrived there from a financial background. They aren't attempting to give people the answers they are looking for, they just want to give them the information to come to their own conclusions. Through it all, there was a notion of humility and the desire to improve the product in every way possible and that bodes well for the future of the software.

Monday, January 4, 2010

Is Javier Vazquez Unclutch?

Throughout his career it has been intimated that Javier Vazquez has the stuff of an ace but the track record of a back of the rotation starter. He has excellent peripheral numbers (3.45 K/BB), a better than league average ERA (107 ERA+) and averages well over 200 innings per season. Above average performance and lots of innings should be a winning combination, but Vazquez has always seemed to underachieve when it comes to his won-lost record. His four postseason appearances haven't been pretty either.

Plenty of Yankee fans are worried that he doesn't have the make up to pitch in New York. Vazquez bombed in the second half of his only season with the Yankees and served up perhaps the most infamous home run in Yankee history. But all of that was more than 5 years ago and the latter was the result of one pitch. Ozzie Guillen ran him out of town in Chicago after calling him out for not being a big game pitcher, but Guillen was supposedly responsible for ousting Nick Swisher too and that worked out pretty well for the Yankees.

Does Javier Vazquez's performance get demonstrably worse under pressure, or has he been a victim of marginal teams and bad luck? Is his reputation as someone who shrinks as the expectations grow deserved, or is it the conflation of a few unrelated events?

As I attempt to take a deeper look into these questions, I'm going to go back to something Tom Tango said during his Q & A with Mike Silva last week. In the process of justifying FIP as a useful statistic, Tango pointed out that there is a strong correlation between the best pitchers in the league over a long period of time when ranked by FIP and ERA.

Those familiar with the two stats might take that for granted at this point, but it's fairly remarkable considering all the things that FIP totally disregards that factor into ERA: singles, doubles, triples, runs allowed, etc. What's more intriguing, he contends, is what the significant gaps between FIP and ERA over the long term can tell us:
Tom Glavine’s career FIP is about 0.50 runs worse than his career ERA. This signals that Glavine does something extra, either he can sequence his events better (leaves alot of runners on base for example), or he has better control on his balls in play. And Javy Vazquez’s ERA is about 0.30 runs worse than his FIP, which signals something different, that perhaps he gives up alot of doubles, or doesn’t sequence his events well, etc.

Overall, two-thirds of pitchers will have their FIP and ERA be within 0.20 runs of each other, and almost all will be within 0.40 runs of each other.

That’s the power of FIP: that it’s designed to tell you one specific thing, and it tells you a second, perhaps even more important, thing.
So what's that thing? That's the million dollar question.

One interesting fact is that Vazquez has been extremely consistent in under performing his FIP throughout his career. He's only had an ERA lower than his FIP twice in his 12 professional seasons and only then by the slimmest of margins.

Even in his career year with the Braves last season, his FIP was lower than his ERA (his low FIP was one of the reasons Keith Law included him on his Cy Young ballot). I think we can agree that this not just the result of random chance: there is something about the way that Vazquez pitches that causes him to have an FIP lower than his ERA.

So what is it about Vazquez's pitching that could be causing this? Let's look at the career FIP/ERA differentials for Vazquez, Glavine as well as Andy Pettitte:

Above, Tango says that almost all pitchers will have an FIP and ERA within .4 of each other, so keep in mind that we are looking at two outliers extreme outliers here in Vazquez and Glavine. This is the main reason I chose to include Pettitte, who is closer to the norm.

If you're a Yankee fan, I bet you're a little surprised that Pettitte's differential is closer to Vazquez's than Glavine's. Pettitte has the reputation of being able to bear down and pitch better when it matters and seems to induce a lot of double plays, both of which would lower his ERA but not FIP. But the numbers indicate that he's not as great at controlling his outcomes as many assume.

As Tango mentioned above, one thing that might inflate ERA independently of FIP is giving up a lot of doubles. (Numbers are normalized over 200 IP to account for different career lengths):

Vazquez and Pettitte give up more two-baggers than Glavine does, but four over the course of 200 innings doesn't seem like enough to account for a shift in 3/4 of a run in ERA. Furthermore, Vazquez and Pettitte give up doubles at a nearly identical rate, but the latter has an FIP significantly closer to his ERA. There's something else at work here.

The other possibility that is mentioned above is "sequencing". Maybe Vazquez gives up hits in a way that hurts him. For example, starting an inning with a walk before a double often results in a run scored and a man on second. Conversely, a double before a walk likely means men on first and second without a run scoring. Vazquez gives up his fair share of home runs as well, and they are obviously more harmful with men on base (something that doesn't show up in FIP but does in ERA).

A quick and dirty way of teasing this out of the data is to look at a pitcher's stats with men on base. The numbers below are shown in tOPS+, which compares a pitcher's performance in a given scenario to his performance in other situations. (A score of 100 means is a pitcher is average in that situation while anything lower means they were better and higher means they were worse)

One caveat: the bases loaded data is by far the smallest sample size of the bunch. Vazquez has faced 163 batters with the bases loaded while Glavine and Pettitte have 428 and 237, respectively. That said, there is a major difference between Vazquez and the other two (particularly Glavine) when it comes to their performance with men on base and especially with the bags packed. And those situations are the ones that separate FIP from ERA the most.

Now let's look at how each pitcher performed based on the score of the game, again using tOPS+:

The smallest sample size in this case is "> 4R", but even for Vazquez that includes 811 plate appearances. Pettitte has the most favorable distribution, pitching his best when the game is close and his worst when it is out of hand. Vazquez, however, is just about average in tight games but much better when the lead or deficit is greater than 4.

Now let's put the last two items together. Baseball-Reference's Leverage stats take into account the occupancy of the bases as well as the score of the game, as does FanGraphs Clutch stat:

This tells a slighty different story about Pettitte but affirms what's becoming a trend for Vazquez - he doesn't perform as well in high leverage situations as he does in lower ones.

Does WPA agree?

Yes. It appears that Vazquez has not given his team as good of a chance to win as either Pettitte over Glavine over the course of his career, but this is to be expected to a certain degree because WPA is tied more closely to ERA then FIP. In a sense, we already knew this.


=====
So what do we make of all this?
=====


It's tempting to use these stats to claim that Vazquez can't handle pitching in New York. His data seem to trend very uniformly (remarkably so) from good to poor as the leverage increases. However, why didn't he cripple under the weight of the New York media in the first half of 2004 when he had a 3.56 ERA in 118 2/3 IP and won 10 games?

To say that he can't handle New York not only gives too much weight to a small sample size but requires a jump that conflates the pressure of in-game situations to be analogous to the demands of pitching for one franchise or another. Does the weight of overall expectations have the same effect on performance that increases in in-game leverage do?

Well, maybe. As 'Duk from Big League Stew pointed out in September, Vazquez's best seasons in terms of ERA+ have come with teams that were out of contention and his worst years came with teams in playoff races:
Three of his top ERA+ years came in the anonymity of Montreal and one came for the 2007 White Sox, who went 72-90. This year's ERA+ of 139 equals his career-best with the 2003 Expos, but while the Braves stuck around as a potential contender for longer than expected, they didn't occupy striking distance space for long.

Meanwhile, Vazquez's worst ERA+ years — with the exception of his first two seasons — all came with contenders: the '04 Yankees, the '05 D'Backs and the '06 and '08 White Sox.
So what is he doing wrong? As we saw when trying to identify some differences between A.J. Burnett's good and bad starts last week, it's extremely difficult isolate any one underlying factor or find a specific reason that Vazquez's numbers go bad as the leverage of the game increases. Javy's K/BB ratio slips from 4.34 to 3.26 to 2.57 as the leverage rises from low to medium to high. His batting average, on-base and slugging percentages all ascend with the gravity of the situation as well.

Even if you grant that Vazquez gets worse under pressure and will pitch worse just by virtue of being a Yankee, he's still likely to be better than league average and throw more than 200 innings. It would be extremely difficult to do that and not add significant value to a team regardless of how his performance is distributed by leverage.

And of course, there's a big difference between "hasn't" and "can't". I'm willing to say that Vazquez certainly hasn't pitched well under pressure in his career, but not that he can't. He clearly had a great year in Atlanta and some of that has been attributed to an improved change up, giving him a second pitch to miss bats with in addition to his curveball. FanGraphs shows that his curveball was what stood out last year, but his change up looked to be improved as well.

It's certainly not impossible that Javy has a good year and surprises his doubters (and even some of his supporters). Vazquez might be frustrating to watch at times in 2010, but the beauty of the situation is that our expectations for him shouldn't be that high to begin with. Not too many teams have the luxury of acquiring a very good pitcher and hoping that he's just average.

Wednesday, December 23, 2009

The Fine Line Between Awesome & Awful

Monday at The Hardball Times, Nick Steiner attempted to figure out what stats (particularly Pitch f/x) could tell us about the difference between a pitcher at their best and at their worst. We continually lean on clichés like "He didn't have his best stuff today" to explain why a pitcher has a bad outing. It seems apparent fairly early in a game, at least in hindsight, when a starter is dealing or is not.

But how much of that is confirmation bias? In other words, how much does the outcome of the start effect how we remember our perceptions of the beginning of the game? Maybe the difference between a 7 inning shutout and a 7 run disaster isn't "stuff". Perhaps, from the pitcher's perspective, there isn't much difference at all. Is it possible that Joba Chamberlain really did "throw a lot of good pitches" in some of his poor outings?

For a subject, Steiner chose A.J. Burnett, because of the stark difference between his best and worst outings. When sorted by Game Score, Burnett's 10 best starts in 2009 added up to an ERA of 1.06 while his 10 worst came out at 9.13. He went (6-2) in his top 10 and (0-6) in his bottom 10.

There should be some major differences between these two groups of starts. You'd expect to see some patterns emerging in terms of velocity, or movement, or location or pitch selection, right?

In short, no. There was almost no difference at all.

Steiner dug through all of Burnett's Pitch f/x data for this year, painstakingly categorizing it by pitch type (4-seam fastball, 2-seam fastball, change up, curveball, slider), movement (horizontal, vertical), location (outside, border out, border in, middle), batter (lefty, righty) and count (pitcher's, hitter's, neutral).

He sliced the data in lots of intuitive ways but found almost no significant differences between Burnett's good and bad starts. And for every directional variation which might explain his better starts (fewer pitches down the middle in good starts), there is another which runs counter to what is expected (better velocity in bad starts). In my own look at the numbers, I found that Burnett actually walked fewer batters (29) is his bad starts than he did in his good ones (32).

So what separates a great start from a terrible one, if not for pitch selection, movement and location?

For one thing, there is a whole lot more luck involved with pitching than we realize. In a span of three starts this year, Burnett bookended a 4 2/3 inning, 7 run outing against the White Sox with two shutouts against the Rays and Red Sox, each at least 7 innings. Three starts, two absolutely brilliant ones and one that a AAA call-up would be ashamed of. (Relax conspiracy theorists, Jorge Posada caught all three of them.) There are few other professions where such wild variations between success and failure are common at such a high level.

Part of this is the fact that it only takes one pitch to alter the outcome of a game. One three run home run can change the complexion of a start entirely. And the difference between it ending up as a round tripper and a fly ball on the warning track is a matter of a fraction of an inch on the bat. That's just one pitch out of 100 or more.

If you look beyond Pitch f/x, some other things turn up in Burnett's starts. While his percentage of strikes looking was almost exactly the same regardless of the type of outing (18.7% to 18.5%), the occurrence of strikes looking was much higher in his better starts (10.4% to 6.6%). He also allowed almost twice as many fly balls and line drives in his 10 worst starts while ground balls were between 7% and 8% in both.

Are we to believe that he is throwing the same quality of pitches in both groups of outings and getting wildly different results just based on luck? If it was a random chance, the swings and misses, line drives and fly balls would be more evenly distributed. I think it's more likely that there is something that Pitch f/x isn't capable of telling us.

Especially in a broad analysis like the one Steiner conducted, it's difficult (maybe impossible) to zero in on the things that separate a curve ball that induces swings and misses from one that results in an opposite field single. It would have a hard time telling a fastball down the middle in a 3-0 count (unlikely to be swung at) from one when the batter was ahead 2-1. It can't tell which locations are preferable to which hitters, given that some like the ball inside while others favor it out over the plate, for example. Mistakes made with men on base are most costly than ones with the bags empty. What about pitch sequencing, or the amount of pitches hitters saw, how often Burnett was working from behind in the count and so on and so on...

One of the great things about baseball is the amount of data available, but it's a double-edged sword. It makes general questions like this one almost impossible to answer because of the endless number or variables. No two outings are exactly alike and something tells me that even if there was a parallel universe where two of the same games began at the same time, they would probably turn out completely differently anyway.

Tuesday, November 24, 2009

Parsing The AL MVP Vote

Good morning, Fackers. To the surprise of essentially no one who followed the 2009 Major League Baseball season, Joe Mauer won the A.L. MVP yesterday. He led the league in batting average, on base percentage, slugging percentage, runs created, wOBA, wRAA, VORP and would have in WAR if Ben Zobrist's UZR's wasn't propped up by small sample sizes or the stat gave any credit for a catcher's defense behind the plate.

With the exception of one writer from a Japanese newspaper based in Seattle (who inexplicably voted for Miguel Cabrera), Mauer was the unanimous choice. He didn't have 30 HR or 100 RBI, but the man from Minnesota was close on both counts. He didn't play in a game until May 1st, but at bat for at bat, he was the best hitter in the American League by a country mile.

While credit should go to the BBWAA for another award winner properly selected, the reality is that, even if you don't understand the concept of positional adjustment, there's no one else that had a legitimate case. And judging by the respective finishes of Derek Jeter and Mark Teixeira, it's apparent that many writers still don't grasp that concept.

Teixeira received 15 second place votes with only 3 voters ranking him lower than 4th. Nine voters ranked Jeter second, 16 others placed him between 3rd and 6th with the remaining three identifying him as the 8th, 9th or 10th most valuable player in the league. In other words, the general consensus was that Teix was a notch above Jeter.

According to Weighted On Base Average or wOBA, the statistic that most accurately measures a hitter's ability to get on base and hit for power, Teixeira (.402) led Jeter (.390) by fairly slim margin. However, wOBA doesn't take into account that Jeter played a much more difficult defensive position and, at least according to UZR, had a much better year in the field.

Perhaps UZR is selling Teixeira short, which most observers would argue is the case. Maybe Teixeira even saved Jeter a few errors by scooping balls in the dirt, although John Dewan's research doesn't seem to indicate that. But even if you grant both of those assumptions, it's unlikely they close the gap from Teixeira's 5.1 wins to Jeter's 7.4.

Of course, most voters don't care about players' wOBA or WAR. The biggest reason that Jeter finished lower than Teixeira on the majority of the ballots was that he only drove in 66 runs while Teix led the AL with 122. Runs Batted In are to the MVP vote what pitcher's wins are to the Cy Young: a context-driven, luck-determinant counting stat that depends largely on the production of one's teammates.

What was the biggest reason that Teixeira was able to drive in 122 runs despite a batting average (.264) and a slugging percentage (.471) with runners in scoring position well below his season marks (.292 & .565)? Derek Jeter and Johnny Damon's on base percentages of .406 and .365, respectively. The same thing happened to Mauer in 2006 when his .429 OBP teed up Justin Morneau for a huge amount of his 130 RBIs. That total was second in the league to David Ortiz and Morneau won the MVP award but Mauer ended up finishing 6th and was behind his teammate on every single ballot.

There were some other oddities within the voting aside from Miguel Cabrera getting a vote for first place but 3 votes for 10th and Teix topping Jeter. Mariano Rivera placed ahead of Zack Greinke although he didn't receive one Cy Young vote and Greinke won the award. Robinson Cano got three votes - all for 7th place. A-Rod netted a third place vote despite being left off 3/4 of the ballots all together.

A commenter over at BBTF took the liberty of compiling a "bizarro ballot", made up of actual selections writers submitted:
1. Miguel Cabrera
2. Kevin Youkilis
3. Alex Rodriguez
4. Jason Bay
5. Aaron Hill
6. Chone Figgins
7. Jason Kubel
8. Michael Cuddyer
9. Placido Polanco
10. Ian Kinsler
Sure, I'm nitpicking a little bit here. The writers have thus far got the 3 major award winners right, but with the exception of Tim Lincecum, they have been absolute no-brainers. When people have to list out 10 players, there are going to be some perceived sleights, but how many of those 10 actual placements do you think you could legitimately justify with statistical evidence?

Maybe we're not quite as far along the road to statistical enlightenment as we thought after Lincecum won the NL Cy Young. Perhaps, as Moshe Mandel from The Yankee Universe contends, we aren't seeing the voters wise up but the ability of sabermetricians (or at least those who are stat savvy) to influence the "buzz" surrounding players. And make no mistake, this is in large part due to the increasing influence of the internet which has given people like Rob Neyer and Joe Posnanski a national voice.

Is anyone going to remember who finished second or third in the voting when next year rolls around? Probably not. But that doesn't negate the fact that many voters (ostensibly "journalists") who have the privilege of voting for these awards are so severely lacking in objective analytical skills when that is one of the most important parts of their job description.

Friday, November 20, 2009

WHIP, FIP & The WAR Against Wins

Good morning, Fackers. In the wake of the senior circuit Cy Young, like the AL version, being awarded to a pitcher not on the basis of his won-lost record but on the quality and number of his innings pitched, we're again going to disagree with the well-respected Tyler Kepner.

As Matt pointed out on Wednesday, Kepner noted that Zack Greinke acknowledged FIP in his post-award conference call but was grasping at straws in an attempt to frame the knowledge of advanced statistics as a key component in Greinke's success. Last night, Kepner tried to connect what was said by Greinke (or more accurately, Brian Bannister) with Tim Lincecum's explanation of his approach and made the same conflation:
Obviously, there is no substitute for pure talent. But in Greinke, Bannister’s teammate, we are seeing what can happen when off-the-charts talent meets sophisticated understanding of numbers.

The same is true of Lincecum, to a degree. His stuff is filthy, but he said he was mainly concerned with how many walks-plus-hits he allows per inning – which was curious, in a way, because Dan Haren, Chris Carpenter and Javier Vazquez all had a better WHIP than Lincecum in the N.L. this season.
First, Lincecum's WHIP was a minuscule 1.05. Haren, Carpenter, and Vasquez? 1.00, 1.01, 1.03. That's a difference of one batter for every 20, 25 and 50 innings, respectively, which Lincecum easily erases with his superior strikeout ratio.

But more importantly, what if Lincecum had simply said that he was trying his best not to allow batters to reach base? It would have been dismissed as a typical cliche. On the offensive side of the ball, we frequently hear batters saying that they were "just trying to get on base" which is just a different way of saying that they were trying to improve their on base percentage.

In most cases, the goals in baseball are pretty obvious. If you are a batter, don't use up outs. If you are a pitcher, try to keep men off the basepaths, preferably via strikeout. There is still some wiggle room regarding the value of sacrifice hits and bunting, but there isn't a whole lot advanced statistics can teach players. Knowing about UZR isn't going to make someone a better defender. They already know they should be trying to field as many balls, as far away from them as possible.

The main function of the more advanced metrics that are steadily gaining in popularity such as WAR, wOBA, VORP, WPA, RE24, UZR, and FRAA is that they allow observers to more accurately compare players to one another. While players citing FIP and WHIP can only increase their popularity, which is certainly a positive thing, understanding them doesn't provide much in terms of strategical, on-field edge.

The real story emerging from the 2009 Cy Young voting is that the voters have begun to value better statistics and in turn, more objective analysis. Which is to say, they're not blindly picking the pitcher who had the most wins.

Kepner demonstrates this by comparing this year's voting to win-skewed results 1990 and 1998 (both of which illustrate the writer's old reliance on wins), but Dave Cameron over at FanGraphs sums it up best:
Congratulations to the members of the BBWAA, who have been willing to adapt as the game changes. They deserve recognition for being willing to accept the shift towards better analytical methods. And getting away from wins as a measure of the value of a pitcher is a big first step.
Of course, Adam Wainwright who left his last game of the season with a 6-1 lead in line for his 20th win still received the most first place votes in the NL. He only finished 10 points behind Lincecum, so maybe if his bullpen had held up, the aftermath of these awards would be slightly less celebratory.

Both Matt and I have taken turns raining on this parade, but I think it's likely that in hindsight, 2009 will be cited as the year that Advanced Stats won the war against Conventional Wisdom. However, I'm more inclined to think that this was the Battle of Saratoga. And considering Bill James penned the hardball version of the Declaration of Independence over 30 years ago, it's probably going to a while before we see any sort of Treaty of Paris.

Monday, October 26, 2009

Value Evaluation: CC or A-Rod

I doubt many people are going to remember who the ALCS MVP was a couple years from now. Postseason series MVPs are even more haphazardly given out than their regular season counterparts the BBWAA gets to vote on. As we had pretty much determined by Game 4, it was going to either A-Rod or CC Sabathia.

Last night as LoHud, Josh Thompson said "In no surprise to anyone, CC Sabathia was named MVP of the ALCS after winning both starts."

Um, I'll admit it. I'm a little surprised. A-Rod had a fantastic series and given how much the media loves stories of redemption, I had thought he would be the slight favorite to win. He put up huge numbers, and had clutch home runs, which I would think the media would value as much ro more than a guy who made two excellent starts.

Like everyone else, I don't really care who won the award, but I thought it would be interesting to look at who was more valuable in the series.

Here go the basic stats. A-Rod hit .429/.567/.952. That's a 1.519 OPS. He walked 8 times and struck out thrice. Three homers, six RBIs and six runs scored, meaning he was at the center of 9 of the Yanks' 33 runs in the series.

Sabathia started and won twice, going 8 innings and allowing one run each time. He had as many strikeouts (12) as walks (3) and hits (9) combined. An ERA of 1.12, a WHIP of 0.750.

Both guys put up numbers in the series that if you extrapolated to a full season would comfortably be the best of all time as a batter and pitcher respectively, I'm willing to say. I'm not sure of a place to get Wins or Runs Above Replacement data for a postseason series, so the best measure of comparing a pitcher to a position player would probably be WPA.

CC takes that one pretty handily which demonstrates the importance of an overpowering starting pitcher. With the Yankees' offense, he virtually assured them of two wins, the only two comfortable victories of the series.

A-Rod, on the other hand, had a hit in every game and was on base twice or more in all except Game 2, when of course he blasted a game-tying home run in the bottom of the 11th when the Yanks were down to their last breath. In Game 5, where he was at .005 in WPA, he still had a hit and two walks.

You can't go wrong with either of these guys obviously, and even A-Rod magnanimously said that CC deserved it. These are nice issues to be able to sort through, aren't they? Everybody wins!

Thursday, October 1, 2009

FanGraphs Salary Values

When trying to quantify a player's contributions to their team, sometimes Matt and I link to FanGraphs because in addition to Runs Above Replacement and Wins Above Replacement, they also translate that player's value into salary dollars. According to their calculations, here are the 15 most notable Yankees and their respective values:
Seems pretty high, doesn't it? That's not even all of them. There are part-time contributors like Chien Ming Wang (yes, he actually had positive value) and Hinske and Hairston and Bruney and Cervelli and Pena that add to that number incrementally as well.

The Yankees' Opening Day payroll was $201.5M, which is roughly what the top 10 most valuable guys on the team add up to (Damon and above on that list). Are the Yanks actually getting that much more than their money's worth?

Well, that depends on your view point.

As David Pinto of Baseball Musings points out, FanGraphs calculates player value based the value of marginal wins, and thereby attempts to valuate all players as if they were free agents. So, when you add up the value of all the batters and pitchers on FG, it comes out to $4.6B, whereas the total payroll of the MLB is roughly $2.7B.

With the obvious disclaimer that the folks behind FanGraphs are much smarter than I am, I would like to respectfully disagree with this methodology.

They use a system that corrects for the artificial forces depressing the salaries of players who are not available to the free market, which makes sense in it's own right. But we are all familiar with these artificial constraints and understand that is the reason why guys like Tim Lincecum are paid a fraction of what they are actually worth.

Instead of creating a system where the value of players is always going to far exceed the payroll, why not base it in reality? When I look at that dollar figure on FanGraphs, I want to know how much a player was actually worth in relation to what other players throughout the MLB are getting paid. I want to be able to tell who is getting their fair share or the pie. Part of that is the fact that guys like Phil Hughes are able to contribute at far beyond their pay grade but someone like CC Sabathia is unlikely to be worth the checks he's cashing, even during a very productive season.

I want to look at salary on a scale that is familiar to me, not one that is based on a contrived scenario in which everyone is a free agent and would make far more than they really do or even would make under those circumstances. It's not like the owners would suddenly shell out an extra $2.1B dollars if everyone hit the market over the next offseason.

Here is that list above, based on the MLB's actual payroll:
  • Derek Jeter -$19.5M
  • CC Sabathia - $16.3M
  • Mark Teixeria - $14.2
  • Robinson Cano - $11.3M
  • A-Rod - $11.2M
  • Jorge Posada - $10.6M
  • Nick Swisher - $10M
  • Andy Pettitte - $9.4
  • A.J. Burnett - $ 8.3
  • Johnny Damon - $7.1M
  • Hideki Matsui - $6.5M
  • Phil Hughes - $6M
  • Mariano Rivera - $5.2M
  • Brett Gardner - $4.9M
  • Joba Chamberlain $4.2M
  • Melky Cabrera $3.9M
  • Alfredo Aceves $3.5M
  • Phil Coke - $59K

  • Total: $152.2M
Seems like a better approximation of their values. At least to me it does.

Friday, September 4, 2009

Delicious

Because everyone loves pie, here is a graphical representation of the Yankees Wins Above Replacement via UmpBump. Dig in!

Quick thoughts:
  • Interesting that Phil Hughes checks in slightly higher than Joba Chamberlain (and Mariano Rivera), isn't it?

  • I didn't think CC would be more valuable than Teixeira, did you?

  • Even if you adjust for the time that A-Rod missed to begin the season, he still wouldn't be as productive as Jeter.

  • If you project How is Chien Ming Wang OVER replacement?

Friday, August 28, 2009

A Few Lines On Coke

I have to admit, it wasn't until all the Twitter feeds came pouring into yesterday's live chat that I realized just how bad Phil Coke has been of late: 17 ER over his last 15.1 IP. His ERA has ballooned from 2.97 to 5.05 in that stretch.

Some may want to blame this on his over-use, as his 59 appearances are the most on the team by a good margin (Rivera is next with 52), and his 51.2 IP in relief is third behind Alfredo Aceves (61) and Rivera (53). And maybe that has something to do with it, but not very much.

Most of what's happened to Coke over these last 20 outings is just good old statistical correction. As I've stated before, particularly when evaluating relievers, I prefer to look at Fielding Independent Pitching (FIP) rather than ERA. As you may recall, FIP is dependent upon the three things pitchers can control: strikeouts, walks, and home runs, and is then is adjusted to an ERA-like scale.

Unfortunately, I can't find game-by-game logs of Coke's FIP, nor do I have the free time to calculate it at the moment. However, at every point this season that I looked at his numbers, his ERA was outperforming his FIP - by a lot. These past twenty outings have served to correct that gap, so much so that after yesterday, Coke's ERA (5.05) is now worse than his FIP (4.86).

So what's the good news/bad news here? The good news is that the numbers seem to indicate that Coke has at worst leveled off, at best is due for a small improvement. The bad news is that where those numbers stand right are not all that good. The good news is the Coke's strikeout (7.14 per 9) and walk (3.14 per 9) rates are slightly better than the league averages. The bad news is his home run rate (1.57 per 9) is a half home run worse than the league average, and that's what is killing both his FIP and his ERA.

Another number to look at is Coke's batting average on balls in play (BABIP). Evidence has shown that a pitcher has little control over balls in play (which is why FIP is a valuable statistic) and that most pitchers end up having a BABIP close to the league average. Currently, Coke is at .234, far better than the league average of .304. This could suggest that we haven't seen the bottom for Coke yet; if his BABIP regresses to the mean over the final five weeks he could be in for a few more rough outings.

However, given Coke's K and HR rates, I would expect his BABIP to be a bit low. His K rate is higher than the league average, meaning there are fewer balls in play against him. Furthermore, 6.1% of batted balls off Coke are home runs, as opposed to 4.0% for the league. As such, hitters are making contact less against him compared to the league, but when they do, the ball is traveling over the fence far more often. While the latter certainly isn't a good thing, I do think that indicates that Coke's BABIP likely won't get close to the league average by season's end. Besides, any BABIP regression to the mean may be a good thing for Coke, as it could indicate that his gigantic HR rate is dropping off a bit.

So what does it all mean? Phil Coke isn't as good as he appeared to be through most of the summer and isn't as bad as he appears to be right now. He gives up way too many home runs, and Yankee Stadium may have something to do with that (18.2 AB/HR at home, 25 on the road). If he can get his home runs down, his K and BB rates suggest he is a relatively effective pitcher.

Relievers are highly volatile due to the relatively small number of innings the pitch. One or two bad outings, like yesterday or Coke's 0.1 IP 6 ER disaster in Chicago on 8/1, can have a major impact on a reliever's statistics. Despite being a former starter with a decent arsenal of pitches, and despite his numbers being good against right handed batters for most of the season, Joe Girardi has insisted upon using Coke as a match-up lefty for most of the year, with 30 of his 59 appearances lasting less than an inning.

Coke has been pretty high up in the bullpen pecking order for most of the season. He'll need to show some improvement over these last five weeks to justify keeping that status in October. If he can keep the ball in the ballpark more often, he has a good chance at making those improvements.

Thursday, August 27, 2009

When Is A Slump A Slump?

Good morning, Fackers. Yes, I just made a Geology picture "joke". Doc Nardacci would be so proud... Now get ready for some logarithms! WAKE UP!

Over at the Freakonomics blog at the NYT, they used some statistics to propose a more solid definition for what actually consititues a slump (or any streak with an absence of a certain event) using A-Rod as an example (h/t BBTF):
It occurred to me that it would be pretty easy to derive a statistical standard for determining when an athlete was having a “statistically significant slump.” For example, Alex Rodriguez recently went through a homerless drought of 72 at-bats. Over his career, A-Rod has averaged one homer for every 14.2 at bats — suggesting there is about a 93 percent chance that he will not homer on any individual at bat. It would be crazy to say that he was in a home-run slump after failing to homer after just a few at bats. But the question is how many homer-less at bats is enough to be a statistically significant drought?

The answer is 42. There is less than a 5 percent chance that Rodriguez would go homerless 42 times in a row — so we can reject the hypothesis (at a 5 percent level of statistical significance) that he is going homer-less merely as a matter of chance.
They are essentially drawing the line at a 95% confindence interval (2 standard deviations), but you can set your own parameters by altering the simple formula:
Total consecutive number of bad events > log(.05)/log(probability of single bad event)
It's a little more difficult because you have to play around with it to find the right number, but you can also figure out what the likelihood of A-Rod going on a 72 at bat homerless streak (beginning in his next at bat) would be. It's about one half of one percent.

Using this method, we can determining the (im)probability that Derek Jeter would go 113 plate appearances without working a walk like he did from July 28th to August 25th. In 9656 career PAs, Jeter has walked 863 times, giving him a walk rate about approximately 8.9%. This makes the odds of him going that long without a base on balls 0.0025% or 1 in 4,000.

Fun stuff, huh? No? Well at least it gives you a way, numerically, to prove that Tim McCarver is an idiot. You're welcome.

Friday, July 24, 2009

Risky Robby

Leading off the bottom of the 5th inning last night, Robinson Cano slashed at the first pitch he saw from Vin Mazzaro, sending it bending down the left field line and into the corner. It hit off the wall, just inside of the line and as the ball bounced directly to Matt Holliday and the cameras panned back to Cano, he was just rounding first, and I knew he was dead to rights at second.

It was part of that innate feel you develop as a fan. If you know who's at the plate, you've got a pretty good sense of how long the ball needs to rattle around before an outfielder gets his hands on it for your guy to get to second or third. You know when a balls rolls up to the wall and Jeter is running, he's thinking about a triple (and so are you). With that clean carom, it would have been a tight play even for Brett Gardner.

Replays showed that Cano paused for a second to see if the ball was fair or foul off the bat (understandable), and started down the first kind of slowly, which would have been fine if he was going to settle for a single. But inexplicably, as he was nearing first base, he broke into a full sprint, only to be gunned down at second about literally two full strides at the least. In the second picture on the right, it looks as if first base coach Mick Kelleher yelling at him to stop. That's probably because Holliday had already released the ball and Cano could have turned back.

As you can see, he was at barely in the fame when the ball arrived to Mark Ellis at 2nd base, and was out by 3 full strides.

This is an isolated incident, and it might seem like I'm dwelling on a baserunning mistake for way too long, but I think lends some insight into his lack of discipline at the plate and in turn his inability to hit with runners in scoring position.

His baserunning mistakes (he's 16 for 34 in SB in his career) say more about his level of risk aversion than his speed on the basepaths. The same can be said for someone with a lack of discipline at the plate. They are willing to swing at pitches that are harder to hit, thus increasing the likelihood of failure.

Cano has the 13th highest swing percentage in the Majors with 51.7%, but makes contact 91.2% of the time, which is 7th highest, where the leaders in that category (Luis Castillo, Marco Scutaro) have some of the lowest swing percentages in the game. Think about how good you have to be at putting the bat on the ball to rank so highly in both of those categories. He makes contact with 79.5% of the pitches he swings at out of the zone.

My contention is that his lack of plate discipline is what is eating away at his production with runners in scoring position and men on base in general. He's at .205 w/RISP this year as opposed to .356 without. For his career, he has a .743 OPS men on base as opposed to .895 with the bags unoccupied, and the former includes 14 intentional walks.

My half-baked theory goes like this: Pitchers are more reluctant to give a batter a pitch to hit with men on base, especially early in the count, but Cano goes up swinging like he always does, and puts the bat on the ball, but makes poorer contact as a result. His BABIP bears this out, as it is .297 with men on as opposed to .338 with the bases empty. We usually cite BABIP as a statistic to explain away fluky performances, but this is over his entire career, 2765 plate appearances. There are no more flukes at that sample size.

It's easy to imagine how good Cano could be if he was just more selective at the plate, but as Bill James has pointed out, it's not easy for a hitter to change his approach:
I think it is easier to learn plate discipline than it is to learn speed or to develop a strong throwing arm—but not much easier. A player who lacks plate discipline at age 18 will usually lack plate discipline at age 30. But not always; some players can adapt well to the challenge of learning to lay off certain pitches.
Robinson Cano is an excellent player as he is. Don't get me wrong. He's probably my favorite Yankees' position player and I love watching him take a ball 6 inches off the plate into the home bullpen as much as the next guy. But he will never be a scary, middle of the order type presence unless he can be more selective and make pitchers throw him balls in the strike zone with runners on base.