Tuesday, February 3, 2009

The age-old NY subway question--unlimited or pay-per-ride?

Before you get onto the New York City subway these days, you have to purchase a metro card. For us daily commuters, the choice would appear to be obvious--purchase an unlimited card. With it, you can get on and off the subway as many times as you like within some period of time. Surely, the MTA prices it to make it worth the money.

However, when I do the calculation for my own behavior, I never seem to get my money's worth from the unlimited. This is because the unlimited card price is always higher than the cost of buying a per ride card if you are only using the card to commute to work during the week. Even if you are using it for one round trip on the weekend, you would still pay less by buying a "pay-per-ride" card unless you buy a 30-day card.

The following table shows the cost of each unlimited ride card, followed by the amount of trips that could be purchased for that same amount. Because the MTA gives a 15% bonus for all "par per ride" purchases over $7, the nominal value in the table shows the value that will be shown on your metro card if you purchase a pay per ride card.
Unlimited Days Unlimited Cost Nominal Value if purchased as a "pay per ride" card Trips if purchased per trip Trips used if going to and from work only, 5 days a week Trips lost if buying unlimited only for work versus purchasing regular card Also one weekend trip each week Trips lost if buying unlimited versus purchasing regular card with 1 weekly fun trip
1 $7.50 $8.63 4.3 na na na
7 $25.00 $28.75 14.4 10 4.4 12 2.4
14 $47.00 $54.05 27.0 20 7.0 24 3.0
30 $81.00 $93.15 46.6 44 2.6 52 -5.4

The first line shows the one-day card, which can be purchased for $7.50. You can use that same $7.50 to purchase $8.63 in value instead, which will be good for 4 trips plus $0.63. Thus, you'd only want to get an unlimited one-day ride if you were making at least 2 round trips.

The 7 day unlimited costs $25. If you use that same $25 to instead purchase a pay-per-ride card, you get $28.75 of value, entitling you to 14 trips (plus $0.75 additional of stored value). If you go to work every weekday during the 7 day period, you'd use just 10 trips (5 round trips). If you also use the card for one round trip during the weekend, you are up to 12 trips, still 2.4 trips short of what you could have purchased with the $25 for a pay-per-ride.

As the table shows, you are always better off purchasing pay-per-ride cards instead of unlimited cards if you are just using your metro card for commuting. Even if you take one round trip in addition to work every week, only the 30-day unlimited would be worth it, and this only if you go to work every weekday during the period and use the card once each weekend. Many people work at home from time to time and there is typically a federal holiday each month, so the 30-day figures are optimistic.

The other issue with the unlimited is a psychological one: I get upset if I forget my unlimited card or end up not taking the subway a couple days when I could have used the card. With the pay-per-ride, you only pay for what you use. Perhaps more annoyingly, the pay-per-ride cards display the amount left each time you enter the subway, but the unlimited cards do not tell you the number of days left on your card when you enter the subway, and thus, if you don't keep track it yourself, you will be jammed in the legs with a locked turnstile at least once a month when you purchase an unlimited card.

I realize there are some who not only commute to work but also very frequently take subway trips to go out or run errands. For those, the unlimited cards may be worth it. For others, stick with the pay-per-ride.

Tuesday, January 20, 2009

Nutty about Peanuts.

Visiting South Carolina this weekend, I picked up an old Southern favorite from Publix: peanut butter cookies (I didn't see my real favorite: boiled peanuts). No sooner had I returned home than my Mom admonished me for buying them, because they were unsafe, and possibly tainted with Salmonella.

Sure enough, there was an article in The State confirming the outbreak. So far, around 500 people around the country have been sickened (and possibly 6 deaths) from what is believed to be contaminated peanuts. USA Today confirms the continuing "epidemic" today. While these figures seem high, 500 people sickened with food poisoning in a period of four months, across the entire U.S., is hardly a risk worth mentioning. According to wrongdiagnosis.com, the number of incidents of food poisoning or sickness is 200,000 a day. OK, you might say, but Salmonella is pretty serious and if you don't take antibiotics you might be laid up for several days. Fine, but the same site says that there are about 1.4 million cases of Salmonella annually, or about 3,835 a day (the CDC says about 40,000 cases are reported annually, but that there are many more unreported).

So why are we getting exercised about a mere 4 cases a day, as with the current outbreak? My best answer is that 1) it makes for interesting news, 2) any problem that affects so broadly a population, even with minuscule or infinitesimal risk, is seen by reporters as being important, and 3) people cannot easily assess their relative risk.

As for me, I explained to my Mom that I'm not too concerned, and quickly had a peanut butter cookie before she could run back to the store. After waiting a day or so to make sure I was Salmonella-free, the rest of the family followed. ;-)

Friday, December 5, 2008

Are we entering "unprecedented" territory?

(click graph for greater resolution)

The reports of gloom and doom are abounding, and, I must admit, I believe most of them.

I am going to focus purely on the stock market, because the data is readily available, and because I believe the broader economic problems are only just beginning. My March blog pointed out that in the stocks versus bonds 20-year view, stocks almost always won, but the results are much more mixed over shorter periods. I also need to point out that I overestimated the results for stocks by assuming dividends were not included in the indices. For the Dow indices, the subject of much of that discussion, dividends are included (see the Dow Jones site), so the graphs in that blog are correct, but the numbers should not be adjusted further for dividends, meaning that stocks' edge over bonds is less impressive.

Today's post, though, is really about the graph above, showing 1 ,10, and 20 year returns on the Dow since 1928 (from December to December). From December 3, 2007 through December 1, 2008, the Dow lost 37% of its value. This horrible run is beaten only once, from December 1930 to December 1931, when the Dow lost 53%. The years 1930, 1937, and 1974 (again, December to December) were the only other years where the 12 month loss was more than 20%.

Thus, historically, though not unprecedented, the yearly drop in the Dow is, well, statistically "improbable" (that is, if you base your probabilities only on history). While the 10 and 20 year numbers are much more in line with history, they are still on the low end of the distribution. The last time the 10-year change was negative, as it is now, was 30 years ago, in 1978, in the waning years of very tough economic times.

The next few months will start to indicate how deep an economic hole we've dug for ourselves, but the stock market numbers are not encouraging, and the extent to which the economy is dependent on the market (in the sense that assets are tied to it) seems much more like the 20s and 30s than like the 70s. Let's hope I'm wrong.

Friday, October 31, 2008

Election Prediction Explained

So here's the explanation.

I am following 3 major websites now:

www.electoral-vote.com – This consolidates polls by state to predict the count. Electoral-vote apparently uses simple averaging to consolidate its data. I prefer this method because it requires little interpretation on their part. Interpretation involves assumptions about bias in the polls, and I believe it is hard to figure out the exact impact of the bias or even the direction. Electoral-vote has Obama at 364. At this time in 2004, they had Kerry at 283 (see this page), whereas his election day total was 252, with the main difference being Florida. More telling, the "strong" Obama States total 264 votes, as opposed to 95 for Kerry at this point.

www.fivethirtyeight.com – This consolidates polls by state to predict the count using some complex weighting system. It’s a neat idea but it’s end result is about the same as averaging, and I am not at all convinced it is better.

They’ve got Obama at 346.5, much more than the 270 needed to win.


www.gallup.com – This well-established survey company is different from the two above in that they actually conduct the polls. Gallup is showing primarily national results, and has Obama significantly up, both in raw percentages and when adjusting for “likely” voters—people Gallup has determined are likely to vote, based on two different models. Gallup's daily tracking polls has Obama's lead almost unchanged since the start of October (never more than the statistical error).

My conclusion from the above---Obama will be the next US President.

So why the change from before, when I said polls are difficult to trust and spoke of biases?

Three reasons:

1) the closer we get to the election, the better correlation between intentions and actions

2) the closer we get to the election, the fewer undecided voters. A recent Reuters poll shows this at about 2%. Even if it is 5% and the undecided break 4 to 1 for McCain, he's going to lose.

3) The biases appear to lean in Obama’s favor: more younger voters likely and more early voters. Very biased reporting from Grandma in S.C. says that lots of young people were out voting early (she spent 2 hours on line to vote early, by the way).

Thursday, October 30, 2008

Election Prediction

Ok, sure I waited until nearly the end, but here's my prediction:
Obama wins, with 401 Electoral votes.

I'll explain why tomorrow.

Wednesday, October 8, 2008

Election Polls

A short note about election polls, which I've been following somewhat religiously for the last few weeks.

Election polls differ in at least four significant ways from actual voting.

First, polls are typically of around 1,000 people or less, which means that at best, they are statistically precise to within plus or minus three percent. This means that a six point difference between 2 candidates may be nothing more than sampling error (i.e., a statistical anomaly).

Second, polls tend to be of the general population and not of likely Electoral College votes, which is how the election is counted (but see electoral-vote.com for a count of Electoral votes, according to polls). As we know from recent elections, the Electoral vote percentages frequently (and seemingly increasingly) do not correspond to popular vote percentages.

Third, polls are snapshots on how people feel on a certain day. Americans seem to be particularly fickle in their opinions recently, perhaps due to the economic turmoil, so don't trust that today's lead won't disappear tomorrow.

Finally, many polls do not remove unlikely voters (though you do see some figures concerning "likely voters"). Polls of people who do not vote are fairly useless, but pollster's haven't been very successful in predicting who will actually vote. Thus, the tendency is to include respondents who are registered and say they plan to vote, without looking at their demographics to see what they've done in the past.

For all these reasons, if you're an Obama supporter, you should be worried and if you're a McCain supporter, you should have some hope. Either way, vote!

Thursday, August 28, 2008

The Atlantic Monthly is criminally misusing statistics

I spent the last week vacationing in South Carolina, where my parent's house seems to have Atlantic Monthly's and Harper's from the dawn of time. What luck, then, that one of the most interesting articles (at least statistically) was in an issue as recent as the July/August 2008 issue of the Atlantic. The article is called "American Murder Mystery" and it's by Hanna Rosin.

The article talks of the recent increase in violent crime in mid-sized cities. In many of these cities, government housing projects (called "Section 8" housing) have been torn down. In their place, the government has provided the poor with rent subsidies so that they can move to private housing. Rosin describes how Phyllis Betts and Richard Janikowski, of the University of Memphis, tie the increase in crime in these cities to the destruction of these projects. A striking quote in the article is from the Memphis police chief: '“It used to be the criminal element was more confined,” said Larry Godwin, the police chief. “Now it’s all spread out."'

The primary statistical evidence given in the article of an association between crime and former Section 8 residents, is a map that shows areas with high incidents of crime correspond to areas with a large number of people with Section 8 subsidies (i.e., former residents of housing projects). As convincing as this might sound, it has a fatal flaw: the map looks at total incidents rather than crime rate. This means that an area with 10,000 people and 100 crimes (and 100 Section 8 subsidy recipients) will look much worse than an area with 100 people and 1 crime (and 1 Section 8 subsidy recipient). However, both areas have the same rate of crime, and, presumably, the same odds of being a victim of crime (see my earlier blog about the safest place to live for some explanation of the use of rates in measuring crime). Yet in Betts and Janikowski's analysis, the area with 10,000 people has a higher number Section 8 subsidy recipients and higher crime, thus "proving" their theory of association.

Of course, there will be both a greater number of Section 8 subsidy recipients and a greater number of crimes in the area with 10,000 people than in the area with 100 people . Thus, while the map presented in the Atlantic article does indeed seem to indicate that there is higher crime in areas where there are more Section 8 subsidies, this differential might be entirely an artifact of population density, and, in fact, the crime rate may be completely unrelated to where Section 8 subsidy recipients reside. Without an adjustment for population density, the inferences made from the association are statistically meaningless.