Monday, March 23, 2009
How big a sample?
The fact is that you would use the same size sample, in almost all practical circumstances, for the two populations above. Ditto if the population were 100,000,000 or 1,000.
The reasons for this, and the concept behind it, go back to the early part of the 20th century when modern experimental methods were developed by (Sir) Ronald A. Fisher. Though Wikipedia correctly sites Fisher in its entry on experimental design, the seminal book, Design of Experiments, is out of stock at Amazon (for $157.50, you can get a re-print of this and two other texts together in a single book). Luckily, for a mere $15.30, you can get David Salsburg's (no relation and he spells my name wrong! ;-) ) A Lady Tasting Tea, which talks about Fisher's work. Maybe this is why no one knows this important fact about sample size--because we statisticians have bought up all the books that you would otherwise be breaking down the doors (or clogging the internet) to buy. Fisher developed the idea of using randomization to create a mathematical and probability framework around making inferences of data. In English? He figured out a great way to do experiments, and this idea, or randomization, is what allows us to make statistical inferences about all sorts of things (and the lack of randomization is what sometimes makes it very difficult to prove otherwise obvious things).
Why doesn't (population) size matter?
To answer this question, we have to use the concept of randomization, as developed by Fisher. First, let's think about the million loans we want to know about at BIGbank. Each of them is no doubt very different, and we could probably group them into thousands of different categories. Yet, let's ignore that and just look at the two categories we care about: 1) good loan or 2) bad loan. Now, with enough time studying a given loan, suppose we can reasonably make a determination about which category it falls into. Thus, if we had enough time, we could look at the million loans and figure out that G% are good and B% (100% - G%) are bad.
Now suppose that we took BIGbank's loan database (ok, we need to assume they know who they loaned money to), and randomly sampled 100 loans from it. Now, stop for a second. Take a deep breath. You have just entered probability bliss -- all with that one word, randomly. The beauty to what we've just done is that we've taken a million disparate loans and with them, formed a set of 100 "good"s and "bad"s, that are identical in their probability distribution. This means that each of the 100 sampled loans that we are about to draw has exactly a G% chance of being a good one and a B% chance of being a bad one, corresponding to the actual proportions in the population of 1,000,000.
If this makes sense so far, skip this paragraph. Otherwise, envision the million loans as quarters lying on a football field. Quarters heads up denote good loans and quarters tails up denote bad loans. We randomly select a single coin. What chance does it have of being heads up? G%, of course, because exactly G% of the million are heads up and we had an equal chance of selecting each one.
Now, once we actually select (and look at) one of the coins, the chances for the second selection change slightly, because where we had G% exactly, now there is one less quarter to choose from, so we have to adjust accordingly. However, that adjustment is very slight. Suppose, G were 90%. Then, we'd have, for the second selection, if the first were a good coin, a 899999/999999 chance of selecting another good one (that's an 89.99999% chance instead of a 90% chance). For smallbank, we'd be looking at a whopping reduction to an 89.9999% chance from a 90% chance. This gives an inkling of why population size, as long as it is much bigger than sample size, doesn't much matter.
So, now we have a sample set of 100 loans. We find that 80 are good and 20 are bad. Right off, we know that, whether dealing with the 100,000 population or the 1,000,000 population, that our best guess for the percentage of good loans, G, is 80%. That is because of how we selected our sample. It doesn't matter one bit how different the loans are. They are just quarters on a football field. It follows from the fact that we selected them randomly.
We also can calculate several other facts, based on this sample. For example, if the actual number of good loans were 90% (900,000 out of 1,000,000), we'd get 80 or fewer in our sample of 100 only 0.1977% of the time. The corresponding figure, if we had sampled from the population of 100,000 (and had 90,000 good loans), would be 0.1968%. What does this lead us to conclude? Very likely, the proportion of "good" loans is less than 90%. We can continue to do this calculation for different possible values of G:
If G were 89%: .586% of the time would you get 80 or fewer.
If G were 88%: 1.47% of the time would you get 80 or fewer.
If G were 87%: 3.12% of the time would you get 80 or fewer.
If G were 86.3%: 5.0% of the time would you get 80 or fewer.
If G were 86%: 6.14% of the time would you get 80 or fewer.
In each of the above cases, the difference between a population of 1,000,000 and 100,000 loans makes a difference only at the second decimal place, if that.
Such a process allows us to create something called a confidence interval. A confidence interval kind of turns this calculation on its head and says, "Hey, if we only get 80 or fewer in a sample 1.47% of the time when the population is 88% good, and I got only 80 good loans in my sample, it doesn't sound too likely that the population is 88% good." The question then becomes, at what percentage would you start to worry?
For absolutely no reason at all (and I mean that), people seem to like to limit this percent to 5%. Thus, in the example above, most would allow that, if we estimated G such that 5% (or more) of the time, 80 or fewer of 100 loans would be good (where 80 is the number of good in our sample), then they would feel comfortable. Thus, for the above, we would say, with "95% confidence, 86.3% or fewer of the loans in the population are good." If we also want a lower bound on the percent of G loans, we could calculate the percent of G such that there is a 5% chance that 80 or more loans in a sample of 100 would be good. This percentage is 72.3%, and we could say that "with 95% confidence, 72.3% or more of the loans in the population are good." We can combine these two 95% confidence intervals into a 90% confidence interval, since the percentages not included, of 5% in each of the two intervals add to 10%. We can thus say: "with 90% confidence, between 72.3% and 86.3% of the loans in the population are good." We can calculate the highest and lowest percent of good loans we estimate there to be in the population, with any level of confidence between 0 and 100%. We could state the above in terms of 99% confidence or in terms of 50% confidence. The higher the confidence, the wider the interval and the lower the confidence the narrower the interval.
Back to sample size versus population. As stated above, the population size, though 10 times bigger, doesn't makes a difference. For a given probability above, we are using the hypergeometric distribution to calculate the exact figure (the mathematics behind it are discussed some in my earlier post).
Here are some of the chances associated with a G of 85% and a sample size of 100 that yields 80 good loans or fewer.
Population infinite : 10.65443%
Population 1,000,000: 10.65331%
Population 100,000 : 10.64%
Population 10,000 : 10.54%
Population 1,000 : 9.49%
Population 500 : 8.21%
This example follows the rule of thumb: you can ignore the population size unless the sample is at least 10% of the population.
A note to the commenter regarding population size:
Anonymous, The reason that population size does barely matters is because the statistical inferences are based on the random behavior of the sample and this behavior does not depend on the population size. Suppose you randomly selected 20 people for a survey regarding preference for a black and blue dress versus a white and gold dress and all 20 preferred black and blue. Whether I told you these people were randomly selected from the state of NY or from the whole US, in either case, you would think that the preferences (state or national) clearly favor the black and blue. That intuitive feeling is because the inference you make in your mind is regarding the behavior of the sample. Statistics takes it a bit further and figures out, if the selection is truly random, how likely such an outcome would be under different scenarios. However, the key is the observed randomness in the sample and its size and not the population size. In other words, the sample size IS important because that is where you make your observations but the population is not as long as the sample is representative of it (and it will be, as long as the sample is random).
Wednesday, March 18, 2009
7 letter scrabble word redux
This, of course, is not correct, as was pointed out in comments to the Numbers Guy's blog (this was also pointed out by my sister). All you need to do to use all your letters with a 7-letter word is find a place to connect that is parallel to an existing word. For example, 'weather' could be connected parallel to a word ending in 'E", since 'we' is a word.
Maybe that's why my sister won so many scrabble games against me when I was a kid.
Wednesday, March 11, 2009
Are same-sex classes better?
In particular, the article states: "Students of both sexes in the co-ed fifth grade did better on last year’s state tests in math and English than their counterparts in the single-sex rooms, and this year’s co-ed class had the highest percentage of students passing the state social studies exam."
In other words, the City is continuing this program, even though the evidence indicates that not only are students in same-sex classes doing no better, they are doing worse! The principal, who has introduced some programs that have achieved material results, said: "“We will do whatever works, however we can get there...we thought this would be another tool to try.” This seems reasonable, but the article states,"...unlike other programs aimed at improving student performance, there is no extra cost." There may not be a monetary cost, but making these students laboratory rats in someone's education research project doesn't help them, and, apparently in this case, hurts them. Not to mention the opportunity cost of not exposing these children to other programs that might actually help.
To be fair, the scholarly literature is not consistent in its conclusions about whether same-sex classes improve achievement. However, many of the U.S. studies showed little or no improvement. See, for example:
Singh and Vaught's study
LePore and Warren
On the other hand, some English and Australian studies indicate that, at least for girls, same-sex classes or schools may result in higher achievement (see, for example, Gillibrand E.; Robinson P.; Brawn R.; Osborn A.) while others indicate that there are no differences (see Harker).
So the literature seems to be mixed, and I would imagine there are numerous confounding factors that make this something hard to measure--for example, typical single-sex classes in New York City consist of low-income minority students, where the boys are seen as being at-risk more than the girls. Contrast with the British and other foreign studies, where the girls are the greater concern for under-achievement.
Despite this, it's questionable how long it is ethical to continue a program, like the one at P.S. 140, where the current known outcome is that boys and girls are doing worse in same-sex classes.
Tuesday, February 3, 2009
The age-old NY subway question--unlimited or pay-per-ride?
However, when I do the calculation for my own behavior, I never seem to get my money's worth from the unlimited. This is because the unlimited card price is always higher than the cost of buying a per ride card if you are only using the card to commute to work during the week. Even if you are using it for one round trip on the weekend, you would still pay less by buying a "pay-per-ride" card unless you buy a 30-day card.
The following table shows the cost of each unlimited ride card, followed by the amount of trips that could be purchased for that same amount. Because the MTA gives a 15% bonus for all "par per ride" purchases over $7, the nominal value in the table shows the value that will be shown on your metro card if you purchase a pay per ride card.
| Unlimited Days | Unlimited Cost | Nominal Value if purchased as a "pay per ride" card | Trips if purchased per trip | Trips used if going to and from work only, 5 days a week | Trips lost if buying unlimited only for work versus purchasing regular card | Also one weekend trip each week | Trips lost if buying unlimited versus purchasing regular card with 1 weekly fun trip |
| 1 | $7.50 | $8.63 | 4.3 | na | na | na | |
| 7 | $25.00 | $28.75 | 14.4 | 10 | 4.4 | 12 | 2.4 |
| 14 | $47.00 | $54.05 | 27.0 | 20 | 7.0 | 24 | 3.0 |
| 30 | $81.00 | $93.15 | 46.6 | 44 | 2.6 | 52 | -5.4 |
The first line shows the one-day card, which can be purchased for $7.50. You can use that same $7.50 to purchase $8.63 in value instead, which will be good for 4 trips plus $0.63. Thus, you'd only want to get an unlimited one-day ride if you were making at least 2 round trips.
The 7 day unlimited costs $25. If you use that same $25 to instead purchase a pay-per-ride card, you get $28.75 of value, entitling you to 14 trips (plus $0.75 additional of stored value). If you go to work every weekday during the 7 day period, you'd use just 10 trips (5 round trips). If you also use the card for one round trip during the weekend, you are up to 12 trips, still 2.4 trips short of what you could have purchased with the $25 for a pay-per-ride.
As the table shows, you are always better off purchasing pay-per-ride cards instead of unlimited cards if you are just using your metro card for commuting. Even if you take one round trip in addition to work every week, only the 30-day unlimited would be worth it, and this only if you go to work every weekday during the period and use the card once each weekend. Many people work at home from time to time and there is typically a federal holiday each month, so the 30-day figures are optimistic.
The other issue with the unlimited is a psychological one: I get upset if I forget my unlimited card or end up not taking the subway a couple days when I could have used the card. With the pay-per-ride, you only pay for what you use. Perhaps more annoyingly, the pay-per-ride cards display the amount left each time you enter the subway, but the unlimited cards do not tell you the number of days left on your card when you enter the subway, and thus, if you don't keep track it yourself, you will be jammed in the legs with a locked turnstile at least once a month when you purchase an unlimited card.
I realize there are some who not only commute to work but also very frequently take subway trips to go out or run errands. For those, the unlimited cards may be worth it. For others, stick with the pay-per-ride.
Tuesday, January 20, 2009
Nutty about Peanuts.
Sure enough, there was an article in The State confirming the outbreak. So far, around 500 people around the country have been sickened (and possibly 6 deaths) from what is believed to be contaminated peanuts. USA Today confirms the continuing "epidemic" today. While these figures seem high, 500 people sickened with food poisoning in a period of four months, across the entire U.S., is hardly a risk worth mentioning. According to wrongdiagnosis.com, the number of incidents of food poisoning or sickness is 200,000 a day. OK, you might say, but Salmonella is pretty serious and if you don't take antibiotics you might be laid up for several days. Fine, but the same site says that there are about 1.4 million cases of Salmonella annually, or about 3,835 a day (the CDC says about 40,000 cases are reported annually, but that there are many more unreported).
So why are we getting exercised about a mere 4 cases a day, as with the current outbreak? My best answer is that 1) it makes for interesting news, 2) any problem that affects so broadly a population, even with minuscule or infinitesimal risk, is seen by reporters as being important, and 3) people cannot easily assess their relative risk.
As for me, I explained to my Mom that I'm not too concerned, and quickly had a peanut butter cookie before she could run back to the store. After waiting a day or so to make sure I was Salmonella-free, the rest of the family followed. ;-)
Friday, December 5, 2008
Are we entering "unprecedented" territory?
The reports of gloom and doom are abounding, and, I must admit, I believe most of them.
I am going to focus purely on the stock market, because the data is readily available, and because I believe the broader economic problems are only just beginning. My March blog pointed out that in the stocks versus bonds 20-year view, stocks almost always won, but the results are much more mixed over shorter periods. I also need to point out that I overestimated the results for stocks by assuming dividends were not included in the indices. For the Dow indices, the subject of much of that discussion, dividends are included (see the Dow Jones site), so the graphs in that blog are correct, but the numbers should not be adjusted further for dividends, meaning that stocks' edge over bonds is less impressive.
Today's post, though, is really about the graph above, showing 1 ,10, and 20 year returns on the Dow since 1928 (from December to December). From December 3, 2007 through December 1, 2008, the Dow lost 37% of its value. This horrible run is beaten only once, from December 1930 to December 1931, when the Dow lost 53%. The years 1930, 1937, and 1974 (again, December to December) were the only other years where the 12 month loss was more than 20%.
Thus, historically, though not unprecedented, the yearly drop in the Dow is, well, statistically "improbable" (that is, if you base your probabilities only on history). While the 10 and 20 year numbers are much more in line with history, they are still on the low end of the distribution. The last time the 10-year change was negative, as it is now, was 30 years ago, in 1978, in the waning years of very tough economic times.
The next few months will start to indicate how deep an economic hole we've dug for ourselves, but the stock market numbers are not encouraging, and the extent to which the economy is dependent on the market (in the sense that assets are tied to it) seems much more like the 20s and 30s than like the 70s. Let's hope I'm wrong.
Friday, October 31, 2008
Election Prediction Explained
So here's the explanation.
I am following 3 major websites now:
www.electoral-vote.com – This consolidates polls by state to predict the count. Electoral-vote apparently uses simple averaging to consolidate its data. I prefer this method because it requires little interpretation on their part. Interpretation involves assumptions about bias in the polls, and I believe it is hard to figure out the exact impact of the bias or even the direction. Electoral-vote has Obama at 364. At this time in 2004, they had Kerry at 283 (see this page), whereas his election day total was 252, with the main difference being Florida. More telling, the "strong" Obama States total 264 votes, as opposed to 95 for Kerry at this point.
www.fivethirtyeight.com – This consolidates polls by state to predict the count using some complex weighting system. It’s a neat idea but it’s end result is about the same as averaging, and I am not at all convinced it is better.
They’ve got Obama at 346.5, much more than the 270 needed to win.
www.gallup.com – This well-established survey company is different from the two above in that they actually conduct the polls.
My conclusion from the above---Obama will be the next
So why the change from before, when I said polls are difficult to trust and spoke of biases?
Three reasons:
1) the closer we get to the election, the better correlation between intentions and actions
2) the closer we get to the election, the fewer undecided voters. A recent Reuters poll shows this at about 2%. Even if it is 5% and the undecided break 4 to 1 for McCain, he's going to lose.
3) The biases appear to lean in Obama’s favor: more younger voters likely and more early voters. Very biased reporting from Grandma in S.C. says that lots of young people were out voting early (she spent 2 hours on line to vote early, by the way).