Monday, May 3, 2010

Reading the Random (or a pedantic rant on conditional probability)

This is something I was supposed to have known and understood from high school.  But I got to appreciate something very key and insightful about this today.  So consider the following classic situation where analyzing for conditional probability becomes important :

Suppose we want to screen people for a disease, that's relatively rare but dangerous (HIV for instance), and it's known that the test is about 99% reliable, that is it goes wrong only once every hundred times it is performed. Then suppose you are the poor guy who tests positive for the virus. Does that mean you are  screwed for all practical purposes now, with only an outside chance of escape? 

Clearly that's what it looks like, but we have to remember there are two independent questions at play here : what does the test say and whether you really have the disease, and both these need to be addressed here.  We of course have the test results with us but need more data to decide the second matter.  Let's say our small country has a population of a million people, and it's not sub Saharan Africa, so only a hundred people actually have this disease (It's quite another story where this data is coming from, because the tests themselves are not absolutely reliable).  When the healthy population of a "million minus hundred" (9,99,900) is tested, 1%  (9999) of them will presumably test wrong, i.e. will test falsely positive for the disease. There will be nearly 10,000 false positive people as opposed to only 99 true positive people, i.e. an overwhelming proportion of people who test positive will actually be disease free! This calls into question the necessity of this test in the first place.  You have  been saved from a prolonged panic attack by conditional probability.

There are many other such illustrations of how failing to account for conditional probability can be fallacious. It is all very simple a posteriori and yet, many intelligent people frequently slip up on this.  This is where my recent realization comes in.  This has got to do with the fact that the idea of conditional probability needs some sort of a priori data input, that is probably obtained through some other means.  The fallacies mostly happen from people assuming two things to be unrelated (statistically independent) when they're not.  Here we need estimates for probabilities of both the isolated event happening as well as the two happening together. This brings us to the somewhat deep (at least I think so) question of how to get these probability values in the first place, since in practice nothing is absolutely reliable or uncorrelated, i.e. different events influence each other in very unpredictable ways. How do I really know for example that the real disease incidence is only 1 in 10,000 when the tests are 1% unreliable. Is there a second test perhaps? Does repetition help get a better handle on things? Can it be inferred from some other related phenomena?

The elementary problems involving probability are all based on certain reasonable assumptions.  A coin unless biased would be equally likely to come out heads or tails, a die would perfectly randomly give a number from 1 to 6 when thrown, and in thermodynamics atoms and molecules in a gas would have some probable energy depending on its temperature.  The latter might puzzle the non-physicist or non-statistician but it's only an extension of the coin or die example where every possible outcome is equally likely.  So perfect randomness is actually a good thing, as we can then relatively safely pinpoint the probability values.

This problem is generic in probability theory.  We would like to believe that all quantifiable data in the natural or social sciences would have some underlying probability distribution, based on which we can analyze or predict things.  But the step before that is to identify this distribution itself!  This in a way is a chicken-and-egg problem since we have to experiment to guess the distribution which we then want to use to predict the outcome of experiment. Most often therefore the distribution is "constructed" using insight and intuition, like Boltzmann did for the energy of a gas molecule. Then there are various fitness tests to tell us how reliable it is. But a large variety of data from very different walks in life actually fit very well to a very small number of standard distributions (is that some sort of unification or universality in statistical theory then?)!  The Gaussian or normal distribution is a case in point.  Statistics is all about intelligent guesswork.

All that I've said above is not specific to conditional probability, but this I believe is the simplest level, where the problems with assuming things to be uncorrelated when they're really not, and the importance of additional data, become apparent, and this can be demonstrated through very simple and interesting fallacies.


I will end with a self concocted example. Consider the small state again with a  mint that has two faulty presses for making coins, one makes head-weighted coins which have a 52% bias of coming out heads in any given toss, and the other makes coins that have 56% chance of coming out tails.  This piece of info can be found for example by the lazy mint official who whiles away his time by tossing coins from both machines and keeping separate records of the two,  instead of fixing the machines.  Or a theoretical physicist could do it by studying the shape and mass distribution of the coin if he was given a grant to (but always trust experiments more, especially when it comes to numbers, not "order of magnitude estimates"!).  Suppose both machines make the same number of coins, i.e. there are equal numbers of both types of coins in the market. How do I guess given a coin which type it is? Of course by throwing it many times and checking whether it ends up more in tails or heads. But suppose I throw it only once, and it comes out heads, then what is the probability that it came from the first machine? Anybody who knows elementary Bayes theorem can calculate the correct answer (54%).  Here it was relatively easy to know the numbers (52 & 56) because we are in good control of things. But imagine a real world with a complicated interdependence of a lot of factors, and then you have some idea  as to why a trained statistician is so highly valued, and how useful or misleading statistics could really be depending on how careful you are being with them .

 I was inspired to think about this from the following pop talk by a renowned statistician (Peter Donnelly):

No comments: