Statistics and Misidentifications - The weeks Detail
-
Strict Scrutiny
- Posts: 38
- Joined: Thu Aug 31, 2006 10:45 pm
Outsider,
In response to your question, I think it is common to have accusations regarding the quality of police work and misconduct during investigations. I can't think of one that exactly matches the McKie case but I am sure they exist. Tell me if I am correctly interpreting the TS fallacy. I am intrigued. Here is what Wikipedia says and it seems you are applying this rule of logic to the McKie case:
One CANNOT use the same information to construct AND test the same hypothesis - to do so would be to commit the Texas sharpshooter fallacy.
You are saying one CANNOT form the hypothesis that McKie lied AND test that hypothesis with the fingerprint information alone because the hypothesis and the test come from the same information. I guess my response is that you are wrong, you are using two sets of information. A fingerprint and a statement. The hypothesis comes from listening to her statement and the test comes in the form of the fingerprint.
Now the fact that the SCRO blew the fingerprint match is another story. But I think that since you have a statement AND an alleged fingerprint match is where the TS fallacy falls apart.
Now if McKie made no statement, she could not have been accused of lying, and if she had been then your fallacy would be accurate. But she would have been fired for not cooperating with the investigation, ouch my head hurts.
In response to your question, I think it is common to have accusations regarding the quality of police work and misconduct during investigations. I can't think of one that exactly matches the McKie case but I am sure they exist. Tell me if I am correctly interpreting the TS fallacy. I am intrigued. Here is what Wikipedia says and it seems you are applying this rule of logic to the McKie case:
One CANNOT use the same information to construct AND test the same hypothesis - to do so would be to commit the Texas sharpshooter fallacy.
You are saying one CANNOT form the hypothesis that McKie lied AND test that hypothesis with the fingerprint information alone because the hypothesis and the test come from the same information. I guess my response is that you are wrong, you are using two sets of information. A fingerprint and a statement. The hypothesis comes from listening to her statement and the test comes in the form of the fingerprint.
Now the fact that the SCRO blew the fingerprint match is another story. But I think that since you have a statement AND an alleged fingerprint match is where the TS fallacy falls apart.
Now if McKie made no statement, she could not have been accused of lying, and if she had been then your fallacy would be accurate. But she would have been fired for not cooperating with the investigation, ouch my head hurts.
-
Strict Scrutiny
- Posts: 38
- Joined: Thu Aug 31, 2006 10:45 pm
I've been thinking about this a bit more. The hypothesis was that McKie was in the house because she was on the investigation team, that hypothesis was tested by comparing her fingerprints to the latent prints in the house. Then that hypothesis was mistakenly confirmed by the erroneous match. This would not be the TS fallacy because she was not picked at random.
or,
The hypothesis is McKie lied. This hypothesis comes from the statement she made. It also comes from the fact that she was on the investigation team and could reasonably be expected to be in the house. It also comes from the erroneous fingerprint match.
I believe that both scenarios to not fit the TS fallacy since McKie was not picked at random, it was already hypothesized that she was in the house.
or,
The hypothesis is McKie lied. This hypothesis comes from the statement she made. It also comes from the fact that she was on the investigation team and could reasonably be expected to be in the house. It also comes from the erroneous fingerprint match.
I believe that both scenarios to not fit the TS fallacy since McKie was not picked at random, it was already hypothesized that she was in the house.
-
Outsider
- Posts: 166
- Joined: Mon Aug 07, 2006 2:15 am
- Location: Scotland
I had to lie down on the couch for a while to think this one out. One thing I realised is that my definition of the TS fallacy is slightly different from the Wikipedia one, but no matter, lets go with this one for now.One CANNOT use the same information to construct AND test the same hypothesis - to do so would be to commit the Texas sharpshooter fallacy.
You are saying one CANNOT form the hypothesis that McKie lied AND test that hypothesis with the fingerprint information alone because the hypothesis and the test come from the same information.
Before I start, I have to remove "individualisation" from the discussion. It may (or may not) be true for the very best fingerprint work, but it will certainly not be true for any work that falls slightly below the highest standards, and the problem is that nobody ever thinks that their work is below the highest standards. A non-individualised identification means that there is a small chance that a random match may occur with someone who did not deposit the latent. For a conviction to be safe, it must be shown that the likelihood of the person the dock being there as a result of a random match is so small that it can be discounted.
I can see that the Texas sharpshooter fallacy applies to nearly all fingerprint work (but I hope to show later that the McKie case is special). There is only one case where TS does not apply. This is where the police say to the fingerprint department "Tell me if John Smith deposited that latent". The prior hypothesis is "John Smith deposited the latent", then you get to work to test the hypothesis.
Now take the case where you have a latent and you search a database. Before you start the search you do not have a prior hypothesis that any named person deposited it. After the database search finds John Smith, the hypothesis is the same as before - "John Smith deposited the latent" but this time the fingerprint match is both constructing the hypothesis and testing it. You would be committing the TS fallacy if you said that the database search provides the same certainty as when the police provided the suspect. This is just common sense. In the suspect case, the chances of a single comparison between a latent and one inked print looking similar is small, so even with imperfect fingerprint work it is still extremely unlikely that there is any other explanation for the match than that the suspect deposited it. In a database search you will be doing millions of comparisons, many of these will look very similar to the latent, so unless you can guarantee individualisation then there is a heightened chance of a random erroneous match. The probability of error is proportional to the "population" of comparisons. The assumption when searching big databases is that fingerprinting has such a fabulously low chance of matching the wrong person that this chance can be multiplied by millions and it is still so low that it can be discounted. You are not committing the TS fallacy to say that a big database search is safe, you are only committing it if you say that it is as safe as a single suspect.
If the police provide a number of suspects and people to be eliminated then we are in a place that is somewhere between the single suspect and the database search. Some would say that bias could enter the picture here and increase the chance of error.
OK, now we are standing in the court room looking at the accused. I can see that we are not really in the territory of the TS fallacy at this point if we use the Wikipedia definition, because we are not hypothesis testing. But all we really have to ask is "is the chance that we are using a misidentification in evidence so low that we can discount it"? The only way we can do that is to limit the risk by limiting the number of identification that were in the inquiry (it only needs one misID in an inquiry because that is the one the police will choose to base the prosecution case on). If we want to return to the Texas Sharpshooter analogy, it is like giving him a number of bullets to hit the target, one for each latent. The chances of a random hit of the target with only a few hundred bullets is small. But if we give the sharpshooter an unlimited supply of bullets and an unlimited time for firing, he is going to hit the target one day.
I don't know were I got my definition of the TS fallacy from but it is simpler than the Wikipedia one (which would arise out of it). If something can happen by chance, and you want to prove that it did not happen by chance, then you cannot start from the event itself (because it could just be happening by chance). This is why we must not let a fingerprint match be the starting point of a whole new case of assumed wrongdoing. I also think it would also be wrong to tie up an old burglary to a fingerprint match that can not be connected with the crime under investigation because the connection would be too circumstantial. My definition can encompass the police selecting process because all we need to ask is "could this have happened by chance".
OK so time has tested fingerprinting using suspects, and using big database searches and it has been found to be safe (although some would argue that the proof of safety is inadequate). But fingerprinting without limiting the risk of misID by allowing any ID at any time to lead to a new suspicion of wrongdoing has NOT been tested because there have not been enough cases. This is why the McKie case looks different to me and I think that this aspect should be discussed any time someone is inquiring into what went wrong in the McKie case.If it aint broke, don't fix it
Now of course when someone suspects that a misID has occurred you must look at the images and determine if anything other than best practise has been employed. But I am talking about being back in 1997 when there was no reason to suspect that anything was wrong with the McKie ID. I would say that someone who understood that an allegation similar to McKie could arise out of any ID in any inquiry at any time would have alarm bells be going off in his or her head.
Steve Horn
Computer Programmer working in the field of statistics for industry
http://www.stevehornsc.pwp.blueyonder.co.uk/pf.htm
Computer Programmer working in the field of statistics for industry
http://www.stevehornsc.pwp.blueyonder.co.uk/pf.htm
-
Outsider
- Posts: 166
- Joined: Mon Aug 07, 2006 2:15 am
- Location: Scotland
I would say that the hypothesis before the fingerprint work starts is that the criminal has deposited a latent in the crime scene. The name of the criminal is added as a result of matches found but it is not fundamentally changed. If a latent is matched to someone on the elimination list they can give an explanation for why they are not the criminal.The hypothesis was that McKie was in the house because she was on the investigation team
The hypothesis that an unauthorised person entered the crime scene was constructed entirely out of the match between McKie and Y7.
If all this is getting too rarefied then just look at the Town Fingerprint Project. That says it all.
http://www.stevehornsc.pwp.blueyonder.co.uk/tfp.htm
Steve Horn
Computer Programmer working in the field of statistics for industry
http://www.stevehornsc.pwp.blueyonder.co.uk/pf.htm
Computer Programmer working in the field of statistics for industry
http://www.stevehornsc.pwp.blueyonder.co.uk/pf.htm
-
Strict Scrutiny
- Posts: 38
- Joined: Thu Aug 31, 2006 10:45 pm
Now this has devolved into "which came first the chicken or the egg?"
I just read your paper again. It appears every disputed ID in the Town Fingerprint Project is the result of a random computerized match in a computer database.
The McKie match is different than the examples you posed. It was made because she was believed to be in the house by the fingerprint examiner. Her print was not matched by computer, nor did the examiner randomly snag her card from the file; there was a belief prior to the ID which then caused the ID.
If McKie was not suspected of being in the house then her fingerprints never would have been compared to Y7. Therefore the hypothesis had to precede the comparison. No TSF, just a bum ID.
I just read your paper again. It appears every disputed ID in the Town Fingerprint Project is the result of a random computerized match in a computer database.
The McKie match is different than the examples you posed. It was made because she was believed to be in the house by the fingerprint examiner. Her print was not matched by computer, nor did the examiner randomly snag her card from the file; there was a belief prior to the ID which then caused the ID.
If McKie was not suspected of being in the house then her fingerprints never would have been compared to Y7. Therefore the hypothesis had to precede the comparison. No TSF, just a bum ID.
-
Outsider
- Posts: 166
- Joined: Mon Aug 07, 2006 2:15 am
- Location: Scotland
I don't think it is unresolveable. In a normal case the crime comes first, the accusation that someone entered Marion Ross' house without permission and that person was McKie came after the ID. When David Asbury won his appeal the question arises, "If Asbury did not kill Marion Ross then who did"?. When Shirley McKie was found not guilty of perjury nobody asked the question "If it wasn't McKie who entered the house without permission then who did"?Now this has devolved into "which came first the chicken or the egg?"
She was known to have been in David Asbury's house. The prior expectation with regard to police officers in the elimination list is that they will give an explanation for why their fingerprint was found so the latent can be eliminated.The McKie match is different than the examples you posed. It was made because she was believed to be in the house by the fingerprint examiner.
Steve Horn
Computer Programmer working in the field of statistics for industry
http://www.stevehornsc.pwp.blueyonder.co.uk/pf.htm
Computer Programmer working in the field of statistics for industry
http://www.stevehornsc.pwp.blueyonder.co.uk/pf.htm
-
Outsider
- Posts: 166
- Joined: Mon Aug 07, 2006 2:15 am
- Location: Scotland
Six logical points that affect the Shirley McKie case (criticism and comments welcome)
1) The people who are accused of lying after being identified from a fingerprint are those who deny having deposited the print. This is a non-random selection and it applies to both good identifications and misidentifications (probabilities are different after non-random selection).
2) We can divide accusations of lying into two groups for analysis (and because we can we should).
A) accusations of committing the crime under investigation, and
B) accusations that are not connected with the crime under investigation.
Group B can be further divided into those accusations where there is evidence of the alternative wrongdoing and those accusations where there is no evidence of the alternative wrongdoing.
3) Every fingerprint identification carries a risk of misidentification. Fingerprint experts claim an error rate of zero for competent work but unless this can be proved and there is a reliable way of dividing the good work from the bad we must assume that every identification carries a risk.
4) At the point of an accusation of lying, the probability that the accusation stemmed from a good fingerprint identification rather than a misidentification is the ratio of two independent probabilities.
A) the probability that the accused deposited the print, and
B) the probability that it is a misidentification
5) In a normal investigation which was set up because a crime occurred, we can expect probability A to be high because the fingerprint team were sent to a location where we know that the criminal attended, and probability B should be low because the crime itself limits the number of crime scene fingerprints in the investigation. This in turn limits the population of identifications which carry the risk (the probability of error after selection is the probability of error for one identification multiplied by the population).
6) For an accusation which is not the crime under investigation and where there is no evidence that the wrongdoing occurred then probability A is low (and might be approaching zero) and probability B is high because the population of identifications from which the accusation could stem is unlimited. Although finding the fingerprint of a police officer in a crime scene is not unusual the probability of finding a fingerprint of a police officer who will deny having deposited the print in the face of a fully verified identification will be low.
1) The people who are accused of lying after being identified from a fingerprint are those who deny having deposited the print. This is a non-random selection and it applies to both good identifications and misidentifications (probabilities are different after non-random selection).
2) We can divide accusations of lying into two groups for analysis (and because we can we should).
A) accusations of committing the crime under investigation, and
B) accusations that are not connected with the crime under investigation.
Group B can be further divided into those accusations where there is evidence of the alternative wrongdoing and those accusations where there is no evidence of the alternative wrongdoing.
3) Every fingerprint identification carries a risk of misidentification. Fingerprint experts claim an error rate of zero for competent work but unless this can be proved and there is a reliable way of dividing the good work from the bad we must assume that every identification carries a risk.
4) At the point of an accusation of lying, the probability that the accusation stemmed from a good fingerprint identification rather than a misidentification is the ratio of two independent probabilities.
A) the probability that the accused deposited the print, and
B) the probability that it is a misidentification
5) In a normal investigation which was set up because a crime occurred, we can expect probability A to be high because the fingerprint team were sent to a location where we know that the criminal attended, and probability B should be low because the crime itself limits the number of crime scene fingerprints in the investigation. This in turn limits the population of identifications which carry the risk (the probability of error after selection is the probability of error for one identification multiplied by the population).
6) For an accusation which is not the crime under investigation and where there is no evidence that the wrongdoing occurred then probability A is low (and might be approaching zero) and probability B is high because the population of identifications from which the accusation could stem is unlimited. Although finding the fingerprint of a police officer in a crime scene is not unusual the probability of finding a fingerprint of a police officer who will deny having deposited the print in the face of a fully verified identification will be low.
Steve Horn
Computer Programmer working in the field of statistics for industry
http://www.stevehornsc.pwp.blueyonder.co.uk/pf.htm
Computer Programmer working in the field of statistics for industry
http://www.stevehornsc.pwp.blueyonder.co.uk/pf.htm
-
Outsider
- Posts: 166
- Joined: Mon Aug 07, 2006 2:15 am
- Location: Scotland
I have put the above on a web page here:
http://www.stevehornsc.pwp.blueyonder.co.uk/logic.htm
It is also available as a link from my main page:
http://www.stevehornsc.pwp.blueyonder.co.uk/short.htm
http://www.stevehornsc.pwp.blueyonder.co.uk/logic.htm
It is also available as a link from my main page:
http://www.stevehornsc.pwp.blueyonder.co.uk/short.htm
Steve Horn
Computer Programmer working in the field of statistics for industry
http://www.stevehornsc.pwp.blueyonder.co.uk/pf.htm
Computer Programmer working in the field of statistics for industry
http://www.stevehornsc.pwp.blueyonder.co.uk/pf.htm
-
Outsider
- Posts: 166
- Joined: Mon Aug 07, 2006 2:15 am
- Location: Scotland
I have thought out another way to illustrate my point which I hope is the clearest yet (it was either spend time on this or dig the garden):
Imagine we know that the error rate of fingerprinting, after verification, is one per million identifications. A fingerprint team go round the streets of towns identifying people and if anyone denies having deposited the print, they will be accused of lying. They know that for every million identifications they make they will wreck one person’s life but take the attitude that nothing in life is risk free and the good done will outweigh the bad.
The first town is Innocentville where very few crimes occur. Nearly everyone is happy to agree that they deposited the print so most of the 999,999 good identifications are used up identifying innocent people. At the end of 1 million identifications they have found only 5 people who deny having deposited the print and they have all been accused of lying. It doesn’t seem much to account for all that work and they are uncomfortable to think that one out of the 5 is innocent.
For the next million identifications the fingerprint team go to Badville. Crime is worse here so instead of finding only 5 of people denying depositing a print they find 25. In Badville the error rate of accusations is 1 in 25 compared with 1 in 5 in Innocentville. Same fingerprint team, same quality of fingerprint work, same per-identification error rate but different per-accusation error rate. What is different between the two towns is the proportion of latents where the person who deposited it will lie because they have something to hide.
Of course the place where you find the highest concentration of this type of latent is in crime scenes. If the fingerprint team restricts itself to only going to crime scenes then the concentration of latents where the person who deposited it will lie will be hundreds or thousands of times higher even than Badville and the error rate per accusation will reduce by the same amount.
But if while in a crime scene the fingerprint team make an identification where the person identified denies having deposited the print but there is no connection with the crime under investigation, this is the equivalent to walking the streets looking for crimes to become apparent out of fingerprint identifications.
I think the principle is that you can never prove that a crime or incident of wrongdoing occurred from a fingerprint.
I have just plucked figures out of the air but I hope it gives an understanding relative safety in different contexts. There are many things in engineering that never get beyond “back of envelope” type calculations, they can give great insight into the mechanisms at play.
Imagine we know that the error rate of fingerprinting, after verification, is one per million identifications. A fingerprint team go round the streets of towns identifying people and if anyone denies having deposited the print, they will be accused of lying. They know that for every million identifications they make they will wreck one person’s life but take the attitude that nothing in life is risk free and the good done will outweigh the bad.
The first town is Innocentville where very few crimes occur. Nearly everyone is happy to agree that they deposited the print so most of the 999,999 good identifications are used up identifying innocent people. At the end of 1 million identifications they have found only 5 people who deny having deposited the print and they have all been accused of lying. It doesn’t seem much to account for all that work and they are uncomfortable to think that one out of the 5 is innocent.
For the next million identifications the fingerprint team go to Badville. Crime is worse here so instead of finding only 5 of people denying depositing a print they find 25. In Badville the error rate of accusations is 1 in 25 compared with 1 in 5 in Innocentville. Same fingerprint team, same quality of fingerprint work, same per-identification error rate but different per-accusation error rate. What is different between the two towns is the proportion of latents where the person who deposited it will lie because they have something to hide.
Of course the place where you find the highest concentration of this type of latent is in crime scenes. If the fingerprint team restricts itself to only going to crime scenes then the concentration of latents where the person who deposited it will lie will be hundreds or thousands of times higher even than Badville and the error rate per accusation will reduce by the same amount.
But if while in a crime scene the fingerprint team make an identification where the person identified denies having deposited the print but there is no connection with the crime under investigation, this is the equivalent to walking the streets looking for crimes to become apparent out of fingerprint identifications.
I think the principle is that you can never prove that a crime or incident of wrongdoing occurred from a fingerprint.
I have just plucked figures out of the air but I hope it gives an understanding relative safety in different contexts. There are many things in engineering that never get beyond “back of envelope” type calculations, they can give great insight into the mechanisms at play.
Steve Horn
Computer Programmer working in the field of statistics for industry
http://www.stevehornsc.pwp.blueyonder.co.uk/pf.htm
Computer Programmer working in the field of statistics for industry
http://www.stevehornsc.pwp.blueyonder.co.uk/pf.htm
-
Outsider
- Posts: 166
- Joined: Mon Aug 07, 2006 2:15 am
- Location: Scotland
Bayes Theorem
I have used Bayes Theorem to put some figures into my statistical observation about the McKie identification. I get a “per accusation” error rate of 1 in 5000 for a normal crime scene and 1 in 2 for a McKie type situation with the following assumptions:
a) identification error rate of 1 in 1 million after verification (errors can have any cause)
b) 1 in 200 identifiable latent prints in crime scenes are deposited by the criminals
c) 1 in 1 million identifiable latent prints in crime scenes are deposited by police officers disobeying orders who will lie if identified.
Fundamental to Bayes Theorem is the concept of “prior probability”. This is the chances - before the fingerprint analysis - that a latent print was deposited by the wrongdoer. In crime scenes the criminal only has to touch something and leave an identifiable print so the probability will be relatively high (I have guessed at 1 in 200). For an allegation which is not the crime under investigation we also have to take into account the prior probability (before the fingerprint analysis) that the suggested act might have happened. Without some independent evidence, the prior probability will be a very very small value, similar, I guess, to the error rate (which means it is not safe to make an accusation).
I have put my calculations in the form of a step-by-step introduction to Bayes Theorem. It is not at all complicated:
http://www.stevehornsc.pwp.blueyonder.co.uk/bayes.htm
“Fingerprinting Innocentville” is a less mathematical way of saying the same thing:
http://www.stevehornsc.pwp.blueyonder.c ... tville.htm
To put it simply, making an accusation of lying where a fingerprint ID is only evidence to suggest that the disputed thing happened is asking for trouble. I hope that the abnormal statistical aspects of the McKie case will not be forgotten in the long-term. Whether or not “individualisation” is a reasonable assumption for competent work, I believe that people involved with quality assurance in fingerprinting should have some statistical awareness. It gives an insight into what to expect when the highest standards are not achieved. In the process industries, developing statistical awareness in operators, supervisors and managers has been the biggest single contributor to the quality revolution of recent years (the software I develop is connected with this).
Ultimately, like everything else, it comes down to a risk/benefit consideration. No human activity is risk free. When serious crimes occur, society requires that all reasonable steps are taken to find the criminal. Accusing Shirley McKie of lying when she said that she had not entered the murder house involved a risk thousands of times greater than when fingerprinting is used to solve a crime, yet this inflated risk was taken without any real crime to solve.
a) identification error rate of 1 in 1 million after verification (errors can have any cause)
b) 1 in 200 identifiable latent prints in crime scenes are deposited by the criminals
c) 1 in 1 million identifiable latent prints in crime scenes are deposited by police officers disobeying orders who will lie if identified.
Fundamental to Bayes Theorem is the concept of “prior probability”. This is the chances - before the fingerprint analysis - that a latent print was deposited by the wrongdoer. In crime scenes the criminal only has to touch something and leave an identifiable print so the probability will be relatively high (I have guessed at 1 in 200). For an allegation which is not the crime under investigation we also have to take into account the prior probability (before the fingerprint analysis) that the suggested act might have happened. Without some independent evidence, the prior probability will be a very very small value, similar, I guess, to the error rate (which means it is not safe to make an accusation).
I have put my calculations in the form of a step-by-step introduction to Bayes Theorem. It is not at all complicated:
http://www.stevehornsc.pwp.blueyonder.co.uk/bayes.htm
“Fingerprinting Innocentville” is a less mathematical way of saying the same thing:
http://www.stevehornsc.pwp.blueyonder.c ... tville.htm
To put it simply, making an accusation of lying where a fingerprint ID is only evidence to suggest that the disputed thing happened is asking for trouble. I hope that the abnormal statistical aspects of the McKie case will not be forgotten in the long-term. Whether or not “individualisation” is a reasonable assumption for competent work, I believe that people involved with quality assurance in fingerprinting should have some statistical awareness. It gives an insight into what to expect when the highest standards are not achieved. In the process industries, developing statistical awareness in operators, supervisors and managers has been the biggest single contributor to the quality revolution of recent years (the software I develop is connected with this).
Ultimately, like everything else, it comes down to a risk/benefit consideration. No human activity is risk free. When serious crimes occur, society requires that all reasonable steps are taken to find the criminal. Accusing Shirley McKie of lying when she said that she had not entered the murder house involved a risk thousands of times greater than when fingerprinting is used to solve a crime, yet this inflated risk was taken without any real crime to solve.
Steve Horn
Computer Programmer working in the field of statistics for industry
http://www.stevehornsc.pwp.blueyonder.co.uk/pf.htm
Computer Programmer working in the field of statistics for industry
http://www.stevehornsc.pwp.blueyonder.co.uk/pf.htm
-
Outsider
- Posts: 166
- Joined: Mon Aug 07, 2006 2:15 am
- Location: Scotland
Probability changes over time
If I approach a National Lottery player at random, I can be very sure that they have not won the Jackpot. If they tell me otherwise, I can confidently accuse them of lying. This is because the odds of winning the lottery are 1 in 14 million and I have chosen a player at random. If I am sitting next to someone on a train who tells me that they have won the lottery I would not dispute it because they came to my attention by a non-random process. If I gatecrash a lottery winners’ party, the probability of choosing a lottery winner from among the guests is very high. So we have three different probabilities all based around the same lottery win. Probability changes after a non-random selection process (random selection does not alter probability).Steve, I am fascinated by your assertion that probability changes over time, could you please give some examples of this?
The police choosing an identification because the identified person cannot give an innocent explanation for being at the location is non-random selection. Only a small proportion of good identifications will be chosen, they will be the criminals. Nearly all misidentifications will be chosen because the most natural thing for a misidentified person to say is that they did not deposit the print.
Steve Horn
Computer Programmer working in the field of statistics for industry
http://www.stevehornsc.pwp.blueyonder.co.uk/pf.htm
Computer Programmer working in the field of statistics for industry
http://www.stevehornsc.pwp.blueyonder.co.uk/pf.htm
-
Daktari
- Posts: 582
- Joined: Fri Aug 18, 2006 2:50 am
- Location: Glasgow
Probability changes over time?
Sorry Steve, but it doesn’t. Probability is the ratio of a specific outcome over all possible outcomes. If one tosses a true coin a specific outcome, say heads, is simply heads over heads and tails i.e. one over two or 50% if you prefer.
If one throws a true die the a specific outcome, say a six, is six over one, two three four, five or six i.e. one in six or 16.67%.
This won’t vary over time. The experiment could have been carried out twenty years ago or in forty years time. The design of the coin may have changed but the probability of tossing a head is still the same 50%.
The probability of tossing a head will not change either no matter where, or by whom, the coin is tossed. Or, for that matter how many times it was heads in previous tosses.
Let me show where you are being confused. Imagine a bag containing three red balls and three blue ones. The probability of a blindfolded person picking out a red ball is, specific outcome, red ball, over possible outcomes, three red and three blue, is three over six viz. 50%. Let’s now suppose the person does pick out a red ball. What is the probability of picking another red ball? It is specific outcome, red ball, over possible outcomes, two reds three blues, is two over five, 40%. What is the probability of picking the last red ball?
Specific outcome, red ball, over possible outcomes, one red three blues, is 25%. So, someone with no understanding of probability theory may think that the probabilities of picking a red ball have changed from 50% to 40% and then to 25%. However someone with a better understanding of statistics would immediately be aware that it is not the probability that has changed it is the event that has changed.
You should also know that the odds of meeting a lottery millionaire are not the same as winning the lottery, circa 14m to 1. The probability, as I explained above, of meeting such a winner is the number of such winners over the entire population. Since the lottery has been running for more than ten years and there are, on average, about three millionaires created each week. That would mean something like 1,500 winners. Assuming they all stayed in the UK then we would have a probability of 1,500 over 60m i.e. 1in 40,000. Q.E.D.
Sorry Steve, but it doesn’t. Probability is the ratio of a specific outcome over all possible outcomes. If one tosses a true coin a specific outcome, say heads, is simply heads over heads and tails i.e. one over two or 50% if you prefer.
If one throws a true die the a specific outcome, say a six, is six over one, two three four, five or six i.e. one in six or 16.67%.
This won’t vary over time. The experiment could have been carried out twenty years ago or in forty years time. The design of the coin may have changed but the probability of tossing a head is still the same 50%.
The probability of tossing a head will not change either no matter where, or by whom, the coin is tossed. Or, for that matter how many times it was heads in previous tosses.
Let me show where you are being confused. Imagine a bag containing three red balls and three blue ones. The probability of a blindfolded person picking out a red ball is, specific outcome, red ball, over possible outcomes, three red and three blue, is three over six viz. 50%. Let’s now suppose the person does pick out a red ball. What is the probability of picking another red ball? It is specific outcome, red ball, over possible outcomes, two reds three blues, is two over five, 40%. What is the probability of picking the last red ball?
Specific outcome, red ball, over possible outcomes, one red three blues, is 25%. So, someone with no understanding of probability theory may think that the probabilities of picking a red ball have changed from 50% to 40% and then to 25%. However someone with a better understanding of statistics would immediately be aware that it is not the probability that has changed it is the event that has changed.
You should also know that the odds of meeting a lottery millionaire are not the same as winning the lottery, circa 14m to 1. The probability, as I explained above, of meeting such a winner is the number of such winners over the entire population. Since the lottery has been running for more than ten years and there are, on average, about three millionaires created each week. That would mean something like 1,500 winners. Assuming they all stayed in the UK then we would have a probability of 1,500 over 60m i.e. 1in 40,000. Q.E.D.
-
g.
- Posts: 247
- Joined: Wed Jul 06, 2005 1:27 pm
- Location: St. Paul, MN
Daktari, Steve...I couldn't resist chiming in.
Only because this is such a classic, well known struggle in statistics, it's sort of fun to watch it play out on this site. This is the classic argument between Bayesians and Frequentists.
In effect, you are both right, but I think Daktari, you are being somewhat unfair to Steve (characterizing his lack of statistical knowledge) because he isn't incorrect and is demonstrating sound knowledge of advanced statistics. Neither are you incorrect though, to a point.
Your statement, probabilities don't change over time isn't entirely accurate. They don't, as you have given several examples, if the events are fixed, the conditions are fixed and known, and all possible outcomes can be tallied. Then you are correct.
In the real world, it is not always true. Perhaps for lotteries and raffle drawings, but a prime example has to do with Poisson estimators.
Let's say the EVENT is a person entering my restaurant. You may find the probability of the event is higher around Noon to 1PM (people on their lunch hour) vs. 10 AM when everybody is working. The event doesn't change, but the probability changes over time to reflect the conditions. So clearly it is a huge issue WHEN the experiment is carried out, say measuring the number of people coming into your restaurant to choose the correct number of staff to accomodate your patrons.
Also the other thing about the Bayesian approach, is while in some instances, the TRUE probability perhaps is a fixed event, but what changes is our knowledge of what that value is. Poker games are classic examples. While the probability of 4 kings is a fixed event, prior to dealing the cards, once the game is under way and you will see other players' cards, or cards in the community flop, your own cards, etc. So you begin to update the probabilities of what everybody else has. So while the original probability didn't change, your KNOWLEDGE of the EVENT does change, and thus can be updated (called the posterior probability).
I don't necessarily agree with all of Steve's approach to the McKie case. I do agree there is a Bayesian aspect to it the way he is now looking at it. The previous approach he gave I thought was lacking (the one Michelle Triplett offered her criticism and started this thread), but now with his inclusion of a Bayesian approach, it begins to make more sense. Essentially with respect to the probability of a false match error given the events that have occured (updated with our knowledge of the event) he is saying it more likely a fingerprint error. I have some issue with some of the estimates and suppositions to arrive at the answer, but the theory and math is known and acceptable.
Just my thoughts on the issue. For more check out the famous "Monty Hall Problem" and other debates between Frequentists and Bayesians. The DNA literature is rife with this age old debate as well.
g.
Only because this is such a classic, well known struggle in statistics, it's sort of fun to watch it play out on this site. This is the classic argument between Bayesians and Frequentists.
In effect, you are both right, but I think Daktari, you are being somewhat unfair to Steve (characterizing his lack of statistical knowledge) because he isn't incorrect and is demonstrating sound knowledge of advanced statistics. Neither are you incorrect though, to a point.
Your statement, probabilities don't change over time isn't entirely accurate. They don't, as you have given several examples, if the events are fixed, the conditions are fixed and known, and all possible outcomes can be tallied. Then you are correct.
In the real world, it is not always true. Perhaps for lotteries and raffle drawings, but a prime example has to do with Poisson estimators.
Let's say the EVENT is a person entering my restaurant. You may find the probability of the event is higher around Noon to 1PM (people on their lunch hour) vs. 10 AM when everybody is working. The event doesn't change, but the probability changes over time to reflect the conditions. So clearly it is a huge issue WHEN the experiment is carried out, say measuring the number of people coming into your restaurant to choose the correct number of staff to accomodate your patrons.
Also the other thing about the Bayesian approach, is while in some instances, the TRUE probability perhaps is a fixed event, but what changes is our knowledge of what that value is. Poker games are classic examples. While the probability of 4 kings is a fixed event, prior to dealing the cards, once the game is under way and you will see other players' cards, or cards in the community flop, your own cards, etc. So you begin to update the probabilities of what everybody else has. So while the original probability didn't change, your KNOWLEDGE of the EVENT does change, and thus can be updated (called the posterior probability).
I don't necessarily agree with all of Steve's approach to the McKie case. I do agree there is a Bayesian aspect to it the way he is now looking at it. The previous approach he gave I thought was lacking (the one Michelle Triplett offered her criticism and started this thread), but now with his inclusion of a Bayesian approach, it begins to make more sense. Essentially with respect to the probability of a false match error given the events that have occured (updated with our knowledge of the event) he is saying it more likely a fingerprint error. I have some issue with some of the estimates and suppositions to arrive at the answer, but the theory and math is known and acceptable.
Just my thoughts on the issue. For more check out the famous "Monty Hall Problem" and other debates between Frequentists and Bayesians. The DNA literature is rife with this age old debate as well.
g.
-
Outsider
- Posts: 166
- Joined: Mon Aug 07, 2006 2:15 am
- Location: Scotland
Frequency and probability
Let’s look at a situation from frequency and probability viewpoints.
Suppose that over a period of time we have thousands of CCTV recordings of criminals entering shops at night to steal goods, and leaving the shops again. In half the incidents we can see that the criminals are wearing gloves as they enter and leave the shop.
In the cases where the criminal was not wearing gloves there is a very good chance that they will leave an identifiable print. Say there are ten thousand such incidents, and we identify the criminal in 5000. As always, the criminal will have been identified out of all the hundreds of elimination identifications in each incident because they deny having deposited the print.
In 10000 cases where the criminal is wearing gloves we would not expect to find a fingerprint of the criminal in the shop. There is a very small chance that the criminal took off the gloves, touched something and put on the gloves again before leaving. In the 10000 gloved incidents, we only find 2 prints where the person identified denies having deposited the print. There is a different frequency of prosecutions between the gloved and non-gloved robberies.
In every robbery there will be large number of eliminations and every elimination has the potential to be misidentified to someone who can not give an innocent explanation for having been at the location. So the frequency of misidentifications will be the same for both gloved and non-gloved robberies. Let’s say one misidentification each for the 10000 gloved and 10000 non-gloved robberies. So the ratio (probability) of good identifications to misidentifications for the non-gloved robberies is 1 to 4999 but for the gloved robberies it is 1 to 1.
In an non-gloved robbery the prior probability of a fingerprint having been deposited by the criminal is high and for the gloved robbery the prior probability of a fingerprint having been deposited by the criminal is low. Until the police select an identification to base a prosecution case on, every identification (including all the hundreds of thousands of eliminations) has the same very small probability of error. The probability of error changes after selection and it depends on the prior probability of a latent in a location having been deposited by the criminal.
The prior probability of a police officer entering a crime scene without permission and denying it when confronted with a fully verified identification and there is no independent evidence of anybody entering the crime scene without permission will be orders of magnitude lower than the prior probability of a (non-gloved) criminal leaving an identifiable print in a crime scene.
Suppose that over a period of time we have thousands of CCTV recordings of criminals entering shops at night to steal goods, and leaving the shops again. In half the incidents we can see that the criminals are wearing gloves as they enter and leave the shop.
In the cases where the criminal was not wearing gloves there is a very good chance that they will leave an identifiable print. Say there are ten thousand such incidents, and we identify the criminal in 5000. As always, the criminal will have been identified out of all the hundreds of elimination identifications in each incident because they deny having deposited the print.
In 10000 cases where the criminal is wearing gloves we would not expect to find a fingerprint of the criminal in the shop. There is a very small chance that the criminal took off the gloves, touched something and put on the gloves again before leaving. In the 10000 gloved incidents, we only find 2 prints where the person identified denies having deposited the print. There is a different frequency of prosecutions between the gloved and non-gloved robberies.
In every robbery there will be large number of eliminations and every elimination has the potential to be misidentified to someone who can not give an innocent explanation for having been at the location. So the frequency of misidentifications will be the same for both gloved and non-gloved robberies. Let’s say one misidentification each for the 10000 gloved and 10000 non-gloved robberies. So the ratio (probability) of good identifications to misidentifications for the non-gloved robberies is 1 to 4999 but for the gloved robberies it is 1 to 1.
In an non-gloved robbery the prior probability of a fingerprint having been deposited by the criminal is high and for the gloved robbery the prior probability of a fingerprint having been deposited by the criminal is low. Until the police select an identification to base a prosecution case on, every identification (including all the hundreds of thousands of eliminations) has the same very small probability of error. The probability of error changes after selection and it depends on the prior probability of a latent in a location having been deposited by the criminal.
The prior probability of a police officer entering a crime scene without permission and denying it when confronted with a fully verified identification and there is no independent evidence of anybody entering the crime scene without permission will be orders of magnitude lower than the prior probability of a (non-gloved) criminal leaving an identifiable print in a crime scene.
Steve Horn
Computer Programmer working in the field of statistics for industry
http://www.stevehornsc.pwp.blueyonder.co.uk/pf.htm
Computer Programmer working in the field of statistics for industry
http://www.stevehornsc.pwp.blueyonder.co.uk/pf.htm