Statistical Model Article
-
Gerald Clough
- Posts: 557
- Joined: Wed Jul 06, 2005 6:27 am
- Location: Lockhart, Texas
- Contact:
Statistical Model Article
First, I got this from Michelle, and by rights, she ought to get to post it here, but I want to comment on it, so I'll start it, but she should get the credit.
http://www.eurekalert.org/pub_releases/ ... 020812.php
Statistical model unlocks barriers to use of fingerprint evidence in court
Potentially key fingerprint evidence is currently not being considered due to shortcomings in the way it is reported, according to a report published today in Significance, the magazine of the Royal Statistical Society and the American Statistical Association. Researchers involved in the study have devised a statistical model to enable the weight of fingerprint evidence to be quantified, paving the way for its full inclusion in the criminal identification process.
Fingerprints have been used for over a century as a way of identifying criminals. However, fingerprint evidence is not currently permitted to be reported in court unless examiners claim absolute certainty that a mark has been left by a particular suspect. This courtroom certainty is based purely on categorical personal opinion, formed through years of training and experience, but not on logic or scientific data. Less than certain fingerprint evidence is not reported at all, irrespective of the potential weight and relevance of this evidence in a case.
Today's Significance paper, which publishes in advance of the full study in the Journal of the Royal Statistical Society: Series A later this year, highlights this subjectivity in current processes, calling for changes in the way such key evidence is allowed to be presented. According to Professor of Statistics Cedric Neumann, "It is unthinkable that such valuable evidence should not be reported, effectively hidden from courts on a regular basis. Such is the importance of this wealth of data, we have devised a reliable statistical model to enable the courts to evaluate fingerprint evidence within a framework similar to that which underpins DNA evidence."
Neumann, from Pennsylvania State University, and his team devised and successfully tested a model for establishing the probability of a print belonging to a particular suspect. After mapping the finer points of detail on a "control print" and "crime scene print", two hypotheses were then tested. The first test, to establish the probability that the crime scene print was made by the owner of the control print (the suspect), compared the control print with a range of other prints made by the suspect. The second test, to establish the probability that the crime scene print was made by someone other than the suspect, compared the crime scene print with a set of prints in a reference database. A likelihood ratio between the two probabilities was calculated; the higher the ratio indicating stronger evidence that the suspect was the source of the crime scene print.
"Current practice allows a state of certainty to be presented which is not justified scientifically, or supported by logical process or data," said Professor Neumann. "We believe that the examiner should not decide what evidence should or should not be presented. Our method allows all evidence to be supported by data, and reported according to a continuous scale."
Their approach was to use a plotting scheme not uncommon if automated FP database searches. They reasoned that they could test for the likelihood of a false positive identification by the examiner by using the US database to check UK identifications, making a reasonable assumption that so few people appear in the US database and have the opportunity to leave an impression at a UK scene that the likelihood was sufficiently small to make the US database a reasonable test.
They then develop a statistically useful measure of similarity and search for the person in the US database that mostly closely matches the proposed UK source. In short, none of the closest US matches were sufficiently close to fall within their threshold for courtroom evidence, which would have deemed them as misleading. In other words, no one would have been charged based on the statistical degree of match between the UK print and the closest US subject. Their model also distinguishes sufficiently between the actual identified UK subject and other UK subjects.
It is hardly a game-changing report, but it is a good demonstration of the way the subject can be approached and a seemingly reasonable measure of reliability. Unfortunately, while it doesn't take away from the principles and value of the work, they make the same error so many make about how they characterize the expert opinion of examiners. That is that they represent examiners as concluding factual absolute certainty. The reality is, of course, that any intellectually honest examiner cannot claim absolute certainty. Rather, the examiner reports that their belief as an expert that the mark is that of the individual. Few scientists who do not perform as expert witnesses understand the place of the expert in the legal system and the courts' view of expert evidence.
They also bemoan what they see as one discipline - DNA - presenting statistical conclusions and another barred from statistical argument. While the observation is essentially correct, the situation is not contradictory when seen within the whole range of expert disciplines, some of which present quantified evidence and some of which do not.
That aside, they are not condemning the use of expert fingerprint testimony. They seem to believe that statistical methods could bring more fingerprint evidence into court. They are, as again so many scientists so often are, overly optimistic in believing this will become a hotly contested issue in court. They both misunderstand how protective are courts to evidence they value and about how long it takes for significant changes in law to take place.
This study is published in the February issue of Significance. Media wishing to receive a PDF of this article may contact physicalsciencenews@wiley.com.
http://www.eurekalert.org/pub_releases/ ... 020812.php
Statistical model unlocks barriers to use of fingerprint evidence in court
Potentially key fingerprint evidence is currently not being considered due to shortcomings in the way it is reported, according to a report published today in Significance, the magazine of the Royal Statistical Society and the American Statistical Association. Researchers involved in the study have devised a statistical model to enable the weight of fingerprint evidence to be quantified, paving the way for its full inclusion in the criminal identification process.
Fingerprints have been used for over a century as a way of identifying criminals. However, fingerprint evidence is not currently permitted to be reported in court unless examiners claim absolute certainty that a mark has been left by a particular suspect. This courtroom certainty is based purely on categorical personal opinion, formed through years of training and experience, but not on logic or scientific data. Less than certain fingerprint evidence is not reported at all, irrespective of the potential weight and relevance of this evidence in a case.
Today's Significance paper, which publishes in advance of the full study in the Journal of the Royal Statistical Society: Series A later this year, highlights this subjectivity in current processes, calling for changes in the way such key evidence is allowed to be presented. According to Professor of Statistics Cedric Neumann, "It is unthinkable that such valuable evidence should not be reported, effectively hidden from courts on a regular basis. Such is the importance of this wealth of data, we have devised a reliable statistical model to enable the courts to evaluate fingerprint evidence within a framework similar to that which underpins DNA evidence."
Neumann, from Pennsylvania State University, and his team devised and successfully tested a model for establishing the probability of a print belonging to a particular suspect. After mapping the finer points of detail on a "control print" and "crime scene print", two hypotheses were then tested. The first test, to establish the probability that the crime scene print was made by the owner of the control print (the suspect), compared the control print with a range of other prints made by the suspect. The second test, to establish the probability that the crime scene print was made by someone other than the suspect, compared the crime scene print with a set of prints in a reference database. A likelihood ratio between the two probabilities was calculated; the higher the ratio indicating stronger evidence that the suspect was the source of the crime scene print.
"Current practice allows a state of certainty to be presented which is not justified scientifically, or supported by logical process or data," said Professor Neumann. "We believe that the examiner should not decide what evidence should or should not be presented. Our method allows all evidence to be supported by data, and reported according to a continuous scale."
Their approach was to use a plotting scheme not uncommon if automated FP database searches. They reasoned that they could test for the likelihood of a false positive identification by the examiner by using the US database to check UK identifications, making a reasonable assumption that so few people appear in the US database and have the opportunity to leave an impression at a UK scene that the likelihood was sufficiently small to make the US database a reasonable test.
They then develop a statistically useful measure of similarity and search for the person in the US database that mostly closely matches the proposed UK source. In short, none of the closest US matches were sufficiently close to fall within their threshold for courtroom evidence, which would have deemed them as misleading. In other words, no one would have been charged based on the statistical degree of match between the UK print and the closest US subject. Their model also distinguishes sufficiently between the actual identified UK subject and other UK subjects.
It is hardly a game-changing report, but it is a good demonstration of the way the subject can be approached and a seemingly reasonable measure of reliability. Unfortunately, while it doesn't take away from the principles and value of the work, they make the same error so many make about how they characterize the expert opinion of examiners. That is that they represent examiners as concluding factual absolute certainty. The reality is, of course, that any intellectually honest examiner cannot claim absolute certainty. Rather, the examiner reports that their belief as an expert that the mark is that of the individual. Few scientists who do not perform as expert witnesses understand the place of the expert in the legal system and the courts' view of expert evidence.
They also bemoan what they see as one discipline - DNA - presenting statistical conclusions and another barred from statistical argument. While the observation is essentially correct, the situation is not contradictory when seen within the whole range of expert disciplines, some of which present quantified evidence and some of which do not.
That aside, they are not condemning the use of expert fingerprint testimony. They seem to believe that statistical methods could bring more fingerprint evidence into court. They are, as again so many scientists so often are, overly optimistic in believing this will become a hotly contested issue in court. They both misunderstand how protective are courts to evidence they value and about how long it takes for significant changes in law to take place.
This study is published in the February issue of Significance. Media wishing to receive a PDF of this article may contact physicalsciencenews@wiley.com.
"Nothing has any value, unless you know you can give it up."
-
Boyd Baumgartner
- Posts: 567
- Joined: Sat Aug 06, 2005 11:03 am
Re: Statistical Model Article
Except for the fact that they are:Gerald Clough wrote:That aside, they are not condemning the use of expert fingerprint testimony.
Aside from all of the ranting I've done on this board about statistics, I've never expressed this opinion. At the end of the day the argument is essentially an analog versus digital argument in the same way that audio has the argument about vinyl vs mp3/cd and the quality of the sound that results. What's at issue is the use of a full complement of information versus discreet units. Traditionally this was done through the use of level 2 detail, but was largely a linguistic convention. For instance, the historical statistical models were based on what types of level 2 features actually existed in combination with a grid size. But if I came along and gave rise to or consolidated the number of types of level 2 features based upon what I called them (for instance bifurcations really comprise byfurcations and byefurcations) based upon some metric of distinction, then the numbers change. Look at early trends to explicitly name every type of ridge events to see what I'm talking about. Philosophically, this approach has never admitted to being arbitrary, but it is.This courtroom certainty is based purely on categorical personal opinion, formed through years of training and experience, but not on logic or scientific data.
Friction ridge examination has qualitative aspects which statistics are very incapable of quantifying. All my past posts including the fingerprint with the question mark core speak to the very fact that friction ridge examiners don't count points for a reason. That reason is that information can act qualitatively in abstract ways on us through representation. There is an aesthetic code we follow in addition to strictly Gestalt organizational perceptions we go through (See Ashabaugh). The red flags that Michele pointed out in this post from 2007 is an example of an explicit recognition of that viewtopic.php?f=2&t=592
Let me be very clear here, the value of information contained in an impression is not directly representational to its rarity in a population. I will go further and make an even more bold statement (think: conjectures and refutations) to say that a comparison is not a determination of rarity. Rarity gets established at the foundation of why comparisons are possible, it is not the diagnostic technique used to reach an evaluation. The determination of source origin is a determination of similarity. The notion of sign representation and signficance (aka value) is known by the study of semiology or visual semiotics. The fact that we call a delta a delta because it looks like a delta is as explicity an iconic representation as one can make (See Google: Visual Semiotics). Combine that with the notion of ACE as pragmatic vindication outlined in my Philosophy of Friction Ridge Examination vid and to say that evaluations are personal opinions not based on logic or scientific data is a reflection of ignorance on the author of that statement moreso than a reflection of a factual statement. These philosophic concepts are hundreds of years old now and flourish in a number of other disciplines than ours. The very real truth is that the so called 'leaders' in this discipline have been impotent in issues that matter and extremely effective in making sure the open bar in the Presidential Suite at the conference is stocked.
I've also got issue with the sole goal of research being repeatability, because repeatability is a complementary virtue of methodology not an end. Do you know what's also repeatable and accurate? Not comparing fingerprints. I will be accurate 100% of the time and you will be able to reproduce the result 100% of the time if I tell you I did not compare the fingerprints and my conclusion is that I did not compare them. Problem solved, now all that grant money can go to me to fund baldness research and the effects of homebrewing beer on a 39 year old man's stomach girth.
-
Gerald Clough
- Posts: 557
- Joined: Wed Jul 06, 2005 6:27 am
- Location: Lockhart, Texas
- Contact:
Re: Statistical Model Article
I give them some slack in the issue of what they say about the current state of examination. The "science" argument is one we have even among examiners. I don't fault a scientist for the opinion that the process doesn't pass scientific muster, that there's no factual scientific certainty involved. I have to agree with that. And, if they mistakenly, as I observed many scientists do, mistake the nature of the conclusion (something well understood where it matters, in court) for a statement of factual certainty, then they understandable see that as running counter to any logic. It plainly would.
You see, I can agree with the argument from scientific rigor that this sort of expertise produces no conclusion that would be accepted as scientifically valid. And at the same time, because I know the value of such expertise lies almost exclusively with in the legal system, I don't share their offense at it being attended to. Fingerprint evidence is very far to the good end of the scale compared to some other expert evidence.
There's a peculiar thing that latent print examination looks at casual glance like it ought to be properly a subject of scientific validation and improvement. If it is, for all the reasons we've talked about for years, it's going to be a very long time, if ever, that science creates any significant change in how it's used. This discipline really suffers in this regard for being observational and therefore subject to examination by students of perception and cognition. And it suffers from its subject being concrete physiological characteristics and their impressions - and further because it looks like the focus is largely on things that look countable. So it feels like something that science ought to be able to do a job on. But science is far short of the tools to take it on with any hope of predictable success at providing any additional value in court.
I think there is a truth that forensic disciplines get attention from hard science proportional to the facility with which scientists perceive they ought to be able to study it. In other words, they don't criticize, unless they think they ought to be able to make it better somehow. I compare to forensic psychology. Aside from some statistical wrangling over things like IQ score distributions and test validations, they don't catch much criticism from science. They deal with things that are rarely if ever concrete and only in some peripheral ways subject to scientific study. It's mostly agreed upon definitions and argument between experts on whether or not a subject fits one or more of the definitions. It's offensive to scientists that the latent print examiner makes statement that sound to the scientist like unsubstantiated truths. But it doesn't offend them that in almost every matter where they appear, two psychologists state opposing absolute conclusions. That, to the scientist, ought to look like at least one of them must be factually wrong. (They aren't, any more than an argument between LPE's is a factual argument.) But they don't complain, because there's so little for them to come to grips with that they couldn't hope to make a scientific study of it. It's off their radar. So far outside the realm that they don't see it.
But I don't worry about science and fingerprint evidence. The courts know, and they have more than once taken care of their fingerprint evidence, even when the fingerprint experts couldn't clearly explain. And I become ever more confident that statistical attacks on fingerprint evidence will continue to make attorneys' and judges' eyes glaze over. These guys haven't changed that. In a peculiar way, most of these efforts really apply to the forensic practice only in that they substantiate the value of fingerprint evidence.
You see, I can agree with the argument from scientific rigor that this sort of expertise produces no conclusion that would be accepted as scientifically valid. And at the same time, because I know the value of such expertise lies almost exclusively with in the legal system, I don't share their offense at it being attended to. Fingerprint evidence is very far to the good end of the scale compared to some other expert evidence.
There's a peculiar thing that latent print examination looks at casual glance like it ought to be properly a subject of scientific validation and improvement. If it is, for all the reasons we've talked about for years, it's going to be a very long time, if ever, that science creates any significant change in how it's used. This discipline really suffers in this regard for being observational and therefore subject to examination by students of perception and cognition. And it suffers from its subject being concrete physiological characteristics and their impressions - and further because it looks like the focus is largely on things that look countable. So it feels like something that science ought to be able to do a job on. But science is far short of the tools to take it on with any hope of predictable success at providing any additional value in court.
I think there is a truth that forensic disciplines get attention from hard science proportional to the facility with which scientists perceive they ought to be able to study it. In other words, they don't criticize, unless they think they ought to be able to make it better somehow. I compare to forensic psychology. Aside from some statistical wrangling over things like IQ score distributions and test validations, they don't catch much criticism from science. They deal with things that are rarely if ever concrete and only in some peripheral ways subject to scientific study. It's mostly agreed upon definitions and argument between experts on whether or not a subject fits one or more of the definitions. It's offensive to scientists that the latent print examiner makes statement that sound to the scientist like unsubstantiated truths. But it doesn't offend them that in almost every matter where they appear, two psychologists state opposing absolute conclusions. That, to the scientist, ought to look like at least one of them must be factually wrong. (They aren't, any more than an argument between LPE's is a factual argument.) But they don't complain, because there's so little for them to come to grips with that they couldn't hope to make a scientific study of it. It's off their radar. So far outside the realm that they don't see it.
But I don't worry about science and fingerprint evidence. The courts know, and they have more than once taken care of their fingerprint evidence, even when the fingerprint experts couldn't clearly explain. And I become ever more confident that statistical attacks on fingerprint evidence will continue to make attorneys' and judges' eyes glaze over. These guys haven't changed that. In a peculiar way, most of these efforts really apply to the forensic practice only in that they substantiate the value of fingerprint evidence.
"Nothing has any value, unless you know you can give it up."
-
ER
- Posts: 351
- Joined: Tue Dec 18, 2007 3:23 pm
- Location: USA
Re: Statistical Model Article
I have a real problem with Cedric's charaterization of our discipline. I think what it really comes down to is that without the foundation of thousands upon thousands of comparisons, he doesn't truly understand what we do. He thinks he does, so he writes papers like this. While I applaud his work on the model and believe that it will eventually be an important tool for a latent print examiner, I don't believe that it can ever take the place of the comparison. As Boyd states, there are qualitative aspects to fingerprint comparison that cannot be accounted for in the model. The infinite variation of a fingerprint combined with the infinite variation of a touch will always leave a statistical model lacking in some regard. The best we can hope for is that the conclusion of an examiner (identification, exclusion, or inconclusive) will have additional information provided from the model. We will never reach a point where the model produces the conclusion. This is what it seems like Cedric has been striving for. His model, however, must always remain a tool to help support the examiner's conclusion because his model will never outperform an examiner. While he derides the examiner's 'subjective' conclusion, he fails to reveal the inherent subjectivity of his model. He also incorrectly describes all of our conclusions as statements of absolute certainty, which they are not. Latent print examiner's state (or at least should be stating) their conclusions in terms of certainty that are supported by the current research, that is, certain to the degree that the compared prints (data) and the comparison process (logical process) dictate. This consistently results in extremely reliable conclusions that Cedric has yet to demonstrate he can surpass.
-
Cedric
- Posts: 12
- Joined: Wed Aug 01, 2007 2:24 pm
- Location: Penn State
Re: Statistical Model Article
ER and Gerald,
Thank you for your (mostly) kind posts.
Between the two of you, I think that you properly nailed most of our intentions:
1) The main purpose of this trend of research is not to create some model that will make decisions instead of the examiners. We have made that statement over and over again: the statistical model is only a tool that aims at providing more objectivity to the examination process and help demonstrate the claims of the community when challenged in court. The aim is definitely not to undermine, belittle, or provide arguments to defense attorneys looking for a cheap thrill.
2) Indeed, most of the efforts are aimed at substantiating the value of fingerprint evidence. None of the research is aimed at designing something that will outperform examiners. It simply cannot in terms of feature detection and comparison. What that tool can do is to process in a few seconds/minutes the lifetime amount of fingerprint records that experienced examiners have in their heads and generate some hard data on the rarity of a given set of friction ridge features. Something that examiners may struggle to quantify objectively.
3) Yes, indeed, the model is an imperfect representation of reality. It has never been intended otherwise. And it will never capture all the features that examiners use. The human brain is, and will remain for the foreseeable future, superior when it comes to pattern recognition and comparison. No doubt about that. Similarly, the model is based on many assumptions that are clearly labelled as such (read the papers). But as we have said it many times, building such model is a step towards more transparency in the way we form our conclusions. It is not meant to form the conclusions by itself, but it is meant to pave some of the way. The human examiner will have to walk the rest by him/herself.
4) And definitely, there is a clear difference between the "weight of the evidence" and the "error rate". An examiner may perfectly associate a latent print with the correct source, while at the same time overstate the evidence by claiming that nobody else than the identified source could have left that latent print. In such cases, the association is perfectly correct, but the weight of the evidence is misrepresented. It is not because a model say that "there is a 1% chance to observe somebody else in the World than the putative source with features similar to the latent print" that "there is a 1% chance that the examiner in the case has done an error". These elements are somewhat connected, but it is not as straightforward. Bottom line, the model is not a measure of the error rate, and I agree with all the data that have shown that fingerprint examiners consistently result in extremely reliable conclusions.
5) And obviously we realize that most examiners (should) express a personal belief when they report a conclusion, as opposed to a scientific fact. That said, even personal believes may overstate the weight of particular comparisons.
6) Finally, of course we realize that change takes time. But that said, if i wasn't slightly optimistic about what I was doing, i wouldn't see any point of doing it... and in order for change to happen, even over the course of a generation, we need to start at some point. Read the FSI paper on the field study done in MN, you will see that the model has some benefits, at least to support current identifications and for quantifying the weight on current inconclusives.
Cedric
Thank you for your (mostly) kind posts.
Between the two of you, I think that you properly nailed most of our intentions:
1) The main purpose of this trend of research is not to create some model that will make decisions instead of the examiners. We have made that statement over and over again: the statistical model is only a tool that aims at providing more objectivity to the examination process and help demonstrate the claims of the community when challenged in court. The aim is definitely not to undermine, belittle, or provide arguments to defense attorneys looking for a cheap thrill.
2) Indeed, most of the efforts are aimed at substantiating the value of fingerprint evidence. None of the research is aimed at designing something that will outperform examiners. It simply cannot in terms of feature detection and comparison. What that tool can do is to process in a few seconds/minutes the lifetime amount of fingerprint records that experienced examiners have in their heads and generate some hard data on the rarity of a given set of friction ridge features. Something that examiners may struggle to quantify objectively.
3) Yes, indeed, the model is an imperfect representation of reality. It has never been intended otherwise. And it will never capture all the features that examiners use. The human brain is, and will remain for the foreseeable future, superior when it comes to pattern recognition and comparison. No doubt about that. Similarly, the model is based on many assumptions that are clearly labelled as such (read the papers). But as we have said it many times, building such model is a step towards more transparency in the way we form our conclusions. It is not meant to form the conclusions by itself, but it is meant to pave some of the way. The human examiner will have to walk the rest by him/herself.
4) And definitely, there is a clear difference between the "weight of the evidence" and the "error rate". An examiner may perfectly associate a latent print with the correct source, while at the same time overstate the evidence by claiming that nobody else than the identified source could have left that latent print. In such cases, the association is perfectly correct, but the weight of the evidence is misrepresented. It is not because a model say that "there is a 1% chance to observe somebody else in the World than the putative source with features similar to the latent print" that "there is a 1% chance that the examiner in the case has done an error". These elements are somewhat connected, but it is not as straightforward. Bottom line, the model is not a measure of the error rate, and I agree with all the data that have shown that fingerprint examiners consistently result in extremely reliable conclusions.
5) And obviously we realize that most examiners (should) express a personal belief when they report a conclusion, as opposed to a scientific fact. That said, even personal believes may overstate the weight of particular comparisons.
6) Finally, of course we realize that change takes time. But that said, if i wasn't slightly optimistic about what I was doing, i wouldn't see any point of doing it... and in order for change to happen, even over the course of a generation, we need to start at some point. Read the FSI paper on the field study done in MN, you will see that the model has some benefits, at least to support current identifications and for quantifying the weight on current inconclusives.
Cedric
Cedric
-
ER
- Posts: 351
- Joined: Tue Dec 18, 2007 3:23 pm
- Location: USA
Re: Statistical Model Article
Cedric,
My original comments were based on the summary article from the Royal Statistical Society: "Fingerprints at the crime-scene: Statistically certain, or probable?". I've now also read your paper: "Quantifying the weight of evidence from a forensic fingerprint comparison: a new paradigm".
I am very happy to see your paper published. It is a tremendous step forward in the science of fingerprint comparisons, and it makes me proud as a latent print examiner to see our discipline represented in such a prestigious journal.
Of particular note are the results that show the scientific foundation of fingerprint evidence. Especially, that there is a distinct difference in the LR of same-source prints when compared to the LR of different-source prints. Also, that this LR gap grows larger as more minutia are compared.
Ever since I first heard of the research into these statistical models, I've always envisioned it as an eventual tool to support the decision of the examiner. An opinion of identification, inconclusive, or exclusion is rendered, and the model gives a LR to support that decision. The LR may even cause the examiner to take another look at a comparison, but the examiner still makes the decision. Your post above seems to agree with me on this. However, your article says the opposite.
But, and this is a big but, the model won't work for the vast majority of these prints. There is a small percentage of inconclusive results where the examiner found corresponding features, but there was insufficient detail for an identification. I come across, maybe, one a month. This is not 30% of comparisons, but rather 1%, at best. The remaining inconclusive results are inconclusive because the exemplars were incomplete or the latent mark was too distorted or too small or lacked any indication of its anatomical origin. Without corresponding features, the model won't work.
Your limitations include: not considering discordant minutiae, feature labeling inconsistencies, not considering all distortion possibilities, quality of the latent print, arbitrariness of the formula, and relying on a limited reference database, and an apparent limit to the LR on higher minutiae comparisons. Considering all these, the model's performance is exciting and promising.
Some of my questions are: Would more difficult comparisons lead to more ambiguous results? The 'close non-matches' in this study didn't seem to give experts that much difficulty. What would it take to ensure close non-matches of lower quantity and quality could still produce the distinct separation of LR values? Why don't the LR values increase exponentially with each added minutiae? Are ridge counts going to replace distance measurements in the model soon?
Please don't misinterpret this (extremely) long post as an attack. I pore over papers like this and devour them with an intense delight. (Even though the equations make my eyes gloss over.) Like I said earlier, I think this is great, except for the implication that the categorical opinion should be eliminated.
My original comments were based on the summary article from the Royal Statistical Society: "Fingerprints at the crime-scene: Statistically certain, or probable?". I've now also read your paper: "Quantifying the weight of evidence from a forensic fingerprint comparison: a new paradigm".
I am very happy to see your paper published. It is a tremendous step forward in the science of fingerprint comparisons, and it makes me proud as a latent print examiner to see our discipline represented in such a prestigious journal.
Of particular note are the results that show the scientific foundation of fingerprint evidence. Especially, that there is a distinct difference in the LR of same-source prints when compared to the LR of different-source prints. Also, that this LR gap grows larger as more minutia are compared.
Ever since I first heard of the research into these statistical models, I've always envisioned it as an eventual tool to support the decision of the examiner. An opinion of identification, inconclusive, or exclusion is rendered, and the model gives a LR to support that decision. The LR may even cause the examiner to take another look at a comparison, but the examiner still makes the decision. Your post above seems to agree with me on this. However, your article says the opposite.
Now, while your article never comes right out and calls for an elimination of the categorical opinion, completely replacing it with a likelihood ratio, the paper does certainly leave the impression that the categorical opinion is antiquated and needs to be eliminated. I hope that you continue to improve upon the statistical model to help make categorical opinions more robust with the understanding that they can never be replaced. As you stated, the model is an imperfect representation of reality. Since this will always be the case, the need for a categorical opinion based on the expertise of the examiner and supported by a statistical model is clear."In the immediate future, we do not see that the current practice of presenting categorical opinions will change. But longer term we expect an evolution towards a framework that is similar to that which underpins DNA evidence.”
“Ultimately, we see this kind of approach (statistics) replacing the existing paradigm.”
I really wish that a statement like this appeared in the paper. The implication from the paper, and especially from the article co-written by Julian Champkin, is that the model will replace the categorical opinion.It is not meant to form the conclusions by itself, but it is meant to pave some of the way. The human examiner will have to walk the rest by him/herself.
While this may be what Saks, Koehler, and Cole are saying, it is far from the truth. I will not claim absolute certainty. While it may have been common just a few years ago, the IAI and SWGFAST have both backed away from this overstatement. But you know this. It's unsettling to see this fallacy repeated over and over in this article. Far from being forbidden, the IAI's resolution allows statements of probability so long as the model has been validated and accepted.“Fingerprint experts are required to claim absolute certainty for their judgements.”
“A fingerprint expert will report that he or she is absolutely certain that an impression from the scene of a crime comes from the finger of the accused.”
“If it is his belief that the mark 'probably does' or almost certainly does' or 'is rather unlikely to' match, he is forbidden to say so in court.”
“A fingerprint expert will tell the court that he or she is absolutely certain that a particular mark was made by the person who provided the control print to the exclusion of all other individuals on Earth.”
I understand the underlying concept behind this quote. Fingerprints are unique on people's fingers, but they can look similar and fool the examiner after being transferred to a surface and processed and compared. However, to state that the uniqueness of fingerprints then becomes irrelevant to the examination is a gross oversimplification of an important underpinning of this science.“The claim of the uniqueness of fingerprints becomes then irrelevant to fingerprint examinations and even more so to the formation of the conclusions.”
That's only part of the truth. While the model is currently limited to the minutiae points and the distances between them, reliable conclusions are formed from comparing the entirety of the print, the minutiae, the unit relationships between minutiae, the spaces in between and the fine details of each ridge unit. Granted, the model is still in its infancy and may expand to encompass all of these areas, but the paper should be clear on the current practices of fingerprint comparisons.“These points of detail are called minutiae, or points, and it is by comparing the minutiae in a print and a trace that examiners form their opinions.”
You go on to state that this is a major reason to begin implementation of the model in casework. The use of statistics would give meaningful numbers to this large proportion of comparisons that are currently on reported as 'inconclusive': or that this “potentially useful evidence must be discarded”.“There is a third class of comparisons: those where the examiner considers, for various reasons, that the evidence is of insufficient quality, or comprises too little detail for an opinion to be formed. This means that there is a class of cases where evidence which might be of corroborative value is not taken further because the examiner considers it to be insufficient for an opinion of complete certainty.”
“This happens is an many as 30% of the comparisons performed in a fingerprint bureau.”
But, and this is a big but, the model won't work for the vast majority of these prints. There is a small percentage of inconclusive results where the examiner found corresponding features, but there was insufficient detail for an identification. I come across, maybe, one a month. This is not 30% of comparisons, but rather 1%, at best. The remaining inconclusive results are inconclusive because the exemplars were incomplete or the latent mark was too distorted or too small or lacked any indication of its anatomical origin. Without corresponding features, the model won't work.
Your limitations include: not considering discordant minutiae, feature labeling inconsistencies, not considering all distortion possibilities, quality of the latent print, arbitrariness of the formula, and relying on a limited reference database, and an apparent limit to the LR on higher minutiae comparisons. Considering all these, the model's performance is exciting and promising.
Some of my questions are: Would more difficult comparisons lead to more ambiguous results? The 'close non-matches' in this study didn't seem to give experts that much difficulty. What would it take to ensure close non-matches of lower quantity and quality could still produce the distinct separation of LR values? Why don't the LR values increase exponentially with each added minutiae? Are ridge counts going to replace distance measurements in the model soon?
Please don't misinterpret this (extremely) long post as an attack. I pore over papers like this and devour them with an intense delight. (Even though the equations make my eyes gloss over.) Like I said earlier, I think this is great, except for the implication that the categorical opinion should be eliminated.
-
Cedric
- Posts: 12
- Joined: Wed Aug 01, 2007 2:24 pm
- Location: Penn State
Re: Statistical Model Article
ER,
We need to get a beer (for you) and a black russian (for me) next time I am in your area of the country (or that you are in mine). This will give us time to talk about all your questions.
In the meantime, one thing: I support the move away from the tripartite categorical opinion scheme, and this move needs to be supported by some sort of method to quantify the weight of fingerprint evidence (e.g. a statistical model, but also charts such as the one in the new SWGFAST standard). But that doesn't mean that the method will make the decisions and form the conclusions instead of the examiners. I still argue that the method is just a tool to provide data to the examiners when they form their conclusions.
The said conclusion can then be anything from the current tripartite categorical conclusion scheme if that's what we want based on QA and operational consideration, through a 174 categories scheme, to a completely continuous reporting scheme. We may indeed end up reporting just a number (the LR), but (a) that's not necessarily the best strategy (we are setting up a group to discuss the various possibilities), (b) this won't happen anytime soon. Nevertheless, it is certain that the method supporting our conclusions will need to be encapsulated into a logical framework, and that so far, the best framework we have is the one that underpins DNA evidence.
If the above seems confusing, the main message is: the reporting scheme and the decision-making process are two different things - we can move away from the categorical reporting scheme, without removing the human participation in the decision-making process.
Cedric
PS: the paper was originally written in 2009, before the IAI and SWGFAST implemented their latest recommendations. The delay in the publication is due to the special peer-reviewing process that we went through.
We need to get a beer (for you) and a black russian (for me) next time I am in your area of the country (or that you are in mine). This will give us time to talk about all your questions.
In the meantime, one thing: I support the move away from the tripartite categorical opinion scheme, and this move needs to be supported by some sort of method to quantify the weight of fingerprint evidence (e.g. a statistical model, but also charts such as the one in the new SWGFAST standard). But that doesn't mean that the method will make the decisions and form the conclusions instead of the examiners. I still argue that the method is just a tool to provide data to the examiners when they form their conclusions.
The said conclusion can then be anything from the current tripartite categorical conclusion scheme if that's what we want based on QA and operational consideration, through a 174 categories scheme, to a completely continuous reporting scheme. We may indeed end up reporting just a number (the LR), but (a) that's not necessarily the best strategy (we are setting up a group to discuss the various possibilities), (b) this won't happen anytime soon. Nevertheless, it is certain that the method supporting our conclusions will need to be encapsulated into a logical framework, and that so far, the best framework we have is the one that underpins DNA evidence.
If the above seems confusing, the main message is: the reporting scheme and the decision-making process are two different things - we can move away from the categorical reporting scheme, without removing the human participation in the decision-making process.
Cedric
PS: the paper was originally written in 2009, before the IAI and SWGFAST implemented their latest recommendations. The delay in the publication is due to the special peer-reviewing process that we went through.
Cedric
-
ER
- Posts: 351
- Joined: Tue Dec 18, 2007 3:23 pm
- Location: USA
Re: Statistical Model Article
Sounds good. Hopefully, I'll see you at the conference this summer.
In very difficult comparisons that rely on clear 3rd level detail with limited 2nd level detail and a little distortion, an LR will always understate the weight of the evidence.
I believe that the examiners will always have to choose between identification, exclusion, or inconclusive. But, we may get to the point of having different flavors of inconclusive (we currently have 3 or 4). Is this what you mean by moving away from a tripartite scheme?
In very difficult comparisons that rely on clear 3rd level detail with limited 2nd level detail and a little distortion, an LR will always understate the weight of the evidence.
I believe that the examiners will always have to choose between identification, exclusion, or inconclusive. But, we may get to the point of having different flavors of inconclusive (we currently have 3 or 4). Is this what you mean by moving away from a tripartite scheme?
-
Cedric
- Posts: 12
- Joined: Wed Aug 01, 2007 2:24 pm
- Location: Penn State
Re: Statistical Model Article
yes, potentially
And yes, a LR that focuses on 2nd level details will not capture the evidential value of 3rd level features, and therefore will probably underestimate the total value of the comparison.
And yes, a LR that focuses on 2nd level details will not capture the evidential value of 3rd level features, and therefore will probably underestimate the total value of the comparison.
Cedric
-
Gerald Clough
- Posts: 557
- Joined: Wed Jul 06, 2005 6:27 am
- Location: Lockhart, Texas
- Contact:
Re: Statistical Model Article
Cedric,
I guess I should explain why I think the effort is largely futile. (Not that any inquiry is inherently futile. You never know what you'll discover.)
If one is seeking a credible method of supporting examiner conclusions, the first requirement is to examine what determines reliability. There are two mechanisms operating in reaching the conclusions. These are perceiving and characterizing minute aspects of the impression and judging whether sufficient characterized detail is properly present in both samples to declare them as from the same source. I as I read the study, it addresses the likelihood of error on account of erroneously declaring a unique source. But in the real world of forensic identification, the possibility of correctly characterizing minutiae but incorrectly concluding as to match is extremely remote and rarely considered.
It is true, of course, that an examiner might be found in any case to oppose a conclusion on the grounds that the examiner was too liberal or too conservative. And if fingerprint identification was a more recently developed field, we would likely see more of that opposition in court. But although we can never defy logic to declare that some specific degree of agreement assures a correct match to the absolute exclusion of all others, the examiner community establishes an undefined but high threshold for agreement. It is apparent that the threshold is sufficiently high that we've never found it insufficient. Nothing says we will never find two impressions from different sources that all examiners will mistakenly identify as a single source, but neither does a comparison run against a large print database preclude that error. And, since the cards have no memory, no statistical result can see that momentous pair of sources approaching.. No method can ever say more than that the community standard is very reliably high.
But there are errors. And they are important. They are important both because they demonstrate that an examiner can be dead wrong. And they are important because they show us that, if an examiner is wrong, it's almost certainly because of incorrect characterization of an impression, almost always the latent impression. If we exclude administrative mistakes, such as comparing to a mislabeled impression, the revealed errors are typically those presumed to have been made in notorious cases like Mayfield. The examiner simply was wrong about what physical skin characteristics he was observing in an impression. That is important, because the threshold for characterization of minutiae is different in nature from the threshold for identification. The identification threshold is essential numerical, setting a vague but high number of points of agreement, however that number might shift from one examination to another as the natures of the impressions and the confidence in the minutiae observations vary.
The threshold for confidence in the character of a detail, though, has several differences. For one, unlike the fairly consistent community threshold for identification, the threshold for confident characterization is subject to substantial differences on account of experience and individual nature of the examiner. I think this explains why errors are often seen to be made by the most experienced and respected examiners. They believe they can "make the hard ones." The less experienced examiner does not possess that confidence born or experience and a long career of successes with difficult cases. And there is a very natural tendency of human observers to built an observation of many details of probably but less than absolutely confident characterizations into an overall confidence higher than any single detail would present in isolation. The longer you're on the trail, the more you believe it's going where you think it's going.
Now, it is never possible, when looking at a single characteristic in an impression, to say it is absolutely the impression of precisely such-and-such skin variation. For that matter, there is no absolute point at which one skin structure becomes another and even less assurance that it was impressed fully and accurately. Every decision as to character of an impressed detail you decide to use is a best guess.
If a statistical argument is made after the fashion of the study, and if it is to be of any use to a legal finder of fact, it must clearly tend to make it more likely that the correct conclusion of two that are offered in opposition to each other will be accepted. It cannot do that if it merely adds an additional layer of potential fog. In other words, it should not simply be something else to argue over. I do not mean argument over the validity of the studies, but rather argument over whether the application has produced a reliable result. Can I trust the application to reveal error? But if you look at the revealed errors, I think it's apparent that, had the latent been correctly interpreted, the proposed method would have suggested confirmation of the erroneous identification. What the examiner believed the latent to be was in complete agreement with the skin of the identified person. The errors were revealed when others made different characterizations of the latent impressions that were not in agreement. The fact that the proposed method supported the initial examination is entirely moot. If applied using the new characterizations, it would then support the revised conclusion of non-identification. In short, the proposed method is itself unreliable as an aid to the fact-finder's task, because it cannot consider the most common source of error.
Fingerprint evidence has some unusual characteristics in practical litigation. In many cases, the fact-finder is not judging whether the expert or their proxy observed correctly. Rather, they as, to put it colloquially, looking at the right stuff. No one doubts that the offender trying to escape the death sentence by claiming to be mentally retarded was a wretched student. No one doubts the teachers accurately recorded his miserable grades. But the fact-finder is more interested in whether he failed because he was incapable of learning or if he failed because he only attended half the school days. One expert is citing the poor grades as evidence he was mentally deficient. The other is ignoring it. With fingerprint evidence, no one contests that each and every detail is meaningful. The contest is over whether they are what each expert says they can reliably be taken as.
The reason I talk about two experts in opposition to each other is that measures such as that proposed do not even enter the legal arena unless there are two experts offering different opinions. Don't mistake this as being similar to DNA. DNA is always a probabilistic conclusion. Even a single expert must inevitably offer their conclusion in absolute numerical terms. An opposing expert may argue assumptions about populations or presumptive likelihoods in paternity. They are argument of what the numerical conclusion means, not whether the process properly characterized he DNA profile. An opposing fingerprint expert will most often argue the accuracy of the observed characteristics of details, not the threshold for identification, because the threshold is set so high that in most cases, lay jurors can understand why the examiner found a match. But opposing fingerprint experts in cases are pretty rate to begin with.
I do think there is some significant value to be had from this general sort of study, though. I do not believe it will ever be possible to avoid the human interpreter's role in characterizing the impressed details. (I think it may be possible that it might one distant day be reliably automated, but I think it's going to take a great deal of analysis of patterns of friction ridge development to bring that automation to the skill level of a human examiner.) But where I think this kind of study may be most useful is in guiding examiners. For example, I think it might be helpful to know how the LR varies when different characteristics are characterized differently from the same portion of the source skin. As I accumulate less than absolutely certain interpretations of details, to what degree am I likely to be affecting the calculated confidence in the conclusion? I can't have absolute numbers here, but I'm interested in the nature of the curve, how quickly the conclusion becomes unreliable. I think there is only the most vague notion of how the likelihood of error accumulates with uncertainties about individual characteristics. Perhaps you can proceed with that analysis. Next week with be fine.
I guess I should explain why I think the effort is largely futile. (Not that any inquiry is inherently futile. You never know what you'll discover.)
If one is seeking a credible method of supporting examiner conclusions, the first requirement is to examine what determines reliability. There are two mechanisms operating in reaching the conclusions. These are perceiving and characterizing minute aspects of the impression and judging whether sufficient characterized detail is properly present in both samples to declare them as from the same source. I as I read the study, it addresses the likelihood of error on account of erroneously declaring a unique source. But in the real world of forensic identification, the possibility of correctly characterizing minutiae but incorrectly concluding as to match is extremely remote and rarely considered.
It is true, of course, that an examiner might be found in any case to oppose a conclusion on the grounds that the examiner was too liberal or too conservative. And if fingerprint identification was a more recently developed field, we would likely see more of that opposition in court. But although we can never defy logic to declare that some specific degree of agreement assures a correct match to the absolute exclusion of all others, the examiner community establishes an undefined but high threshold for agreement. It is apparent that the threshold is sufficiently high that we've never found it insufficient. Nothing says we will never find two impressions from different sources that all examiners will mistakenly identify as a single source, but neither does a comparison run against a large print database preclude that error. And, since the cards have no memory, no statistical result can see that momentous pair of sources approaching.. No method can ever say more than that the community standard is very reliably high.
But there are errors. And they are important. They are important both because they demonstrate that an examiner can be dead wrong. And they are important because they show us that, if an examiner is wrong, it's almost certainly because of incorrect characterization of an impression, almost always the latent impression. If we exclude administrative mistakes, such as comparing to a mislabeled impression, the revealed errors are typically those presumed to have been made in notorious cases like Mayfield. The examiner simply was wrong about what physical skin characteristics he was observing in an impression. That is important, because the threshold for characterization of minutiae is different in nature from the threshold for identification. The identification threshold is essential numerical, setting a vague but high number of points of agreement, however that number might shift from one examination to another as the natures of the impressions and the confidence in the minutiae observations vary.
The threshold for confidence in the character of a detail, though, has several differences. For one, unlike the fairly consistent community threshold for identification, the threshold for confident characterization is subject to substantial differences on account of experience and individual nature of the examiner. I think this explains why errors are often seen to be made by the most experienced and respected examiners. They believe they can "make the hard ones." The less experienced examiner does not possess that confidence born or experience and a long career of successes with difficult cases. And there is a very natural tendency of human observers to built an observation of many details of probably but less than absolutely confident characterizations into an overall confidence higher than any single detail would present in isolation. The longer you're on the trail, the more you believe it's going where you think it's going.
Now, it is never possible, when looking at a single characteristic in an impression, to say it is absolutely the impression of precisely such-and-such skin variation. For that matter, there is no absolute point at which one skin structure becomes another and even less assurance that it was impressed fully and accurately. Every decision as to character of an impressed detail you decide to use is a best guess.
If a statistical argument is made after the fashion of the study, and if it is to be of any use to a legal finder of fact, it must clearly tend to make it more likely that the correct conclusion of two that are offered in opposition to each other will be accepted. It cannot do that if it merely adds an additional layer of potential fog. In other words, it should not simply be something else to argue over. I do not mean argument over the validity of the studies, but rather argument over whether the application has produced a reliable result. Can I trust the application to reveal error? But if you look at the revealed errors, I think it's apparent that, had the latent been correctly interpreted, the proposed method would have suggested confirmation of the erroneous identification. What the examiner believed the latent to be was in complete agreement with the skin of the identified person. The errors were revealed when others made different characterizations of the latent impressions that were not in agreement. The fact that the proposed method supported the initial examination is entirely moot. If applied using the new characterizations, it would then support the revised conclusion of non-identification. In short, the proposed method is itself unreliable as an aid to the fact-finder's task, because it cannot consider the most common source of error.
Fingerprint evidence has some unusual characteristics in practical litigation. In many cases, the fact-finder is not judging whether the expert or their proxy observed correctly. Rather, they as, to put it colloquially, looking at the right stuff. No one doubts that the offender trying to escape the death sentence by claiming to be mentally retarded was a wretched student. No one doubts the teachers accurately recorded his miserable grades. But the fact-finder is more interested in whether he failed because he was incapable of learning or if he failed because he only attended half the school days. One expert is citing the poor grades as evidence he was mentally deficient. The other is ignoring it. With fingerprint evidence, no one contests that each and every detail is meaningful. The contest is over whether they are what each expert says they can reliably be taken as.
The reason I talk about two experts in opposition to each other is that measures such as that proposed do not even enter the legal arena unless there are two experts offering different opinions. Don't mistake this as being similar to DNA. DNA is always a probabilistic conclusion. Even a single expert must inevitably offer their conclusion in absolute numerical terms. An opposing expert may argue assumptions about populations or presumptive likelihoods in paternity. They are argument of what the numerical conclusion means, not whether the process properly characterized he DNA profile. An opposing fingerprint expert will most often argue the accuracy of the observed characteristics of details, not the threshold for identification, because the threshold is set so high that in most cases, lay jurors can understand why the examiner found a match. But opposing fingerprint experts in cases are pretty rate to begin with.
I do think there is some significant value to be had from this general sort of study, though. I do not believe it will ever be possible to avoid the human interpreter's role in characterizing the impressed details. (I think it may be possible that it might one distant day be reliably automated, but I think it's going to take a great deal of analysis of patterns of friction ridge development to bring that automation to the skill level of a human examiner.) But where I think this kind of study may be most useful is in guiding examiners. For example, I think it might be helpful to know how the LR varies when different characteristics are characterized differently from the same portion of the source skin. As I accumulate less than absolutely certain interpretations of details, to what degree am I likely to be affecting the calculated confidence in the conclusion? I can't have absolute numbers here, but I'm interested in the nature of the curve, how quickly the conclusion becomes unreliable. I think there is only the most vague notion of how the likelihood of error accumulates with uncertainties about individual characteristics. Perhaps you can proceed with that analysis. Next week with be fine.
"Nothing has any value, unless you know you can give it up."
-
ER
- Posts: 351
- Joined: Tue Dec 18, 2007 3:23 pm
- Location: USA
Re: Statistical Model Article
I think this is the key to your post. While the confidence in each individual minutiae may vary depending on a wide variety of factors, a model that can produce a reliable numerical representation of that 'vague but high number of points' would be a damn useful tool. Not just for guiding the examiner in calibrating things like, how many more points do you need on an ID in the tail of a loop, but in providing more detail to the court on the clear, small latent that is now just reported as inconclusive.The identification threshold is essential numerical, setting a vague but high number of points of agreement, however that number might shift from one examination to another as the natures of the impressions and the confidence in the minutiae observations vary.
Now, there is separate research into measuring the reliability of the minutia position. That seems to be headed in a direction where only the most reliable points are used in a comparison. To me, that seems to be falling to the least common denominator where you can't use a feature in your comparison unless everyone in the room can see it. I don't like the idea of giving up the ID's that I can reliably make just because other people can't see the points.
In any case, I think the point of this research (and Cedric, correct me if I'm wrong) is not to give the court a number to tell them which of two opposing conclusions is the right one, but to describe the ID in relation to other ID's. The weight of the evidence. Each comparison and each ID is different. You would, in general, have extreme confidence in your average 12 point ID. You would, in general, have even more confidence in a 24 point ID of the same general clarity. The court seems to want some numerical representation of how good the ID is. Currently, we can tell them a number of points, but that, obviously, is only part of the story. The research already shows that the number generally gets better as you add more points, but that it can vary a lot with the same amount of points. We already knew that... but it's nice to see statistics to back it up.
-
Gerald Clough
- Posts: 557
- Joined: Wed Jul 06, 2005 6:27 am
- Location: Lockhart, Texas
- Contact:
Re: Statistical Model Article
I admit to a recurring impulse to insist that no detail can be used unless it can be characterized with the same confidence required to declare an identification. In those moments, it seems to me that if you have less than that kind of certainty in the details, you're accumulating likelihood of errors in what you believe to be the skin features that impressed them. But I then consider two things. One is that neither actual skin features nor the impressions they can make are always one thing. They are all really defined as tendencies. A bifurcation is two ridges approaching one ridge with a tendency to cojoin. At some degree of close inspection, it might equally be a bifurcation or a ridge and an adjacent ending ridge. And the impressions may more often be even less sure. So, it can just as well be called something like a "conjunction-like structure," if you had to name it. Not that we have to name it. We see it and can easily point it out and describe it. The point is that as you continuously vary the form of such a structure, you can't say where it crosses the line between bifurcation and ending ridge. So you're making a judgment that the skin that made it will be a lot like the impression. And it may not. They may just have common tendencies. The point is that you can never assume you've made no error at all and have seen in the impression exactly what you would see in the skin, down to the unit level. Some degree of error is inevitable. The question is how much is tolerable.
The other thing is that we don't really know much about the relationship between accurate characterization of details and the effect on conclusions of different degrees of certainty. For that matter, we don't understand much about the degree of certainty in the conclusion itself and the truth. There certainly are different degrees of certainty. We conclude by saying, in effect, we're going to bet the farm on this conclusion as reflecting truth, but it's still a bit, a guess. And we make it without knowing the odds. But what matters is that, so far as anyone knows, WE ALWAYS WIN THAT BET. So far as anyone ever determines the truth, they agree we won and get to keep the farm. And that's because we bet by the seat of the pants, but we bet so conservatively that, without knowing the odds, our conservatism puts us way on the right side of chance. We are proven to be, by virtue of training, experience, and community experience, safe guessers. I suppose we're not necessarily extraordinarily good guessers. If we were all that good at guessing, we be guessing right and making more identifications with less to go on. But our virtue and the key to our success is that we KNOW we don't know enough to make close call guesses. We win by never chancing a loss.
We're doing it right. Or we certainly know how to do it right, even though once in a great while, one of use gets sloppy about how we interpret an impression. But with the right data in hand, we're fine. The real question about research of this kind in my mind is how and if it can be applied. The only application I can see is some potential to influence the degree of certainty an examiner adopts before declaring an identification. If you show me that some situation that I find I can recognize when it appears in an examination is very close to probability 1.0, will I use that information to declare an identification that I wouldn't before. Just for a dumb example, what if the research shows that cascading series of four ending ridges virtually assures a match?
________________________________________
________
___________________________________________
______________
_________________________________________
____________________
__________________________________________
____________________________
___________________________________________
And all I can see in the latent is that portion containing those ending ridges. Do I declare an identification, supported by the statistical study? I frankly don't know. That's a real paradigm change.
The other thing is that we don't really know much about the relationship between accurate characterization of details and the effect on conclusions of different degrees of certainty. For that matter, we don't understand much about the degree of certainty in the conclusion itself and the truth. There certainly are different degrees of certainty. We conclude by saying, in effect, we're going to bet the farm on this conclusion as reflecting truth, but it's still a bit, a guess. And we make it without knowing the odds. But what matters is that, so far as anyone knows, WE ALWAYS WIN THAT BET. So far as anyone ever determines the truth, they agree we won and get to keep the farm. And that's because we bet by the seat of the pants, but we bet so conservatively that, without knowing the odds, our conservatism puts us way on the right side of chance. We are proven to be, by virtue of training, experience, and community experience, safe guessers. I suppose we're not necessarily extraordinarily good guessers. If we were all that good at guessing, we be guessing right and making more identifications with less to go on. But our virtue and the key to our success is that we KNOW we don't know enough to make close call guesses. We win by never chancing a loss.
We're doing it right. Or we certainly know how to do it right, even though once in a great while, one of use gets sloppy about how we interpret an impression. But with the right data in hand, we're fine. The real question about research of this kind in my mind is how and if it can be applied. The only application I can see is some potential to influence the degree of certainty an examiner adopts before declaring an identification. If you show me that some situation that I find I can recognize when it appears in an examination is very close to probability 1.0, will I use that information to declare an identification that I wouldn't before. Just for a dumb example, what if the research shows that cascading series of four ending ridges virtually assures a match?
________________________________________
________
___________________________________________
______________
_________________________________________
____________________
__________________________________________
____________________________
___________________________________________
And all I can see in the latent is that portion containing those ending ridges. Do I declare an identification, supported by the statistical study? I frankly don't know. That's a real paradigm change.
"Nothing has any value, unless you know you can give it up."
-
cchampod
- Posts: 16
- Joined: Tue Jul 19, 2005 9:52 am
- Location: Lausanne, Switzerland
- Contact:
Re: Statistical Model Article
Dear Gerald,
I believe that this all discussion is leading to the following: we need to be careful not confusing an assignment of a probability on an issue (e.g. the proposition that this is an identification) with the decision reported by a fingerprint expert that: "I consider this as is an identification". The probability remains a probability (and will never reach 1 for the above proposition) and the decision is indeed carried out according to an additional system of values (or utility functions) amounting to a cost/benefit analysis.
That process has been perfectly described (and that is valid for fingerprint identification as well) in the following paper (if anyone need a copy, please drop me a line: christophe.champod@unil.ch) : Biedermann, A., Bozza, S., and Taroni, F., "Decision Theoretic Properties of Forensic Identification: Underlying Logic and Argumentative Implications", Forensic Science International, vol. 177 (2-3), pp. 120-132, 2008.
The arguments of that paper have been taken up in the chapter 3 of the recent report: Expert Working Group on Human Factors in Latent Print Analysis, Latent Print Examination and Human Factors: Improving the Practice through a Systems Approach. Washington DC: U.S. Department of Commerce, National Institute of Standards and Technology, 2012. Now available on http://www.nist.gov/oles/human-factors- ... alysis.cfm
It reads (p. 68): Bayesian decision theory offers a more complete, formal model of the intuitive decision-making process. According to this model, a latent print examiner should consider both the source probability and the costs and benefits (utilities) of correct and incorrect decisions. More specifically, the Bayesian examiner:
Very kind regards
Christophe
I believe that this all discussion is leading to the following: we need to be careful not confusing an assignment of a probability on an issue (e.g. the proposition that this is an identification) with the decision reported by a fingerprint expert that: "I consider this as is an identification". The probability remains a probability (and will never reach 1 for the above proposition) and the decision is indeed carried out according to an additional system of values (or utility functions) amounting to a cost/benefit analysis.
That process has been perfectly described (and that is valid for fingerprint identification as well) in the following paper (if anyone need a copy, please drop me a line: christophe.champod@unil.ch) : Biedermann, A., Bozza, S., and Taroni, F., "Decision Theoretic Properties of Forensic Identification: Underlying Logic and Argumentative Implications", Forensic Science International, vol. 177 (2-3), pp. 120-132, 2008.
The arguments of that paper have been taken up in the chapter 3 of the recent report: Expert Working Group on Human Factors in Latent Print Analysis, Latent Print Examination and Human Factors: Improving the Practice through a Systems Approach. Washington DC: U.S. Department of Commerce, National Institute of Standards and Technology, 2012. Now available on http://www.nist.gov/oles/human-factors- ... alysis.cfm
It reads (p. 68): Bayesian decision theory offers a more complete, formal model of the intuitive decision-making process. According to this model, a latent print examiner should consider both the source probability and the costs and benefits (utilities) of correct and incorrect decisions. More specifically, the Bayesian examiner:
- (1) Assesses the prior probability that source of the exemplar left the latent print. An examiner who purports to rely solely on the information in the prints rather than the context of the case or the other evidence against this individual would have no reason to distinguish this individual from anyone else on the planet capable of being where the latent print was found. An examiner who considers more details of the case might treat the source of the exemplar as equivalent to a random person drawn from a smaller population of conceivable suspects. Such reasoning leads to the reciprocal of the size of the relevant population as the prior probability.
(2) Assesses the weight of the evidence as a function of the similarities and differences observed between the latent print and the exemplar. Formally, the weight is a likelihood ratio, as discussed below and in Chapter 1. Examiners may arrive at the weight intuitively (using their knowledge and experience) or by consulting data-driven likelihood models. However, these models currently use fewer features than a human examiner would, and they have other limitations.
(3) Computes the posterior probability of the proposition that the source of the exemplar left the latent print by combining the prior probability and the likelihood ratio according to Bayes’s rule (see Chapter 1).
(4) Combines the posterior probabilities with the utilities of the possible correct and incorrect decisions to reach and report the optimal decision.
Very kind regards
Christophe
C. Champod
ESC
University of Lausanne
Switzerland
ESC
University of Lausanne
Switzerland
-
Boyd Baumgartner
- Posts: 567
- Joined: Sat Aug 06, 2005 11:03 am
Re: Statistical Model Article
Don't recent likelihood models depend on the 'additional system of values' applied by examiners when captured via PiANos anyway? Considering this 'additional system of values' is ill or non defined, how would you compensate for iterated expectation in value rating by Examiners in such a model? After all, when a model only considers ridge events, you've basically got a Keynesian Beauty Contest in which examiners attempt to pick the the most 'beautiful' points to satisfy others expectations of sufficiency in the Verification. If what I am saying is true, such a model actually contributes to a 'confidence bubble' the same way a derivatives/commodity market would bubble in the stock market based upon expectation of future (verification/consensus) value. Thus resulting in a confident, but erroneous conclusion.cchampod wrote:I believe that this all discussion is leading to the following: we need to be careful not confusing an assignment of a probability on an issue (e.g. the proposition that this is an identification) with the decision reported by a fingerprint expert that: "I consider this as is an identification". The probability remains a probability (and will never reach 1 for the above proposition) and the decision is indeed carried out according to an additional system of values (or utility functions) amounting to a cost/benefit analysis.
Thoughts?
Edit: I'd also appreciate it, if you answered this yourself as opposed to referring to some article/book/etc. I'd like to hear what Christophe thinks, not somebody else; considering you're a fingerprint examiner by your own admission.
-
cchampod
- Posts: 16
- Joined: Tue Jul 19, 2005 9:52 am
- Location: Lausanne, Switzerland
- Contact:
Re: Statistical Model Article
Hi b..
If you take some time to read what I referred to, you will realize that I am not so alien to the material....
To respond to your specific questions:
(1)
(2)
(3)
Christophe
If you take some time to read what I referred to, you will realize that I am not so alien to the material....
To respond to your specific questions:
(1)
Utility functions refer to something different than the probabilities computed by fingerprint models (at least the ones I am aware of). What the recent model captures (some through PIANOS indeed) is the variability examiners may display when annotating minutiae.Don't recent likelihood models depend on the 'additional system of values' applied by examiners when captured via PiANos anyway?
(2)
No way. Models are not substitutes for proper fingerprint comparison. The process should not offer the possibility of picking and choosing the minutiae that pleases or suits the examiner (or the probabilistic value derived thereof). it should be based on an proper analysis. I like the concept of a consensus analysis in complex cases (meaning not dependent on one examiner only).After all, when a model only considers ridge events, you've basically got a Keynesian Beauty Contest in which examiners attempt to pick the the most 'beautiful' points to satisfy others expectations of sufficiency in the Verification.
(3)
. Well, as I said both in my testimony and in my written reports to the Fingerprint Inquiry, probabilistic models is not THE ultimate safeguard against wrong identification. However they can bring another layer of transparency to the identification process.Thus resulting in a confident, but erroneous conclusion.
Christophe
C. Champod
ESC
University of Lausanne
Switzerland
ESC
University of Lausanne
Switzerland