Zero Error Rate vs. No Error Rate

Discuss, Discover, Learn, and Share. Feel free to share information.

Moderators: orrb, saw22

Michele
Posts: 384
Joined: Tue Dec 06, 2005 10:40 am

Zero Error Rate vs. No Error Rate

Post by Michele »

I’d like to start a different conversation regarding the error rate of ACE-V. I'm making a new thread since this is so different than the other thread.

What is a method? I think it’s an explanation of HOW to do a procedure (specific directions). Ace-V doesn’t specifically tell us HOW to do a comparison, it doesn’t tell us what rules to use and which not to, and it doesn’t set any standards of when we’re done. I’m thinking that ACE-V is a vague outline but not specific instructions.

Here’s an example:
What if we said that we wanted to wash the cars in a car dealership by ‘cleaning them’. This is very vague because there are several ways to ‘clean the cars’ and some ways will produce a better result than others. You can drive cars through several types of car washes (water only, brushes, long plastic strips) or you can hand wash the car by several different methods (brush, towel, power sprayer). All of these specific methods will produce different results. If we washed 100 cars then some people might want to produce a test and determine how many of the cars passed the test for cleanliness. This number would be an overall number (error rate) and it might not apply to any individual method. Other people might think it’s more logical to look at the different ways that people clean cars and then test each method differently.

With regard to ACE-V, it doesn’t specifically tell us HOW to compare prints...... it tells us WHAT to do but it doesn’t describe HOW to do it. Some people may count points, others may use ridgeology, some people might use level one to exclude, while others might do a full analysis. Maybe Ace-V isn’t a methodology……Maybe we don’t know the methodology someone uses because we don’t specifically ask them??? If we want to find an error rate for a method then we need to establish what method was used (counting points, ridgeology, guessing, etc) and we should know the rules of that method.

I’m wondering if others might agree that ACE-V tells us WHAT to do but it doesn’t describe HOW to do it. It might not be a SPECIFIC METHOD but a VAGUE DESCRIPTION that could include several methods inside of it.

For years people have recommended doing more research but don't we need to define certain things before we can research them?
Michele
The best way to escape from a problem is to solve it. Alan Saporta
There is nothing so useless as doing efficiently that which should not be done at all. Peter Drucker
(Applies to a full A prior to C and blind verification)
Michele
Posts: 384
Joined: Tue Dec 06, 2005 10:40 am

Post by Michele »

After getting a few emails from friends telling me how they interpreted my post, I thought I should expand on something.

I didn’t realize some people might think that stating ‘ACE-V wasn’t specific enough’ would be interpreted as meaning that ‘it lacked reliability’. Reliability isn’t dependent on how specific the process is, it’s dependent on the how often the results are reproduced. From the empirical data over the last 100 years, I think it’s safe to say that our overall results are extremely reliable. Anyone claiming something different would be hard pressed to find any data from casework to support their view.

The scientific method of hypothesis testing isn’t specific either but it’s been used as an accepted scientific technique for hundreds of years. We don’t need for something to be specific to be useful or reliable. The advantage of having a specific method would be that it would help us speculate on a potential error rate and it would also help us determine the probability of an error occurring. For instance, I might exclude someone from leaving a latent because of the pattern type alone. Someone else might exclude this person by looking for a specific target group of minutia. And a 3rd examiner might exclude by looking for a target group and then checking themselves by looking for another target group. In this example, the probability of me making an error is higher than the 2nd and 3rd examiner. And the probability of the 3rd examiner making an error is extremely low because they have a specific method with quality assurance controls in place. I know this example isn’t about ID’s but it was easy to differentiate between different methods.

I just thought I’d clarify this so it’s not implied that I think the results of our professions ID’s aren’t reliable. On the contrary, I think they are extremely reliable.
Michele
The best way to escape from a problem is to solve it. Alan Saporta
There is nothing so useless as doing efficiently that which should not be done at all. Peter Drucker
(Applies to a full A prior to C and blind verification)
Pat A. Wertheim
Posts: 872
Joined: Thu Jul 07, 2005 6:48 am
Location: Fort Worth, Texas

Post by Pat A. Wertheim »

We all admit errors occur. We always have, both out of court and in court under oath. There are several problems I see, and I believe Michele has put her thumb on one big factor in "error rate." While we discuss "ACE-V" as our methodology, the actual way we work can vary. For example, one examiner may prefer to start a comparison with a target exactly on a delta or core, another may look for a recognizable combination of "points" a few ridges away. A third examiner may begin the comparison with a more critical glance at pattern and never even get to the stage of trying to find a target to search. Would the "error rate" be different for the three examiners? I could see where it might be, but who knows? We've never tried to analyze it that way, and as Michele points out, there are far greater variations in our application of the ACE-V methodology.

Another problem with "error rate" is that it could never be calculated precisely. The biggest issue in definining "error rate" is to define an "error." In the parallel thread, Charles differentiates between Type I and Type II errors. But even more to the point, if we are only considering erroneous identifications (Type I), when exactly does an error become an error? If Examiner A brings a latent to Examiner B and says, "Check the #2 finger on this one," and B has a look and says "Whoa, what about this short ridge way up here? I see the half dozen or so points in the middle that I believe you are looking at, but this is an exclusion, not an identification." So A agrees it is an exclusion and moves on to other suspects. I think we would all agree that no erroneous identification occurred there to be added to the error rate calculations. But what if A had said, "I think this is ident to #2. Would you verify it for me?" and B's response is the same? Would we use that in calculating error rate? What if A said, "I'll stake my career on this one. It's #2." and B's response was the same? My point is, the prints were the same, B's observations and response was the same, only A's statement in seeking a verification was different. Are they all erroneous identifications?

Or, are any of them erroneous identifications? What if your lab policy only considers an erroneous identification when the entire ACE-V breaks down and the error is verified? Until then, maybe no erroneous identification occurred at all. Certainly, if we are computing error rate for entire ACE-V, you cannot count an error until the V is in error, too.

Then there is the problem of individual error rate versus overall error rate. Let's throw in laboratory error rate, too. Most examiners never make erroneous identifications in their entire careers. Some examiners make multiple errors. Should the error-free examiner be painted with the same brush as the examiner who makes several errors during his career?

Yet another problem is that an error rate would not be used correctly if it were used by opposing counsel to suggest that was the chance there was a mistake in a specific case at trial. As Lisa Steele (defense attorney who posts on this board) has said many times, she is not worried about the clear prints with lots of points. She is concerned with the badly distorted prints with few points. If error rate were calculated, I think we would all understand that it was based on those badly distorted prints with only a few points. So how could that error rate be appled to a clear print with twenty or thirty points?

I agree that research should be done in the area of error rate. But I believe we need to agree first on exactly what an "error" is, and wee need to understand and be able to explain the meaning and limitations there are on "error rate."
Pat A. Wertheim
P. O. Box 150492
Arlington, TX 76015
George Reis
Posts: 147
Joined: Wed Jul 27, 2005 1:00 pm
Location: Orange County, CA - USA
Contact:

Post by George Reis »

When I read Charles' first post on this, I almost replied with something similar to what Michele did here. The issue is that a method may or may not include a test within that method. If there is no test, then an error rate cannot be applied.

Here's an example:

To check to see if a solution is acid the method is used to dip pH paper into the solution, report the result, verify. Notice that there is no test here, so the method doesn't tell us what criteria the examiner uses to determine the acidity. Change the method to: dip the pH paper into the solution, compare the result to the standard, if the color is X it is acid, report the result, verify. Here, we have a test - a comparison is made and a criteria must be met. In the first case, it would be silly to attach an error rate because there isn't a test - in the second case we can attach an error rate for both the test and for the examiner.

ACE-V doesn't have an objective test of this sort. Without it, we can't apply an error rate to it. Maybe that's fine, and we simply state that this is an opinion based science. Maybe that's not adequate and we find a way to quantify it. I'm not sure of the solution - but it seems that we've spent a lot of time addressing the wrong question.

George
I can resist anything except temptation - Oscar Wilde
Michele
Posts: 384
Joined: Tue Dec 06, 2005 10:40 am

Post by Michele »

George,

That's so true. If we had a test, what would we be testing it against (since we never know the ground truth)?

With blind verification, we are testing the conclusion against the conclusion of someone else. Some people may consider this a test (which I don't have a problem with) but we need to determine when a test passes and when a test fails. Then determine what to do if a test fails.

For example, what if the conclusions of blind verifications aren't the same? Does that invalidate the ID? If it doesn't why, or when should it? If it does invalidate the ID, I'd ask the same question, why? I think we should establish make the rules prior to having tests instead of having tests and just assuming the results will always turn out the way we want them to.
Michele
The best way to escape from a problem is to solve it. Alan Saporta
There is nothing so useless as doing efficiently that which should not be done at all. Peter Drucker
(Applies to a full A prior to C and blind verification)
Charles Parker
Posts: 586
Joined: Mon Jul 04, 2005 6:15 am
Location: Cedar Creek, TX

Post by Charles Parker »

Michele, the following is from my personal point of view:

1. ACE is an acronym for the examination model of Analysis, Comparison, and Evaluation. This term is generally used by the Latent Print Community to simply describe the process they use to reach conclusions.

2. There have been prior models and there are even a number of versions of ACE that have been offered in the past 10 years. All of these prior models and current model/versions have one thing in common. They use comparison as one of the steps in their proposed model.

3. The real “Methodology” is the mental process of comparison of two or more items.

3. B.C. Bridges used the term “Comparative Examination”. I adopted the term “Comparative Analysis” after reading an article in the 1980’s on the subject as it was being applied to foot wear comparisons.

4. Comparative Analysis is the major factor in any pattern and even some non-pattern type evidence. (shoeprints, bullet/casings, hair/fiber, etc.)

What is a method? I think it’s an explanation of HOW to do a procedure (specific directions). Ace-V doesn’t specifically tell us HOW to do a comparison, it doesn’t tell us what rules to use and which not to, and it doesn’t set any standards of when we’re done. I’m thinking that ACE-V is a vague outline but not specific instructions.
5. The model (method) that I have seen that comes close to specific instructions is Robert Olsen’s comparison flow chart. I think it is still published in Lee and Gaensleen book. Even though it is step by step, try to tell it in a court of law or to bunch of 5th graders. But it is more specific than anything and if someone was to add standards and principles to it, they would be getting a lot closer than ACE.
With regard to ACE-V, it doesn’t specifically tell us HOW to compare prints...... it tells us WHAT to do but it doesn’t describe HOW to do it. If we want to find an error rate for a method then we need to establish what method was used (counting points, ridgeology, guessing, etc) and we should know the rules of that method.
6. I believe if an error rate was developed it should be determined based upon the conclusion and not how one reached that conclusion (you and Pat both make arguments along that line). My reasoning is that it is subjective and there is no accurate way to determine the correct path to reach a conclusion. The conclusion is an opinion (court definition not lay definition).
For years people have recommended doing more research but don't we need to define certain things before we can research them?
7. I agree with more research but what do you think would need better definitions?

8. Pat, from my POV errors occur when they are marked or it is in writing. Seen too many over-eager LPE claim someone made an error by just being vocal. If someone came to me and stated they would bet their career on it, then I am going to ask them to mark it first. If someone comes to me and states I think this is correct but I am having a problem and can you help me, then I help them. An error could occur before or after “V” but not before there is any documentation of the conclusion. You need documentation in the process to avoid the finger pointing personalities.
Then there is the problem of individual error rate versus overall error rate. Let's throw in laboratory error rate, too. Most examiners never make erroneous identifications in their entire careers. Some examiners make multiple errors. Should the error-free examiner be painted with the same brush as the examiner who makes several errors during his career?
9. I think the courts are looking for the error rate of the discipline, but that could always be reinforced in court by that examiners individual error rate as well. It would be like machines that are different make and models. The more advanced would have a lower rate.
Yet another problem is that an error rate would not be used correctly if it were used by opposing counsel to suggest that was the chance there was a mistake in a specific case at trial. As Lisa Steele (defense attorney who posts on this board) has said many times, she is not worried about the clear prints with lots of points. She is concerned with the badly distorted prints with few points. If error rate were calculated, I think we would all understand that it was based on those badly distorted prints with only a few points. So how could that error rate be applied to a clear print with twenty or thirty points?
10. Pat, interesting! The way I would think it would go would be like any of the other tests (CLPE, Competency, Proficiency, etc) a range of Simple, Moderate, and Complex prints would be the best. But I could see where the critics would say it was too many Simple and Moderate level prints so the error rate would be less. They would like it if it consisted of just complex prints as the error rate would be much higher. Either way it is SBT. Figuring error rate (if the discipline goes that way) is going to be very tough thing to do.
I agree that research should be done in the area of error rate. But I believe we need to agree first on exactly what an "error" is, and we need to understand and be able to explain the meaning and limitations there are on "error rate."
11. No problem. Something for SWGFAST to do?
George, That's so true. If we had a test, what would we be testing it against (since we never know the ground truth)?

With blind verification, we are testing the conclusion against the conclusion of someone else. Some people may consider this a test (which I don't have a problem with) but we need to determine when a test passes and when a test fails. Then determine what to do if a test fails.

For example, what if the conclusions of blind verifications aren't the same? Does that invalidate the ID? If it doesn't why, or when should it? If it does invalidate the ID, I'd ask the same question, why? I think we should establish make the rules prior to having tests instead of having tests and just assuming the results will always turn out the way we want them to.
12. Michele, trying to get my head around this one. I have to think some more on it.
Knuckle Draggin Country Cousin
Cedar Creek, TX
mdavis
Posts: 154
Joined: Mon Jan 02, 2006 6:07 am
Contact:

Post by mdavis »

As I've mentioned elsewhere, how do you determine error rate if you don't know the correct answer with absolute certainty? We can devise proficiency tests from known controls, but in the real world with actual casework, there is very seldom any way to know the correct answer. Without that absolute answer, all bets are off and any pseudo-results of accuracy are essentially worthless. My contention is that there is no way to measure error rates in actual casework. Only when a reported individualization is "proven" by other evidence to be in error, do we accept that there is one. Honest, ethical blind verification by another qualified examiner is the best hope for reliability.

We can't do much better than the ACE-V guidelines and procedure. Once we attempt to specify "how" to effect an analysis and individualization, then we stray into the realm of quantification rather than qualification. We are forced to revert back to a perhaps more sophisticated form of "point counting", and I don't think many of us want to revisit those ruins.
Michele
Posts: 384
Joined: Tue Dec 06, 2005 10:40 am

Post by Michele »

Without that absolute answer, all bets are off and any pseudo-results of accuracy are essentially worthless.
The ground truth is rarely known in science. Error rates are established by comparing the answer given to the expected answer. This is overly simplified, but that’s the basic idea. If we have no way of determining if an error was made then many of the past cases of erroneous ID's couldn't be considered to be erroneous.
Honest, ethical blind verification by another qualified examiner is the best hope for reliability
What if the blind verification doesn’t produce the same result as the first examiner? Then what happens?
Once we attempt to specify "how" to effect an analysis and individualization, then we stray into the realm of quantification rather than qualification.
I think we can have basic principles within ACE-V that lead us to better results without giving up the critical thinking aspect.

For instance, do you have to do a full analysis prior to moving on to the comparison phase?
Can you identify a person with level 1 detail?
Can you exclude using level 1 detail?
Is the 1-Discrepancy rule a mandated rule?
Does the overall size of the latent matter?
Can you overlay images to make an ID?
Can you overlay images to exclude?
Do the ridges widths need to be the same?
Does an ID need to be verified?
How many people need to agree?
Are you supposed to be confirming the conclusion or scrutinizing it?

Having basic principles to follow doesn’t take away from ACE-V, it should make it stronger.

I’m not saying ACE-V doesn’t have objective rules of how to use it. From reading the list above I’m sure everyone knows which of these are standard rules and which aren’t as standard. Once we know if someone used the rules appropriately, then we can determine the error rate accordingly.
Michele
The best way to escape from a problem is to solve it. Alan Saporta
There is nothing so useless as doing efficiently that which should not be done at all. Peter Drucker
(Applies to a full A prior to C and blind verification)
mdavis
Posts: 154
Joined: Mon Jan 02, 2006 6:07 am
Contact:

Post by mdavis »

The stock answer to all of your questions, Michelle, is "it depends on the print." You can't make generalizations fit the range of latents encountered, can you?

And what is "the expected answer?" Expected by whom? If we compare agreement between examiners, whose opinion counts? If there is no agreement, does that mean an error, or does that show that one examiner has a different threshold than another?

Comments submitted for consideration, not as a rebuttal.

mdavis (Devil's Advocate) :evil:
Amy Hart
Posts: 43
Joined: Tue Oct 11, 2005 7:00 am

Post by Amy Hart »

In reading these discussions, I become extremely frustrated that we, as latent print examiners, are trying to create an error rate based on how many bad identifications have ever been made. No other discipline has an error rate based on operator error. I think the problem that we are experiencing is twofold.

First, pretty much everyone agrees that fingerprints are unique. Because of this, the fingerprint community has not been compelled to determine what the probability that someone else has a similar fingerprint would be. This is (more or less) how DNA has calculated their error rate. Based on a finite sample, the frequency of alleles at particular loci was measured. Statistics were then calculated based on those measurements. The error rate in DNA is, therefore, the probability of another person having the same alleles at the same loci. Where in that description does it come up whether or not someone followed the procedure correctly? Where is the possibility of mistaken interpretation of the electrophoresis? Where is the confirmation bias? These are certainly factors that come up, but they are separate from the error rate.

The second problem is that if we were to calculate an error rate, it might introduce the concept of a threshhold. We would be calculating a probability based on the presence of 1st, 2nd, and 3rd level detail. You would have a starting probability of pattern type/location, then you would have a probability of particular points occuring in a particular sequence, then you would have a probability of the occurrence of 3rd level detail. Once you multiplied those out, you would have the error rate of your sample. However, in order to convince a jury that you have made an identification, your error rate would most likely have to exceed the population of the earth. So, for example, if you have a print that has an apparent whorl pattern (probability 1/3) with opposing bifurcations two ridges directly above the core (probability of 1/x) and then an ending ridge 2 ridges above the bifurcation (probability 1/y). Your error rate would be 1/3xy, and the population of the earth is 1/3x. You may have made a case for your ID, but you have just wiped the quality portion of our usual equation and gone completely over to quantity. Juries will be much more impressed with the probability of 1/3xy than they would be with just 1 in the earth's population. Thus, you have just created an arbitrary threshhold...

In any event, I think that the fingerprint community needs to find a statistician who will marry a latent print examiner, and have a latent print examiner statistician baby, who can calculate a logical and meaningful error rate for our discipline. Otherwise, I am being held accountable for other people's lack of training, lack of ability, headache, hangover, bad day, or sinus infection without receiving any credit for the good, ethical work that we're doing most days.
Charles Parker
Posts: 586
Joined: Mon Jul 04, 2005 6:15 am
Location: Cedar Creek, TX

Post by Charles Parker »

Amy, I must say you have made some very good statements which were very easy for me to understand. and after reading them, you have convinced me.

Rock On Amy.
Knuckle Draggin Country Cousin
Cedar Creek, TX
Outsider
Posts: 166
Joined: Mon Aug 07, 2006 2:15 am
Location: Scotland

Post by Outsider »

The type of error commonly discussed with DNA profiles is the Random Match Probability, the chances that someone in the population, picked at random, will have the same profile. This is not operator error, it is stable and predictable. The other type of error a false positive or lab error is different. I cited a paper recently on another thread about these two types of error with DNA evidence. The authors are concerned that while the courts require the RMP when DNA evidence is used, the court does not ask for a rate for false positives. When the RMP is very low, false positives will be the dominant source of errors.

http://www.bioforensics.com/conference/ ... %20Pos.pdf

The same logic will apply to fingerprint evidence except that the term “individualization” means an RMP of zero. But that still leaves false positives. I know from my work in industry that the frequency of this type of error will almost certainly be unstable, meaning that it will change unpredictably over time. Prediction is only possible with stable systems. So the idea of an error rate that can be used for precise calculations is probably a non-starter (but I still think that research would be useful). Even if we could measure the error rate of identifications for an examiner or department, and it was stable (which it won’t be), this would be very different from the chances that a particular identification is erroneous after the police selection processes, particularly when the identified person challenges the forensic evidence. The error rate here will be higher, and possibly much higher because any forensic error will look suspicious to the police, with the subject is unable to give an innocent explanation. This has been discussed this at length on the Statistics and Misidentification thread.

The paper highlights a particular problem when the false positive rate is unknown and the other evidence in the case is very weak. All evidence is a balance of explanations ( the figure of “Justice” holds a set of scales). When the rest of the case against the accused is very very weak (very low prior probability) the error rate of the forensic evidence, even if it is very low, becomes critical (particularly if a very large number of forensic tests have been used to select the accused). With an even balance small differences can tip the scales one way or the other. That is not to say that it is right to keep any evidence off the scales. There are no easy answers to this.

I understand the official position to be that the method of fingerprint identification has a zero error rate but humans can make mistakes if they do not follow the correct method. I would take this to mean that that the RMP is thought to be zero, but I would not completely rule out the possibility of a false positive. I would consider this with the other evidence in the case. But it could also be interpreted as meaning if the expert has impressed the jury that he or she has taken care, they should consider the evidence to be infallible.
Steve Horn
Computer Programmer working in the field of statistics for industry
http://www.stevehornsc.pwp.blueyonder.co.uk/pf.htm
mdavis
Posts: 154
Joined: Mon Jan 02, 2006 6:07 am
Contact:

Post by mdavis »

In comparison to other types of forensic "evidence" that gets full credit in the courtroom, I think latent print individualization is on extremely firm ground. A recent conviction was overturned after 26 years based on a "eye witness" identification that was proven false bymodern DNA testing. How many convictions are made by these eye witness "experts" to the crime scene in question? How many people are convicted by a few items of circumstantial evidence (yes, even latent print evidence can be circumstantial in a few cases).

As a judge, I would think you could put more credibility on a latent print ident as potentially accurate than any other single piece of evidence (a few eye witness idents by known persons or detailed image capture of the crime in progress notwithstanding).
Gerald Clough
Posts: 557
Joined: Wed Jul 06, 2005 6:27 am
Location: Lockhart, Texas
Contact:

Post by Gerald Clough »

Eyewitness identification is an interesting area in which the law is rapidly evolving in response to research. It was, until recently, routine for trials courts to refuse to allow experts to testify to the reliability of eyewitness identification in general. The rejection was based upon the notion that jurors could apply their common experience and knowledge to judge the credibility of eyewitness, the lack of peer-reviewed research, and the fact that the expert would not be testifying about the specific eyewitness identification situation. In many appeal jurisdictions, it is now being held to be error to exclude such expert testimony. This is because there is now a substantial body of formal research, in fact there are a number of eyewitness identification laboratories at universities, and the findings run contrary to the jurors assumptions and lie outside their common knowledge. There are some significant differences among states and the federal circuits, but the law seems clearly to be moving in a new direction.

The investigative use of eyewitness evidence is becoming much more subject to expert review and is subject to being criticised if it does not reasonably conform to published standards, mainly the USDOJ guidelines and those published by the New Jersey Attorney General, both of which are reasonably consistent but should be taken together to get a complete treatment of all the procedural issues.

Latent print examination may one day be subject to similar expert opinion on the nature of LPE in general, but it's going to take a body of research which, as we well know, seems very difficult too design. At least it is not now so amenable to the same kind of modeling done in the eyewitness labs.
"Nothing has any value, unless you know you can give it up."
mdavis
Posts: 154
Joined: Mon Jan 02, 2006 6:07 am
Contact:

Post by mdavis »

I don't know if it is still being done or not, but in one of the criminal justice classes taught at the university that housed our lab, they had a rather sobering lesson for would-be LEOs. Sometime, during a routine subject lecture, one of us would come in from the crime lab and hand the professor a note, then leave. Since most of us did not teach classes, few of the students knew us by name or face. The professor would look briefly at the note, place the note on the podium and continue with the lecture for a minute or two. He would then ask students to get out a piece of paper and write down what they had observed during the previous several minutes, including the sex, height, weight, facial features, hair color, eye glasses, clothing, shoes and demeanor of the "messenger." Comparing these "eye witness" observations with the actual messenger showed that eye witnesses are seldom good at describing or recognizing observed events. Place a victim under duress during commission of a criminal act, and the chances diminish. These were "trained" LEOs, not civilians, yet their powers of observation were routinely very poor.

The main strength that latent print identification has going for it is a long history of success with an incredibly high percentage of accuracy. Of the handful of cases brought to national or international attention as suspect idents (and you can bet that essentially every bad ident is publicized by the media), there are tens of thousands of valid idents that are never contested by the accused, consistent with other forensic and circumstantial evidence, and as accurate as humanly possible. No one claims that a bad ident might fall through the cracks and never be caught and never publicized, but the chances of that happening is arguably less than any other form of forensic evidence presented in a courtroom.
Post Reply