Page 2 of 2

Re: British judge rules that Bayes Theorem can no longer be used

Posted: Mon Oct 31, 2011 11:56 am
by Cedric
Michele, All,

I must clear some confusion: while Saks and others have often criticized the community for the 'training and experience' answer to all challenges, they have also done some very interesting work on juries to see how they would react to various ways of communicating results in a court settings. So I was referring to their constructive work, which applies directly to GrayMatter's question, as opposed to the publications where they criticize us, which was not the topic of today's conversation.

Thanks

Re: British judge rules that Bayes Theorem can no longer be used

Posted: Mon Oct 31, 2011 12:29 pm
by ER
While I hope that some sort of statistics or LR can back up our 'training and experience', I'm still in a wait and see mode.
A group of us believe that it is indeed appropriate, and even necessary, to use a LR.
While LR's have shown us a lot of valuable things in regards to latent print comparisons, they may not be usable as a basis for every conclusion. There will probably ALWAYS be a comparison where the identification decision comes down to 'training and experience' because a LR cannot account for every detail.
One thing for sure is that we need to move away from categorical conclusions
Moving away from categorical conclusions is not a certainty. Whatever LR or statistical form of natural language that is developed must be shown to be at least as precise, accurate, and repeatable as the categories that we currently use. The current three categories may turn out to be the best way to report our results. We need research before tossing them out.
I think ‘training and experience’ is a weak basis for a conclusion. I believe it because it puts too much weight into elements that can't be looked at and tested.
'Training and experience' may have its weaknesses, but it's not untestable or unknowable.

Re: British judge rules that Bayes Theorem can no longer be used

Posted: Mon Oct 31, 2011 2:30 pm
by GrayMatter
While I enjoy Saks and Dawn’s more constructive work http://legalpsychology.asu.edu/Lab/Publications.html particularly relevant to this thread is the paper: What Expert Witnesses Say and What Factfinders hear. Given that judges were part of their experimental group.

BUT my question was really more related to the policies of FSS at the time the examiner in R v. T conducted his analysis. Given that you worked there I hoped you could answer.

1. If a footwear examiner had a “match” did they bother using Bayes’ theorem after they had determined the match?
2. Do you think it is possible examiners were intentionally leaving out the use of Bayes theorem in their report because it had been banned since 1996?

Re: British judge rules that Bayes Theorem can no longer be used

Posted: Mon Oct 31, 2011 3:08 pm
by Michele
ER,

If an examiner is asked, "When do you know when you have an identification?" and the examiner answers with, "it's based on my training and experience" or "when I'm confident based on my training and experience", how do you test the conclusion against the standard of 'the examiners training and experience'? I don't understand how it can be done.

Re: British judge rules that Bayes Theorem can no longer be used

Posted: Mon Oct 31, 2011 3:29 pm
by ER
First of all, I hope that no one is leaving the answer at "I'm confident based on my training and experience." At the very least, I hope that they would say that the two prints share a sufficient amount of unique features that, based on my training and experience, would not be present in another person's print. (Or something to that effect.) A LR would be saying basically the same thing, but would assign a number to the uniqueness of the features that the two prints share.

The real question for me is: Which is the better method?

You can test examiners in general (or examiners at a specific lab) against known samples to determine how often they have false positives, false negatives, or missed ID's. (Glenn's studies, FBI Black Box, etc.) If a LR model can be shown to reduce these errors, then we should start using it right away. That's obviously a few steps down the road. My point is that just because you develop a model that gives you a LR, doesn't mean that anything was improved. I need to be shown that using an LR in conjunction with ACE-V produces more reliable results than using ACE-V with only my training and experience. I hope it does. I want to produce the most reliable results possible.

But to answer your specific question, you test a wide group of examiners against a known set and then extrapolate that test to the specific print in this case. It's not a perfect test, but that's how it's being done. Again, my point is that we don't yet know whether these LR models are just going to add to our general understanding of uniqueness, or if they will actually improve our ability to make accurate conclusions. I'm hoping for both.

Re: British judge rules that Bayes Theorem can no longer be used

Posted: Mon Oct 31, 2011 4:29 pm
by Cedric
GrayMatter,

1) having a match merely means that two objects/individuals have identical features at some analytical level (and explainable differences). This is obviously the first question of importance. The second important question is to know how many other objects/individuals unrelated to the first one also would also share those characteristics. Answering those two questions is broadly equivalent to doing a LR. Therefore, an examiner who finds features in agreement between two shoe impressions (e.g. a trace and a control) will want to know how many other impressions from other shoes would also show these similarities to the trace. If he find none, then he can declare an identification. If he finds some, then depending on the number found, the examiners will have more or less strong evidence.

So yes, footwear examiners with a "match" should compute a LR to quantify the weight of the evidence.

2) As mentioned earlier in the thread, a LR is different than the Bayes theorem, therefore, FSS examiners were perfectly correct to compute LR even though the Bayes theorem was banned since 1996.

ER,

Yes, you are right, there are 2 different issues: 1) providing supporting data that forensic scientists can do what they claim they can do, and 2) improving the accuracy and precision of the process. In both cases, data provided by statistical, error rates, bias and other studies are critical.

Re: British judge rules that Bayes Theorem can no longer be used

Posted: Mon Oct 31, 2011 8:00 pm
by Michele
ER,

This isn’t really what the thread is about but since we’ve started the conversation (about training and experience) I’d like to bring up a few more points. About the ability to test training and experience, I think a lot of it may be semantics. With the studies you mention, I wouldn’t feel comfortable saying these tested T and E or tested the ACE-V process. My opinion would be to say that they may just test the ability of the person to get the right answer and/or tested a person’s thoroughness.

You mentioned needing “a sufficient amount of unique features”. For me, that’s even more important than training and experience. In the errors I’ve seen, the errors were caused because too much weight was given to T and E and not enough weight was given to having a sufficient amount of unique features. If we want to diminish errors, I think putting weight in the quality and quantity is more important than putting weight in T and E.

Re: British judge rules that Bayes Theorem can no longer be used

Posted: Tue Nov 01, 2011 9:40 am
by ER
Michele,

We have strayed a little off topic, but I do enjoy the journey.

The studies that we're talking about tested the examiners ability to get the right answer with their T&E using the ACE-V process (or at least ACE). I don't think these terms can be separated. The studies have shown that examiner's are generally very good at using ACE with their T&E to get the right answer. That's a valuable thing to know. And I wouldn't dismiss a study just because it tested someone's ability to get the right answer. When you boil away all the acronyms and protocols and methodologies, isn't 'getting the right answer' the essence of what we're all trying to do?

Errors have been made by examiners that are too confident in their T&E. They probably need more training on the limitations of ACE-V. Diminishing errors is the goal for everyone in this discipline. All I'm saying is that when we get a working model that will generate LR's for comparisons in casework, that might reduce errors, but it might not. Using LR's in casework might increase bad ID's, missed ID's, bad exclusions, or all of the above. We can't know until it's done and tested. Then we can have the debate about whether it's better to just use T&E or to supplement with LR's.

One thing I don't understand is how do you propose to put more weight in quality and quantity except by training people to do so?

Re: British judge rules that Bayes Theorem can no longer be used

Posted: Tue Nov 01, 2011 1:29 pm
by Michele
ER,

From what I’ve seen, many words are being used very casually by some researchers. Lets start with the word expert.

I’ve been involved in several studies that have said they were studying ‘experts’ but the studies never established me as an expert. They didn’t ask me about my education, training, experience, years in the field, competency testing, proficiency testing or certification. From what I’ve seen, it appears that people are using the term ‘expert’ to mean ‘employed’. To be honest, I’m not even sure if being employed was a qualification to be involved in their research. I’ve never seen a study that competency tested anyone to establish their expertise (or asked if someone was competency tested).

As I’ve stated earlier, the studies seem to look at the reproducibility of the conclusion but they don’t seem to be looking at if the conclusion is valid (even though some claim that they are studying validity).

The next term I think that is being used too casually is ACE-V. If people are saying this is the method we’ve always used then is ACE-V just a fancy name for doing a visual comparison and having someone else do a visual comparison and arriving at the same conclusion? Is research claiming to be testing the use of ACE-V with no rules on how it should be used? Are there any rules you have to go by to claim to be using ACE-V? I read in a transcript that an examiner was using ACE-V even if they didn’t realize it. Does that mean that ACE-V is really just saying someone looked at the two impressions?

As for your question, I think more weight can be put in the quality and quantify by stating rules that they must adhere to. As an example for quantity, can you use scars, level 3 details, pores or creases as characteristics? Even if we can’t quantify the number of characteristics needed to make an ID, shouldn’t we at least state what the characteristics that we can count include? If I had a scar (or maybe it’s a crease) that cuts across 4 ridges (you can clearly see how it cuts across the ridges), can I count each side of the it as a characteristic (giving me 8 characteristics)? This may sound ridiculous to some but others are waiting to hear how others answer. How are discrepancies counted? What is a discrepancy? How can you tell if something is a discrepancy or distortion? How can we label something as a discrepancy if we can’t describe what a discrepancy is? The answer of “it’s based on my T and E” really just says, I know it when I see it (I think that’s the antithesis of science).

Re: British judge rules that Bayes Theorem can no longer be used

Posted: Wed Nov 02, 2011 3:29 am
by Pat
I think the original topic of this thread had to do with testimony to “probability” of one sort or another. It seems to have shifted to whether we should testify to an opinion based or our training and experience, or on the quality and quantity of corresponding detail in the prints. I would suggest that testimony should include components of both.

There are five basic areas of expert testimony:
1) Witness Qualifications (trained, experienced, and tested for competencey);
2) Foundations of the science (biology: permanence and uniqueness);
3) Origin and chain of custody of the evidence;
4) Examination process (ACE-V, including quality and quantity of features);
5) Examiner’s conclusion (opinion).

The thread started with discussion of 5) above, but has shifted to a question of whether 1) or 4) is more important. The truth is that all five are essential components of testimony.

But to bring the discussion back to the original topic, the question we are struggling with is whether a conclusion can be stated in probabilities of some sort, or whether we should stick with the traditional three conclusions: 1) Absolute certainty of identification, 2) absolute certainty of exclusion, or 3) Can’t tell.

The trick is to separate good science from dogma. I would suggest that science includes probabilities. The requirement of absoluteness in testimony is a dogma that served us well in the early days, but one which should be reexamined in the light of evolving needs of science in the courtroom.

The problem is that we have not learned to adequately and appropriately articulate what we mean when we are not absolutely certain. I, for one, am not afraid to admit that I have had latent prints that I thought were “probably” made by a person to whose inked prints I had compared the latents. But I was not “absolutely” certain. In such cases, I followed the old dogma and rendered a conclusion of “inconclusive.” The latent print is of “no value.” I can’t tell a thing about it. Could be him, could be somebody else. Ho hum.

But isn’t that denying some shred of evidence to the jury? The problem is that if I testify to something like “I find some matching details, not enough to be sure, but I think it is probably him,” then we are afraid of the risk that the court will understand us to be making a positive identification. So we just say “inconclusive.”

I remember a violent rape. Woman alone at home while father has the kids on a weekend outing. Woman awakes in the dark, aware of the presence of someone in the bedroom. She reaches and turns on the lamp on the nightstand, but as she does so a blanket is thrown over her head. She does not see the man. He turns the light back off and has his way with her.

The responding officer fingerprinted the lamp. He found a latent print that was a whorl pattern with several clear points. All of the family, woman, husband, and all kids, are strictly loops. A suspect is developed on other evidence (prior to DNA) and a comparison is requested. So far as Level 1 pattern, the whorl on the lamp was as exact an overlay as you can imagine. The few points plus a little Level 3 was enough for the first examiner, Glenda, to make the identification. But with only two or three points and, in my opinion, insufficient Level 3 to support a positive identification, I declined to verify the identification. The other examiners in the department also declined to verify. The final report said only that the latent print was “of no value.”

Really? No value? What about combined with the other evidence that had led to identifying the guy as a suspect in the first place? Was the jury given all of the evidence? Was it correct to deny telling the jury that the print matched in all respects, but was insufficient for absolute certainty?

As a discipline, we are still struggling with the question of how to adequately explain corresponding features that fall short of absolute certainty. But if we are to present ourselves as good science, we cannot blindly subscribe to old dogma without at least honestly considering refinements.

Re: British judge rules that Bayes Theorem can no longer be used

Posted: Wed Nov 02, 2011 8:59 am
by GrayMatter
Boyd:

Thank you for posting the articles despite your dissatisfaction with them. I only read the first one but I rather enjoyed it. You could tell ultimately what side the authors fell on but I thought they gave decent treatment to limitations of likelihood ratios and the benefit of experience in qualitative testimony.

I am interested in this
The ethical reality however that lives in stark contrast to these idealized views of the likelihood ratio's use is that statistical calculations in arriving at evaluative weight becomes a crutch and a black box that is abused.
I didn't see an ethical argument in the decision nor the article. What is the ethical argument you see in either the R v T case or in forensic statistics in general?
As for the second part of the sentence I feel like Pat has made a good argument for why the "crutch" might be preferential to nothing.

What do you mean about the black box? Do you mean like in R v T that the examiner didn't disclose how he arrived at the conclusion of "moderate support'?

For everyone: Do we have professional forensic ethicists?

Re: British judge rules that Bayes Theorem can no longer be used

Posted: Wed Nov 02, 2011 9:25 am
by ER
Michele,

A few things:

Just because a study is flawed, doesn't mean it's worthless. Every study has inherent flaws and limitations that should be discussed and taken into consideration. Hopefully, future studies will look more closely at what constitutes an expert. However, a study should look at a broad selection of people employed in LPU's. These are the people that are testifying on conclusions in court, so these should be the people that are studied.

Each study I've seen validates their conclusions by creating ground truth samples or determining the valid answer by a panel of experts. How else would you propose to validate the conclusions? A LR model? They're not quite ready for prime-time yet.

What q&q rules would you propose? The newest computer models are trying to assign uniqueness values to quantity. There's ongoing research into quality-mapping. They're yet to be proven and accepted into casework. I've been trained to count everything. Flow, path, shape, crease, scar, pore, edge, stop, split, dot, angle, width, etc. Each of these features is assigned a variable weight depending on how unique I think it is and it's clarity in the latent and the known. It doens't matter if you count it as 1 crease, 4 breaks, or 8 ridge endings. It is what it is. There are no equivalencies in LP comparions. A crease across 4 ridges doesn't equal 8 points. So you can't make a rule that a crease is worth an exact amount. The only rule that you can make is 'look at everything'.

As for how to count discrepencies... I think our discipline is struggling with this question right now. And I'm not sure if anyone has a completely satisfactory answer yet. We've found a problem with bad exclusions that suggests that our current philosophy of 'one discrepency equals exclusion' needs some tweaking.


Pat,

I completely agree that all those aspects are important in what we do.

Things have changed dramatically in the past 10-12 years. As things continue to change, we will probably find through extensive studies that some of our dogma needs to be changed, while other dogma may be the best way to continue operating.

Re: British judge rules that Bayes Theorem can no longer be used

Posted: Wed Nov 02, 2011 11:17 am
by Michele
ER,

I don’t think the studies are worthless, I just think we need to acknowledge the factors of a study in order to know how much weight to put in the results. This will be important in moving toward a statistical basis as well, people need to know what the numbers mean so they can relay the information to the courts appropriately. This seems to have been a big issue in the case that started this thread.

As for rules that I would propose, I don’t think that’s for me to judge. All I’m saying is that if you want conclusions to meet an expectation then you need to state what that expectation is. When I was a new examiner, I was told the expectation was to get the right answer. That’s kind of vague and doesn’t give examiners enough information to meet the expectations. I’ve found that science has rules that are well researched and accepted. If we use these protocols then I think our conclusions will be on solid footing. The more protocols we use, the stronger our conclusions will be.

As for the validity, I think we’re using the term in different ways which is confusing the conversation. Comparing conclusions to ground truth answers is great but it’s looking at the accuracy (not the validity behind the conclusion). Comparing conclusions to panels of experts is another measure but it’s measuring reproducibility of the conclusion. I’m using the term valid to look at if the conclusion was arrived at in an appropriate way, which is different than if the conclusion is correct or reproducible. Carl Speckles wrote an article earlier this year called ‘Can ACE-V be Validated’ in the JFI. It’s a great article that expounds on the confusion behind all these concepts (accuracy, validity and reliability).

Re: British judge rules that Bayes Theorem can no longer be used

Posted: Wed Nov 02, 2011 1:36 pm
by ER
Michele,

I don't disagree with anything that you're saying here. I just think that there are some assumptions that are being made. You suggest that with a more 'statistical basis' or 'the more protocols we use, the stronger our conclusions will be'. I don't think these things have been proven yet. I know of some protocols (ahem, FBI, cough) that, I believe, lead to errors and produce weaker conclusions.

What I hear you saying (and correct me if I'm wrong) is that a valid answer is always right, but a right answer isn't always valid. That's a very interesting concept. How about this? For a specific comparison, can there be more than one valid answer? Not opposing answers, obviously. But couldn't 'exclusion' and 'inconclusive' both be valid answers? In some cases, couldn't 'ID' and 'inconclusive' both be valid answers?

My immediate question is, "What IS a valid answer?" How is a valid answer defined. I didn't really like Carl's definition of validity as 'the ability to measure the data'. Maybe it's just sematics, but I would define a valid answer as accurate and reliable but also well-founded and appropriate. We can show that an answer is accurate by using ground truth samples. We can show that an answer is reliable if most examiners reach the same conclusion. I think what you're suggesting is that a well-founded answer is based on q&q in some sort of objective and measurable way. While that's true, I think that an answer can also be well-founded if it's based on q&q in a somewhat subjective but demonstrable way.

If this is how we define a valid answer, then we could make that the expectation for examiners. We no longer expect the 'right' answer. We expect a 'valid' answer. The result of your comparison should be accurate (no bad ID's or bad exclusions), reliable (the same answer that most examiners would reach), and well-founded (you should be able to demonstrate and defend your decision based on the q&q of the print).

Re: British judge rules that Bayes Theorem can no longer be used

Posted: Wed Nov 02, 2011 9:37 pm
by Michele
ER,

I agree that valid means well-founded, using accepted protocols. Being able to demonstrate the basis behind a conclusion is how we show the validity. What is an ‘accepted’ protocol may change over time when a protocol is found to be deficient but I think it’s important to be able to establish what is ‘accepted’ and also why something is not accepted. As an example, many people put weight in the one discrepancy rule. I’ve seen many examples of discrepancies that I couldn’t explain but there’s still plenty of information to make an identification. Another example is excluding based solely on level 1 detail. I think it works the majority of the time but I’ve seen lots of examples where this protocol doesn’t work. For me, I feel like I could show that this protocol isn’t as valid as many people claim. I think showing the validation (the justification) behind each conclusion is much stronger than saying a conclusion is valid based on the appropriate use of ACE-V or based on my training and experience.

I don’t think that showing something is valid means you will always have an accurate conclusion. In casework, we don’t know the ground truth so we can’t establish the accuracy of the conclusion, the best we can do is show the validation behind the conclusion. If the explanation is valid (well-founded) then the conclusion can be accepted as accurate for the time being. It’s possible that the validity behind an idea can change over time and therefore give a conclusion less weight than it originally had.
We no longer expect the 'right' answer. We expect a 'valid' answer.
I think that’s a great way to put it!! I think that people who understand this concept will arrive at better conclusions. I’ve looked at quite a few of the reported erroneous identifications and the one thing they all had in common was that the examiner wasn’t able to demonstrate the validation behind their conclusion (they weren’t able to show a well-founded conclusion). I just realized that this applies to erroneous exclusions as well.