Fallacy of using Bayesian likelihood ratios in court

Discuss, Discover, Learn, and Share. Feel free to share information.

Moderators: orrb, saw22

Post Reply
Dr. Borracho
Posts: 157
Joined: Sun May 03, 2015 11:40 am

Fallacy of using Bayesian likelihood ratios in court

Post by Dr. Borracho »

NIST Urges Caution in Use of Courtroom Evidence Presentation Method

Use of 'Likelihood Ratio' not consistently supported by scientific reasoning approach, authors state.


October 12, 2017

Two experts at the National Institute of Standards and Technology (NIST) are calling into question a method of presenting evidence in courtrooms, arguing that it risks allowing personal preference to creep into expert testimony and potentially distorts evidence for a jury.

The method involves the use of Likelihood Ratio (LR), a statistical tool that gives experts a shorthand way to communicate their assessment of how strongly forensic evidence, such as a fingerprint or DNA sample, can be tied to a suspect. In essence, LR allows a forensics expert to boil down a potentially complicated set of circumstances into a number—providing a pathway for experts to concisely express their conclusions based on a logical and coherent framework. LR’s proponents say it is appropriate for courtroom use; some even argue that it is the only appropriate method by which an expert should explain evidence to jurors or attorneys.

However, in a new paper published in the Journal of Research of the National Institute of Standards and Technology(link is external), statisticians Steve Lund and Hari Iyer caution that the justification for using LR in courtrooms is flawed. The justification is founded on a reasoning approach called Bayesian decision theory(link is external), which has long been used by the scientific community to create logic-based statements of probability. But Lund and Iyer argue that while Bayesian reasoning works well in personal decision making, it breaks down in situations where information must be conveyed from one person to another such as in courtroom testimony.

These findings could contribute to the discussion among forensic scientists regarding LR, which is increasingly used in criminal courts in the U.S. and Europe.

While the NIST authors stop short of stating that LR ought not to be employed whatsoever, they caution that using it as a one-size-fits-all method for describing the weight of evidence risks conclusions being driven more by unsubstantiated assumptions than by actual data. They recommend using LR only in cases where a probability-based model is warranted. Last year’s report(link is external) from the President’s Council of Advisors on Science and Technology (PCAST) mentions some of these situations, such as the evaluation of high-quality samples of DNA from a single source.

“We are not suggesting that LR should never be used in court, but its envisioned role as the default or exclusive way to transfer information is unjustified,” Lund said. “Bayesian theory does not support using an expert’s opinion, even when expressed numerically, as a universal weight of evidence. Among different ways of presenting information, it has not been shown that LR is most appropriate.”

Bayesian reasoning is a structured way of evaluating and re-evaluating a situation as new evidence comes up. If a child who rarely eats sweets says he did not eat the last piece of blueberry pie, his older sister might initially think it unlikely that he did, but if she spies a bit of blue stain on his shirt, she might adjust that likelihood upward. Applying a rigorous version of this approach to complex forensic evidence allows an expert to come up with a logic-based numerical LR that makes sense to the expert as an individual.

The trouble arises when other people—such as jurors—are instructed to incorporate the expert’s LR into their own decision-making. An expert’s judgment often involves complicated statistical techniques that can give different LRs depending on which expert is making the judgment. As a result, one expert’s specific LR number can differ substantially from another’s.

“Two people can employ Bayesian reasoning correctly and come up with two substantially different answers,” Lund said. “Which answer should you believe, if you’re a juror?”

In the blueberry pie example, imagine a jury had to rely on expert testimony to determine the probability that the stain came from a specific pie. Two different experts could be completely consistent with Bayesian theory, but one could testify to, say, an LR of 50 and another to an LR of 500—the difference stemming from their own statistical approaches and knowledge bases. But if jurors were to hear 50 rather than 500, it could lead them to make a different ultimate decision.

Viewpoints differ on the appropriateness of using LR in court. Some of these differences stem from the view that jurors primarily need a tool to help them to determine reasonable doubt, not particular degrees of certainty. To Christophe Champod, a professor of forensic science at the University of Lausanne, Switzerland, an argument over LR’s statistical purity overlooks what is most important to a jury.

“We’re a bit presumptuous as expert witnesses that our testimony matters that much,” Champod said. “LR could perhaps be more statistically pure in the grand scheme, but it’s not the most significant factor. Transparency is. What matters is telling the jury what the basis of our testimony is, where our data comes from, and why we judge it the way we do.”

The NIST authors, however, maintain that for a technique to be broadly applicable, it needs to be based on measurements that can be replicated. In this regard, LR often falls short, according to the authors.

“Our success in forensic science depends on our ability to measure well. The anticipated use of LR in the courtroom treats it like it’s a universally observable quantity, no matter who measures it,” Lund said. “But it’s not a standardized measurement. By its own definition, there is no true LR that can be shared, and the differences between any two individual LRs may be substantial.”

The NIST authors do not state that LR is always problematic; it may be suitable in situations where LR assessments from any two people would differ inconsequentially. Their paper offers a framework for making such assessments, including examples for applying them.

Ultimately, the authors contend it is important for experts to be open to other, more suitable science-based approaches rather than using LR indiscriminately. Because these other methods are still under development, the danger is that the criminal justice system could treat the matter as settled.

“Just because we have a tool, we should not assume it’s good enough,” Lund said. “We should continue looking for the most effective way to communicate the weight of evidence to a nonexpert audience.”


Paper: S.P. Lund and H. Iyer, Likelihood Ratio as Weight of Forensic Evidence: A Closer Look. Journal of Research of National Institute of Standards and Technology, Published online 12 October 2017. DOI: 10.6028/jres.122.027(link is external)
See the article: https://www.nist.gov/news-events/news/2 ... ion-method
"The times, they are a changin' "
-- Bob Dylan, 1964
Boyd Baumgartner
Posts: 567
Joined: Sat Aug 06, 2005 11:03 am

Re: Fallacy of using Bayesian likelihood ratios in court

Post by Boyd Baumgartner »

Here's the paper referenced in the article.

Hari Iyer has been needling the L/R's appropriateness for some time now. I saw him at the Sacramento IAI conference and was intrigued by what he had to say. This paper seems to be directly confronting the Camp 2 people from this thread

From the article
“We’re a bit presumptuous as expert witnesses that our testimony matters that much,” Champod said. “LR could perhaps be more statistically pure in the grand scheme, but it’s not the most significant factor. Transparency is. What matters is telling the jury what the basis of our testimony is, where our data comes from, and why we judge it the way we do.”
I don't disagree with the sentiment, but I would argue the L/R is less transparent only in the sense that A) it leaves out information that it cannot detect (level 3 & population frequency) B) There's going to be a fair amount of Examiners that will just want the 'tell me what to say' answer and won't even understand what the L/R is doing anyway C) People arguably don't even know their own priors let alone the concept of a prior anyway. So I would argue that a chart accompanied by an explanation does the work of transparency a thousand fold over that of a convoluted, arguably non applicable calculation.
You do not have the required permissions to view the files attached to this post.
josher89
Posts: 509
Joined: Mon Aug 21, 2006 10:32 pm
Location: NE USA

Re: Fallacy of using Baysian likelihood ratios in court

Post by josher89 »

From what I've gathered in talking to those that have done research into L/R for FP comparisons, it seems that at least the common theme among them is when to apply L/R. What I mean is, non-complex or basic comparisons don't need L/R if they are easily demonstrable (as Boyd put it). This is for both ID and EXC.

I think where L/R may be of use is in inconclusive decision situations or extremely complex prints. Given what little I know about it, those L/R values can differ with the way that minutiae are plotted (sometimes significantly). So, what value does it hold for crappy prints that we cannot say yea or nay or if we cannot repeat the plotting across multiple examiners to get L/R values that are close to each other that could be statistically significant? I am not smart enough to answer.
"...he wrapped himself in quotations—as a beggar would enfold himself in the purple of emperors." - R. Kipling, 1893
Dr. Borracho
Posts: 157
Joined: Sun May 03, 2015 11:40 am

Re: Fallacy of using Baysian likelihood ratios in court

Post by Dr. Borracho »

Boyd Baumgartner wrote: Thu Oct 12, 2017 9:48 am Here's the paper referenced in the article.

Hari Iyer has been needling the L/R's appropriateness for some time now.
I had tried to read and understand Iyer's earlier article you referenced when it came out. Even though I have a B.S. strong in math and physics, trying to follow those equations and keep track of the symbols turned my brain to stone. Not only that, but if even if I were able to follow and understand those equations, it would turn a the jurors to stone if I tried to derive a L/R in court to explain my identification.

The new NIST article is plain English, which not only can I understand, I could explain myself in that language so a jury could understand. That is why I think this new article is better -- it has more value for my understanding and it will help me more in court.
josher89 wrote: Thu Oct 12, 2017 12:46 pmI think where L/R may be of use is in inconclusive decision situations or extremely complex prints. Given what little I know about it, those L/R values can differ with the way that minutiae are plotted (sometimes significantly). So, what value does it hold for crappy prints that we cannot say yea or nay or if we cannot repeat the plotting across multiple examiners to get L/R values that are close to each other that could be statistically significant? I am not smart enough to answer.
You have very succinctly stated the issue. A solution to the question you pose? I don't there there is one. Thank you!
"The times, they are a changin' "
-- Bob Dylan, 1964
Boyd Baumgartner
Posts: 567
Joined: Sat Aug 06, 2005 11:03 am

Re: Fallacy of using Bayesian likelihood ratios in court

Post by Boyd Baumgartner »

And here I thought you had a PhD, with your title and whatnot... :roll:
Dr. Borracho wrote: Fri Oct 13, 2017 2:50 am The new NIST article is plain English, which not only can I understand, I could explain myself in that language so a jury could understand
It is and has always been an issue of garbage in/garbage out. Arguably, it's the same issue that plagues black box experimental design, the quality of candidates you get back on an AFIS search and decision making in general. Another way to say it would be, 'You are what you eat'.

So if you feed an equation junk, guess what? it gets bloated. And if you starve your equation, guess what? it gets gaunt.

This in essence, is what the paper says when it says "but one could testify to, say, an LR of 50 and another to an LR of 500".

Hari Iyer's initial critique (the original paper I referenced) the way I understood it was 'You're standardizing what you're feeding the equation, but that's unwarranted because it's not based on anything'. Ideally, what you feed the equation is something like population frequency data. So, for us it would be how prevalent are certain feature types or combinations of feature types (See DNA/Hardy Weinberg). Absent that, it becomes more subjective and is something like the combination of what you believe about the rarity, significance and objectivity of the points in a print, which is logically more consistent anyway, but how are you going to convert that into a number? It's why pain charts are measured in smiley faces and not millimeters or amperes.

And none of this even gets into the notion of separating the examiner's ability to interpret what's there correctly vs ground truth, which is basically what PCAST gets into on the applicability of black box studies/proficiency testing to casework. Because to Josh's point and as g. has said for some time, the application is probably most useful in the inconclusive domain where things are going to be more complex due to the limited amount of data, the amount of interpretation, the reliance on features that not everyone agrees should be used, or some combination of those factors. That is to say, even inconclusive prints exist on a spectrum, so the applicability question gets even more thorny, more quickly.

Considering all the hype about overstating and certainty, etc. why even attempt to put a number on it? Why not just show your hand, let the attorneys attempt to frame it within the context of their narrative and let the jury do what the jury does and make render a verdict? It seems that so much energy is being placed into such a narrow focus with no real clear advantage in doing so. Or to rephrase what Josh is saying, the point at which the L/R is arguably most needed, it's the least powerful.
Dr. Borracho
Posts: 157
Joined: Sun May 03, 2015 11:40 am

Re: Fallacy of using Bayesian likelihood ratios in court

Post by Dr. Borracho »

Here's another analysis of yesterday's article by Lund and Iyer:
https://www.forensicmag.com/news/2017/1 ... 3dheadline

Likelihood Ratios—Can Objective Odds Come From Subjective Decision Theory?
by Seth Augenstein

Can human judgment and decision making, adjusting constantly to a changing set of circumstances, be quantified for juries and judges?

Likelihood ratios, a complex statistical model, attempts to do just that in forensic science, by putting numbers to the complexities of crime solving and detection. The LR concept is increasingly being used in European courts and elsewhere.

But they are an intensely-debated phenomenon in much of the world—especially in the U.S., amid the growing pressures to overhaul forensic science in favor of reproducible, quantifiable evidence.

A new study by two statisticians at the National Institute for Standards and Technology (NIST) contends LRs should not be used in American courtrooms—since the attempt at objectivity is actually only a mathematical model of an expert’s subjective reasoning process.

“We are not suggesting that LR should never be used in court, but its envisioned role as the default or exclusive way to transfer information is unjustified,” said Steve Lund, one of the statisticians, in a NIST statement on the work. “Bayesian theory does not support using an expert’s opinion, even when expressed numerically, as a universal weight of evidence. Among different ways of presenting information, it has not been shown that LR is most appropriate.”

The LR concept was proposed by Reverend Thomas Bayes in England in the 18th century, a way to marry probabilities to observations and decisions. The modern breakthrough in its usage was by Alan Turing, who used LRs to crack the Germans’ Enigma Code during World War II, providing a key advantage to the eventually-victorious Allies.

A simplified example of the use of LR was presented in the new NIST paper. Two urns have 100 balls inside. Urn 1 has 99 red balls and one green ball; Urn 2 has 99 green balls and one red ball. One of the urns is picked, and the balls are thoroughly mixed. The observer, if guessing which urn was picked at this point, would have a 50-50 chance of getting it right. But then a ball is selected, and it is red. The single ball selection means it is very likely that the randomly-chosen urn is Urn 1, where 99 percent of the balls are red (as opposed to Urn 2, where selecting the single red ball was very unlikely on the first pull). The observer would thus update their belief of which urn had been picked, based on the new information. This is a kind of Bayesian decision making, accounting for an altering flow of information.

(But the use of LR often involves a much more complicated set of circumstances—for instance, where each urn had an equal 50 red balls and 50 green balls, and three balls were selected).

Amid increasingly complex situations, an expert presenting such decision-making theories in a courtroom would only be presenting their own solitary logic process—and not a universal language that the jury could justly interpret, the NIST numbers experts contend.

In the wider vantage point of a criminal trial, involving other experts and detectives and numerous other players and factors, putting a number to a sole brain’s workings does not warrant a quantification, the pair argue.

“Thus, while decision theory may have a normative role in how a 'decision maker' processes information presented during a case of trial in accordance with his or her own personal beliefs and preferences, it does not dictate that a forensic expert should communicate information to be considered in the form of an LR,” they conclude.

The LR debate has potential ramifications in the increasingly complex world of DNA mixture interpretation. (NIST last week announced it would be conducting tests and analyses of the various mixture methods currently available to DNA forensic sciences, including probabilistic genotyping software programs like STRmix and TrueAllele.)

Some argue that it is simply the translation and communication aspects of LR that need to be refined for understanding by non-mathematicians. Some experts have suggested ways to express an LR “in plain English” to laymen, using the strength of the match between the evidence and a suspect as a numerator, and the possibility for a coincidental match as the denominator.

The two NIST statisticians indicate that DNA mixtures could indeed be an exception where LR could help parse out incredibly complex evidentiary situations.

“Forming a lattice of assumptions and uncertainty pyramid, including explicitly identifying what data will be considered, for applications in the field of high-template, low-contributor DNA evaluations could help to provide clarity to other forensic disciplines seeking to demonstrate or develop a basis for using a similar LR framework,” they write.

The debate over LR has been years in the making, and has only accelerated recently. A paper in the journal Forensic Science International in 2014 found the mere mention of numbers strengthened the persuasive power of experts. An entire issue of the journal Science and Justice was even devoted to the debate last year.
"The times, they are a changin' "
-- Bob Dylan, 1964
Bill Schade
Posts: 243
Joined: Mon Jul 11, 2005 1:46 pm
Location: Clearwater, Florida

Re: Fallacy of using Baysian likelihood ratios in court

Post by Bill Schade »

So we have heard from some who support this point of view.

I would be interested to hear the opinion of those who have supported the move to L/R

What is the average practitioner to do? Who do we follow? What direction would be the safest path forward?

Life sure was simpler in the 70's and 80's
Post Reply