Page 1 of 1
#Article: The Mismeasure of Science
Posted: Wed Jun 15, 2011 6:23 am
by Boyd Baumgartner
http://www.plosbiology.org/article/info ... 071-Gould2
Summary: An interesting discussion on the Steven Jay Gould / George Morton controversy and the discussion of bias.
I particularly liked this section:
Biased Scientists Are Inevitable, Biased Results Are Not
Samuel George Morton, in the hands of Stephen Jay Gould, has served for 30 years as a textbook example of scientific misconduct [12]. The Morton case was used by Gould as the main support for his contention that “unconscious or dimly perceived finagling is probably endemic in science, since scientists are human beings rooted in cultural contexts, not automatons directed toward external truth” [1]. This view has since achieved substantial popularity in “science studies” [2]–[4]. But our results falsify Gould's hypothesis that Morton manipulated his data to conform with his a priori views. The data on cranial capacity gathered by Morton are generally reliable, and he reported them fully. Overall, we find that Morton's initial reputation as the objectivist of his era was well-deserved.
That Morton's data are reliable despite his clear bias weakens the argument of Gould and others that biased results are endemic in science. Gould was certainly correct to note that scientists are human beings and, as such, are inevitably biased, a point frequently made in “science studies.” But the power of the scientific approach is that a properly designed and executed methodology can largely shield the outcome from the influence of the investigator's bias. Science does not rely on investigators being unbiased “automatons.” Instead, it relies on methods that limit the ability of the investigator's admittedly inevitable biases to skew the results. Morton's methods were sound, and our analysis shows that they prevented Morton's biases from significantly impacting his results. The Morton case, rather than illustrating the ubiquity of bias, instead shows the ability of science to escape the bounds and blinders of cultural contexts.
What does this say about the application of ACE-V? Is it constructed in a way that prevents bias? That is to say, is it normative?
The same questions also apply to studies that claim to study ACE-V or decisions in general.
Re: #Article: The Mismeasure of Science
Posted: Wed Jun 15, 2011 12:14 pm
by briano
Re: #Article: The Mismeasure of Science
Posted: Thu Jun 16, 2011 12:34 pm
by Gerald Clough
I don't think it's possible, in any real practical operation, to design a process entirely free of bias. ACE-V is, as I think we pretty well thrashed out a while back, not necessarily strictly linear. "A" and "C" can interact back and forth. But I think the worry about bias in examinations largely grows out of what I think is a mistaken analogy to scientific experiment. Hypothesis testing is simply not what's going on in examiners' heads. There's not really a "hypothesis" for an examiner to attach to and to bias in favor of.
But we began this with reference to an article on Morton, and we should probably see if it can relate to LPE. Morton's objective metric observations were no doubt highly accurate. And unbiased. There was no bias reason to exaggerate or minimize anatomical metrics. (It really appears Gould's own prejudice to presume Morton's common 19th century bias would infect his metrics biased Gould's study of Morton.) He didn't need to, because he was observing a physical trait that varied according to his preconception that some such trait would be found to fit the bill. His prejudice was that there were categories of racial performance, his own white race being superior and other races being "crafty," or having merely selfish motives for caring for their children, or eating disgusting food. It suited his prejudice to select cranial capacity as an indicator. He, of course, has what sounded like a rational explanation, that larger crania accommodated bigger and therefore better brains. (Counter-examples like as Anatole France being conveniently forgotten.) We might well wonder whether, if cranial capacity wasn't so variable, Morton wouldn't have emphasized frontal development alone or some other trait. But the point is that Morton didn't decide it just had to be cranial capacity that mattered and fudged his measurements to make it appear so. He picked a correctly measured trait and warped its interpretation to fit his prejudice. It wasn't anything new. Pieter Camper had already done essentially the same thing with his "facial angle." If you set out to find some data to fit into your prejudice, you'll inevitable find it - and maybe found a field of science in the process, no matter how mistaken your conclusions.
But in our terms, Morton Analyzed the metric data. Morton could be said to have accurately Compared the metric data with his peculiar "data" from observing races. His metrics seemed to him consistent with the racial categories, because he made conclusions about intellectual and cultural performance based on his prejudiced interpretations. His Evaluation, then, was that cranial capacity correlated with his flawed categorization of relatively "better" or "worse" characteristics of the races. A similarly prejudiced peer would have Verified by misinterpreting the same characteristics as data.
Morton's observations were rather like an examiner viewing two pristine inked impressions. Simple metrics. Very little room for prejudice to bias the measurements, which in the LPE case are the decisions of the nature of various details and the nature of their interrelationships. Bias, in that case, is limited in potential to the "E"valuation stage where it's a question of a judgment of sufficiency, an individual judgment that will probably not form internally in just the same way in two examiners viewing the same case, even if they agree on conclusion.
Of course, in latent work, far more individual judgment comes into play in interpreting the impressed details as to what physical reality they represent. Now, I was about to say that if ACE were strictly linear, there would be little room for bias, because Analysis would take place and be finalized before any comparison. But, aside from the fact that the examiner is almost always exposed to the record impression before beginning the examination, there is a more subtle potential bias. Let use just presume that some prejudice is at work under some pressure to make an identification from this examination. Especially in a sparse latent, I know that finding some quantity of detail that I judge clear enough to reliably characterize makes it more likely that I will make an ID. I need some good points. My tolerance may be more slack. I essentially guess. I'm guessing what's most likely (not what I an sure of), but that's okay with my prejudice, because (1) that's most likely what that detail is, and (2) I still get to interpret the record and infer how the physical skin the record represents can present in the latent as what I said it was. That's a long way around to explain why I think even linear ACE cannot proof against bias. Obviously, another mode, "recursive" or "circular," or whatever, allows bias more freedom to alter interpretations.
However, I do NOT think prejudice is a significant routine operator among competent examiners. It can operate, of course, as we well know, and it can operate rather discouragingly among the incompetent. And when we add the "V," we further reduce the potential. And when we address the more global process issues and cultural issues (FBI response), we reduce it further. But all observation of the scientific/technical sort is subject to bias and prejudice, even if it's merely the desire to observe meaningfully. I suspect there's not such a great difference between Morton's "disgust" about one race's supposed food preferences and his assumption that it spoke to intelligence and a negative attitude toward a Muslim suspect in a terror crime. What's most important is that people cannot see their own operative prejudices. If they could see them, they wouldn't operate.
And to the other level, whether you can study ACE-V without bias. Of course not. No one is purely neutral. At the very least, you want data. And what are you really studying? If the examination mode is not ACE-V, what mode is it? What do you compare ACE-V to? I don't think there's anything to compare it to. So all you can do is try to see whether an examination conducted under the only rational way it can be conducted produces correct results. That's FACTUALLY CORRECT results, not the conclusion an examiner should reach. And I don't think you can say that in any test situation that uses any kind of realistic data you can say what any given examiner should conclude, only what you or your committee or your majority think the examiner should conclude. How on Earth do you study bias when you can realistically say that one individual's conclusion is correctly concluded and another not? If in your study, you only consider a response correct if it's factually correct, you've already biased your study toward the invalid notion that all examiners will conclude alike and consistent with factual reality. If you try to prejudge what conclusions were correctly produced, regardless of reality, you're necessarily prejudging correct tolerances to conform to your or your advisors' individual guesses. Back to Gould studying Morton. Nope. You do the best you can, and never presume you can't be biased.
Re: #Article: The Mismeasure of Science
Posted: Fri Jun 17, 2011 2:55 pm
by Neville
Hi Gerald,
You stated 'aside from the fact that the examiner is almost always exposed to the record impression before beginning the examination,' I have been pondering this point for some time and can make no sense of it. Surely you must examine the latent first before entering it into AFIS. If you do not have AFIS (and I am sure you do) or if you are handed the ten print set and the latent, then you would have to look at the latent first to decide which finger you are going to compare, you can't surely be suggesting you remember the target points of all ten fingers before looking at the latent.
Hi Boyd
This guy you are talking about must have had less knowledge than I do about history, a quick scan of the brain I think of the Ottoman empire, the Byzantine Empire, the Persian-Mede and Egyptian eras all highly advance cultures, lets face it the Arabs were performing complexed medical procedures long before his forebears had stop an existence living scratching about in the dirt. Am I correct in saying 1844 the Victorian era not long after the slave traders stopped it is not surprising really, brain washed rather than being biased.
Re: #Article: The Mismeasure of Science
Posted: Mon Jun 20, 2011 1:53 pm
by Gerald Clough
Boyd asked if ACE-V could be structured to prevent bias. My observation was simply that only a strictly linear examination had any anti-bias built into it, and that, aside from the impracticality of working with strict linearity, there was the practical issue of having seen the record print, during a prior comparison or some other way. I doubt that many examiners strictly isolate themselves from the record print before analysis of every latent. And few would prevent themselves from reanalyzing a latent after having seen something in the record print that gave them insight, insight that always has potential for being contaminated by bias. The larger point was that some processes are amenable to a largely bias-proof design. Photo line-up procedures recognize the potential harm and recommend that an officer with no knowledge of the case administer the line-up. The classic double-blind drug trial is another. I just don't think there's much way to proof a latent print examination against bias potential without crippling a significant amount of its potential for discovering identifications. There are many, many other scientific and technical processes aimed at gathering data and concluding from it that are also potentially biased. It's just the way it is. And our tolerance is so low that the process absorbs most of the bias effects.
Re: #Article: The Mismeasure of Science
Posted: Mon Jun 20, 2011 2:30 pm
by Tazman
I'm sorry, but in spite of all of the efforts the British are putting into proving bias is a serious problem, I have yet to see any significant number of cases in which bias has led to an erroneous identification. Okay, "Brandon Mayfield" raises the number up to "1." Can anyone enlighten me differently?
Re: #Article: The Mismeasure of Science
Posted: Tue Jun 21, 2011 1:51 pm
by Boyd Baumgartner
I would make the assertion that no one has substantively demarcated or shown bias (with the exception of Ken Moses).
Since Tazman brings up Mayfield, and there's arguably a root cause analysis performed in the OIG report, I'd say let's start there.
A Point to consider:
The OIG report brings up both an unusually similar print and Mayfield's status as a practicing Muslim. On top of this, I've heard Ken Moses speak back in 2005 where he says he was biased by the results of the original examiner.
Discussion:
How can you untangle bias and/or measure it in the form of 'unusual similarity' (known to unknown bias), from preconceived notions of guilt (assuming that's what the mention of Mayfield's religion can be classified as), from confirmation bias from an independent examiner (Ken Moses)?
While these are extensive topics of which I'm only scratching the surface by asking the questions, I think it boils down to lack of norms (call them standards if you like). I'd have to say that I agree with Gerald in the regards of never eliminating bias, but on the flip side of that, seeing bias everywhere like the boogeyman under your bed is a bit on the other extreme of the spectrum.
It's conceivable that if the FBI had norms in place that culturally drove them to base Evaluations on demonstrable, objective, justified data to the satisfaction of consensus, making use of QA measures such as consultation and/or blind testing along the way, they might not have made the error. Instead, as the other thread indicated, we get the implementation of a strict methodological adherence based on what I would call a mistaken assumption.
Edit: I would also point out that the reference to unusual similarity is not warranted due to the fact that A) it met a threshold of sufficency for Individualization to another person, and B)other than the subjective probabilities that underlie the identification process, there are no accepted probabilistic measures from which to make, or by which to evaluate this claim.
Re: #Article: The Mismeasure of Science
Posted: Wed Jun 22, 2011 8:40 am
by Gerald Clough
Tazman wrote:I'm sorry, but in spite of all of the efforts the British are putting into proving bias is a serious problem, I have yet to see any significant number of cases in which bias has led to an erroneous identification. Okay, "Brandon Mayfield" raises the number up to "1." Can anyone enlighten me differently?
Bias always operates. It's an essential part of the universal process of working efficiently in decision-making. The negative connotations of various perceived undesirable biases not withstanding, everything that operates in human interaction with the environment evolved to a useful purpose. The bias that evolved to increase the odds of survival in more primitive situations is, however, not altogether compatible with more deliberate processes toward modern purposes. I think this has to be put in perspective, so that bias is not seen as a lurking monster that will always hijack decision-making, nor should it be seen as something that can be eliminated simply on account of awareness. We are not helpless in the inevitable presence of bias.
I think the standards and practices in latent print examination are such that biases indeed don't produce any sort of general unreliability. (Nor does it in the many other realms of human study and practice.) In other words, there's no reason to consider every result less reliable on account of factoring in bias. Best practices now in place tend to establish the fundamental data within an examination without much risk of it being corrupted by bias. And the high threshold almost universally in place in the LPE community dramatically reduces the potential for any bias affecting the outcome. This does not mean that there must be an automatic presumption that no examination can be seriously affected. All criminal forensic opinion evidence should be subject to critical review. Bias is just one potential reason. So it plain administrative error, a major feature in Cowans. So is plain incompetence or insufficient competence, which is always a possibility. These are true for every field. Peers in astronomy don't presume each other to have been corrupted by bias, but they consider the possibility.
A corollary to this notion is that we can never really know how many erroneous identifications were made. The nature of criminal investigation, in which a latent print is frequently not the sole evidence, makes it quite possible that the actual criminal actor may be erroneously identified as the source of a latent print. But logically, it would seem that if such a thing were more frequent than rare, that there would be more Mayfield-class errors provable by knowledge of the factual truth that the person simply could not be the source. We would be even more confident if latent evidence was more often subjected to critical review.
I would further suggest that we know far too little of human perception and memory to imagine that we can identify pure bias errors or even craft a relatively realistic experiment that can reliably distinguish bias. All that we have observed of the workings of the human brain, often from studying subjects with various defects, makes it clear that practically all working cognition takes place "behind the curtain," out of the view of the observing conscious. Very persuasive argument suggests that interpretation of observed experience and the formation of memories are highly creative and the result of some very fluid processes over which we have very little, perhaps no, control. Even a slight extension of what is reliably known from brain imaging studies makes it almost certain that the analysis that we consciously imagine to occur during the instant that we intentionally observe is really produced independently and unconsciously in the non-conscious and presented to the conscious for adoption. (Imaging studies show that the response to a signal is unconsciously formulated in the appropriate portion of the brain shown to become active well before the subject is conscious of observing the signal. We only think we are doing the responding. Consciously, we are just along for the ride.)
That non-conscious brain, though, is not unlearned. Is, in fact, the only part that learns. So awareness of the hazards of bias and other potential factors becomes part of its learning and is accounted for in its processes where such factors should be a part of the process, and thereby their potential is lessened. But because it is all happening on the other side of consciousness, we're not going to be able to get a clear view of the works. We can only judge the effects by noting, as you say, the apparent low incidence of fatal error.
Re: #Article: The Mismeasure of Science
Posted: Wed Jun 22, 2011 9:17 am
by Michele
Gerald,
I think the standards and practices in latent print examination are such that biases indeed don't produce any sort of general unreliability.
Best practices now in place tend to establish the fundamental data within an examination without much risk of it being corrupted by bias.
What standards, practices, and best practices are you referring to? I haven't seen any of these things to be consistent enough between examiners or agencies to make such a claim.
Re: #Article: The Mismeasure of Science
Posted: Wed Jun 22, 2011 10:39 am
by Gerald Clough
Michele wrote:Gerald,
I think the standards and practices in latent print examination are such that biases indeed don't produce any sort of general unreliability.
Best practices now in place tend to establish the fundamental data within an examination without much risk of it being corrupted by bias.
What standards, practices, and best practices are you referring to? I haven't seen any of these things to be consistent enough between examiners or agencies to make such a claim.
I think there are general practices that apply widely among all but the most lame excuses for examiners and units. However, keep in mind that one reason I would never want the potential for a biased result to be ignored is that the standards and practices are really not entirely universal.
First and foremost, the threshold for interpreting an impressed detail as having a particular character and the threshold for the amount of consistent data required for identification, particularly the identification threshold, are so high that I think the effect of normally present biases is merely to nudge the thresholds a bit. They are already so high that it requires a major cascade of effects to lower a threshold so much as to produce an error. The known error cases suggest it takes a whole array of different influences to produce plain error. Whether individual examiners think about it or not, we run our thresholds so high because we don't have an objective measurable threshold. We so much want to avoid error that we "tune" the whole examiner community to this kind of high threshold. We probably could cut things much finer and still make good ID's, but we don't know where the danger point is, so we stay high. I think this level of threshold is so ingrained that it's relatively high even in the least effective training.
ACE-V, which we have to view as a model from which we diverge in practical examination, is still motivated by our awareness that our analysis of latent features can be influenced by a record print, under any number of biasing influences. We don't fanatically adhere to linear ACE, and we don't have to, but it has been woven into our approach to examination. Not only do we always operate with that idea of bias-immune analysis in mind, but it disciplines us in general.
Verification, while not so specifically built in to avoid biasing influences, still inserts another examiner into the process, so you have to have not one examiner so affected as to push his thresholds fatally low, but it takes another similarly malfunctioning to put the erroneous ID through. It can happen, but it's terribly rare and requires a rather extraordinary circumstance.
Also, examination in general isn't treated as an off-the-cuff exercise in mere guesswork. Even the least trained takes it seriously and approaches it as a deliberate operation aimed at accuracy and truth. That gives biases an immediate significant hurdle to overcome. This treatment of examination as a serious matter of a demanding technical sort is something we take for granted, but among all human decision making, it has to seen as part of the examination "standard."
What I find interesting is that if I were forced to guess whether biasing influences were more likely to produce one of these rare errors among the most highly trained and experienced examiners or among the inadequately trained and lightly experienced (or maybe just badly experienced) examiners, I could not decide which to choose. I am kind of inclined to expect that sort of error more from the higher class of examiner, because they more confidently deal with much more difficult complex latents where there are many more interpretive decisions to make and therefore more exposures. I suspect the lesser lights might not even recognize the potential for a match in those iffy, complex latents and therefore never undertake the comparison. And my feeling is that the threshold for how poorly impressed a detail may be and still be characterized for comparison gets lower as experience becomes vast, and that that can happen to the degree that other influences reach error-producing potential. I can't back that up - it's just a notion. But the errors I recall from poorly trained or poorly managed examiners have tended to be simply missing rather obvious points in rather plain and uncomplicated latents.
(I formulated that comparison of high-level examiners versus inadequate examiners with the U.S. situation in mind, there being almost no effective standard, other than talking a judge into letting you testify. There are many, many people doing examinations who we never hear from and who have no conception of the issues we discuss and who have never heard of the famous error cases, ACE-V, etc. Obviously, the situation may be very different in other places. But there may still be differences between the most experienced and those lightly experienced. And no matter where or how well trained, there are always people more and less conscientious, thoughtful, and ethical.)