As I promised, I will share my response to David Kaye's criticisms of my DNA study.
For all of those who were more than happy to jump and attack my work based on a one-sided debate, I only ask that you also read and consider my response, so you have a two-sided debate before you make your mind up.
Also, if I may, can I ask those who widely used and distributed Kaye's criticism of my work, to also please distribute my response, so everyone can hear both sides.
Here is my response (it can also be viewed at:
http://cci-hq.com/Dror_SJ_Cognitive_For ... search.pdf --where the tables and formatting is better):
Kaye [1] raises a number of important issues regarding the conclusions and experimental design that we report in a recent study on interpretation of mixture DNA [2]. Our study compares the decisions made by 17 forensic DNA analysts who examined a DNA mixture without biasing contextual information with two forensic examiners who examined the identical DNA mixture within biasing contextual information of a real criminal case. Only 1 out of the 17 examiners in the control non-biasing condition reached a decision of 'cannot be excluded', whereas both of the two examiners in the biasing condition reached a 'cannot be excluded' decision that was consistent with the biasing extraneous information. We conclude that bias may have played a role in the interpretation of the DNA mixture by the two examiners who were subject to the contextual information.
Kaye questions the strength of the data reported in our study, and based on his statistical analysis he concludes that the data "cannot easily reject the null hypothesis that the extraneous information had absolutely no effect."
I think that Kaye may be under estimating the weight properly assignable to the difference between the results of the control test group and the real world group exposed to biasing information.
Let us consider the data, and what is justified to conclude. Our data show the following:
Computing the exact p-value and the associated effect size (r-equivalent) based on the Fisher Exact Test [3], gives a p-value of .018 (one tail) and an r-equivalent effect size of .49
This comfortably and confidently, at the very least, justifies what we concluded in our study: that the data are "suggesting that the extraneous context of the criminal case may have influenced the interpretation of the DNA evidence" (p. 204, emphases added).
I think the correct statistic to apply to the data is the Fisher Exact Test [3], which is computed above. However, Kaye suggests a different statistical approach. Based on the control data (no biasing context) that 1 out of 17 examiners reached a decision of 'cannot be excluded' --what Kaye refers to as "the observed unbiased inter-examiner variability" --he computes 1/17 as the probability of an examiner reaching the decision of 'cannot be excluded' without the effect of bias, p = .058. And concludes with a concern that "with p> 0.057, one cannot easily reject the null hypothesis that the extraneous information had absolutely no effect on the Georgia group."
However, Kaye seems to be mistaken in his analysis, as the data is not from one examiner, but from two examiners, both reaching the same 'cannot be excluded' decision. Therefore, the correct calculation is not the probability that a single examiner will reach a 'cannot be excluded' decision without bias, but the probability that two examiners will both reach a 'cannot be excluded' decision without bias. This, of course requires computing 1/17 * 1/17, giving the probability of p= .003, not p= .058.
Such a low p-value "easily" rejects the null hypothesis. This sheds quite a different light on Kaye's criticism that our "study offers little experimental support for the claim that exposure to extraneous information was the cause of the disparity between the Georgia analysts and the 17 others who participated in the experiment".
The above computation indeed requires that the two examiners were totally independent, but even if they did not reach their decision totally independently, it still reduces the probability of .058, of only a single examiner. Furthermore, the examiners may have not been totally independent because one of the examiners knew the decision of the other --a common bias in forensic work, or because both were affected by extraneous information --the exact topic of the study.
The data even supports the conclusion from our study if I adopt a more conservative data analysis. Rather than analyzing the 'cannot be excluded' vs. 'other decisions', I can reduce the 'other decisions' by removing the 'inconclusive' decisions'. Thus, only using direct 'excluded' decisions when comparing to 'cannot be excluded' decisions (thereby reducing the n= 16 to only n= 12) --see table below.
Even this more conservative approach, the statistical analyses provide comparable results. Using the statistic that I think is most appropriate, the Fisher Exact Test, computes a p-value of .029 (one tail) and an r-equivalent effect size of .50. Using Kaye's statistic (now the "the observed unbiased inter-examiner variability" is reduced to 1/12), computes a p-value of .007.
Computing both statistics, and using even a more conservative data analysis approach, all yield comparable results: All clearly supporting the conclusions of our study. If anything, we have understated our findings, and took extra care and caution in the conclusions and claims we made.
Kaye has further concerns about our study, however, he criticizes us for conclusions we did not make. For example, he incorrectly attributes to us as claiming that "extraneous information was the cause" (italics emphasis added). We did not conclude or infer that that was the cause, but that it appears to have been the cause, that our study is about potential DNA bias, and conclude that it may be susceptible to bias. We clearly erred on the side of caution, and can be criticized for excessively playing down and understating our conclusions too much.
Kaye makes good suggestions for additional control conditions, which may allow isolation of the single causal factor without alternative explanation. This idealized suggestion is nice in theory, but is far from being realistic in studies into contextual bias in forensic work. Kaye would have preferred, as would I, a design with a larger number of participants examining test materials as part of routine casework, with all participants working under the same protocols, while experimentally controlling and manipulating potentially biasing information, all without their knowledge and the participants actually believing it is routine casework --these are all important design components in experiments on bias that perhaps allow isolation of a single causal factor without alternative explanations.
That would be wonderful. However, in practice, it is a long way from reality. For example, conducting research with such a design requires that involuntary participation in such “real casework” studies becomes a condition of employment in forensic laboratories (and perhaps part of the laboratory accreditation requirements). Of course, such deception of participants may not be possible if the ethical IRB approval demands that participants sign an informed consent form.
Studies that are based on data collected in the field, on data from real casework, offer great insights and opportunities. They more accurately reflect the reality of forensic work, and the data is from examiners who were actually within the contextual information and believe it (just as when we study the effects of alcohol on decision making we must actually give alcohol to the participants, we must make sure that when we study bias the participants truly believe the contextual information, along with all the pressures/reality of forensic work). These are critical for collecting data on bias in forensic work [4]. However, such studies, in contrast to experimental laboratory studies, do not allow freedom to control and manipulate conditions, do not provide all details, and ground truth is usually unknown.
Of course Kaye is right that our design leaves open questions about the contribution of other factors that could not be controlled for. His concern is that, in contrast to our 17 control examiners, the examiners in the criminal case may have examined (or 'peeked' at) the suspect's electropherograms before characterizing the profiles potentially present in the mixtures. Although Kaye's alternative explanation is plausible and might account for the decisions in the real case, it constitutes nothing more than another form of biasing information irrelevant to the characterization process, and therefore does not undermine any conclusion about the possible effects of bias.
Nevertheless, there is always the possibility of other potential confounds and alternative explanations for data from real casework, we recognized this in our study. That is why we qualify and only state that our data is "suggesting that the extraneous context of the criminal case may have influenced the interpretation of the DNA evidence." We need to think of how further studies can address the limits of the previous studies, how we can expand the literature on bias and other areas of cognitive forensics, that together, a whole set of studies provide a growing understanding of the issues, and how to deal with them.
Understanding the nature of cognitive bias and experimental design, especially within an ecologically valid forensic study, entails that a perfect study with all possible control conditions is not possible. Rather many studies, together, each making a small step forward and a contribution to better understanding of these cognitive phenomena are warranted. The first few steps are most difficult.
Therefore, we need to move forward with the small studies that are possible (what Kaye characterizes as "small observational" studies). The study we report on [2], includes both elements of a laboratory experiment and of casework, and shows:
1. That in an experimental setting, 17 DNA analysts examining identical data, using identical protocols, reached different conclusions. Thus reflecting the subjective nature and lack of consistency among examiners.
2. That in casework, two DNA examiners, within contextual information that suggested a suspect 'cannot be excluded', both concluded that a suspect cannot be excluded (in contrast to only 1 of the 17 analysts who performed an identical task, but without the potentially biasing contextual information). This suggests that they may have been influenced by the context --see statistical analysis, including effects sizes, above.
Kaye [1] raises good and valid questions, which we welcome. There are many factors at play in forensic work, and we cannot exclude that they too have contributed to the decisions made by the examiners. Nevertheless, the data from this study raises important concerns, and issues to be investigated. The issue of bias is complex and difficult, and there has been a lack of research in this area. Both because of the sensitive nature of this topic as well as the experimental challenges it presents.
The study we report [2] is particularly important, as DNA, with all its variations, has been clumped together and viewed as a Gold Standard, objective and not susceptible to contextual influences and bias. Our study, the first published study to examine bias and contextual influences in DNA interpretation, raises important questions about how realistic this view of DNA interpretation is, and calls for more research in this area.
Acknowledgements:
I would like to thank David Kaye for his thoughtful comments, they were insightful and thought provoking. I would also like to thank Michael Risinger, Reinoud Stoel, Jennifer Mnookin, and Robert Rosenthal for stimulating discussions about some of these issues. However, the opinions, findings, and conclusions expressed in this paper are the sole responsibility of the author and do not necessarily reflect the views of any one else.
References:
[1] D.H. Kaye, The design of “the first experimental study exploring DNA interpretation”, Science & Justice (2011), doi:10.1016/j.scijus.2011.10.003
[2] I.E. Dror, G. Hampikian, Subjectivity and bias in forensic DNA mixture interpretation, Science & Justice (2011), doi:10.1016/j.scijus.2011.08.004
[3] R. Rosenthal, D.B. Rubin, r-equivalent: A simple effect size estimator, Psychological Methods 8 (2003) 492-496.
[4] I.E. Dror, On proper research and understanding of the interplay between bias and decision outcomes, Forensic Science International 191 (2009) e17-e18, doi:10.1016/j.forsciint.2009.03.012