Page 2 of 6
Re: Recent News - Articles and Discussion
Posted: Sat May 16, 2009 10:12 am
by Pat A. Wertheim
Thanks. We are on the same page here. I think the research is valuable and should continue, that's the main thing I'm saying. I remember ten years ago when Christophe Champod told me he envisioned the day we would have a button on the computer that we could push after we made an identification in AFIS, and the computer would calculate a statistical probability to support the identification. I laughed and told Christope that I wished him luck because that would be a valuable thing, but I didn't think I would see it in my lifetime. Well, taking the class with Cedric and Glenn showed me that I may just be wrong -- I may see it in my lifetime after all. The model Cedric demonstrated is a major step in the right direction, but it is still a long way from validation and acceptance. If you ever get the chance to take that class with Cedric, don't miss the opportunity. You would enjoy it. And who knows -- you may even be encouraged by what you see, as I was.
Re: Recent News - Articles and Discussion
Posted: Mon May 18, 2009 9:10 am
by Gerald Clough
mdavis wrote:
We have gotten into more trouble with this statistical modeling than we will ever get out of. We are headed down a narrow path and the vines are getting closer. What was wrong with staying on the pavement, using all three levels of detail and years of training and experience, and calling our results "expert opinion?" A newborn can quickly recognize with essentially 100% accuracy its own mother. The fact that a comparison discipline does not lend itself to "scientific validation" does not render it "invalid."
That last sentence is entirely true. There is no rational person genuinely involved in the questions who holds that fingerprint identification is invalid. That would be to hold that it cannot be done. But the problem with calling a conclusions of identification "expert opinion" is that it represents taking refuge in clinical guessing. In saying that, I had to chose between calling it "clinical opinion" and "clinical guessing." "Opinion" seems to me to imply a confident belief. "Guess" sounds a bit too casual and imprecise. I find "guess" more defensible, for all that it implies little weight. If I state an "opinion," I should be able to defend that belief in precise, concrete argument. I should have a proof. But, given the current state of knowledge, the critical components of my "proof" turn out to be mostly arguments from experience.
But the real problem is that the legal component of the system is demanding a more solid foundation for expert conclusions that carry the weight of scientific certainty. As I have said before, clinical proposals of likely fact have value. Most medical practice is exactly that kind of highly experienced "guessing." Psychological and psychiatric testimony, so prominent in legal questions of competency, sanity, and retardation, is entirely clinical. That is why it is almost always contested by experts offering contrary opinions. If I say that my opinion in latent print examination should be accepted on account of my experience and my knowledge of the experiences others, then anyone could and should consider another expert's opinion to the contrary, according to their judgments of relative credibility. If neither of us can demonstrate that we applied knowledge gained through an acceptable scientific inquiry that has shown that, presuming that the characters of all details have been unambiguously determined, that identifying a quantity of relationships among details exceeds a tested threshold means there is a specific, very high probability of identification, we are making clinical guesses. (We would be decidedly unhappy to accept a rate of erroneous criminal convictions similar to the overall rate of medical misdiagnosis.)
I think the reason psychologists are not held to the standard of scientifically validated conclusion is that it is difficult to even conceive of a way to rigorously quantify their observations and a way to establish any set of knowns to test against. But it's all there is to assist in answering the legal issues. There being no better way, we use what's there. If kind of works, simply because the legal issues are stated in clinical terms to start with. When the court says a mentally retarded person may not be executed, it means "mental retardation" in clinical psychological terms. And the courts accept that they will often hear two contrary opinions and that they must judge which of the two to accept.
That is NOT what you want to see happen in disciplines like ours. You do not want the better rhetorician to win. You do not want the better resume to win. You do not want the win to go to merely whoever can make the jury feel they understand the argument. Or even to the eldest or the most impressive figurehead. Those are all factors in who wins contests of psychological testimony. The closed shop that has kept latent print examiners as a whole within similar nonspecific thresholds for conclusion is opening, and awareness of the issues is growing. Without scientific validation, which means probabalistic conclusions, you have no answer to another examiner who believes, from the same kind of experience and knowledge, that a more conservative threshold is required for identification.
We need not be able to quantify every human factor of latent print examination to show its conclusions valid and quantified. Yes, we will validate conclusions using datasets that are clear impressions, without artifact or significant distortion. In fact, we may do it with idealized data generated by proven statistical models of friction ridge skin. Such validation will cause much of the bias argument to drop away. Issues of perception, cognition, and bias will remain to be argued in how the data is identified in individual examinations. The credibility of that observation and interpretation can and should be argued. The OIG argued exactly those interpretations in their Mayfield analysis in a very convincing way. I don't think there's any doubt of what effect would have been on a jury had that kind of presentation been made in a trial.
The desire to maintain the status quo of rendering opinion from true expertise exerts a powerful pull, especially since we feel the conclusions are accurate. But fingerprint identification is better than that. It can and will be scientifically validated. And admitting questions of accuracy of interpretation of features to create data for analysis will at once protect against error and demonstrate why examiners made those observations with confidence. And it may well bring in issues of likelihood that the individual could have been the actual source. And that is appropriate and in no way a bad thing in the process of finding truth. In the vast majority of fingerprint cases, there will be no credible argument over what the impression represents and, when validation studies have been accepted, little argument over whether or not the conclusion of identification is both accurately stated and powerfully persuasive.
I don't know that newborn recognition of the mother has much to say about the issue. Their recognition process is poorly understood. They seem to work well off of scrambled features and a study suggests that there is a significant, perhaps even very powerful auditory component..
Re: Recent News - Articles and Discussion
Posted: Mon May 18, 2009 3:46 pm
by Les Bush
Greetings from down under,
Thanks Pat and Gerald, I agree the door once closed is being opened to examine the basis of how fingerprints are individualised. The fingerprint community by virtue of recent events such as Mitchell, McKie and Mayfield has been challenged to clarify to the courts and scientific community that reliability of the conclusions reached remains trustworthy. By entering the arena of science our procedures require the ability to be validated with an appropriately defined methodology and technique. The one issue that has been a constant pain is the definition of sufficiency. But put in context with the nature of the physical reality that is fingerprinting then by having a blurred standard does coincide with the highly variable parts of fingerprint pattern formation and transfer onto receiving surfaces. Being practitioners we know and accept that each fingerprint examination is not going to repeat a previous comparison experience. In this sense we remain flexible while ensuring that every examination is done with accurate application of the principles. The design of algorithms incorporating subjective probabilities about the occurence rates of combinations of fingerprint features is not flexible and ultimately requires a tolerance that adjusts the likelihood of supporting the conclusion. It must be remembered that those proposing statistical models are not advocating individualisation but a new subjective measure of identification using arbitrary weightings from a limited sample population. So how do we practitioners hold our own belief that no two fingerprints are the same. Using the word 'forensic' from its latin origins we must argue logically and scientifically to both courts and science forums. Without a presentation program that captures all the decisions of a comparison and allows them to be discussed and negotiated our work remains in the 'black box' of the experts mind. We also need validation studies about each aspect of fingerprint identification science and considerable investment by all practitioners to elevate the understanding of how we individualise. Pushing a button on a computer and getting a mathematical result is not dissimilar to the way we use AFIS with the difference being that having tested AFIS we still apply the human mind to sort the plethora of variables in order to reliably conclude a result. The science successfully does this for both tenprints and latents. If Gerald is inclined to lead us I'd invite him to spell out a simple program of testing about how we can validate our practices as a counter argument that we have a basis for confidence in our results. Cheers from oz. les
Re: Recent News - Articles and Discussion
Posted: Tue May 19, 2009 7:57 am
by Gerald Clough
Les Bush wrote: If Gerald is inclined to lead us I'd invite him to spell out a simple program of testing about how we can validate our practices as a counter argument that we have a basis for confidence in our results. Cheers from oz. les
I don't know how simple it is. The idea is simple. The devil is definitely in the details. I frankly don't know how slippery those devils might be and won't, until someone tackles the project, but a number of folks are working on it.
There are two components. One is fundamental validation of the proposal that fingerprints are very powerful identifiers. Power, as it applies to such things, means the ability to declare that there is a very, very high probability that a particular individual is the source.
The second component of identification is the ability to say with high confidence that any potential sources (the individuals who might exist according to the statistical likelihood) bear a portion of friction ridge skin that produces an impression that is in no way observably different in the character of its pore unit ridges from the latent under examination.
The two components may sound like the same thing. But the first is a statement of validity. The second is a statement of reliability. Validation is done by by comparing a very large set of ideal representations of fingerprints. To do this, you must develop an unambiguous way to code the characteristics you wish to use, the characteristics you propose to use in actual examinations. (These need not be what you use now. For instance, you may wish to define a single coding that takes in two characteristics that have similar flows but would be called by different names because of vagaries in how they were impressed, even when they are impressions of the same skin. Or you might not. See below.) And you must develop a coding for relationships among features, because those will be used in examinations.
Having those codings developed sufficiently to characterize any ideal impression, you can run the comparisons. You must, of course, also have a quantifiable measure of how much contiguous impression is in agreement between two impressions. This may be defined in terms of area, number of features, or both. After you run this for each member of the set against the whole set, you will have the numbers telling you how many members of the set might match a single members of the set, that probability being a function of something like what we now call the quantity of detail. So, you don't have a single threshold for identification. You have various thresholds that can each put the probability past a particular likelihood.
There is a great deal more to it, of course, and I am pretty sure it will require many validation studies to validate many variations in what we use as accumulating data toward conclusions. (For example, you can validate accumulating data across simultaneous impressions.) But that's the idea. That's validation of what we proposed we could do with fingerprints. Now, we confront the reality that, while we had to use idealized data in order to get solid numbers from our validation study, impressions are often not ideal, or perhaps it's more accurate to say that there are often features that are not ideally impressed but which we can, with confidence, assign to one of the specific coded characteristics we used in validation. We do that now, when we have a continuous ridges impression of a bifurcation in one image and a small gap at the junction in another. This not a fatal ambiguity, and here's the solution. And the thresholds may end up being defined by quantities of different types of details that occur with different frequencies in the population, or by some measure of complexity, or both, or something I haven't thought of.
Why did we use idealized data in validation? Because we must have that benchmark established from idealized data before we introduce the ambiguities. But the data we will accumulate to see which thresholds we can cross in an examination are made by OBSERVING and by INTERPRETING. We could make errors in either. We don't all have the same ability to judge gradations of density. We have varying degrees of experience and insight. When it comes to less than ideally represented details, we can differ in how we code the data. We don't want that sort of thing running loose in validation, because it's just not part of the fundamental validation issue.
How are we to settle the interpretation issue? There is no absolute way to determine which interpretation is correct. We can say this. If one of us is wrong, if the two impressions are of ridge details that truly have different characters, it may bar identification. How do we decide? Again, there are two factors. One is plain credibility. We each demonstrate and explain how we decided on our particular interpretations. The other is that we invoke the results of the validation studies to determine how likely it is that the actual sources of the rest of the impression are different. There's something going on at that point in the impression that may or may not be an impression of a particular detail. I can be guided in judging how likely it is that they are the same genuine feature by the weight of the other details. It's akin to what we are talking about when we say we use Quality to judge how much detail is sufficient. I do NOT mean one or the other interpretation is automatically defeated. I mean the likelihood supports one or another position.
This indeed admits the possibility of testimonial contests over the interpretation of details. That's the reality of a world in which not all the details are clear. I would point out that these interpretations are NOT SUBJECTIVE. They are not made through external bias. They are studied, expert judgments. That is objective.
Now, I know this tends to make it sound like it's a major change in paradigm. It's not. It allows us to state, once we settle on our interpretations of details, a conclusion with great confidence, because our conclusion, which must be probablistic, has been scientifically validated. And there will be relatively few contests over interpretations if we craft our detail coding scheme well. We could, in fact, do our validation using very broad definitions of details, so broad that there would be almost no arguments. But that might leave us with thresholds higher than we would practically like to be bound to, but it demonstrates that this whole thing is far more dynamic than it first appears. In most cases, the interpretations will be very strong, and the probabilities will lock in nicely and will be very powerful.
The issue of bias will remain. More study is needed to determine where in the process, to what degree, and under what conditions bias may influence interpretations in the investigative phase of a case, so that we better avoid error. And litigation is another phase in which others (with different potential biases) get to reexamine the evidence.
The big changes? The issue of all-or-nothing zero tolerance in identification will evaporate. The position that research supports no conclusions of likelihood will become invalid. The vocal (and largely correct) critics will have to find other topics for their papers. All those forensic science students now crowding into universities will have no end of possibilities for their dissertations. And latent print examination will have moved from clinical practice to valid scientific analysis.
Re: Recent News - Articles and Discussion
Posted: Tue May 19, 2009 3:50 pm
by Les Bush
Thanks Gerald
Your posting is very encouraging as the picture painted is not bleak but promising that with effort our community can meet the expectations of scientific validation. I particularly liked the consequence of several examiners producing different thresholds in their comparison conclusions, this is an example of the blurred definition of 'sufficiency' but the validation test creates meaningful data provided the test design is correct. The other aspect I liked was the selection of samples for testing, we know from experience that latents can be a real challenge for interpretation and decision, in a validation test the science as much as the examiner is being evaluated so my thoughts went in the direction that 'normal' reproductions of fingerprint patterns should be reasonable samples for any testing since the degree of latent difficulty is not under test. I've printed out your posting to study further since it forms a roadmap to where we need to be positioned. The validation test is a global issue and any readers from geographic regions may like to come on board, contribute and provide influence. The timing of this issue is appropriate in our world where digital technologies and communications are well advanced.
Gerald you've given us the criteria of what needs to be included is there a chance you could spell out the elements needed for each criteria? If a validation test was conducted on the sample pattern ( Medium count Loop ( Right or Left slope)) single core and delta, comprising an arrangement of ridge endings, bifurcations and dots as primary features. How would the test design incorporate the criteria and elements to produce results?
Regards. Les
Re: Recent News - Articles and Discussion
Posted: Wed May 20, 2009 7:29 am
by Dr. Dror
I agree with most of what is said here, but think you need to use these new opportunities not only to 'validate' what you are doing. Yes, you believe in what you are doing (some say for 'good' reasons, some say for 'not so good' reasons); but regardless, as you/I/we, do more and more 'science', why not use it not only to 'validate' what you already believe in, but to learn new things, to better understand, and to *change* and to *improve* things. 'Change' and 'improvement' does not mean 'admitting' that anything is terribly wrong'.
What I urge for is a bit less of "what we are doing is right, and now we need to validate it and further underpin it in scientific studies", but more let's use these new opportunities brought about by the NAS report (or other reasons), to do more experimental scientific studies so as to improve the field (regardless of how already wonderful it is, or not). This, I believe is not only a better scientific approach, but moves us from divisive discussions how good (or not good) this area is, to more constructive discussions, moving together to make improvements. This may be only a nuance to some people, but to me it is important, as sometimes I get a feeling that some just want to 'validate' what is done already, not a great willingness to hear criticism or that anything is wrong, and want to do research with a sole aim to show that the existing procedures, protocols, and methods are working just fine.
Itiel
Re: Recent News - Articles and Discussion
Posted: Wed May 20, 2009 8:20 am
by Gerald Clough
I should point out that I do not anticipate that I will actually be doing any validation studies and that all this is just from thinking about the problem. I expect that those who are serious about doing genuine studies are holding their ideas pretty close and are probably shooting that the current round of NIJ grant applications.
I think that's the toughest part, settling on how to code the characteristics as data. I really don't have an answer I'm happy with. I do think anyone or any group attempting it should try to be open to thinking about it without being too constrained by conventions. At the same time, the definition of characteristics used now is pretty clear. I always have to make an effort to keep the interpretation issues out of it, so I suppose the coding of the dataset should be done by interpreting very strictly. But I then start thinking that it might be done in terms of how ridges are trending at a given point. For instance, a ridge changes direction toward an adjacent ridge and ends very close to that ridge. A little more, and it would be a frank bifurcation. I can't settle on whether that and the completed bifurcation should be coded as the same feature or if they should be coded in the conventional way. I usually retreat to the conventional characterizations, and I think that is the least ambiguous situation for coding. Perhaps studies done both ways will reveal the most meaningful results. I'm not sure it matters much. Either way will produce validation that will apply to any examination done using the same characterization scheme.
The more difficult one, I think, is how to code relationships among details. I might be more clear on that if I knew more details of AFIS coding. But however relationships are coded, it must be in a way that examiners can use. An examiner must be able to argue that a particular relationship is as the examiner describes it. Again, some experimental validation projects may well reveal what works. The ultimate question will be if a particular examination can be shown to be one that is addressed by a particular accepted validation. That's why there will be multiple validations. I can do one validation for examination comparing one contiguous latent impression. I must do another to validate comparison of, for instance, two simultaneous impressions in presumed anatomical order to two record impressions, if I am using characteristics of both impressions in one operation. I probably must do other validations for situations in which the useful characteristics in a latent are separated by a void or such a poorly impressed area that it is no better than a void.
And how does it change the probabilities when you can determine LI? How does that change the outcome for particular numbers of characteristics? And what if two or more of these situations are combined? It will probably pretty interesting work. (If you're not on the examiner end holding things down until it's done.)
One then has to work out how to weight data. This takes in things like frequency of occurrence of different characteristics and L1 classes and maybe different relationships. The work being done to computer-model friction ridge skin may be very helpful here. To make an extreme example, if you saw one ridge and four adjacent ridges, and each adjacent ridge diverted, one after another, to form a row of bifurcations with the first, even if that's all you saw, you'd likely believe you were seeing something so unlikely as to be an identification in itself. What would validation say about that? What would the the modeling algorithm say about how likely that was to have happened? We have some pretty good data on frequency of characteristics and can easily obtain more. We need that, too, to help refine modeling so we can generate as large a dataset as we wish that is statistically identical to the sample we have of the real human population.
But before we get lost in the more exotic, we have to come back to the fundamental, which is plain validation of the proposition that fingerprints can identify. Just the plain confirmation that the fundamental theory that no one really doubts now can be applied by observation to produce a very, very small number of possible matches in a population. The lack of that is the great criticism of fingerprint identification and applies in a directly and completely to the vast majority of examinations. That validation can be done with ideal, even computer-modeled impressions, not even impressions. It can be done with skin models without reference to impressions, although the first will likely be done with actual impressed/scanned records. That validation alone provides the essential critical support of all the extensions to the less common situations of things like simultaneous impressions and latents with voids. That void situation is interesting and important. We have held that a contradicting feature kills the ID, but we just work around voids that, for all we know, is where the contradiction would be found. Of course, we also assume that no such contradiction appears in ANY of the skin not impressed. Select some of the published "close calls," and blot out the contradictions and ask if they would have been concluded to identify. Validation helps settle that issue.
Considering that when validation studies have been done in this and other comparative disciplines, most or all conclusions will be in terms of probabilities, I would not be at all surprised if the expert title Forensic Statistician becomes very common. We see their issues already with DNA evidence. They will virtually have their own forensic field. Fingerprint validation will produce functions that can provide likelihoods for different quantities of data in a latent. What if the actor is clearly left-handed? How does the frequency of left-handedness in the population combine with the fingerprint conclusion? What if the pistol was .45 caliber Glock, and the individual addressed by the latent print owns one? Or drives an 1992 red Audi? Or is Vietnamese? Or wears size 10 shoes? Or owns size 10 Adidas rock-climbing shoes? How much statistics will a latent print examiner need to know, just to talk about the fingerprints? (The horror....The horror...)
Re: Recent News - Articles and Discussion
Posted: Thu May 21, 2009 3:53 pm
by Les Bush
Thanks again Gerald, and Itiel,
The picture is forming that we do need validation of the power of fingerprints as well as other research. Relevance and resources are always high on the appreciation list for what is being invested. Having research grants certainly helps in promoting the cause even down to using images subject to copyright which all comes with fees (as i've just discovered). The concept of validating the power of fingerprints has a two edged sword. One side is kind and will provide a scientific basis that supports all the previous examinations for the past century, tenprints and latents. The other side is unknown as the design of the validation test can result in the power of fingerprinting being regulated by statistical probability. From reading Geralds last post it appears the consequence of validation is linked to a result that is weighted with probabilities. Both the design of the validation test and the examiners chosen to complete it are two very important issues for fingerprint science. I would have thought that given ideal test samples, using a coding system that is already familiar, and confirming the result against the source skin while employing competent experts to observe and interpret them would reveal the power of fingerprints for what it is, individualisation to the exclusion of all others. The issue of exclusion is well established through the large number of examinations conducted on both ten prints and distal phalange latents, none of which have exposed a duplication of skin feature arrangements sufficient for individualisation. None. The whole body is a composite of variables and individual parts are highly variable some with special sensory needs. Friction skin is a special sense organ combining sensory function with grip. It doesnt surprise me that the pattern of features represents uniqueness and we have been detecting that for a long time. Through validation we should not have to compromise the variables of skin development and transfer of patterns within the design of a statistical test. We should also not be expected to accept a very low threshold of feature arrangement as a standard for reporting a probable result. The reliability of the science is also a test of human reasonableness coming from considerable experience and by which community trust is given. Any statistical model based on a validation study has several levels of checking and understanding to be investigated and verified particularly if the impact is a recommendation of a new method for presenting our conclusions. Knowing how the validation test and study were completed is the first part. Regards again from oz. Les
Re: Recent News - Articles and Discussion
Posted: Fri May 22, 2009 8:05 am
by Gerald Clough
I think validation will show fingerprints to be as unique as any randomly generated thing can be, but there's really very little that can be proven absolutely unique. I could make up special cases that would, with the right limitations, be unique, but with the kind of objects we consider for forensic use, we can only test with a finite set, so we just have to go with likelihoods. I don't that that's a killer. It's not now. If the question is asked of us if it's possible there could be another source for a latent that we couldn't discriminate, we have to answer that it's strictly possible. Validation tells us about how very small that possibility is for a given quantity of detail. I'm much more happy to contemplate a situation in which we don't have to worry about the challenges to the fundamental principles and don't have to say, "I don't know." quite so much. We're actually pretty fortunate that the research activity is ramping up well ahead of well-structured challenges becoming common in trial courts.
You mention the examiners who might be coding the test data. That's really not to much of a problem. First, the impressions are idealized. Details can be absolutely unambiguous, to the point of being represented schematically. Maybe even as non-image data, if it can be shown that a generator can produce data statistically identical to a large sample of actual impressions. That would be nice, since the dataset could be very, very large indeed. If you can argue that analyzing a very large set of real impressions produces a reliable statistical model, you might even realize that "all the fingerprints of all people who have ever lived" that's sometimes attached to the "have you examined?" question. But interpretations doesn't come into the validation issue. That is strictly another issue. That's where credibility, demonstration, and bias are argued. In validation, each impression does actually represent a single source. You're not coding "latents." There aren't really pairs of impressions to "examined." You just code every fingerprint in your set. You are trying to see what it takes to tell them apart. That tells you how likely it is that if you perfectly impress two of them, you couldn't tell the difference. You're not validating "examination." You're validating the proposal that some kind of determination of details in fingerprints has particular power to discriminate. We happen to be thinking about doing what we do in examination, but validation doesn't care at all how we detect the details. It just confirms that the theory can be applied and how you can state the meaning of impressions found to be alike to certain degrees. It's all about fingerprints, not fingerprint examiners.
Seen in that light, the bias issue is actually directed at internal workings of examinations and trying to produce the best product. Sometimes, that product will be contested, but you want the products of an identification unit to be as solid as possible so you don't have to defend every product. We want to cover all those bases in-house first. But it's nothing to do with validation. Validation just addresses the potential discriminating power of fingerprints. And "reliability" fades to almost nothing as an issue. If you don't think the product is reliable, challenge it, and see if you can be convincing. Just like any other evidence. And even if you can pick away at some details, what we will know from validation still leaves most fingerprint evidence not significantly less powerful in its ability to find fact.
I haven't thought too much about exclusion. I think that's something where expertise will play a big role, just as it does now, if you don't claim to have excluded every portion of an individual's skin as the source of an impression. As a legal issue, exclusion shares some issues with Inconclusive. What does "inconclusive" mean in a validated discipline? Well, it could mean you just can't interpret any details from this smudge. But some cases that are today inconclusive on account of small quantity of detail might still produce probabilities of common source. That's going to generate some interesting problems for attorneys and judges. In light of the cultural status of fingerprints, judges are going to have to make some decisions on whether the probability presented by an examination lends too much prejudicial weight to something that a jury might misinterpret as too powerful. How that works out depends a lot on just how the validations studies work out.
If you or anyone else responds to this, know that I will be away next week. My step-son, who I raised from the age of six (his mother helped), will graduate from the U.S. Air Force Academy and be commissioned, and we will be in Colorado for that. He will be going to Canaveral to work on space launches, and I will soon turn things around and start sending HIM emails every couple of weeks to send ME $100.
Re: Recent News - Articles and Discussion
Posted: Sun May 24, 2009 3:10 pm
by Les Bush
Thanks again Gerald, hope all goes well with your step son's graduation and the money flows in, my experience with two daughters is the opposite but love conquers all.
I agree that our science has a very strong basis for coming through the validation test and as an expert I wish to understand the conditions of a reliable test design so that any publication of results or any proposed statistical model can be assessed. The issue of absolute conclusions remains a sticking point since our science has two very objectively strong strings to its bow. The first is ten print examinations where yours and my fingerprints against the rest of the world would find that no two sets of ten prints are duplicated. Given the ability and reliability of AFIS to sort ten prints and experts to examine them the science of fingerprint individualisation is axiomatic. The second string is the quality latent where the pattern is presented in three levels of detail. Once again AFIS can successfully sort the pattern of a single digit against all the holdings and then refer results to the expert to complete the examination. Our fingerprint history is dotted with numerous examples of latent individualisations. The range of quality of detail and its interpretation support the spatial, sequential and sufficiency criteria of the technique. The diversity of data available in a quality latents provides the strong basis for an absolute conclusion of individualisation. A validation test of our science must start with our strengths before progressing to the more subjective examinations of less quality latents. This would be a logical approach and not dissimilar to how a new trainee is exposed to familiarisation of fingerprint patterning. An attempt to propose a new technique for examining fingerprints based on validation testing that focusses on the more subjective conditions and not incorporating the strengths of fingerprint science should meet resistance. Our principle of unique detail is capable of being falsified through any fingerprint examination (ten print and latent)conducted anywhere in the world. It is evidence based examinations that are repeatable. To date there is no evidence to falsify the principle. Cheers Les
Re: Recent News - Articles and Discussion
Posted: Mon May 25, 2009 5:42 am
by mdavis
There have been some excellent posts on this topic. But the devil is in the details, to be sure. As Gerald says:
Validation is done by by comparing a very large set of ideal representations of fingerprints. To do this, you must develop an unambiguous way to code the characteristics you wish to use, the characteristics you propose to use in actual examinations. (These need not be what you use now. For instance, you may wish to define a single coding that takes in two characteristics that have similar flows but would be called by different names because of vagaries in how they were impressed, even when they are impressions of the same skin. Or you might not..... And you must develop a coding for relationships among features, because those will be used in examinations.
Indeed. One step at a time. What is "very large?" What are and where do we obtain "ideal representations" of fingerprints? Not, surely, from our AFIS and IAFIS collections which abound with pathetic excuses of "controlled" known impressions. What are "unambiguous" characteristics? What "characteristics you wish to use, the characteristics you propose to use in actual examinations" do we attempt to code? This sounds as if we will cherry-pick only characteristics we might be able to code someday while ignoring those that we cannot. And yes, we do use L2D in actual examinations, and a whole lot more. We must first develop a foolproof method of capture that shows friction ridge skin in accurately reproduced and accurately reproducible detail. But in the process of designing such a capture method, we must first decide what detail matters. My earlier post touches on the issues of limiting the validation study to simple L2D with no regard for L1D or L3D, assuming we can agree on what constitutes a commonly agreed upon definition of observable and codable details for each level. Can we validate using only L2D? We know we can't use only L1D. Is L2D good enough without L3D and how do those stats vary with L1D taken into consideration? What happens to my example of the plain arch with no L2D? Any so-called validation using only L2D would render that a totally useless impression.
But I've been down this road before. I don't mean to rain on everyone's parade. I'm simply being the critical scientist that I was trained to be who looks for the flaws in research that could undermine or even invalidate conclusions of a particular study or result. Peer review is the last step in the scientific method. If you don't design a logical study with controlled parameters, your study is invalid and your peers will eat you alive. So here are a few hurdles to which we must agree in simply setting up:
How large a database is needed?
How do we design a consistent capture method that faithfully reproduces the detail needed for the study, including but not limited to control of contact, surface area and pressure?
How much surface area needs to be captured from the finger? Is 2nd and 3rd joint detail to be included?
How do we define adequate capture of tip area and how do we control that with our capture method?
How do we determine the thresholds in including/excluding friction ridge edge detail from digital images (grayscale boundaries for example)
What detail is needed for the currently proposed study? What might we need for future studies?
Do we go with the recent trends toward segregation of Level 1,2,3? How are these defined? Where are the isolation boundaries?
Once we've solved these issues, we "simply" need an accurate, unbiased, reproducible computer algorithm to isolate, record, and compare details. Then, having come up with statistical probabilities of two friction ridge impressions having 12 "details" in common, what do we make of that? Is it reasonable to include all known fingerprints in a database when your suspect list is limited to a dozen local burglars? Isn't a statistical probability circumstantial? Can the courts deal with the information in a meaningful way or are we overwhelming them with irrelevancies to please a few creative defense attorneys and the NAS?
Everyone has given "lip service" to the idea of validating our work. But in reading each post, I'm seeing the same "if's" and "buts" and "whens" with no scientifically valid solutions -- yet. No one would be more grateful than I for such work to be designed and successfully completed, if only to placate the courts. What I'm seeing is that we are light years away from a valid "validation" that meets scientific rigor. And if we do not come up with a rigorously designed study, we run into the scientific concept of "significant digits." Can we truly say that a study limited to, say, L2D is sufficiently accurate and representative of the incredibly complex comparisons performed by the human brain? It's like trying to weigh a tractor-trailer to the nearest gram when your scales are calibrated to the nearest 100 kg.
My view is that such a rigorous scientific study is overkill in today's sloppy, contentious, adversary court system in which human emotion rules untrained-juror decisions. (Perhaps we could marshal a counter movement requiring accreditation and proficiency testing of prospective jurors, judges and attorneys with required courses in math, statistics, chemistry, physics and biology.) : - )
Perhaps the proposed "lightweight" study of L2D probabilities is sufficient, like the 100 kg. scales.
Re: Recent News - Articles and Discussion
Posted: Thu May 28, 2009 7:50 pm
by Gerald Clough
mdavis wrote:
How large a database is needed?
How do we design a consistent capture method that faithfully reproduces the detail needed for the study, including but not limited to control of contact, surface area and pressure?
How much surface area needs to be captured from the finger? Is 2nd and 3rd joint detail to be included?
How do we define adequate capture of tip area and how do we control that with our capture method?
How do we determine the thresholds in including/excluding friction ridge edge detail from digital images (grayscale boundaries for example)
What detail is needed for the currently proposed study? What might we need for future studies?
Do we go with the recent trends toward segregation of Level 1,2,3? How are these defined? Where are the isolation boundaries?
How large a database is needed?
As large as can be produced. How large that is depends on just how it is generated. The size of the dataset will determine in part the probability thresholds.
How do we design a consistent capture method that faithfully reproduces the detail needed for the study, including but not limited to control of contact, surface area and pressure?
Variations in reproduction is not a validation issue. Validation does not compare two impressions of the same source. It works with one representation of each portion of skin used or emulated. One would not, of course, use impressions in which details could not be unambiguously coded on account of fog. But some of that is part of the coding problem, developing codable descriptors that can be used by examiners so that the results of validation apply to examinations. I strongly suspect that computer simulation of friction ridge skin can be developed so that its product is statistically indistinguishable from a large set of actual skin portions. That, of course, would be a dataset of unambiguous detail.
How much surface area needs to be captured from the finger? Is 2nd and 3rd joint detail to be included?
If one uses a dataset taken from actual skin impressions, it need only be large enough. Not a very satisfactory answer, because it's one of those things that depends on the kinds of details and relationships being used to define thresholds. It may well turn out that the area of the skin being examined becomes part of the threshold definitions.
How do we define adequate capture of tip area and how do we control that with our capture method?
I don't think that's an issue in the simple validation of L2 detail/relationships.
How do we determine the thresholds in including/excluding friction ridge edge detail from digital images (grayscale boundaries for example)
Again, we have to leave perception/interpretation issues out of validation. Validation is after confirmation of the simple proposition that aspects of friction ridge skin can discriminate among individuals. It uses a completely idealized version of that skin.
What detail is needed for the currently proposed study? What might we need for future studies?
A single validation study only validates a single defined scheme. For instance, L2 detail and the relationships among details. Adding L3 is another study. Using L3 alone is another. Using L1 along with the L2 is another. Any combination of observed phenomena proposed to be used in an examination requires validation. By this, I mean NATURAL phenomena. Matter of scars, distortions, interpretation of incipient ridges, etc. are not validation matters. They are for examiners to interpret.
Do we go with the recent trends toward segregation of Level 1,2,3? How are these defined? Where are the isolation boundaries?
As above.
mdavis wrote: Can we truly say that a study limited to, say, L2D is sufficiently accurate and representative of the incredibly complex comparisons performed by the human brain? It's like trying to weigh a tractor-trailer to the nearest gram when your scales are calibrated to the nearest 100 kg.
My view is that such a rigorous scientific study is overkill in today's sloppy, contentious, adversary court system in which human emotion rules untrained-juror decisions. (Perhaps we could marshal a counter movement requiring accreditation and proficiency testing of prospective jurors, judges and attorneys with required courses in math, statistics, chemistry, physics and biology.) : - )
Perhaps the proposed "lightweight" study of L2D probabilities is sufficient, like the 100 kg. scales.
Validation does not have to address all the interpretive issues of examinations. The "weighing" is for the examiner who makes, as required, unambiguous determinations of the reality that is represented. If an examiner can't determine the nature of the actual skin that was impressed, no conclusion is possible.
If one thing is abundantly clear, it is that validation is required, if we want to do more than merely point out similarities without conclusions. The courts will be tolerant for a while, since they will be loathe to throw out a valuable tool. And not all courts will be that tolerant. The reason jurors, judges, and attorneys need not be qualified experts in the sciences is that they depend on actual experts to deliver accurate assistance in matters of fact. They are perfectly willing to have any particular expert limit testimony to observations and clinical guesses, as they do with psychologists. But that's not what we want for fingerprint identification. It is exactly because the court is not expert that they will require proven valid conclusions from those who pretend to deliver them.
Validation is really not so difficult. It's not trivial, but it becomes more clearly straightforward when, as you must, you make a clear distinction between the fundamental proposition of fingerprint identification and all the examiner issues of how to interpret what is observed.
Re: Recent News - Articles and Discussion
Posted: Fri May 29, 2009 4:58 am
by mdavis
My point is, and it has been made increasingly clear by Gerald's post, that such gross oversimplification of the complexity of friction ridge skin by limiting "validation" to an "idealized" database of simple L2D in which "variations in reproduction is not a validation issue" cannot survive serious scientific vetting by defense experts. I repeat my example of a "simple" plain arch in which there are no L2D, or even better, one L2D. This would be at the very least a highly improbable yet possible occurrence that would receive a near 0% probability of being unique based on L2D. Every latent print in the database (with the exception of my plain arch with no L2D) would share at least one L2D, so the probability of uniqueness would be essentially zero, especially if we ignore the area of friction ridge skin captured in this ideal database and it's rotational orientation. The statistical probabilities would be meaningless and the "validation" study a sham if it cannot be applied to real world examples, or if we go from point counting thresholds to probability thresholds.
How many idents could not be called because they contain an "inadequate" probability threshold percentage. How many latent impressions would be "pushed" over the threshold by an overzealous examiner using the "eye of faith" struggling to reach the "magic" percentage to "validate" the ident? This is precisely the problem we have with the few bad idents hung out on the town square for the world to see -- issues not with a "valid" number of L2Ds, but with forcing sufficient detail to exist where there is inadequate clarity of the impression. The number of L2Ds obtained from an idealized database needed to meet some threshold probability of uniqueness is effectively meaningless until we reach large numbers of L2D, then it becomes simply dangerous.
I think the issue here is that someone comes up with SOME percentages, somehow. This seems to be an ugly trend led on by the relentless tide of "accreditation." Do something, write it down to satisfy the inspectors and make it part of your SOP. Dumb down the study to simple Galton ridge ends and bifurcations, throw in a few thousand "perfect" impressions and give 'em a handful of numbers to satisfy the courts and the critics that we did the validation studies, and then go back to some serious casework. If that makes someone happy, go for it. It does nothing to address the real issue of the reason for those extremely rare but highly damaging bad idents made by those examiners who live too close to the edge.
Re: Recent News - Articles and Discussion
Posted: Sun May 31, 2009 8:30 am
by Gerald Clough
We are at a very early stage of this. It is very true that a simple validation study of strictly L2 will indeed scientifically validate the theory that friction ridge skin has features that can identify an individual with a very, very small possibility of it being someone else. That has not been done. Fingerprint identification has only ever been presumed valid, entirely based upon experience. It is not acceptable to go on in that mode when validation is so doable. Do not make too much of the thresholds thing. That just comes along with any validation, because it's impossible to examine every possible variation in skin. A substantial part of validation is showing that the likelihood of being mistaken for another is extremely small. Until that is shown, it's plain guessing. You can dress it up with all the accumulated experience, and it's still guessing. Guessing isn't evil. Your doctor is almost always guessing when he makes a diagnosis and prescribes for you. But he has no choice. We do.
I will take specific issue with this: Idealizing friction ridge skin is not over-simplification. It isn't simplification at all. It is exactly the skin itself. And it is the skin we are concerned with in validation, not the examination. This is the hardest thing for some examiners to deal with. Two issues. One is what skin CAN tell you about who someone is. That's the one that validation addresses. The other is the issue of HOW one way of using impressions of that skin works. If a machine was to do the examination, the validation studies would be the same. If examiners decided to do quick looks and roughly guess as what details were, validation would be the same. If a committee of twelve examiners had to unanimously agree on the character of each detail, validation would be the same. Validation has nothing at all to do with examination, EXCEPT that examination uses the same skin features and therefore can cite validation results. In a sense, it is validation that applies to the real world, the world of actual skin. Examination issues apply to the world of impressions of skin. That's real, too. But it's addressed in a different way.
And validation really has nothing much to say about most of the notorious erroneous identifications. All of them I recall would have been incredibly high-confidence conclusions, had the examiners not erroneously characterized details. Validation doesn't affect that at all. If the analysis of those latents are presumed accurate, they are rock solid identifications, as certain as they could be. The fix for those errors lies in issues of proper procedure and proper review. They are not just oddities. The examiners made specific errors. They are not just the bad you have to take with the good. And you can't even say they are rare, except that you can say it is rare that errors of that kind are discovered. But they are not validation issues.
Validation is not a "sham," and it is totally meaningful. It is arguably the most meaningful thing we must deal with, at least the most fundamental. It's the one aspect of fingerprint identification that the examiner will not have to defend. It's the one thing that you don't have to admit to guessing about or depend on a large body of experience, experience in which actual absolute determination that the correct identification was made was rarely done or even could be done. That body of experience is scientifically meaningless. No one doubts in the least that that experience reflects a reality about fingerprints, but general agreement that we all fully expect that to be validated in no way establishes it as fact. Astrologers are all in agreement that their data is meaningful, too, and lots of people who aren't adepts accept it as meaningful.
The rest is up to the examiner, to demonstrate credibly analysis of the latent impression. And you're exactly right about SOP's and such. They don't insure accurate analysis. Nothing does. Any analysis based on perception and interpretation may always be questioned, or at least reviewed, and seeing that access to independent reexamination is a legal issue, not even an LPE issue.
Bottom line here is that fingerprint identification can remain stuck in clinical practice, educated guessing, or it can move on to solid scientific ground and do the best it can to try to make the analysis as reliable as possible. If it remains clinical, its evidence will end up being subject to exactly the same kind of "I say" - "You say" contests that are nearly always seen in other clinical opinion fields, like psychology. If, right now, you are confronted with an opposing expert who says, "That's not a sufficiently reliable identification.", you can't answer it with anything but "He says, but I say." You're both citing experience. And like the psychologists, you can't both be right, but neither can either of you be wrong. You just guess differently. The fact that you're almost never confronted with such an expert witness has more to do with the historical professional culture and structure than anything else. That's changing, and you can wait until it happens, or you can begin now to work to lock down what can be scientifically known. Or, if you're on the down slide of a career, you can just wait and leave it to others to worry about. The law evolves slowly, and I would guess that half the examiners working today won't be seriously confronted with the validation issue in their own cases.
Re: Recent News - Articles and Discussion
Posted: Sun May 31, 2009 10:47 am
by mdavis
I'm bowing out of this discussion. IF a
simple validation study of strictly L2 will indeed scientifically validate the theory that friction ridge skin has features that can identify an individual ....
and the courts buy that, then what? It doesn't matter if I don't buy it for reasons given elsewhere. If you send me to court with a probability percentage, so what? What does the court do with the numbers? What do I do with the numbers? What does it mean? Are we to allow the court to draw the line based on a "validation" study, or is that my job? I don't know what the numbers mean as applied to any given comparison and neither does anyone else. The "strength" of probability would depend on the comparison at hand. I restate my example of the plain arch. Generic "validations" will put such a comparison out of reach. If a L2D "validation" so mis-fits such a case, how many other cases does it mis-fit on both ends and the middle of the spectrum? Gerald seems to think we can do a L2D validation, put it in the bank and go back to business as usual with the blessing of statistics. All I'm saying is that if and when we come up with probability percentages for a given number of L2Ds, we are back to point counting using different semantics. I think this will open up a whole new can of worms because we are using numbers based on half truths and inadequate attention to detail because the complexity of our analog craft is too great to digitize. A simple "validation" study of L2D is a trap, not a lifeline.