Miami-Dade Research Project with NIJ Grant

Discuss, Discover, Learn, and Share. Feel free to share information.

Moderators: orrb, saw22

Post Reply
josher89
Posts: 509
Joined: Mon Aug 21, 2006 10:32 pm
Location: NE USA

Miami-Dade Research Project with NIJ Grant

Post by josher89 »

Shout out to Brian and Igor (and the examiners that assisting in this study).

Just saw that this was posted and wanted to be sure and recognize their hard work on this research.

https://www.youtube.com/watch?v=tHaEqj6MjV4

I met with them in Sacramento and was excited to discover they aren't done with this research; they hope to expand it and I am looking forward to seeing it as well!

Link to NIJ Study site:

http://nij.gov/multimedia/Pages/video-m ... deo-122015
Boyd Baumgartner
Posts: 567
Joined: Sat Aug 06, 2005 11:03 am

Re: Miami-Dade Research Project with NIJ Grant

Post by Boyd Baumgartner »

Feel free to correct my math, but in the video they reported a 3% false positive rate for the ACE only portion and by my calculation that is a 2900% increase in errors over the black box study false positive rate at .1%.

I just used a normal % change calculation where V1 is Black Box false error rate and V2 is the Miami Dade false error rate

Code: Select all

% Change = ( V2 - V1/V1 ) * 100
% Change = ( 3 - .1/.1 ) * 100
% Change = ( 2.9/.1 ) * 100
% Change = (29 ) * 100
% Change = 2900
2900%
 
Black Box Study


Anyone want to take a stab at how they'd explain a 2900% change in the false error rate on the stand or in a defense interview? Alternatively, does this study reflect poorly on the question posed by the title of the video (How Reliable Are Latent Fingerprint Examiners?).
g.
Posts: 247
Joined: Wed Jul 06, 2005 1:27 pm
Location: St. Paul, MN

Re: Miami-Dade Research Project with NIJ Grant

Post by g. »

Sure,

1) I wouldn't use the math that you used. If the rate increased to 0.2% from 0.1% (a pretty insignificant increase), using that math, you would say it was a 100% increase. I'd use more of a log scale and note that the increase was basically an order of magnitude (i.e. a 10-fold increase). Approx. 1 in 1000 to approx 1 in 100. It's a log relationship and not a linear one (as you have calculated). There are other methods to compare these values too (using a test for significance for differing proportions, etc.)

2) It's easily explained in each study's definition of a "false positive". In Black Box a false positive was "an identification decision given that the paired images were non-mates". In Pacheko, et al. it was essentially "the listed identification did not match the ground truth answer for a specific source/finger or palm". In other words, it allowed that "clerical errors" would be counted as false positives since the researchers had no way to distinguish a clerical error from a true erroneous ident. This is always a tricky thing when computing error rates from 1 to N (many) comparisons versus more controlled study designs of 1 to 1 comparisons. Set ups like Igor's (or mine and Kaseys in 2006, Gutowski 2006, Langenburg 2009 or CTS tests) are closer to case work design, but suffer from this draw back, unless you require participants to chart out, annotate, or document the steps of ACE-V, it's difficult to parse out the clerical errors from true decision errors. [Kasey and I were strongly criticized for attempting to do so by Haber and Haber, 2006; Kasey and Dror ran into a similar problem in 2010 in their AFIS study, they reported two error rates: one with all the false positives and one with the false positives removed that were "likely" due to clerical error.]

One thing I have seen in 3 separate studies and fairly consistently in CTS tests over 15 years is that clerical errors happen around 1% of the time. It's pretty robust. I used to predict before class would start when I did comparison training based on the number of students, the number of comparisons how many clerical errors would be committed by Friday afternoon. It was a perfect 95% confidence interval. Out of 20 classes, 19 of them fell in the 95% confidence interval. This demonstrates to me it is very random (a Poisson distribution was used) and not dependent on the difficulty of the exam, the skill of the examiner, the clarity of the latent print, etc. It's not dependent on all the things that we know that the false + error rate for true erroneous ID decisions is clearly dependent on.

3) A much more challenging question is what denominator do you use for the number of comparisons? Without boring people with the math, ask yourself, when you make an exclusion to an individual is it 1 exclusion to that person? or is it 10 finger exclusions to that person? plus 2 palms? plus 10 finger joints? How you answer the question means you made 1, 10, 12, or 22 exclusions (another factor of 10 on the log scale). And that can be the difference between 0.1% and 1% error rates. The interesting thing is NONE of the critics/pundits have sounded off on this and there is no standard way to calculate it, but it's a pretty critical question to answer. In Pacheco et al., they chose "1" exclusion per person, so each exclusion in the study resulted in "3" people excluded. So each exclusion counted as 3 and not 30 (or 36 or 66 exclusions). That was their call and they were explicit about it. There's no right way or wrong way, it just shows that it is important to understand where the numbers come from in these studies.

...which I think was Boyd's point... how do you explain the differences between studies? Different designs, different definitions, different limitations.

g.
Bill Schade
Posts: 243
Joined: Mon Jul 11, 2005 1:46 pm
Location: Clearwater, Florida

Re: Miami-Dade Research Project with NIJ Grant

Post by Bill Schade »

g

I read the whole post and my "take away" is:

".. how do you explain the differences between studies? Different designs, different definitions, different limitations."

This brings up the point that perhaps I shouldn't cite any research studies I don't thoroughly understand. (and that seems to be most of them)

Although I consider myself informed, these discussions are way over my head. My only consolation is that it will be beyond the average attorney too and I'm pretty confident that if the jury has to sit through this kind of discussion they will probably base their decision on "Bill was wearing a nice tie and looked awfully professional, I believe him" :P
Boyd Baumgartner
Posts: 567
Joined: Sat Aug 06, 2005 11:03 am

Re: Miami-Dade Research Project with NIJ Grant

Post by Boyd Baumgartner »

Thanks for the reply, g, it makes sense. A comparison of absolutes is less informative than order of magnitude when comparing studies especially those that aren't apples to apples in terms of design and method.

That being said, it still seems off that there would be a 3% (or 2% considering your clerical example) false error rate. In a larger shop that does upwards of 6000 IDs a year that would be a 120 - 180 errors a year that make it to a verifier. Have you (or anyone else who wants to answer) been asked to translate studies into case work, or do you have an opinion on the utility of these studies.
Shane Turnidge
Posts: 81
Joined: Thu Jul 21, 2005 11:55 am
Location: Canada

Re: Miami-Dade Research Project with NIJ Grant

Post by Shane Turnidge »

I'm with Bill on this one.
Whether it be a juror or a judge, studies like these tend to muddy the waters and can detract from the message. I.e. is it him? Are you really sure, is there anything in any process you used that could affect the outcome of your decision? Etc. That's why I always wear a nice tie. 8)

That said, within our own community, studies occasionally help us evolve our processes. Studies that help examiners understand their thresholds and other limitations are important.
With actual casework, I tend to support g.'s numbers although they are very difficult to objectively be demonstrated because it would require 100% transparency on behalf of all agencies conducting the work.
In casework, the full effect of the constraints placed on an examiner becomes apparent.
I've always like the Oliver Shroeder quote that sums up some of the limitations;
"The cornerstone of all ethical thinking including professional ethics is private morality. On this foundation of morality is placed a second layer of responsible performance, the ethics of the profession. A third and final layer is then added in the form of public law. All three of these conditions act to constrain each professional practitioner. A forensic scientist must be moral, ethical and lawful simultaneously."
When you add the knowledge, skills and abilities of an examiner you have the underpinnings of why the number of false positives in actual case work are so low.
In contrast, In a training environment, where we need examiners to test and occasionally exceed their thresholds, my experience is that the numbers of false positives are significantly higher.

Shane Turnidge
ER
Posts: 351
Joined: Tue Dec 18, 2007 3:23 pm
Location: USA

Re: Miami-Dade Research Project with NIJ Grant

Post by ER »

Other possible math options:

Relative difference based on mean: 181%
Relative difference based on max: 97%
Relative difference based on min: 2900%
Relative difference based on log mean: 153%
g.
Posts: 247
Joined: Wed Jul 06, 2005 1:27 pm
Location: St. Paul, MN

Re: Miami-Dade Research Project with NIJ Grant

Post by g. »

In a larger shop that does upwards of 6000 IDs a year that would be a 120 - 180 errors a year that make it to a verifier. Have you (or anyone else who wants to answer) been asked to translate studies into case work, or do you have an opinion on the utility of these studies.
Boyd, I'll take the 2 issues separately:

Issue 1 - Yes, I'd say that's about right. 6000 IDs is a LOT. Does KC keep track/stats of the number of IDs per year? I don't think many agencies would come close to 1000 IDs per year, let alone 6000 IDs in a year. I think it would be difficult for examiners to make 6000 IDs in their career as well. That's 1 ID per day (5 day work week assuming 48 weeks of work per year over 20-25 years).

We kept some rough stats for awhile and found that we only reported about a 400 +/- 100 per year for 6-7 examiners. We may have done 10s of 1000s of COMPARISONS, but reported IDs only occurred in 1 in 3 or 4 cases or so. You'd have to have a huge number of examiners (let's say LAPD, LASD, NYPD, FBI, etc.). where you have 30-50 examiners to get 6000 IDs per year.

That said, yeah, if you have 6000 IDs per year, I'd expect 60-120 or so of those to be clerical (transposed fingers) errors that get caught during verification. I am not worried about that high a number of clerical errors b/c 1) they are relatively easy to spot, 2) they nearly always have the right source named, just wrong finger.

When I was working for the State of MN as an examiner, I found at least 1-2 per year during verification. So yeah... I'd say those numbers seem plausible (based on empirical observations and the reported literature).

Issue 2 - The utility of error rate studies has been mostly in answering questions in court about error rates, particularly in admissibility hearings. While I do not advocate using the studies to generalize to case work error rates, I think the studies tell us quite a bit about the accuracy of the process (generally speaking). It also gives at least some sense of the magnitude of error and what conditions might increase or decrease the error rate. There is no such thing as "THE ONE and ONLY error rate"....there are "error rates" and they change depending on the conditions. But the fact that the error rates in 5-6 different studies, conducted by different researchers with different test objectives and study designs, all keep hovering around the same approximate error rates tells me quite a bit. And I think some very general, but useful, statements in admissibility hearings can be made. I find them to be much more helpful and useful than statements like "zero error rate"; "we don't have any way of knowing/calculating", "It can't be done", "It's always changing so you can't calculate it", etc. These answers aren't helpful, informative, and are not true (in my opinion).

That's how I would use error rate studies. How would you/do you use them, Boyd? Or do you think researchers should be spending efforts on a different question at this point?

g.
Boyd Baumgartner
Posts: 567
Joined: Sat Aug 06, 2005 11:03 am

Re: Miami-Dade Research Project with NIJ Grant

Post by Boyd Baumgartner »

Does KC keep track/stats of the number of IDs per year?
Yes, and we publish them in our Annual Report. The combined total last year was 6381 (pg 10). Our AFIS program is levy funded and includes King County, Seattle PD and Bellevue PD. King County represented about 90% of that number which is approximately 6000.

That said, yeah, if you have 6000 IDs per year, I'd expect 60-120 or so of those to be clerical (transposed fingers) errors that get caught during verification.
I think you are right on in your estimates of clerical errors.
When I was working for the State of MN as an examiner, I found at least 1-2 per year during verification. So yeah... I'd say those numbers seem plausible (based on empirical observations and the reported literature).
I guess my reaction to this depends on where the errors are coming from and the policies of the agency. I would think an error would require an examiner to be pulled from case work and their previous (6 mos/year?) work scrutinized very heavily as well as future workload being reduced. If the errors were distributed evenly among examiners, that could bring an agency, especially a small one to their knees. Also, did the same people make the same errors year after year? That could point to a need for retraining or policies that limited the ability of an examiner to call an ID. In the Miami Dade research scenario with 120 errors (3% - 1% clerical), that would mean roughly 6 errors per person in case work at our agency. We'd be shut down if that happened. I guess ultimately I believe mis-identifications should be more on the order of the black box study than the Miami Dade study. 2% seems unstable in terms of keeping an office running or a public perception of accuracy.
That's how I would use error rate studies. How would you/do you use them, Boyd?
I've only ever had to use them in pretrial interviews. It usually starts with a defense attorney or investigator asking if I've ever heard of the NAS report to which I reply that I have and summarize it by saying that the report recommended regulation and research in the discipline and let them direct the questioning further. They will undoubtedly ask about error rates, to which I will reply there have been a few studies and produce a copy of the black box study with the .1% and 7.5% nicely highlighted for them and the white box study. When pressed further I say that in the same way that climate science data shows trends but cannot be used to predict the weather, fingerprint performance studies show trends but say nothing about the case at hand. (pretty much the same thing as Bill and Shane were saying) When pressed further, I usually remind them that according to the study I'm more likely to have missed identifying their client on a piece of evidence than mis-identified their client. That usually redirects the line of questioning.
Or do you think researchers should be spending efforts on a different question at this point?
With regards to research, I think that performance studies are fine and agree that they tend to show a trend, but ultimately are hard pressed to define recommendations for improvement in the discipline. I personally would like to see more research along the lines of trying to apply machine learning to eye tracking data and see someone come up with a matching algorithm that mimics the routes of comparison that an expert would take and see if they could incorporate some configural processing algorithms to the eye tracking data as well.
g.
Posts: 247
Joined: Wed Jul 06, 2005 1:27 pm
Location: St. Paul, MN

Re: Miami-Dade Research Project with NIJ Grant

Post by g. »

Sorry if I wasn't clear. I found 1-2 CLERICAL errors per year. I agree. 1-2 erroneous identifications per year would be a problem for an ID unit!!!

I gave the 1-2 clerical errors suggesting that if we assume all the other examiners in my lab found the same, then our lab probably has about 6-12 per year, which is about 1-2% of our total # of IDs per year before verification.

I was still surprised to see that your agency had 6000 IDs in a year! That's a lot (for what? 15-18 case working analysts?). You guys must have a very efficient recovery and AFIS entry process. That's pretty good!

I also agree that if you use Miami Dade data, you are likely to overestimate the number of expected erroneous IDs. But they include TRUE ERRONEOUS ID + CLERICAL ERRORS = TOTAL # of FALSE POSITIVES. After removing the "probable" CLERICAL ERRORS that are part of Miami Dade's study, they are left with around 2 or so IDs that are to different people than the ground truth, and likely not a transposed conclusion. The participant associated the wrong person with the latent. That would put their False Positive Error Rate closer to the black box study false positive error rate.

g.
Bill Schade
Posts: 243
Joined: Mon Jul 11, 2005 1:46 pm
Location: Clearwater, Florida

Re: Number of identifications

Post by Bill Schade »

I have to respond to the "surprise" about the number of identifications made in a year in this discussion

I am assuming we are talking about impressions identified since each one could be a potential error, either clerical or otherwise

Seven examiners and a working manager who still looks at a case now and then in a traditional police latent print shop. Looking at our numbers for the 11 months of 2015.

Submissions received is 4,944 containing 23,235 lifts and photos

Subjects identified by request is 629 and by AFIS search is 1,557 for a total of 2,186
This does not include impressions identified against eliminations submitted which are documented but not tracked.

I can assure that most of those subjects were identified on more than one impression, so 6000 impressions identified in a year sounds about right to me.
Post Reply