Page 1 of 1

Validation

Posted: Wed Sep 26, 2007 5:56 pm
by Charles Parker
Validation

I wanted to broach a subject that I have been thinking about the last couple of weeks because I am currently involved in the Validation of a test. I went to SWGFAST to their glossary but the term validation is not defined there. However SWGFAST has a complete approved document on Validation. But after reading the document it is primarily concerned with the Validation of latent processes and not the Validation of a test (competency, proficiency, etc) designed to test human ability and knowledge.

I then went surfing and wound up at Wiki, but the following I thought were interesting to what I wanted to know in the area of Validation of a Test of Human Ability.
In a quality management system, validation should not be confused with verification. Validation usually relates to meeting the needs of an external customer or user of a product, service, or system: Verification is usually an internal quality process of determining compliance with a regulation or specification. An easy way of recalling the difference between validation and verification is that validation is ensuring "you built the right product" and verification is ensuring "you built the product right." Verification is testing to confirm that a product complies with its requirements and specifications. Validation is testing to confirm that it satisfies stakeholder needs.
Verification and Validation (V&V) is the process of checking that a product, service, or system meets specifications and that it fulfills its intended purpose. These are critical components of a quality management system such as ISO 9000.
Now that we have an idea on what type of validation I am pursuing I want to provide you with a scenario. Now this is not a hypothetical scenario because as my friend Pat W. said when he hears the word hypothetical that it probably came from real life. I do not want you to think that this might be real as I do not know but this scenario is strictly an extrapolation on my part. Let us just call it an Another World Scenario or AW Scenario.

Now I am the developer of a test of human ability. I have a client group of say 1,000 individuals. In this group 750 are full time widget makers with varying experience levels ranging straight out of widget school to 35 years making widgets. Of the remaining 250 only 200 make widgets less than 50% of the time again with varying levels of experience. The remaining 50 only make widgets less than 25% of the time.

Now that I have the test made up I need to get it validated so I then try to determine who I can get to validate it. Now I have no idea of the level of experience or the amount of time that my client base has. I could take a guess at it, but I really do not know for sure what the experience base is. REMEMBER now that widget making is a subjective area of human endeavor. Also that in AW Scenario you can make good widgets, bad widgets, maybe widgets, and you can even say that it is not a widget at all. Four responses on your widget making ability.

In thinking about validation of my test I contact 10 widget makers that I know. They have been making widgets for a very long time and most are from the federal widget factory and some from state widget factory and a couple from city widget factories. They all do widgets full time. Now I do not take any information from my 10 friends. I just know they have been doing it a long time and they are all certified widget makers. They all agree to take my test. I tell them that on my test they can only respond with it’s a good widget or bad widget. They cannot say if it is a maybe widget or that it is not a widget at all.

I get the test back and I am amazed. Nine of the ten have all the answers correct and one of them only missed one (I strike this one from my future widget taking test list). Six of the ten said that widget 14 and 15 were tough but they were finally able to get them.

Now if I send out my test to my client of 1,000 widget makers have I done a valid Validation Test.? Have I developed a test to measure the needs of my client base or have I just created a test that measures the top 10%?

Any thoughts or ideas?

My point is this---the next time someone wants to give you a widget test of your abilities you might want to ask to see the Validation Protocol? How many people were used in the validation test? What are their experience level. Comments they made on the test? etc. Should you be able to respond with the same answers on the test as you are allowed on the widget factory floor? Should your 10 years of experience really be compared with those that have 25-30?

Posted: Thu Sep 27, 2007 6:32 am
by Pat A. Wertheim
Hi Charles

I think one of the first things you need to do in the planning stage is to discuss with the client exactly what they need for you to measure and exactly where they want the threshold score drawn. I don't know too much about widgets, but I have designed one or two latent print proficiency tests, both for my agency and for a few other agencies. One of the first tasks, especially for an outside agency, is to define what the proficiency test is supposed to accomplish in all regards. For example, does the client want the test to measure excellence, competence, or just adequacy? On the job, do the examiners ever compare color reversals? Position reversals? Is the test for 10-print examiners who testify to prior conviction documents? If the test is a promotional test for unit supervisor, maybe the level of difficulty needs to be much greater than for journey level or even senior level latent print examiner, or maybe it just needs to measure a basic level of understanding for someone being promoted from a related field to supervise several disciplines.

Once I know exactly what the client expects the test to measure and where the client wants the threshold, then I design the test based on my perception of those criteria. Sometimes I'm right on target, sometimes I'm not. Subjectivity enters the equation, both in designing the test and in taking it. But once I have designed and prepared the test, then I decide on a validation group of people who are not a part of the group who will eventually be measured by the test. The validation group, as you point out, should be people whose ability you know, or think you know, well enough that they set the standard for the level the client wants measured as the threshold.

After you have administered the test to the group for whom it was designed, then you have to evaluate the results. I have been unpleasantly surprised on one or two occasions and I have had to revise my threshold after the fact. That is a difficult thing to do, but because there is that pesky subjective component to both the design and taking of a test, it is sometimes better to tweek the test after the fact than insist on the cold, rigid interpretation you set before the test was administered. Now, before some readers get up tight about this, let me point out that we all had professors in college who did this -- they called it "grading on the curve." Ideally, you can design your test so you don't need to use a curve. But sometimes, not using a curve is unfair to both the client and the test takers.

There are some of my thoughts on test validation this morning. Hope that helps. Maybe a few other things will occur to me with my second cup of coffee.

Posted: Sat Sep 29, 2007 5:50 am
by Charles Parker
Pat, very good protocol. If I could paraphrase and add some points of my own.

1. Determine Clients Need (Desire)
2. Define Test (Expected Outcome with Test Subjects)
3. Design Test
4. Decide Validation Group (Similiar to Test Subjects Ability/Experience)
5. Validate Test and Document V-Group A/E and Comments
6. Develop Formal Procedure for Test Complaints
7. Adminster Test
8. Assessment of Results
9. Modify Result Matrix (If Outside Expected Outcome)
10. Provide Validation Group Matrix if Requested

I see two types of tests. Those that are used one time for a large group and not used again. Then those that are used continually over a long period of time but are used only once per person.

The one-shot tests I would think would be the most difficult to maintain the same level of acceptance since they are so variable from testing period to testing period. A well defined and broad base validation group would be the key to the best successful implementation of this type of test with the least amount of variance. That and well designed. Also one-shot tests should have the widest range of responses available to those being tested.

The static test is good but difficult to maintain the security of the test over a period of time without strict controls over its administration (Proctor--No Imagining Capibility--etc.). Also over a period of time the client base may change and the test should be re-validated every five years with a more up to date validation group OR use the the responses of the test subjects to modify the results. Like you said gradng on the curve.

Those are my thoughts for this morning. Like you I need a second cup of coffee.

Can we apply this to training as well?

Posted: Tue Oct 02, 2007 4:46 pm
by Lauren Cooney
Along similar lines - do you think there should be some form of validation with regard to widget making training to competency programs? Maybe I'm off-base here since it isn't an actual test, but, how does one determine that all the widget factories have sufficient training programs?

Of course, the SWGWMST (Scientific Working Group on Widget Making, Study, and Technology) has guidelines, but they are somewhat generic and leave a great deal of room for interpretation. How do we know all the widget makers at all the different factories have comparable and sufficient training? :?

Posted: Tue Oct 02, 2007 5:19 pm
by Charles Parker
Ahhh, Ms Cooney------You have asked the $64.00 Question.

A question I have been asking for awhile.

I think yes to a validation test in regard to training to competency.

You are right SWGWMST is kind of generic. And the ASWMD (American Society of Widget Maker Directors) is also very generic.

My 2 cents question is what is COMPARABLE and SUFFICIENT for training?

Your definition? My definition? The SWGWMST definition yet to be developed?

Would a 2 year training program do? Would a 1 year training program do?
Would a 6 month training program do?

How about a 40 hour widget making wonder OR widget making out of a box.

Who decides what is comparable and sufficient?

I personally do not think a non-profit volunteer group could do it. It will take each state or the feds to make it LAW. And then tax money to hire people to administer it and test the applicants and make sure the training is proper much like each state does now does with its Comminisoned Officers. Each state has different needs and different amounts of money to spend.

We are now where my state was in the 60's when it came to comminisoned officers and the state created a TCLOSE or in California it is called POST.

Only by law and with tax funds to operate and manage will that $64.00 question be solved.

My 2 Cents.

Posted: Tue Oct 02, 2007 6:01 pm
by Pat A. Wertheim
Validation of training may not be that difficult. First, if the goals and objectives of the training are clearly defined for the audience for whom the training is to be given, and if the audience itself is clearly defined, then a good test measures each objective based on the needs of the "client."

But designing good training does not ensure education. We have all seen outstanding classes in which some students learned much and others learned little. We all have experience as well with classes we took and passed, yet we are not proficient in the subject matter. In other words, we passed the tests but did not achieve the training objectives.

Likewise, we all have seen people who took substandard classes but emerged with more knowledge than the instructor or, indeed, more knowledge than the course provided.

My point is that training is highly variable. Excellence emerges and failure exists no matter how well or how poorly the training is designed.

But a proficiency test! That must be designed independent of training goals and objectives. Therein lies the need for validation, whether in widget making or latent print proficiency testing.