NIST’s new draft framework for evaluating AI systems, TEVV-Athlon, should provide a useful reframe for leaders who rely on checklist-style governance, writes Larry Marks, a GRC adviser and consultant. Demonstrating completion of a process isn’t the same as understanding the evaluation, and those who are tempted to turn the framework itself into the latest checklist are making Goodhart’s Law manifest.
One of the things that caught my attention in NIST’s new TEVV-Athlon framework draft for evaluating AI systems is what it does not do. Namely, the framework does not give organizations a standard list of tests every AI system must pass. Instead, NIST recognizes that AI systems, their uses and their encompassing risks vary too much for a single evaluation methodology to work in every situation. Instead, the framework is designed to be adaptable and customizable to the organization’s objectives and the particular AI system being evaluated.
I have spent much of my career working with cybersecurity and technology risk assessments. One lesson I learned is that an organization can become very good at demonstrating that an assessment was completed without necessarily demonstrating that the right risk was assessed. Once a framework becomes established, there is a natural tendency to standardize it. Questions become checkboxes, evidence collected and management approvals documented. Eventually, completing the process can become almost as important as understanding what the process was intended to tell us.
NIST appears to be trying to avoid that problem in this new draft. TEVV-Athlon starts by asking organizations to articulate their objectives and organize the evaluation around them. Only then does the organization determine what should be measured, how those measurements should be performed and what the resulting evidence means. The four stages move from articulate and organize to define and construct, to apply and measure and finally to synthesize and interrogate.
For compliance and risk professionals, this requires a somewhat different mindset. The objective should not be to demonstrate that an AI system has “passed TEVV.” The more important question is whether the evaluation was designed to tell the organization what it actually needed to know about that AI system in its intended use. That may sound like a small distinction, but, in practice, it is a significant one.
A passing score can create the wrong kind of confidence
One of the more interesting parts of the NIST draft is its discussion of Goodhart’s Law: When a measure becomes a target, it can cease to be a good measure. NIST applies this concept to AI evaluation and warns that optimizing a system to perform well against a particular benchmark may not tell us how well that system will perform in the real world. This should get the attention of compliance and risk professionals.
We are accustomed to measurements. We use risk ratings, control effectiveness scores, key risk indicators and other metrics to help management understand risk. These measurements are very useful. They change complicated information into something that can be compared and acted upon. The danger comes when the score becomes the objective.
The same problem can occur with AI. An organization may establish performance thresholds, conduct testing and conclude that an AI system has successfully met its evaluation criteria. Management then sees a passing result and assumes the risk has been addressed. But what exactly passed?
NIST recommends using multiple complementary evaluation approaches and recognizes the importance of real-world testing rather than relying exclusively on benchmarks. That is an important distinction for compliance. Likewise, the purpose of TEVV should not be to produce a score that makes management comfortable but to produce evidence that helps management understand the AI system, including where uncertainty and residual risk remains.
A good evaluation may therefore produce an uncomfortable answer. It may tell management that the AI performs well under certain conditions but that there is insufficient evidence to reach the same conclusion under others. This may be exactly what management needs to know.
5 Structural Barriers Breaking Your Cybersecurity Compliance Framework
Compliance challenges rarely stem from a lack of intent, but are often rooted in how systems and processes are designed.
Read moreDetailsCompliance needs to ask a different question
The practical challenge for compliance and risk professionals is resisting the natural desire to reduce TEVV to a simple question: Did the AI pass? I would ask something different: What did this evaluation actually prove? Instead of simply confirming that testing occurred, I want to understand what the organization was trying to learn, why particular measurements were selected, what assumptions influenced the evaluation and what was outside its scope.
This is consistent with the structure NIST has proposed. TEVV-Athlon starts with organizational objectives and ends with “synthesize and interrogate,” where the results are interpreted to provide information that can support organizational decisions. For compliance professionals, that last step may ultimately be the most important.
We do not need to become data scientists to challenge an AI evaluation. We do need to understand the reasonable conclusions management can draw from it.
I would want four questions answered before relying on the results:
- What were we trying to learn?
- Why did we choose these measurements?
- What assumptions or limitations affected the results?
- What didn’t we test?
Those questions are different from asking whether the required testing was completed. They require us to examine the relationship between the evidence and the business decision being made. This is particularly important as AI becomes embedded in third-party products. Organizations may receive evaluation reports from vendors rather than conduct every test themselves. A vendor may be able to show impressive benchmark results. The compliance question remains the same: Do those results provide evidence about the way our organization intends to use the AI?
That is where TEVV can become more than another control requirement. Used properly, it can improve the quality of the information management uses to make decisions about AI.
Don’t standardize the value out of TEVV
After years of working with technology risk assessments, I have learned that completing an assessment and understanding risk are not necessarily the same thing. That is why I find NIST AI 200-2 interesting and useful.
NIST is not offering organizations another universal AI test. TEVV-Athlon provides a structured way to determine what an organization needs to know about an AI system and then to construct an evaluation capable of producing evidence relevant to that objective. Its flexibility is not a weakness that compliance departments need to correct. It is part of the framework’s value.
Inevitably, pressures will mount to standardize the process. Organizations want consistency. Auditors want evidence. Executives want understandable results. Regulators want organizations to demonstrate that appropriate controls exist. All of those expectations are reasonable.
The problem begins when demonstrating completion of the process becomes more important than understanding what the evaluation tells us. NIST’s inclusion of Goodhart’s Law should serve as a useful warning. The moment organizations begin managing AI evaluations primarily to achieve acceptable scores, they risk losing sight of why those measurements were selected in the first place.


Larry Marks is a GRC adviser and consultant. He most recently served in a variety of roles at BDO and IBM. 







