The Challenge
NIST released its TEVV-Athlon framework (NIST AI 200-2) to prevent organizations from turning AI system evaluation into a checkbox exercise. Instead of providing a standard list of tests, it asks organizations to define their objectives first, then organize evaluations around those objectives through four stages: articulate and organize, define and construct, apply and measure, and synthesize and interrogate.
However, organizations quickly began doing what NIST warned against. They converted the framework's flexibility into standardized templates, turning "what did we learn about this AI system's actual performance?" into "did we complete the TEVV process?"
This issue isn't just theoretical. It appears when compliance teams receive vendor evaluation reports showing impressive benchmark results, document that testing occurred, and move on without questioning whether those benchmarks measure how the organization intends to use the AI. It also shows up when risk professionals produce passing scores that make management comfortable, rather than providing evidence that helps management understand where uncertainty remains.
The Environment and Constraints
The pressure to standardize comes from legitimate needs. Organizations require consistency across business units. Auditors need documented evidence. Executives want risk information they can compare and act on. Regulators expect appropriate controls.
The constraint is that AI systems, their uses, and their risks vary too much for a single evaluation methodology to work in every situation. A chatbot handling customer service inquiries operates under different risk conditions than an AI system making credit decisions. The same AI model deployed in different contexts requires different evaluation approaches.
NIST recognized this reality by building adaptability into the framework. But adaptability creates documentation challenges. It's harder to demonstrate a customized evaluation than to show you've checked every box on a standard list. The natural response is to create that standard list anyway, which defeats the framework's purpose.
The framework also discusses Goodhart's Law: when a measure becomes a target, it ceases to be a good measure. NIST warns that optimizing an AI system to perform well against a particular benchmark may not indicate how well that system will perform in the real world. Organizations can't simply establish performance thresholds, conduct testing, and conclude that passing scores mean risk has been addressed.
The Approach Taken
The framework's design prevents checklist reduction by structuring evaluation as an investigation rather than a test. It starts with organizational objectives, not measurements. Only after articulating what the organization needs to know about the AI system does the framework move to defining what should be measured and how.
The final stage, "synthesize and interrogate," positions evaluation results as evidence for organizational decisions rather than as pass/fail outcomes. NIST recommends using multiple complementary evaluation approaches and emphasizes real-world testing over exclusive reliance on benchmarks.
This structure forces different questions. Instead of "did the AI pass TEVV?", the framework asks "what did this evaluation actually prove?" It encourages organizations to examine the relationship between the evidence collected and the business decision being made.
For third-party AI products, this approach shifts the compliance question. A vendor may demonstrate impressive benchmark performance. The evaluation question isn't whether the vendor completed testing. It's whether those results provide evidence about how your organization intends to use the AI.
Results and Metrics
NIST's draft doesn't claim to solve the standardization problem. The framework acknowledges that a good evaluation may produce an uncomfortable answer. It may indicate that an AI performs well under certain conditions but that insufficient evidence exists to reach the same conclusion under others. That uncomfortable answer may be exactly what management needs to know.
The framework's value shows up in the quality of information available for decision-making, not in the completion of a process. Organizations using TEVV properly should be able to articulate what assumptions or limitations affected their evaluation results and what scope boundaries they didn't test.
What They Would Do Differently
The draft framework doesn't include retrospective analysis, but the design choices suggest NIST learned from watching other frameworks become standardized. The emphasis on Goodhart's Law and the warning against optimizing for benchmarks indicate awareness of how measurements get misused.
If organizations are already converting TEVV into checklists, the framework might benefit from more explicit guidance on what interrogation looks like in practice. The "synthesize and interrogate" stage is conceptually sound but operationally vague. Compliance professionals need concrete examples of how to challenge an AI evaluation without becoming data scientists.
Takeaways for Your Team
Don't ask whether your AI evaluation passed. Ask what it proved. Before relying on TEVV results, get answers to four questions:
What were we trying to learn? This isn't "did we meet regulatory requirements?" It's "what specific uncertainty about this AI system's performance did we need to resolve?"
Why did we choose these measurements? Every measurement selection involves assumptions about what matters. Make those assumptions explicit. If you're testing for bias, which definition of bias did you use and why does it match your use case?
What assumptions or limitations affected the results? Real-world conditions differ from test environments. Document what you controlled for and what you couldn't control. If your evaluation used synthetic data, state how that affects your conclusions about production performance.
What didn't we test? Scope boundaries matter more than coverage percentages. An evaluation that thoroughly tests 60% of relevant conditions while clearly documenting the untested 40% provides better decision support than one claiming 100% coverage through definitional gymnastics.
Resist the pressure to turn TEVV into a control requirement that you simply demonstrate completion of. The framework's flexibility isn't a weakness your compliance program needs to correct. It's the feature that makes the evaluation useful.
When you receive AI evaluation reports from vendors, don't just verify that testing occurred. Examine whether the vendor's test conditions match your deployment context. A language model that performs well on academic benchmarks may produce different results when exposed to your organization's domain-specific terminology and edge cases.
The moment you start managing AI evaluations primarily to achieve acceptable scores, you've made Goodhart's Law manifest. The score becomes the target, and you lose sight of why you selected those measurements. NIST built TEVV-Athlon to prevent exactly that outcome. Don't standardize the value out of it.





