Section 9 of 15
Measure
National Institute of Standards and Technology · about 5 minutes
5.3 Measure
¶The MEASURE function employs quantitative, qualitative, or mixed-method tools, techniques, and methodologies to analyze, assess, benchmark, and monitor AI risk and related impacts. It uses knowledge relevant to AI risks identified in the MAP function and informs the MANAGE function. AI systems should be tested before their deployment and regularly while in operation. AI risk measurements include documenting aspects of systems’ functionality and trustworthiness.
¶Measuring AI risks includes tracking metrics for trustworthy characteristics, social impact, and human-AI configurations. Processes developed or adopted in the MEASURE function should include rigorous software testing and performance assessment methodologies with associated measures of uncertainty, comparisons to performance benchmarks, and formalized reporting and documentation of results. Processes for independent review can improve the effectiveness of testing and can mitigate internal biases and potential conflicts of interest.
¶Where tradeoffs among the trustworthy characteristics arise, measurement provides a traceable basis to inform management decisions. Options may include recalibration, impact mitigation, or removal of the system from design, development, production, or use, as well as a range of compensating, detective, deterrent, directive, and recovery controls.
¶After completing the MEASURE function, objective, repeatable, or scalable test, evaluation, verification, and validation (TEVV) processes including metrics, methods, and methodologies are in place, followed, and documented. Metrics and measurement methodologies should adhere to scientific, legal, and ethical norms and be carried out in an open and transparent process. New types of measurement, qualitative and quantitative, may need to be developed. The degree to which each measurement type provides unique and meaningful information to the assessment of AI risks should be considered. Framework users will enhance their capacity to comprehensively evaluate system trustworthiness, identify and track existing and emergent risks, and verify efficacy of the metrics. Measurement outcomes will be utilized in the MANAGE function to assist risk monitoring and response efforts. It is incumbent on Framework users to continue applying the MEASURE function to AI systems as knowledge, methodologies, risks, and impacts evolve over time.
¶Practices related to measuring AI risks are described in the NIST AI RMF Playbook. Table 3 lists the MEASURE function’s categories and subcategories.
¶Table 3: Categories and subcategories for the MEASURE function.
¶Categories Subcategories MEASURE 1: MEASURE 1.1: Approaches and metrics for measurement of AI Appropriate risks enumerated during the MAP function are selected for implemethods and metrics mentation starting with the most significant AI risks. The risks are identified and or trustworthiness characteristics that will not – or cannot – be applied. measured are properly documented.
¶MEASURE 1.2: Appropriateness of AI metrics and effectiveness of existing controls are regularly assessed and updated, including reports of errors and potential impacts on affected communities.
¶MEASURE 1.3: Internal experts who did not serve as front-line developers for the system and/or independent assessors are involved in regular assessments and updates. Domain experts, users, AI actors external to the team that developed or deployed the AI system, and affected communities are consulted in support of assessments as necessary per organizational risk tolerance.
¶MEASURE 2: AI MEASURE 2.1: Test sets, metrics, and details about the tools used systems are during TEVV are documented.
¶evaluated for 2.2: Evaluations involving human subjects meet ap- MEASURE trustworthy plicable requirements (including human subject protection) and characteristics. are representative of the relevant population.
¶MEASURE 2.3: AI system performance or assurance criteria are measured qualitatively or quantitatively and demonstrated for conditions similar to deployment setting(s). Measures are documented.
¶MEASURE 2.4: The functionality and behavior of the AI system and its components – as identified in the MAP function – are monitored when in production.
¶MEASURE 2.5: The AI system to be deployed is demonstrated to be valid and reliable. Limitations of the generalizability beyond the conditions under which the technology was developed are documented.
¶Continued on next page Table 3: Categories and subcategories for the MEASURE function. (Continued) Categories Subcategories MEASURE 2.6: The AI system is evaluated regularly for safety risks – as identified in the MAP function. The AI system to be deployed is demonstrated to be safe, its residual negative risk does not exceed the risk tolerance, and it can fail safely, particularly if made to operate beyond its knowledge limits. Safety metrics reflect system reliability and robustness, real-time monitoring, and response times for AI system failures.
¶MEASURE 2.7: AI system security and resilience – as identified in the MAP function – are evaluated and documented.
¶MEASURE 2.8: Risks associated with transparency and accountability – as identified in the MAP function – are examined and documented.
¶MEASURE 2.9: The AI model is explained, validated, and documented, and AI system output is interpreted within its context – as identified in the MAP function – to inform responsible use and governance.
¶MEASURE 2.10: Privacy risk of the AI system – as identified in the MAP function – is examined and documented.
¶MEASURE 2.11: Fairness and bias – as identified in the MAP function – are evaluated and results are documented.
¶MEASURE 2.12: Environmental impact and sustainability of AI model training and management activities – as identified in the MAP function – are assessed and documented.
¶MEASURE 2.13: Effectiveness of the employed TEVV metrics and processes in the MEASURE function are evaluated and documented.
¶MEASURE 3: MEASURE 3.1: Approaches, personnel, and documentation are Mechanisms for in place to regularly identify and track existing, unanticipated, tracking identified and emergent AI risks based on factors such as intended and ac- AI risks over time tual performance in deployed contexts.
¶are in place. 3.2: Risk tracking approaches are considered for MEASURE settings where AI risks are difficult to assess using currently available measurement techniques or where metrics are not yet available.
¶Continued on next page Table 3: Categories and subcategories for the MEASURE function. (Continued) Categories Subcategories MEASURE 3.3: Feedback processes for end users and impacted communities to report problems and appeal system outcomes are established and integrated into AI system evaluation metrics.
¶MEASURE 4: MEASURE 4.1: Measurement approaches for identifying AI risks Feedback about are connected to deployment context(s) and informed through efficacy of consultation with domain experts and other end users. Apmeasurement is proaches are documented. gathered and 4.2: Measurement results regarding AI system trust- MEASURE assessed. worthiness in deployment context(s) and across the AI lifecycle are informed by input from domain experts and relevant AI actors to validate whether the system is performing consistently as intended. Results are documented.
¶MEASURE 4.3: Measurable performance improvements or declines based on consultations with relevant AI actors, including affected communities, and field data about contextrelevant risks and trustworthiness characteristics are identified and documented.