DOI QR코드

DOI QR Code

Applying the NIST AI RMF for an AI Security Platform: A Conceptual Reliability Framework Integrating MITRE ATT&CK with Pilot Empirical Evaluation

AI 보안 플랫폼에 NIST AI RMF를 적용하기: MITRE ATT&CK와 파일럿 실증 평가를 통합한 개념적 신뢰성 프레임워크

  • Received : 2026.01.27
  • Accepted : 2026.04.13
  • Published : 2026.06.30

Abstract

This paper presents a conceptual framework for applying the NIST Artificial Intelligence Risk Management Framework(AI RMF 1.0) to support reliability and security evaluation in an AI security platform. The framework operationalizes governance-driven risk management using adversarial threat scenarios derived from the MITRE ATT&CK knowledge base, bridging the persistent gap between high-level governance principles and technical adversarial threat modeling. The four RMF components (Govern, Map, Measure, and Manage) are aligned with operational components and supplemented by structured evidence templates implemented as an Excel-based prototype. To demonstrate operational feasibility, a pilot empirical evaluation was conducted across three ATLAS-mapped adversarial scenarios, achieving mitigation effectiveness rates of 75%, 100%, and 25% respectively.

본 논문은 인공지능 보안 플랫폼의 신뢰성 및 보안 평가를 지원하기 위해 NIST 인공지능 위험 관리 프레임워크(AI RMF 1.0)를 적용하는 개념적 프레임워크를 제시합니다. 이 프레임워크는 MITRE ATT&CK 지식 기반에서 도출된 공격자 위협 시나리오를 활용하여 거버넌스 기반 위험 관리를 구체화하고, 고수준의 거버넌스 원칙과 기술적 공격자 위협 모델링 간의 지속적인 격차를 해소합니다. RMF의 네 가지 구성 요소(거버넌스, 매핑, 측정, 관리)는 운영 구성 요소와 연계되며, 엑셀 기반 프로토타입으로 구현된 구조화된 증거 템플릿으로 보완됩니다. 운영 가능성을 입증하기 위해 ATLAS에 매핑된 세 가지 공격자 위협 시나리오에 대한 파일럿 실증 평가를 수행했으며, 각각75%, 100%, 25%의 완화 효과를 달성했습니다.

Keywords

I. INTRODUCTION

Artificial Intelligence (AI) systems are now integrated in various critical sectors, including industries, healthcare, manufacturing, and finance, making their reliability and safety crucial for trust and performance [1]. At the core of this conversation is reliability engineering, which ensures that components perform without failure when applied to complex engineered systems. However, AI increases complexity because an AI system is made up of many subsystems, including data, models, infrastructure, and human interactions. These diverse subsystems may fail either intentionally (due to adversarial attacks) or unintentionally (due to human errors, design flaws), both of which may affect the system’s reliability and dependability [1]. Recently, new Generative Agentic AI systems have made decisions on their own by combining large language models along with reasoning, planning, and external tool integration, which enables autonomous actions across enterprise environments. They identified nine major threat categories in the analysis, such as cognitive dimensions, temporal dimensions, tool integration, trust boundary, identity fluidity, and governance complexity, indicating that such autonomy requires structured evaluation through frameworks like NIST AI RMF and MITRE ATT&CK [2]. In this study, we define reliability as an ability of an AI system to maintain expected operational behavior throughout its lifecycle. A reliable system resists adversarial manipulation, minimizes hallucinations, and compiles with governance and safety policies. We distinguish reliability from several closely related terms to avoid confusion as follows. Security focuses on the protection of AI system components-such as data pipelines, model inference layers against unauthorized access, adversarial exploitation, and integrity violations. Trustworthiness is the overarching, multi-dimensional property which encompassing reliability, safety, fairness, explainability, privacy, and accountability[3]. Risk defines the likelihood of a negative outcome from the interaction of a threat, a vulnerability, and the consequence of exploitation within an AI system[3]. Additionally, resilience is the AI’s capacity to bounce back from disruptions and stay online [1] while robustness is its ability to handle manipulated or unusual inputs without falling [1]. While trustworthiness is the primary goal, this framework specifically addresses its measurable, technical foundations such as reliability, security, robustness, and resilience.

There are several frameworks to support AI risk and reliability evaluation, and among them, the NIST AI Risk Management Framework (AI RMF) and MITRE ATT&CK are notable. The U.S. National Institute of Standards and Technology (NIST) adopted AI Risk Management Framework (AI RMF 1.0)in 2023 [3]. The framework is supported by the AI RMF Playbook and the guide application [4], also augmented by key 2025 companion resources such as Adversarial Machine Learning Taxonomy (AI 100-2E2025) and cybersecurity overlays.

The AI RMF provides a voluntary, yet increasingly influential approach for managing AI risks to individuals, organizations, and society. The framework consists of four key functions that together provide a lifecyle-oriented approach to AIrisk management.

∙ Govern: Establishes the direction for organizational policy, culture, and accountability related to the management of AI-related risks.

∙ Map: Mentions the context, objectives, and possible impacts of AI systems as well as the associated risks.

∙ Measure: Develops metrics, assessment approaches, and knowledge on system performance and risk factors.

∙ Manage: Prioritizes and monitors risks using mitigation plans, documentation, and continuous tracking [3].

In late 2025, the AI RMF's importance has risen significantly. The framework is serving as a de-facto reference for trustworthy AI amid evolving global regulations. It aligns with international standards (e.g., ISO/IEC 42001), supporting compliance in regions enforcing risk-based AI rules, such as the EU AI Act (with phased implementation ongoing, including prohibited practices effective February 2025 and high-risk obligations progressing toward 2027). To enable auditable, proactive risk management—bridging policy with technical safeguards many organizations adopt the AI RMF, while fostering innovation in high-stakes applications. As shown in Figure 1, the NIST AI RMF operates as an iterative cycle.

JBBHCB_2026_v36n3_905_3_f0001.png 이미지

Fig. 1. NIST AI RMF as an iterative cycle.

Although this framework gives a strong governance structure, it offers limited technical guidance to handle adversarial or cybersecurity threats. On the other hand, the MITRE ATT&CK framework catalogs adversarial tactics and techniques observed in real-world attacks, offering valuable information for threat detection and response [5]. However, it provides extensive technical elaboration of cyberthreats, without the coverage of governance or reliability management of AIsystems. In addition, “most organizations are in the early stages of AI RMF adoption maturity” as Dotan et al. [7] argue that the NIST AI RMF is mostly used throughout the documentation process, not for risk control implementation. This indicates a growing gap between policy-level governance models and technical adversarial frameworks—highlighting the need for an approach that operationalizes RMF-based governance processes using ATT&CK-based adversarial threat intelligence to comprehensively assess the reliability of the AI system.

Many other governance and threat modeling frameworks exist for AI and cybersecurity; however, the NIST AI RMF and MITRE ATT&CK were selected in this study because they complement each other across organizational and technical layersof AI security. To identify, measure and manage risks AI RMF gives a lifecycle-oriented governance structure. However, it intentionally remains implementation-agnostic and does not prescribe specific adversarial tactics or testing scenarios.

On the other hand, MITRE ATT&CK and its AI-focused extension ATLAS offer an empirically grounded taxonomy of adversarial tactics, techniques, and procedures derived from real-world attack observations. ATT&CK does good job at technical threat modeling, but lacks a governance-oriented risk management lifecycle.

Therefore, the proposed framework does not attempt to merge the two frameworks. Instead, we describe a conceptual approach that operationalizes the governance processes defined in the AI RMF by using ATT&CK adversarial knowledge to instantiate measurable threat scenarios and evaluation procedures. To further illustrate the operational feasibility of this conceptual approach, a pilot empirical evaluation based on a real RAG-based AI security platform is included in this study and providing preliminary evidence of the framework’s practical applicability.

1.1 Contributions

This paper demonstrates a framework that operationalizes the NIST AI RMF governance model using MITRE ATT&CK adversarial threat intelligence within an AI security platform. The main contributions are:

∙ Our framework systematically maps RMF functions with ATT&CK adversarial behavior in order to support structured evaluation of AI system reliability and security posture.

∙ We address the current gap between normative risk alignment and technical adversarial behavior.

∙ We map the four RMF functions—Govern, Map, Measure, and Manage to align with AI platform components including Risk Register, Vector Database Query, Adversarial Testing, and ReAct Flow.

∙ We develop an Excel-based prototype evaluation pipeline to capture risk scenarios, evidence, performance tests, and target thresholds for each workstream.

∙ The framework provides a tangible, measurable, threat-informed approach to the evaluation of the reliability and resiliency of the AI system that connects the NIST guidance with the operational security defenses based on MITRE ATT&CK.

∙ We validate the operational feasibility of the proposed framework through a pilot empirical evaluation across three ATLAS-mapped adversarial scenarios(AML.T0051, AML.T0020, AML.T0015), illustrating mitigation effectiveness rates of 75%, 100%, and 25% respectively.

II. RELATED WORK

The NIST AI Risk Management Framework (AI RMF) [3] is a resource that guides organizations in the design, development, deployment, and useof trustworthy AI systems. The core of the NIST AI Risk Management Framework(AI RMF) has four key components —Govern, Map, Measure, and Manage —that facilitate responsible risk assessment throughout the AI lifecycle. The framework emphasizes cultivating a culture of risk awareness in its Govern section, recognizing system context, identifying threats, and prioritizing mitigation actions. Although this provides an excellent starting point for high-level AI governance based on practical and satisfactory fundaments, its focus remains largely on organizational processes rather than technical adversarial risks or security testing.

However, researchers [7] note that the RMF lacks operational tools to evaluate the maturity of AI risk management because it remains mostly conceptual. Similarly, Swaminathan et al. [8] observed that the framework gave a valuable governance structure when applying to real-world surveillance systems, but required measurable criteria for technical reliability.

On the other hand, MITRE ATT&CK [5] is a widely used and always updated knowledge base of adversarial tactics and techniques based on real-world observations. MITRE ATLAS [6] is an extension of the AI-specific context of the MITRE ATT&CK, covering attacks such as data poisoning, evasion, and prompt injection, providing a threat intelligence matrix tailored to AI systems. Various research, such as [9], [10], analyzed threats to large language model (LLM) powered applications, emphasizing possible attack vectors including data poisoning, jailbreaking, compositional injection, and prompt-based exploitation, and suggested execution to be used in unison with STRIDE and DREAD models to offer a comprehensive model for AI threat modeling, along with related countermeasures. Others also developed a practical AI threat library, which has 63 structured security and privacy questions designed for developers and engineering teams. However, they also observed that such libraries have a hard time capturing rapidly changing attack surfaces, which highlights a limitation of static threat-mapping strategies. Despite these progresses in AI-accommodating threat modeling, not much has been linked adversarial intelligence with top-level governance frameworks such as the NIST AI RMF to create a unified reliability or security evaluation model.

Some studies have attempted to connect legal or high-level governance frameworks with technical threat models. Ruohonen et al. [11] proposed a mapping between the European Union Cyber Resilience Act(CRA) and the MITRE ATT&CK framework to show regulatory controls with adversarial threat tactics. Schroder and Breier [12] introduced a quantitative risk-scoring model based on dimensions such as damage potential and attack effort for evaluating AI system reliability, but their work does not integrate with any governance frameworks like the AI RMF. Meanwhile, another researcher [13] identified the lack of usability of policies with security tools for high-risk AI. Reliability and threat modeling are separated in some recent models, such as CORTEX[14]. Together, these works suggest an increasing awareness of policy with threat alignment; however, no consolidated frameworkties NIST AI RMF to adversarial threat models. Surveys of responsible AI frameworks [15], [16] consistently highlight the absence of unified methods that merge policy-driven governance with technical adversarial intelligence.

In summary, even though policy frameworks could be discussed with threat modeling separately by some researchers, there is no unified approach that connects the NIST AI RMF's governance principles with MITRE ATT&CK for AI reliability. To fill this gap, our work presents a unified approach to map RMF core functions with ATT&CK-based adversarial behaviors, validated through a pilot empirical evaluation across three ATLAS-mapped scenarios. The approach illustrates a conceptual AI security platform which combines RMF functions (Govern, Map, Measure, Manage) with operational components and documents risk evidence through an Excel-based prototype pipeline, demonstrating that governance-driven risk management can be operationalized in practice.

III. PROPOSED FRAMEWORK

In previous sections, the gaps between current AI risk governance and threat modeling approaches were discussed. In this section, we first provide an overview of the conceptual architecture used in our study. Then, we describe how the proposed framework operationalizes the NIST AI RMF by leveraging MITRE ATT&CK within an AI security platform through a structured Excel-based prototype used to document risk scenarios, evidence logs, and adversarial threat mapping.

3.1 Framework Overview

In Figure 2, an overview of our proposed framework is presented, describing how the AI security platform applies trustworthy AI principles by connecting the NIST AI RMF with MITRE ATT&CK.

JBBHCB_2026_v36n3_905_6_f0001.png 이미지

Fig. 2. Conceptual architecture of the proposed AI security framework integrating NIST AI RMF with MITRE ATT&CK/ATLAS.

At the top of the main pipeline, the core functions of the NIST AI RMF guide how governance, mapping, measurement, and management influence the system. The middle layer represents a generalized architecture of an AI security platform used as the evaluation target within the proposed framework. The workflow begins with data collection, which is fragmented and embedded before being stored in a vector database. After that, the stored vectors are retrieved and then combined with a fine-tuned language model using a retrieval-augmented generation (RAG) setup. Then the ReAct pattern is followed by the reasoning stage, and the model can alternate between thought and action based on the context. At the bottom layer, MITRE ATT&CK and its AI-focused extension ATLAS provide adversarial tacticsand techniques that inform the definition of threat scenarios and adversarial testing strategies. In this study, three adversarial scenarios derived from the MITRE ATLAS taxonomy (prompt injection, retrieval poisoning, and adversarial paraphrasing) are used to empirically validate how threat-informed evaluation can be conducted within the proposed framework.

Together, the three layers form a structured viewpoint.

∙ NIST AI RMF provides governance and reliability mandates.

∙ AI security platform provides target components for assessment.

∙ MITRE ATT&CK/ATLAS provides adversarial knowledge to simulate or anticipate potential threats.

This three-layer architecture ensures that governance requirements are not treated in isolation from technical threat realities, but are instead operationalized through a unified evaluation workflow that connects policy mandates to measurable adversarial testing outcomes.

3.2 RMF-Platform Functional Alignment

To operationalize NIST guidance in a technical context, each RMF function is mapped to a corresponding functional block within the AI security platform. This ensures that governance and technical workflows are aligned rather than treated as separate processes.

∙ Govern → Accountability and Policy Controls:

The Govern function informs data provenance rules, module ownership, oversight intervals, and documentation requirements. These governance constraints apply across the data pipeline, vector database maintenance, and model reasoning processes.

∙ Map → Context and System Boundaries:

The Map function facilitates identification of data sources, preprocessing steps, embedding strategies, and interaction boundaries, ensuring traceability between system context and potential risk exposure.

∙ Measure → Performance and Threat Evaluation:

The Measure function aligns with technical evaluation modules. MITRE ATT&CK/ATLAS techniques are used to define adversarial evaluation scenarios such as prompt injection, data poisoning, and evasion attacks, illustrating how system robustness and reliability can be assessed within the framework.

∙ Manage → Mitigation and Continuous Oversight:

The Manage function outlines how corrective actions, risk thresholds, and monitoring cycles would be executed. The Manage function defines how mitigation strategies, risk thresholds, and monitoring cycles can be applied based on the outcomes of adversarial evaluation and system performance observations.

Explainability assurance is identified by the framework as a priority area for futureintegration, aligned with the Measure function’s evaluation scope. As shown in Figure 3, this component remains partially addressed and is marked for full-scale deployment.

3.3 RMF-ATT&CK Threat Mapping

We provide an explicit mapping between the NIST AI RMF functions and representative MITRE ATT&CK adversarial techniques to clarify how governance-driven risk management processes can be operationalized using adversarial threat intelligence. Rather than directly merging the two frameworks, the suggested method uses AI RMF as the governance structure for risk management while using MITRE ATT&CK techniques as technical indicators for adversarial threat scenarios.

This mapping shows how each RMF function relates to specific risk contexts and technical threat scenarios that may affect AI system reliability.

Table 1 presents the mapping between RMF governance functions, AI Platform components, and representative MITRE ATT&CK techniques. The Govern function is associated with adversarial tactics related to unauthorized system access or governance violations such as Initial Access since it focuses on accountability structures and policy enforcement. The Map function is relevant to data-centric threats such as Data Poisoning that may affect training or knowledge sources because it identifies system context, data sources, and interaction boundaries. The Measure function evaluates system robustness and model behavior, and therefore incorporates adversarial testing scenarios such as Prompt Injection that target LLM reasoning processes. Finally, the Manage function focuses on mitigation, monitoring, and corrective responses, which aligns with ATT&CK techniques related to defense evasion or safety bypass attempts during system operation.

Table 1. Mapping between NIST AI RMF functions and MITRE ATT&CK adversarial techniques for the proposed AI security platform

JBBHCB_2026_v36n3_905_8_t0001.png 이미지

3.4 Excel-Based Risk Mapping Pipeline

To illustrate how the RMF–platform alignment translates into a measurable workflow, we designed a five-sheet Excel-based pipeline. This prototype provides a structured documentation framework for capturing risk scenarios, adversarial evaluation evidence, and reliability metrics during AI system assessment. While the initial design served as a pre-evaluation planning artifact, the pilot empirical evaluation conducted in Section 4.2 demonstrates that the pipeline can be populated with real experimental observations, confirming its operational viability.

3.4.1 Risk Register (MAP + MEASURE):

The Risk Register sheet serves as the foundation of the Excel-based risk mapping pipeline. It also serves as a tracking mechanism for the AI- related risks across the AI security platform. Each row defines a specific workflow, including data collection, vector database, and adversarial testing. The Risk scenario, quantitative metric, baseline value, target threshold, impact severity, root cause reference, and mitigation owner are also described. The register was initially designed as a pre-evaluation planning artifact with baseline and target thresholds defined conceptually. Following the pilot empirical evaluation described in Section 4.2, selected metrics-specifically prompt injection compliance, retrieval poisoning resistance, and adversarial paraphrasing detection were updated with empirically observed values. The Map and Measure functions of the NIST AI RMF are supported by this structure, which converts conceptual risk categories into quantifiable evidence-based parameters. A pre-evaluation version of the Risk Register is provided in Table 2. Again, every risk entry is also linked to a mitigation action, which specifies the corresponding preventive or corrective control. These mitigation elements are maintained separately for clarity and traceability, as summarized in Table 3.

Table 2. Risk Register (NIST-GUIDED AI RMF)

JBBHCB_2026_v36n3_905_9_t0001.png 이미지

Table 3. Mitigation Action For Identified Risks

JBBHCB_2026_v36n3_905_9_t0002.png 이미지

Together, the Risk Register and the mitigation mapping support continuous monitoring, policy alignment, and proactive management of AI system reliability.

Furthermore, each risk scenario is conceptually linked to relevant MITRE ATT&CK tactics or techniques; however, the detailed ATT&CK mappings are documented separately in the Attack Posture Sheet (Table V), rather than inside the Risk Register itself.

3.4.2 Governance Sheet:

For the Govern function of the NIST AI RMF 1.0, we have a governance sheet that provides a structured way

to define accountability, policy control, and oversight frequency. This supports the NIST principle that AI governance required transparent assignment of roles, documented policies, and consistent monitoring across technical and organizational layers. As shown in Table 4, every entry aligns with the relevant NIST AI RMF(GOVERN) sub-function for data provenance or red-team testing.

Table 4. Governance Sheet

JBBHCB_2026_v36n3_905_9_t0003.png 이미지

3.4.3 Attack Posture Sheet:

The Attack Posture Sheet implements the Measure and Manage functions of the NIST AI RMF by connecting adversarial events to corresponding ATT&CK techniques and control actions inside the AI security platform. This sheet records each adversarial test, the targeted system component, the observed impact, and the corresponding mitigation or defense response. By documenting these relations, the framework links technical vulnerability evidence with risk metrics and allows traceable updates within the Excel-based risk pipeline.

Each attack entry follows the MITRE ATT&CK and ATLAS taxonomy, where theTactic column identifies the adversarial goal (e.g., data poisoning, model evasion, prompt injection), and the Technique ID provides the unique ATT&CK reference. Table 5 summarizes an excerpt of this mapping.

Table 5. Attack Posture Sheet

JBBHCB_2026_v36n3_905_10_t0001.png 이미지

3.4.4 Root Cause Analysis Log:

The Root Cause Analysis Log serves as the closing component of the Excel-based reliability pipeline, linking directly to the Manage function of the NIST AI RMF. Its purpose is to ensure that every detected failure or abnormal event identified during adversarial testing, fine-tuning, or runtime monitoring is analyzed for its underlying cause and recorded with corresponding corrective and preventive actions. Table 6 summarizes an illustrative excerpt of this analysis log.

Table 6. Root Cause Analysis Log

JBBHCB_2026_v36n3_905_11_t0001.png 이미지

Every entry in the Root Cause Analysis Log is traceable to its corresponding Governance Sheet owner and evidence directory, ensuring transparency and reproducibility of the remediation process. This systematic linkage between incident detection, root-cause analysis, and corrective action closes the reliability management loop defined by the NIST AI RMF and completes the AI security platform’s threat-informed reliability evaluation framework.

IV. RESULTS AND DISCUSSION

In this section, we present the implementation and evaluation workflow of the proposed framework. The study focuses on demonstrating how an AI security platform can operationalize reliability and security assessment through the integration of the NIST AI RMF and MITRE ATT&CK.

To support this, we prepared five structured sheets, which together form the core of the assessment pipeline.

In addition to the conceptual pipeline, we present a pilot empirical evaluation conducted across three adversarial scenarios-prompt injection (AML.T0051), retrieval poisoning (AML.T0020), and adversarial paraphrasing (AML.T0015) derived from the MITRE ATLAS taxonomy and performed on a RAG pipeline built from a corpus of AI security-related documents. Together, the pipeline and the experimental scenarios demonstrate how governance-driven risk management can be translated into practical, measurable evaluation procedures within an AI security platform.

4.1 Framework Initialization and Alignment Outputs

The designed Excel-based pipeline serves as the foundation of our proposed framework. Each Excel sheet was constructed to capture risk context, policy frequency, technical metrics, and adversarial mapping. These initialization components establish the documentation and monitoring infrastructure required for systematic reliability assessment within the AI security platform.

To evaluate how effectively the proposed framework connects policy-driven governance with technical threat modeling, we examined the alignment between the four NIST AI RMF functions and the operational components of the AI security platform. The analysis highlights how each RMF function maps to a concrete module—risk registration (Govern), contextual evidence via Vector DB (Map), adversarial testing (Measure), and the ReAct-based mitigation loop (Manage).

As shown in Figure 3, this alignment reveals both areas of strong correspondence and areas requiring future refinement. Governance and measurement exhibit complete alignment, while explainability assurance and automated oversight remain partially addressed. These insights confirm that the framework can support risk evaluation both conceptually and operationally, as further validated by the pilot empirical evaluation presented in Section 4.2, while also identifying improvements-particularly in explainability assurance and automated oversight- needed for full-scale deployment.

JBBHCB_2026_v36n3_905_12_f0001.png 이미지

Fig. 3. RMF-AI security platform Alignment and Gap Identification Overview.

4.2 Small-Scale Adversarial Evaluation

We conducted a small-scale adversarial evaluation targeting three representative attack scenarios drawn from the MITRE ATLAS taxonomy to provide empirical grounding for the proposed framework. The evaluation was performed on a retrieval-augmented generation (RAG) pipeline constructed from a corpus of AI security-related documents (2136 chunks indexed via FAISS with cosine similarity). For each scenario there were two configurations: a baseline configuration with no defensive controls, and a mitigated configuration incorporating the countermeasures specified in the framework’s Mitigation Actions sheet. An automated Groq-based policy auditor (llama-3.1-8b-instant) was used to evaluate LLM responses for determining whether safety policy violations occurred, ensuring reproducible and objective judgment. We selected the three scenarios based on three criteria. First, representativeness: they address the three primary attack surfaces of the AI security platform – the inference layer (Prompt Injection, AML.T0051), the knowledge base(Retrieval Poisoning, AML.T0020), and the safety filter layer (Adversarial Paraphrasing, AML.T0015). Second, cross-functional RMF coverage: each scenario corresponds to a different RMF function, allowing for validation across multiple governance layers rather than a single functional area. Third, practical feasibility: we can implemented all three in a RAG pipeline without accessing model weights which make them candidates for pilot-scale evaluation. Other ATLAS methods such as model inversion or membership inference are reserved for further large-scale evaluation as they require full model access.

4.2.1 Scenario 1 (Prompt Injection)

To conduct the test first we submitted ten adversarial prompt injection attempts to the RAG pipeline targeting the LLM inference layer. In the baseline configuration, where no system-level defense instructions were applied, 8 out of 10 injection attempts successfully bypassed the model’s default behavior, yielding a bypass rate of 80%. This high baseline vulnerability confirms that unprotected LLMs integrated into RAG pipelines are highly susceptible to instruction override attacks, consistent with ATLAS. Following the application of the mitigation defined in the framework’s Mitigation Action sheet-specifically, a system-level defense prompt enforcing instruction anchoring and context distrust- the bypass rate was reduced to 2 out of 10, representing a 75% reduction. The post-mitigation dashboard indicator was classified as YELLOW, suggesting partial effectiveness and flagging this component for further remediation within the Manage function cycle.

4.2.2 Scenario 2 (Retrieval Poisoning)

Ten queries were submitted alongside adversarially crafted documents injected into a temporary clone of the FAISS knowledge base to simulate a data poisoning attack targeting the vector database layer. In the baseline configuration, we observed only 1 out of 10 queries resulted in the poisoned document influencing the model’s response (10% influence rate). In another 5 additional cases, the poisoned document was retrieved into the top-k results but the LLM demonstrably resisted its influence, producing factually grounded responses. In the remaining 4 cases, the poisoned document did not rank within the top-k results at all. Following the application of provenance-based filtering as the mitigation control, the influence rate was reduced to 0 out of 10 (0%), representing a 100% reduction. The post-mitigation dashboard indicator was classified as GREEN, confirming that provenance filtering is a highly effective control for this threat category within the Map and Measure functions.

4.2.3 Scenario 3 (Adversarial Paraphrasing)

Ten policy-violating queries were paraphrased using semantically equivalent butsyntactically varied language to test evasion of the safety filter layer. In the baseline configuration using a keyword-based filter, we found 4 out of 10 paraphrased queries evaded detection and elicited policy-violating responses from the LLM. This confirms that surface-form variation is a viable and practical evasion strategy against keyword-based moderation, consistent with ATLAS taxonomy. Following substitution of the keyword filter with a semantic similarity filter, we reduced the evasion rate to 3 out of 10 (30%), representing a 25% reduction. The post-mitigation dashboard indicator was classified as YELLOW.

4.2.4 Summary and Framework Validation

A small sample size (n=10 per scenario)was given to perform the empirical testing. We reported the results as descriptive proportions and we did not apply formal statistical inference. We planned larger-scale validation for future work. The risk reduction percentage number reported in this study are a relative measure of adversarial attack success rate between the baseline and mitigated configurations calculated as: Reduction(%) = (Baseline hits – Mitigated hits) ÷ Baseline hits × 100. In the case of scenario 1, for example, the reduction was calculated as (8-2) ÷ 8 × 100 = 75%. These figures reflect the effectiveness of the specific mitigation controls under controlled pilot conditions.

Table 7 summarizes the results across all three scenarios.

Table 7. Small-Scale Adversarial Evaluation Results

JBBHCB_2026_v36n3_905_14_t0001.png 이미지

These results serve two important purposes within the scope of this study. First, they demonstrate that the proposed framework’s evidence templates- specifically the Attack Posture Sheet and Mitigation Action Sheet can be populated with real experimental observations and used to drive meaningful updates to the reliability dashboard. Second, the results confirm that the ATT&CK-to-RMF mapping defined in Table 1 is operationally valid: each scenario maps cleanly to the corresponding RMF function, and the mitigation controls derived from the framework produced measurable improvements across all three scenarios.

4.3 Conceptual Reliability Dashboard

The conceptual reliability dashboard was designed as part of the framework prior to empirical evaluation. The pilot evaluation conducted in Section 4.2 subsequently validated three of its monitored metrics with empirically observed values. The Excel-based pipeline converges governance controls, risk register (evidence + metrics), and threat scenarios (adversarial tests) into a reliability dashboard reporting Red/Yellow/Green risk levels and baseline vs target comparison in Figure 4. This traffic-light system, inspired by NIST AI RMF Playbook examples, provides measurable visualization of reliability posture.

JBBHCB_2026_v36n3_905_14_f0001.png 이미지

Fig. 4. Conceptual Reliability Dashboard Integrating Risk Register and Threat Scenarios

Table 8 presents an illustrativere liability dashboard example showing how evaluation metrics can be mapped to governance monitoring indicators. Each workflow is associated with a measurable metric, a target threshold, and a corresponding R/Y/G status indicator. This structured representation allows security analysts to quickly identify reliability deviations and link them to their corresponding ATT&CK technique and root cause within the RMF-aligned evaluation pipeline. Following the pilot empirical evaluation presented in Section 4.2, three dashboard metrics were updated with empirically observed values: prompt injection compliance(post-mitigation bypass rate 20%, YELLOW), retrieval poisoning resistance(post-mitigation influence rate 0%, GREEN), and adversarial paraphrasing detection (post-mitigation evasion rate 30%, YELLOW). The remaining metrics retain their conceptual threshold definitions pending full-scale system evaluation.

Table 8. Reliability Dashboard With Pilot Evaluation Updates

JBBHCB_2026_v36n3_905_15_t0001.png 이미지

4.4 Identified Gaps and Improvement Areas

While the pilot empirical evaluation confirms the operational feasibility of the proposed framework, the following gaps remain before full-scale system deployment.

Specifically, the current reliability thresholds are derived from conceptual policy definitions rather than long-term operational measurement. As the system is evaluated under real workloads, these thresholds will need calibration using empirical metrics such as retrieval accuracy, hallucination rate, and adversarial bypass frequency. Another limitation concerns the current reliance on manually populated monitoring sheets. Future work will address this through API-based automation, enabling real-time ingestion of system telemetry into the risk management pipeline, as further elaborated in the Future Experimental evaluation section. Finally, additional modules, such as defense orchestration components or advanced monitoring layers may introduce new categories of operational risk that are not yet fully represented in the conceptual model. Rather than weakening the framework, these observations highlight the areas where the governance-driven evaluation pipeline can evolve to support practical AI security deployments.

4.5 Future Experimental EvaluationandContinuous Monitoring

Building on the pilot empirical evaluation presented in Section 4.2, future system-level evaluation will extend the framework by applying operational metrics collected from diverse AI components suchas RAG pipelines, language models, and associated security modules. Additionally, the Excel-based prototype will be progressively migrated toward an API-integrated architecture, enabling real-time ingestion of system telemetry into the risk management pipeline through structured log events and automated threshold monitoring. During these evaluations, metrics including retrieval accuracy, hallucination rates, bypass attempts, and adversarial interaction traces will be recorded and stored within the monitoring pipeline, initially via the Excel-based sheets and progressively through the API-integrated architecture.

These observations will allow the reliability threshold defined in the conceptual stage to be refined through empirical measurement. Different controlled experiments, such as prompt-injection attempts, unsafe tool-calling checks, or real-time model behavior analysis will refine the thresholds and mitigation pathways established in this study.

Following the NIST AI RMF Manage function, the pipeline will shift into a continuous monitoring loop upon completion of full-scale experiments. The system will continuously evaluate metrics against defined thresholds, generating updated reliability indicators in real time. If any metric goes above or below the defined threshold, the dashboard automatically updates the R/G/Y indicators and activates the corresponding mitigation action and review cycle.

As a result, reliability evaluation transitions from a static, point-in-time assessment into a continuous, governance-aligned monitoring process– consistent with the iterative lifecycle model defined by the NIST AI RMF.

V. CONCLUSION

This paper presented a conceptual reliability and security evaluation framework for application within a general AI security platform, integrating the NIST AI Risk Management Framework (AI RMF) with the MITRE ATT&CK knowledge base. To validate the operational feasibility of the proposed framework, a pilot empirical evaluation was conducted across three ATLAS-mapped adversarial scenarios, demonstrating consistent mitigation effectiveness and confirming that the conceptual design can be executed in practice. To achieve this operational structure, five interconnected Excel sheets - Risk Register, Governance Sheet, Mitigation, Attack Posture Sheet, and Root Cause Analysis Log were developed as the backbone of a practical evaluation pipeline. Together, these components demonstrate how governance policies, adversarial testing evidence, and reliability metrics can be integrated into a unified monitoring structure for an AI security platform.

5.1 Key Contributions

The proposed approach introduces several novel aspects. First, it demonstrates an integrated workflow that connects organizational accountability with technical threat intelligence, allowing AI reliability to be assessed through a unified structure. Second, the Excel-based implementation provides an interpretable and reproducible mechanism for capturing, tracking, and updating risk metrics in alignment with RMF guidance. Third, the use of color-coded (Red/Yellow/Green) risk indicators enables visual traceability across policy, performance, and mitigationlevels. Fourth, by embedding MITRE ATT&CK tactics within the testing and monitoring pipeline, the system allows adversarial risks to be quantified alongside governance outcomes, bridging a long-standing divide between AI security and AI reliability domains. Finally, the pilot empirical evaluation conducted across three ATLAS-mapped adversarial scenarios provides preliminary evidence that the proposed framework is operationally viable, achieving mitigation effectiveness rates of 75%, 100%, and 25% for prompt injection, retrieval poisoning, and adversarial paraphrasing respectively.

5.2 Limitations and Future Work

Although the proposed framework successfully connects NIST AI RMF principles with MITRE ATT&CK adversarial mappings, several limitations remain. First,while the pilot empirical evaluation demonstrates operational feasibility across three adversarial scenarios, larger-scale empirical experiments involving diverse AI workloads will be required to statistically validate the reliability thresholds and monitoring indicators defined in this framework. Second, the threat modeling component relied on a selected subset of ATT&CK / ATLAS techniques. As adversarial behaviors evolve rapidly, the framework may require frequent updates to remain aligned with new attack vectors such as multimodal jailbreaks, agentic exploitation, or system prompt manipulation. Third, some modules in an AI security platform, such as quantum machine learning defenses or blockchain-based components, may behave differently under high computational load or real-world latency conditions. These factors could influence the threshold values defined in the Risk Register and introduce complexity that is not yet captured in the current framework.

Lastly, the framework assumes that all observations and incidents can be recorded consistently in the Excel sheets. While this is effective during early prototyping, large-scale deployments will require migration toward an API-integrated architecture, enabling automated ingestion of system telemetry, version-controlled incident logging, and integration with enterprise SIEM and SOAR platforms to ensure long-term governance scalability.

References

  1. S. Mishra, A. Rao, R. Krishnan, B. Ayyub, A. Aria, and E. Zio, "Reliability, resilience and human factors engineering for trustworthy AI systems," arXiv:2411.08981, Nov. 2024.
  2. V. S. Narajala and O. Narayan, "Securing agentic AI: A comprehensive threat model and mitigation framework for generative AI agents," arXiv:2504.19956, Apr. 2025.
  3. National Institute of Standards and Technology (NIST), "Artificial intelligence risk management framework (AI RMF 1.0)," NIST AI 100-1, Jan. 2023.
  4. National Institute of Standards and Technology (NIST), "NIST AI risk management framework playbook," 2023. [Online]. Available: https://airmf.csrc.nist.gov/, accessed Jan. 2026.
  5. MITRE Corporation, "MITRE ATT&CK framework," 2023. [Online]. Available: https://attack.mitre.org/, accessed Jan.2026.
  6. MITRE Corporation, "MITRE ATLAS: Adversarial Threat Landscape for Artificial-Intelligence Systems," 2023 [Online]. Available: https://atlas.mitre.org/, accessed Jan. 2026.
  7. R. Dotan, B. Blili-Hamelin, R. Madhavan, J. Matthews, and J. Scarpino, "Evolving AI risk management: A maturity model based on the NIST AI RMF," arXiv:2401.15229, Jan. 2024.
  8. N. Swaminathan and D. Danks, "Application of the NIST AI risk management framework to surveillance technology," arXiv:2403.15646, Mar. 2024.
  9. R. Tete, "Threat modelling and risk analysis for LLM-powered applications," arXiv:2406.11007, Jun. 2024.
  10. J. Dev, N. B. Akhuseyinoglu, G. Kayas, B. Rashidi, and V. Garg, "Building guardrails in AI systems with threat modeling," Digital Government: Research and Practice, vol. 6, no. 1, pp. 1-18, Feb. 2025. https://doi.org/10.1145/3674845
  11. J. Ruohonen, E.-Y. Kang, and Q. Ramadan, "An alignment between the CRA's essential requirements and the ATT&CK's mitigations," in Proceedings of the IEEE 33rd International Requirements Engineering Conference Workshops (REW), Valencia, Spain, 2025, pp. 209-214, doi: 10.1109/REW61221.2025.00033.
  12. J. Schröder and J. Breier, "RMF: A risk measurement framework for machine learning models," in Proceedings of the 19th International Conference on Availability, Reliability and Security (ARES), Vienna, Austria, Jul. 30-Aug. 2, 2024, pp. 1-6.
  13. A. M. Barrett, D. Hendrycks, J. Newman, B. Nonnecke, and others, "Actionable guidance for high-consequence AI risk management: Toward standards addressing AI catastrophic risks," arXiv:2206.08966, Jun. 2022.
  14. A. E. Muhammad, K. C. Yow, J. Baili, Y. Cho, and Y. Nam, "CORTEX: Composite overlay for risk tiering and exposure in operational AI systems," arXiv:2508.19281, Aug. 2025.
  15. E. Papagiannidis, P. Mikalef, and K. Conboy, "Responsible artificial intelligence governance: A review and research framework," Journal of Strategic Information Systems, vol. 34, no. 2, Jun. 2025, Art. no. 101885, doi: 10.1016/j.jsis.2024.101885.
  16. A. Batool, D. Zowghi, and M. Bano, "AI governance: A systematic literature review," AI and Ethics, vol. 5, no. 3, pp. 3265-3279, 2025. https://doi.org/10.1007/s43681-024-00653-w