SG-logo-white
  • Product
    • The Platform

      Design a security program that builds trust, scales with your business, mitigates risk, and empowers your team to work efficiently.

      • Our technology
      • Built for AI
      • Why Strike Graph
      • All frameworks
    • Features
      • Atlas AI
      • Action Items (POA&M)
      • AI Security Assistant
      • Audits & certifications
      • Customizations
      • Dashboards & reporting
      • Integrations
      • Pen testing
      • Questionnaires
      • Risk management
      • SBOM Manager
      • Self Assessments
      • System Security Plan
      • Third-Party Risk Management
      • Trust Center
      • Verify AI
      • Vulnerability scanning
      • Workspaces
  • Solutions
    • Solutions
      For industries
      • Data Centers
      • Life Sciences
      • Manufacturing
      • Medical Devices
    • Frameworks
      • CCPA/CPRA
      • CMMC
      • DORA
      • GDPR
      • HIPAA
      • SOC 2
      • HIPAA
      • ISO 27001
      • All frameworks
      • HITRUST CSF
      • ISO 27001
      • ISO 27701
      • ISO 42001
      • NIST CSF
      • NIST 800-53
      • NIST 800-171
      • PCI DSS
      • SOC 1
      • SOC 2
      • TISAX
      • All frameworks
  • Pricing
  • Company
    • Strike Graph
      • About us
      • Careers
      • News
      • Partner
      • Press
    • FEATURED

      Cybersecurity is evolving — Strike Graph is leading the way.

      security-compilance
      February 9, 2023
      Security Compliance: Why It’s A Business Accelerator
    • Thought leadership
      It’s your technology and your security controls: Don’t let an auditor become your CTO
      Cybersecurity compliance that is unique to your organization
      Constant compliance is security theater
  • Resources
    • categories
      • Blog
      • Case studies
      • Guides
      • Secure Path events
      • Secure Talk podcast
      • Webinars
      • All resources
    • WHITE PAPER

      The future of compliance AI is already here

      smbr
      How small AI models outperform in compliance
      Download white paper
    • SEARCH

      Find answers to all your questions about security, compliance, and certification.

    • Sign In
    • Schedule a demo
    • Sign In
    • Schedule a demo

    Ready to see Strike Graph in action?

    Find out why Strike Graph is the right choice for your organization. What can you expect?

    • Brief conversation to discuss your compliance goals and how your team currently tracks security operations
    • Live demo of our platform, tailored to the way you work
    • All your questions answered to make sure you have all the information you need
    • No commitment whatsoever

    We look forward to helping you with your compliance needs!

    Fields marked with a star (*) are required

    Find out why Strike Graph is the right choice for your organization. What can you expect?

    • Brief conversation to discuss your compliance goals and how your team currently tracks security operations
    • Live demo of our platform, tailored to the way you work
    • All your questions answered to make sure you have all the information you need
    • No commitment whatsoever

    We look forward to helping you with your compliance needs!

    Article graphic

    The AI Governance Gap: Why ISO 42001 and NIST AI RMF Don't Prove Your AI Works

    How AI governance frameworks certify the organization's process, not the AI itself. 

    WHITE PAPER
    By: Justin Beals

    Executive summary

    AI is the first enterprise technology wave where the governance frameworks arrived with the technology, not years later. But those frameworks only govern how an organization manages AI, not what an AI system actually does, and that leaves a gap.

    The AI governance gap: fast frameworks, persistent risk

    NIST released its AI Risk Management Framework 57 days after ChatGPT launched. ISO/IEC 42001 followed within a year, and the EU AI Act took effect in 2024. By comparison, cloud computing waited four years for its first control framework, and open source waited twenty-three.

    According to IBM's 2026 Cost of a Data Breach Report, 68% of breached organizations lack an approved AI governance policy. Among large US public companies, where nearly all have a policy, EY found that 47% have bypassed it for urgent deployments and 36% have already suffered an AI incident with material impact.

    Ten enterprise AI risks, tested against four frameworks

    To understand why the gap persists, this white paper identifies the ten enterprise AI risks with the highest expected loss. It then tests each risk against the four frameworks organizations most often cite as AI assurance: ISO/IEC 42001, the NIST AI Risk Management Framework, the EU AI Act, and SOC 2. The risks are ranked by measured incidence and cost, using data from the Verizon 2026 Data Breach Investigations Report, IBM, EY, McKinsey, and the OWASP Top 10 for LLM Applications. The ten risks are:

    • ungoverned (shadow) AI adoption

    • excessive agent autonomy

    • prompt injection

    • assurance methods that cannot test non-deterministic systems

    • data and IP leakage

    • regulatory instability

    • model supply chain drift

    • deployer liability

    • failure to realize value

    • AI-enabled attacks

    Each risk is scored on a single question: could an auditor find an organization deficient under the framework?

    Why AI governance frameworks stop short of AI behavior

    The analysis shows that all four frameworks govern how an organization manages AI, and none of them tests what an AI system actually does. An organization can pass an ISO/IEC 42001 audit or a SOC 2 examination by showing that it defined and ran an evaluation, even if that evaluation found the model performing poorly.

    Four different institutions drew this same boundary for two reasons. First, reliable methods for measuring AI behavior do not exist yet. NIST acknowledges this inside the AI Risk Management Framework itself, which instructs organizations to document risks that cannot be measured. Second, standards bodies write within the formats they already own. ISO writes management system standards, which certify an organization's processes. The AICPA writes control attestations, which confirm that controls operated. Neither format is designed to certify how a system behaves.

    AI regulation retreated in 2026

    Instead of closing the gap, AI regulation in 2026 moved away from enforceable requirements. The EU delayed its high-risk AI rules to December 2027 because the technical standards companies need to prove compliance were not finished in time. Colorado repealed its AI Act before it took effect. US bank regulators removed generative and agentic AI from their model risk guidance.

    Two AI risks with no framework coverage

    Two of the ten risks have no framework coverage at all: shadow AI, meaning employee use of unapproved AI tools, and technical evaluation of the third-party models companies build on. Shadow AI is also the highest-loss risk in the paper. Overall, what organizations can be audited against today covers roughly half of what they should be doing.

    Four actions that close the gaps the frameworks leave open

    The actions that address the highest-loss risks are the ones the frameworks either don't require or leave the standard to the organization, so organizations have to take them on their own initiative. Organizations should discover how AI is actually being used before writing more policy. They should give every AI agent an owner with the authority to stop it, and set their own pass-or-fail thresholds for model evaluations. Finally, they should design AI systems so no single agent can access private data, read untrusted content, and communicate externally at the same time.

    For the industry, this means AI certifications should be read as evidence of good process, not proof of safe AI.

    The bottom line: AI certifications prove process, not performance

    An organization can hold an accredited ISO/IEC 42001 certificate and a clean SOC 2 report and still be unable to answer the first questions after an AI incident: what AI was running, who could have stopped it, and how accurate it was at the time. Today's frameworks prove an organization manages AI on paper. They don't prove the AI works.

    Why AI's governance gap is not a missing-rules problem

    AI is the first enterprise technology wave where governance frameworks arrived alongside the technology rather than years afterward. The historical record supports this claim by a wide margin.

    Tech Wave

    Commercial inflection

    First cross-industry control framework

    Lag

    Outsourcing / service providers

    1990s BPO and ITO boom

    SOC 2, via SSAE 16, effective Jun 2011

    ~15–20 years

    Open source in the enterprise

    OSI founded 1998

    NTIA SBOM Minimum Elements, Jul 2021

    ~23 years
    Commerical web /e-commerce

    ~1995

    PCI DSS v1.0, Dec 2004

    ~9 years

    Cloud / laaS

    AWS EC2, Aug 2006

    CSA Cloud Controls Matrix v1.0, 27 Apr 2010

    ~4 years

    Cloud, government

    AWS EC2, Aug 2006

    FedRAMP OMB memo, 8 Dec 2011

    ~5.5 years

    Generative AI

    ChatGPT, Nov 30, 2022

    NIST AI RMF 1.0, 26 Jan 2023; ISO/IEC 42001, 18 Dec 2023

    ~2-13 months

     

    NIST published AI RMF 1.0 on January 26, 2023, fifty-seven days after ChatGPT's release. ISO/IEC 42001 followed on December 18, 2023, and the EU AI Act entered into force on August 1, 2024. In the same period, the Open Worldwide Application Security Project (OWASP) published a Top 10 for LLM Applications, which it has since revised twice.

    Most technologies get governance frameworks only after years of failures force them. Open source is the most extreme example. It waited 23 years for its first framework, and repeated major security failures along the way did not speed it up.

    The warning signs appeared early, and the failures kept coming:

    • 1984: Unix co-creator Ken Thompson showed in his Turing Award lecture that a compiler could secretly insert a backdoor into every program it built.
    • 2014: Heartbleed, a flaw in the OpenSSL encryption library, exposed passwords and private keys on roughly half a million secure web servers. At the time, OpenSSL received roughly $2,000 a year in donations. The incident led to the Core Infrastructure Initiative.
    • 2017: Equifax was breached through an Apache Struts vulnerability (CVE-2017-5638), even though a patch had been available since March.
    • 2021: Log4Shell, a flaw in the widely used Log4j logging library, landed in December. In July 2022, the Cyber Safety Review Board called it an endemic vulnerability expected to persist for a decade.

    The same failure mode, an unfunded critical dependency with no inventory, occurred three times across seven years before a mandate appeared.

    AI did not repeat that pattern. By the time most enterprises ran their first production AI deployment, they already had a NIST framework, a certifiable ISO management system standard, an OWASP threat list, a European regulation in force, and every existing obligation from HIPAA to PCI DSS to SOC 2 still applying to the systems underneath.

    AI's governance gap cannot be blamed on missing rules. In 2006, the industry did not yet know what controls to write for cloud. In 2026, the controls for AI are written, published, and in some cases legally binding, and the gap is still open.

    What the gap looks like

    Nearly every large company now has an AI governance policy, but survey data shows that many bypass it, many have already suffered AI incidents, and even formal certification does not test whether AI works.

    The EY US AI Risk and Governance Survey, published in September 2026, surveyed 202 senior AI decision-makers at publicly traded US companies with $1 billion or more in annual revenue. Results found that 98% have formal AI governance policies. Yet 47% have bypassed those policies for urgent deployments, and 36% have experienced an AI incident that caused materially negative impact, such as data loss, financial damage, brand damage, or operational disruption. Furthermore, 91% use agentic AI in pilot or production, and 49% have not updated their governance framework to cover it.

    The EY survey also shows that governance works when organizations actually use it. Among organizations that run formal AI reviews, 64% modified at least a quarter of their AI systems as a result, 29% paused at least a quarter, and 25% stopped at least a quarter. The reviews work, yet most organizations are not running them.

    IBM's 2026 Cost of a Data Breach Report, which studied 602 breached organizations across 16 countries and regions, found a wider gap. 68% do not have an AI governance policy in place: 35% have no policy at all, and 33% are still developing one. The share of organizations with a policy fell from 37% to 32% in a year, and only 19% said their AI governance and security teams work together.

    Certification does not close this gap. Roughly 350 organizations had publicly announced ISO/IEC 42001 certification as of April 2026, but an ISO/IEC 42001 certificate confirms only that an organization follows a documented AI management process. Compare FedRAMP, the US government's cloud security program. FedRAMP assessments include technical testing of the cloud service itself, which is why federal agencies have relied on existing FedRAMP approvals more than 3,000 times instead of running their own reviews. ISO/IEC 42001 includes no equivalent test of the AI system. So what does this mean in practice?

    AI governance frameworks govern the organization, not the AI

    AI governance frameworks arrived early, but they govern only one thing: the organization. Every framework published between January 2023 and August 2024 describes what a company must do around an AI system. None of them describes what the AI system itself must do. This distinction is the foundation of the rest of this paper, so an example helps make it concrete.

    Consider a mid-size insurer deploying an AI model to triage insurance claims. To meet ISO/IEC 42001, the insurer:

    • documents an AI policy (A.2.2)
    • assigns roles (A.3.2)
    • documents its data sources and their known bias (A.4.3)
    • assesses the impact on affected individuals (A.5.4)
    • defines verification and validation (V&V) measures and criteria (A.6.2.4)
    • records data provenance (A.7.5)
    • monitors the model in operation (A.6.2.6)
    • assigns responsibilities between itself and its model vendor (A.10.2)

    An accredited certification body reviews samples of that evidence in a two-stage audit. It then issues a certificate valid for three years and returns for surveillance audits in years one and two.

    Now suppose the insurer's quarterly test shows that the model wrongly denies 40% of valid claims from one group of policyholders. The auditor asks three questions: Is a V&V criterion defined? Was it applied on schedule? Were the results retained? The answer to all three is yes. The auditor records conformity, raises no finding, and the certificate is issued.

    The certificate is issued because control A.6.2.4 requires the insurer to define and apply V&V criteria, but does not say what those criteria must be. The auditor has no basis to call them too weak. The same is true of A.7.4 on data quality and A.6.2.6 on monitoring, which also leave the criteria to the organization being audited. No control in ISO/IEC 42001 is one a model can objectively fail, because the model is never the thing being tested. The insurer's process is.

    The insurer's certificate is real. It confirms that the insurer runs an AI management system that conforms to ISO/IEC 42001. But the insurer wrote the certificate's scope itself, under clause 4.3. It also chose which controls to apply, in a document called the Statement of Applicability, under clause 6.1.3. Both documents are usually kept confidential between the insurer and its certification body.

    As a result, a buyer holding the certificate number usually cannot tell whether the claims triage model was even covered. The exception is a certificate issued under an accredited ISO/IEC 42006 scheme, which now requires the certificate to name the AI systems in scope.

    The certificate describes the wall, not what's inside it

    Every CISO already knows this failure mode under a different name. Perimeter security built a hard shell around a soft center: a firewall, a DMZ, and controlled entry points, protecting an interior that nobody inspected because the wall was supposed to make inspection unnecessary. The approach worked until an attacker got inside. Then the lack of visibility into the interior became the whole loss, which is what drove the industry's move to zero trust.

    ISO/IEC 42001, the NIST AI RMF, and SOC 2 build the same kind of perimeter around AI. Policies, roles, documented data sources, defined V&V procedures, vendor due diligence, logging, and management review form a shell of process around the model. That shell is real, and it does useful work. But nothing in any of the three frameworks inspects what is inside it.

    The insurer's claims model sits at the center of a well-built perimeter, and its 40% error rate for one group of policyholders is a fact about the interior. The certificate only describes the wall.

    Scope and exclusions: two choices most certificates don't show

    An organization makes two choices that shape what its ISO/IEC 42001 certificate covers. Clause 4.3 lets it declare the scope of its management system. Clause 6.1.3 lets it exclude any control in the standard's Annex A, as long as it justifies the exclusion with its own risk assessment, which it also wrote. An organization running twelve AI systems can therefore certify the management system for just two of them. Unless the certificate was issued under an accredited ISO/IEC 42006 scheme, it looks identical either way.

    This is not a drafting error. It is how every ISO management system standard built on Annex SL works. Annex SL is the common structure ISO uses for all its management system standards, and this flexibility is why ISO/IEC 27001 certificates have always been read alongside their scope statements. The difference is how buyers read them. An AI standard reads to a buyer as a statement about AI safety, while an ISO/IEC 27001 certificate never read as a statement about whether the software was correct.

    The NIST AI RMF was not built to be certified against

    The clearest evidence comes from the NIST AI Risk Management Framework itself. A full-text search of the framework document (NIST AI 100-1), across all 72 subcategories, finds zero uses of "conformity," "accredit," or "attest." The word "certification" appears once, in reference to credentials held by operators and practitioners, and "audit" appears once, in passing, in a discussion of explainability.

    A framework that never uses the language of conformity, accreditation, or attestation was not designed to be certified against, and NIST has said so consistently. The framework states that its actions "do not constitute a checklist, nor are they necessarily an ordered set of steps." The companion Playbook calls itself "neither a checklist nor set of steps to be followed in its entirety." NIST's own AI RMF Roadmap lists ten priorities for the framework's future, and none of them involves conformity assessment, certification, or accreditation.

    As a result, a claim of "NIST AI RMF alignment" has no defined criteria, no accredited assessors, no assessment procedure, no standard way to report findings, and no authority that issues it. When such a claim appears alongside an ISO/IEC 42001 certificate or a SOC 2 report, only those two documents contain anything an auditor actually tested.

    Taken together, the three frameworks give an organization a well-documented process around its AI, and no evidence about the AI itself. That raises a practical question: which enterprise AI risks does a process-only approach actually leave exposed?

    The ten enterprise AI risks with the highest expected loss

    Below are the ten enterprise AI risks that cost organizations the most. The risks are ranked by expected loss to the organization deploying the AI, based on how often each risk occurs and what it costs when it does.

    1. Ungoverned (shadow) AI adoption. Shadow AI is employee use of AI tools the organization has not approved or cannot see. The Verizon 2026 Data Breach Investigations Report found that employee AI use on corporate devices rose from 15% to 45% in twelve months, and 67% of that use runs through personal rather than corporate accounts. Source code was the most common type of data employees uploaded to AI tools, followed by images and structured data. Shadow AI is now the third most common non-malicious insider action, a fourfold increase year over year, and more than 15% of users have unauthorized AI browser extensions installed, many of which quietly collect whatever is on the page. IBM's 2026 report found that shadow AI was involved in 43% of breaches, up from 20% the year before, and those breaches averaged $5.39 million, up from $4.63 million. Organizations cannot govern AI they cannot see, which is why discovery comes first in 'What to prioritize'.

    2. Excessive agent autonomy. Excessive agency is the risk that AI agents take actions no human has reviewed or approved. The EY US AI Risk and Governance Survey found that among organizations using agentic AI, 85% have at least some AI systems taking actions without real-time human involvement, and 26% cannot detect unauthorized AI agents operating inside their environment. A separate August 2026 report from Enterprise Management Associates found that AI agents had acted outside their intended scope at 65% of surveyed enterprises, and 29% reported measurable organizational impact. Open Worldwide Application Security Project (OWASP), the nonprofit that tracks top security risks for AI applications, moved excessive agency from sixth to third in its 2026 Top 10 for LLM Applications. Gartner predicts that 40% of enterprises will scale back or shut down autonomous agents by 2027 because of governance gaps. Every autonomous agent needs an owner with the authority to stop it.

    3. Prompt injection. Prompt injection is an attack that hides malicious instructions inside content an AI system reads, such as an email or web page, so the AI follows the attacker's instructions instead of the user's. It has ranked first on the OWASP list for two consecutive editions. Security researcher Simon Willison, who named the attack in 2022, identified the core problem: an AI model reads instructions and data the same way, so it cannot reliably tell them apart. Willison also identified what he calls the "lethal trifecta," the combination that makes data theft possible: an AI system that holds private data, reads untrusted content, and can communicate externally. EchoLeak (CVE-2025-32711) proved the risk in a real product, stealing data from Microsoft 365 Copilot without the user clicking anything. Because prompt injection cannot be fully prevented, organizations should limit the damage through system design.

    4. Assurance methods that cannot test AI. Traditional audits cannot reliably verify AI systems, because AI does not behave the same way every time. Every control audit since the 1990s rests on the same logic: a control is designed, it operates consistently, an auditor tests a sample, and the sample represents the whole. AI breaks the last step, because the same input can produce different outputs. Research by Casper et al. (FAccT 2024) shows that testing an AI system from the outside cannot cover its range of possible behavior, and that when it fails, testers cannot tell which component caused the failure. Every commercial AI assurance product sold today tests AI this way, from the outside. Since no framework defines what passing looks like, organizations have to set their own performance thresholds.

    5. Data and IP leakage, and training-data provenance. This risk covers two problems: sensitive data leaking into AI tools, and not knowing what data an AI model was trained on. On the leakage side, source code was the most common data type employees uploaded to AI tools in Verizon's data, and Samsung's 2023 source-code leak through ChatGPT led the company to ban generative AI tools company-wide. Provenance is the harder problem. Open source took twenty-three years to develop a standard inventory of software components, and no equivalent yet exists for AI training data. The cost of that gap became clear when Bartz v. Anthropic settled for $1.5 billion in September 2025, in part because the training data's origins could be traced in ways that hurt the defendant but not in ways that helped it. Provenance is the one area where the frameworks set real requirements, so organizations should meet them in full.

    6. Regulatory instability. Regulatory instability is the risk that the rules an organization is preparing for change before they take effect. AI compliance requirements have shifted three times in eighteen months, each time toward less obligation. The section "AI regulation retreated in 2026" covers this in detail. Organizations that build governance on today's deadlines alone risk building for rules that change, so they should manage the risks directly rather than waiting for regulation to require it.

    7. Model supply chain drift. Model supply chain drift is the risk that the third-party AI models a company depends on change without warning. AI providers retire model versions every few weeks or months, which invalidates any testing a company did on the previous version. A vendor's risk profile can also change between annual reviews, and third-party risk programs built on yearly questionnaires will not catch it. The risk is concentrated: a December 2025 study by the Cloud Security Alliance and Google Cloud found that 70% of enterprises use OpenAI's GPT models, 48% use Google Gemini, 29% use Anthropic's Claude, and 20% use Meta's Llama. OWASP ranks supply chain risk fourth on its list. Organizations should retest vendor models after every version change and ask vendors exactly what their certifications cover.

    8. Deployer liability. Deployer liability is the legal principle that a company is responsible for what its AI tells customers, even when the AI is wrong. In Moffatt v. Air Canada (February 2024), a tribunal held the airline liable for its chatbot's false statement about refund policy. In May 2026, two German courts extended the pattern. The Higher Regional Court of Hamm held a clinic liable for false claims its website chatbot made about its physicians' qualifications, and the Munich Regional Court held Google directly liable for incorrect AI Overviews, treating them as Google's own statements rather than third-party content. None of these rulings required new law. They apply the same principle that made Target responsible for the breach that came through its HVAC contractor in 2013. Companies are responsible for what their AI tells customers, so they should know when they've taken on provider obligations and keep the records needed to defend their decisions.

    9. Failure to realize value. Value realization failure is the risk that AI investments never deliver measurable business results. McKinsey's 2025 State of AI survey found that 88% of organizations use AI in at least one function and 62% are experimenting with or scaling AI agents. Yet 10% or fewer have scaled agents in any single function, and only 39% report any effect on earnings (EBIT), most of it under 5%. Roughly 6% qualify as high performers. A Federal Reserve analysis of US Census Bureau data puts firm-level AI adoption at only about 18% at the end of 2025. Organizations certified to ISO/IEC 42001 already have a formal review process for stopping AI projects that aren't delivering.

    10. AI-enabled attacks. This risk covers attackers using AI to breach organizations faster and at greater scale. IBM's 2026 Cost of a Data Breach Report found that one in four malicious breaches was AI-enabled, a 56% increase year over year. Those breaches cost an average of $6 million, compared with a $4.99 million global average for all breaches. More than 20% of organizations had breaches that targeted their AI models or applications directly. In November 2025, Anthropic reported an AI-driven espionage campaign against roughly 30 targets, in which the AI carried out 80–90% of the work and humans stepped in at only four to six decision points. The International AI Safety Report 2026, led by Yoshua Bengio with more than 100 experts from over 30 countries, reached a more measured conclusion: AI is not yet running cyberattacks fully on its own, but it is speeding up the preparation stages. The gap between one company's incident report and a multinational expert assessment is itself worth noting. Organizations should assume attackers are using AI and design systems so a compromised agent cannot reach sensitive data.

    These ten risks fall into two groups. Some are failures of process, such as data provenance and failure to realize value, which governance frameworks were designed to address. Others are failures of system behavior, such as prompt injection and model drift, which no process control can see. Underneath both groups sits a shared problem: most organizations lack the people to manage either kind.

    Most organizations lack the expertise to run their AI policies

    The EY US AI Risk and Governance Survey found that 69% of respondents are concerned their organization lacks the internal expertise to evolve its AI governance controls, and 63% are concerned it lacks the expertise to implement or design them. The policy exists, but the capability to operate it often does not.

    How the four frameworks cover the ten risks

    The table below tests each of the ten risks against ISO/IEC 42001, the NIST AI RMF, the EU AI Act, and SOC 2. Each framework is scored on one question: could an auditor find an organization deficient for failing to manage this risk? The final column gives the overall verdict across all four frameworks.

    Risk

    ISO/IEC 42001

    NIST AI RMF

    EU AI Act

    SOC 2

    Verdict

    1. Ungoverned adoption

    A.9.2, A.9.4, A.4.2: in declared scope only

    GOVERN 1.x, MAP 1.x

    Art 4 AI literacy, softened by omnibus to "support development of"

    Out of scope by definition

    None

    2. Excessive agency

    A.9.4, A.6.2.6

    MANAGE 4.x: no agent content, drafted Jan 2023

    Art 14 human oversight: strong text, applies 2 Dec 2027

    CC6 endpoint access

    Thin, and not yet in force

    3. Prompt injection

    Named in no control

    AI 600-1; Cyber AI Profile still preliminary draft

    Art 15(5) names the attack classes

    CC7 monitoring

    EU only, 2 Dec 2027

    4. Assurance / non-determinism

    A.6.2.4 defines V&V, prescribes nothing

    MEASURE 1.1 permits "cannot be measured"

    Art 15(3): declare metrics of your choosing

    Sampling tests controls, not inferences

    Delegated by all four

    5. Data leakage and provenance

    A.7.3, A.7.4, A.7.5 provenance

    MAP 2.x, MEASURE 2.x

    Art 10 data governance; Art 53 training-data summary

    CC6, C-series

    Best covered

    6. Regulatory instability

    Not addressed

    Not addressed

    Not addressed

    Not addressed

    Meta-risk

    7. Model supply chain

    A.10.2, A.10.3: relationship only

    GOVERN 6.x

    Art 25 role reversal; Art 53 downstream docs

    CC9 + carve-out, chain terminates

    Relationship yes, technical no

    8. Deployer liability

    A.10.2

    Not addressed

    Art 26, Art 27 FRIA

    Not addressed

    EU only

    9. Value realization

    Cl. 6.2, 9.1, 9.3 measurable objectives and management review

    Not addressed

    Not addressed

    Not addressed

    Covered and unused

    10. Adversarial asymmetry

    A.6.2.6

    Cyber AI Profile, draft

    Art 15(5)

    CC7, CC4

    Partial

     

    Seven of the ten risks have partial coverage as process. Regulatory instability sits outside any framework's reach, and two risks have no coverage of the risk itself: ungoverned adoption and technical evaluation of third-party models.

    Where the frameworks deserve credit

    Three parts of the table show the frameworks doing real work.

    • Data provenance has binding requirements with real content. ISO/IEC 42001 control A.7.5 requires organizations to record the origin of their data and maintain that record for as long as the data is kept. Control A.7.3 governs how data is acquired, including rights to personal data and copyrighted material. The EU AI Act's Article 10 goes further, requiring that training, validation, and testing datasets be examined for bias and be relevant, representative, and as free of errors as possible. Provenance is the one risk in the table where an organization can point to a binding requirement with specific content.
    • SOC 2 does more than skeptics grant. SOC 2 genuinely tests access controls over AI model endpoints and training data (CC6), change management over prompts, models, and model weights (CC8), vendor due diligence (CC9), logging and monitoring (CC4 and CC7), and incident response. A well-scoped SOC 2 covering AI infrastructure is worth having. Its coverage simply stops where AI-specific risk starts.
    • ISO/IEC 42001 includes a way to stop failing AI projects. Clauses 6.2, 9.1, and 9.3 require measurable AI objectives, defined monitoring, and a management review that must consider a required list of inputs. Together, these give leadership a formal mechanism to shut down an AI project that isn't working, written into a certifiable standard. Almost no organization uses it that way. Failure to realize value is the only risk in the table where the tool exists and the practice does not.

    Where all four frameworks fail together: third-party AI models

    Third-party AI models expose the biggest shared weakness in the four frameworks. None of them requires a company to test the AI models it buys from vendors.

    Consider a company that builds its product on a third-party AI model, accessed through the vendor's API. Under SOC 2, the model provider counts as a subservice organization, meaning an outside vendor that performs part of the service. Companies can either include that vendor in their SOC 2 examination or "carve it out," and with AI model providers, companies carve them out in essentially every case, because the providers do not participate in customers' audits.

    When the vendor is carved out, the SOC 2 report names the vendor, describes its service in a line, and lists the controls the company assumes the vendor performs. The auditor does not test those controls. Instead, the auditor checks that the company reviewed the vendor's own SOC 2 report once a year, by looking for evidence that the review happened.

    The vendor's own SOC 2 report, in turn, covers its hosting platform, not the model itself. The AI Provider Trust Registry, for example, rates Meta's Llama as only partially covered when accessed through AWS Bedrock or Azure AI, because the certification applies to the cloud platform rather than the model. And nearly all such vendor reports sit behind a nondisclosure agreement or a trust portal, so the company's own customers usually cannot read them.

    ISO/IEC 42001 handles the same dependency through controls A.10.2 and A.10.3 and clause 8.1 on outsourced processes. All three govern the vendor relationship: due diligence, contract terms, and ongoing oversight. None requires anyone to technically evaluate the model.

    For a company whose entire AI risk sits in a third-party model, that is the decisive gap. A company can hold a clean ISO/IEC 42001 certificate and a clean SOC 2 Type II report without ever having run a single test against the model it ships.

    Why all four frameworks stop at the same boundary

    Four institutions with different mandates, drafting traditions, and legal force arrived at the same boundary: none of their frameworks tests how AI systems behave. Two causes explain why.

    The measurement basis does not exist, and NIST says so

    AI RMF 1.0 section 1.2.1 states that "AI risks or failures that are not well-defined or adequately understood are difficult to measure quantitatively or qualitatively," and that "the current lack of consensus on robust and verifiable measurement methods for risk and trustworthiness" is itself an AI risk measurement challenge. MEASURE 1.1 instructs organizations to document risks that will not or cannot be measured as such.

    NIST is saying, inside the framework, that the measurement basis for its own MEASURE function is unavailable. A requirement cannot be written to be assessable against a measurement that does not exist. So the frameworks arrived early because they governed process, and process governance is the part that can be written without waiting for the science.

    The EU AI Act shows the same sequencing under legal pressure. Article 15(2) makes it the Commission's duty to encourage the development of benchmarks and measurement methodologies. That duty has not been discharged. Article 15(1) therefore requires an "appropriate level" of accuracy, robustness and cybersecurity, and the only fixed obligation in the article is 15(3): declare the accuracy metrics in the instructions for use. The provider picks the metric and picks the level.

    Standards bodies write from the templates they own

    ISO/IEC 42001 is an Annex SL management system standard. Its clause structure is the harmonized structure shared with ISO 9001 and ISO/IEC 27001, which is why an auditor trained on 27001 tests it the same way and why certification under ISO/IEC 17021-1 takes the management system as its object. A management system standard cannot certify a product; that is a property of the instrument, decided before anyone wrote a word about AI.

    SOC 2 is an attestation under SSAE 18 against Trust Services Criteria issued in 2017. The 2022 update revised only the points of focus, which are non-mandatory illustrative considerations, so it added no testable requirement. There is no AI criterion, no AI point of focus, and no AICPA AI attestation standard as of September 2026. The architecture is control-existence-and-operation, and it has no reporting vehicle for a substantive finding about a system's properties. A control can say management performs quarterly model evaluations, and a Type II can test that the evaluation occurred and was reviewed. The results are not part of the opinion.

    Processing Integrity is where this is most easily misread. PI tests conformance to specification, and management writes the specification. If the specification says the model produces a summary, PI can test that a summary was produced, delivered completely to the right party, and stored correctly. PI does not test whether the summary is true. As LBMC, a top 35 U.S. accounting and business consulting firm that performs SOC 2 examinations, puts it, "third-party auditors cannot provide reasonable assurance on a Generative AI's output, but can provide reasonable assurance on the controls around the development, maintenance, and use of AI by the organization.”

    The gap between what "processing integrity" means in plain English and what PI1.x requires is the most exploitable ambiguity in SOC 2 for AI marketing.

    AI certifications are repeating the SAS 70 mistake

    SAS 70 was issued in 1992 as a financial-controls attestation letting a user entity's financial statement auditor rely on controls at a service organization relevant to internal control over financial reporting. Through the 2000s it was marketed and consumed as a data center security certification. The AICPA's response was SSAE 16, effective for periods ending on or after 15 June 2011, which split SOC 1, SOC 2 and SOC 3 to correct exactly that.

    The pattern is that a narrowly scoped attestation gets consumed as a broad-spectrum trust signal, because the badge is legible and the report is not, because buyers lack the expertise to read scope, and because no competing signal exists.

    It is happening again, with one qualification worth stating. SAS 70's misuse was a category error. SOC 2's AI problem is an over-extension: SOC 2 genuinely is a security report, and the security of AI systems genuinely is within its reach. What falls outside is model behavior, output quality, bias, provenance and evaluation. The badge does not communicate where that boundary falls.

    The parallel for ISO/IEC 42001 is sharper, because it is an AI standard and therefore reads as an AI safety certification while certifying a management system.

    The standard-setters have not claimed that SOC 2 or ISO/IEC 42001 assures model behavior. The AICPA's own guidance points to NIST, ISO/IEC 42001 and the EU AI Act for AI substance, and Schellman and LBMC both disclaim output assurance in writing. Comp AI, a compliance-automation vendor, states outright that SOC 2 includes no controls unique to artificial intelligence or machine learning systems. The overclaim sits downstream, in procurement practice and in the consultancies answering it.

    And the profession is already worried about SOC 2 credibility for an adjacent reason. In May 2026 the AICPA issued peer review guidance targeting firms whose SOC 2 engagements "are not designed to respond to the unique risks associated with the service organization, resulting in engagements that have identical reports, risk assessments, sample sizes, and testing procedures." Carl Mayes, AICPA Vice President for Ethics and Firm Quality, said some firms are leaning too heavily on third-party SOC platforms without applying the professional judgment the standards require. Peer reviewers now examine roughly five engagements per firm for templating.

    If templated engagements are already a documented quality problem for conventional systems, a templated engagement is unlikely to meet a novel AI risk surface.

    AI regulation retreated in 2026, in the EU, US states, and US banking

    Three regulatory changes in 2026 moved away from enforceable AI requirements, and the certification system has its own weakness.

    The EU deferred because the standards did not arrive

    Regulation (EU) 2026/1744, the Digital Omnibus on AI, reached provisional political agreement 6 May 2026, was confirmed by Member State representatives 13 May, signed 8 July, published in the Official Journal 24 July, and entered into force 27 July 2026. It moved standalone Annex III high-risk obligations from 2 August 2026 to 2 December 2027, and AI embedded in regulated products under Annex I from 2 August 2027 to 2 August 2028.

    The stated reason was standards unavailability, and the record supports it. The standardisation request C(2023)3215 went to CEN-CENELEC JTC 21 in May 2023 with a delivery deadline of 30 April 2025, which was missed. EN 18286, on quality management systems for AI Act regulatory purposes, published 31 July 2026 as the first European standard for the Act. It has not been cited in the Official Journal, so it confers no Article 40 presumption of conformity. Eight further deliverables remain in draft, including prEN 18229-2 on accuracy and robustness and prEN 18282 on cybersecurity, which are the two that would give Article 15 content.

    As of September 2026, zero harmonised standards have been cited in the Official Journal and no common specifications have been adopted under Article 41.

    Most high-risk AI self-assesses anyway

    This is the part most commentary gets wrong, and it matters for anyone budgeting against the AI Act.

    Article 43(2) provides that for Annex III points 2 through 8, the provider follows the conformity assessment procedure based on internal control in Annex VI, "which does not provide for the involvement of a notified body." That is seven of the eight Annex III categories, including critical infrastructure, education, employment and worker management, access to essential public and private services including credit scoring and insurance pricing, law enforcement, migration and border control, and administration of justice. It covers the overwhelming majority of commercially deployed high-risk use cases.

    Annex III point 1, biometrics, is also eligible for Annex VI self-assessment under Article 43(1), on condition that the provider has applied harmonised standards. Because zero have been cited, that option is unavailable in practice and Annex VII notified-body assessment applies by default. So third-party assessment for biometrics is right today, for a contingent reason that changes the moment EN standards are cited.

    Product-embedded high-risk AI under Annex I follows the existing sectoral notified-body route under Article 43(3).

    The notified-body ecosystem is substantially unbuilt. DEKRA announced on 10 March 2026 that it became the first body accredited for AI biometric systems under the Act, and that announcement is explicit that this is accreditation by the Dutch Accreditation Council, a precondition for notification under Articles 29 to 31 and not a designation. The omnibus gave sectoral notified bodies until 28 January 2028 to apply for AI Act designation, and permitted Annex I Section A bodies to assess AI Act conformity for 18 months from 27 July 2026 without it. Neither accommodation would be necessary if the pipeline were populated.

    Colorado repealed and replaced

    The Colorado AI Act was signed 17 May 2024 with an effective date of 1 February 2026, delayed once to 30 June 2026. SB 26-189, signed 14 May 2026, repealed it and substituted the Automated Decision-Making Technology Act, effective 1 January 2027 contingent on the attorney general completing rulemaking.

    The replacement removed the duty of care to mitigate algorithmic discrimination risks, the annual impact assessments, the risk management programs, and every mention of discrimination. It kept developer documentation duties, deployer notice before use, disclosure of adverse outcomes within 30 days, three-year record retention, and a consumer right to meaningful human review as commercially reasonable.

    A December 2025 executive order directed the attorney general to establish an AI Litigation Task Force to identify and challenge state AI laws, and directed Commerce to explore withholding broadband funding from states enacting laws deemed onerous. States still enacted 109 AI laws in the first half of 2026, against 121 by the same date in 2025.

    Banking carved AI out of the one regime that had teeth

    On 17 April 2026 the Federal Reserve, OCC and FDIC issued SR 26-2 and OCC Bulletin 2026-13, replacing SR 11-7 on model risk management. SR 11-7, issued 2011, is the most substantive model-governance regime in US regulation: independent validation, effective challenge, ongoing monitoring, with real supervisory weight behind it.

    The revision narrowed the definition of a model to require complexity and to exclude simple arithmetic calculations and deterministic rule-based processes. It raised the asset threshold to $30 billion, replacing the FDIC's prior $1 billion. It removed prescriptive requirements such as annual validation. It states that non-compliance alone with the guidance will not result in supervisory criticism. And it excludes generative and agentic AI models from scope as novel and rapidly evolving, recommending banks apply broader risk management practices and signalling a future request for information.

    The best-developed model validation regime in American financial regulation was rewritten in 2026 and the rewrite took AI out.

    AI certifiers operated without a standard for nineteen months

    The certification system has not filled the gap regulators left behind. ISO/IEC 42001 was published on December 18, 2023. ISO/IEC 42006, which sets the requirements for the bodies that audit and certify against ISO/IEC 42001, was not published until July 7, 2025.

    For those nineteen months, organizations bought and received ISO/IEC 42001 certificates with no standard defining what a competent certifier was. The major accreditation bodies, ANAB in the US and UKAS in the UK, now accredit certifiers against ISO/IEC 42006. The new standard added auditor competence requirements, a formula for how long audits must take, and a requirement that each certificate state which products, business units, locations, AI systems, and organizational roles it covers.

    Even so, an unaccredited ISO/IEC 42001 certificate is still legal to issue, and a buyer cannot tell it apart from an accredited one.

    What to prioritize

    Priorities are ranked by how severe each risk is and how little the frameworks cover it. They fall into two groups: the first four actions address the highest-loss risks, where the frameworks say nothing or leave the standard to the organization, and the subsequent five actions put existing framework requirements to real use. A final principle applies to both.

    Act first: close the gaps the frameworks leave open

    These four priorities address the risks that cost organizations the most. No framework requires them, or the frameworks require only a process and leave the standard to the organization, so organizations have to act on their own initiative.

    1. Discover AI use before writing policy

    Organizations should find out how employees actually use AI before writing more AI policy. Shadow AI is the most common AI risk, and no framework covers it, because every framework applies only within the scope an organization declares, and shadow AI sits outside that scope by definition.

    Discovery is a technical capability, not a policy exercise. It means monitoring outbound network traffic to AI services, keeping an inventory of browser extensions, tuning data loss prevention (DLP) rules to flag uploads to AI tools, and tracking which user accounts are calling which AI models. An organization that writes an AI policy before it can see its AI usage has governed only the AI it already knew about.

    The cloud era offers a warning about how large that blind spot can be. In 2015, Cisco found that IT departments estimated about 51 cloud services in use, when the actual number was roughly 730, later revised to 1,220. Organizations should expect a similar gap with AI.

    2. Match agent autonomy to real authority

    Organizations should classify AI agents by how much independence they have, and give each level an owner with the authority to stop it. Gartner's four-level model is a useful starting point:

    • Level 1, observe: the agent can only read information.
    • Level 2, advise: the agent makes recommendations, which require quality testing.
    • Level 3, act with approval: the agent takes actions a human approves, which require audit trails.
    • Level 4, act autonomously: the agent acts on its own, which requires monitoring and guardrails.

    Applying the same controls to all four levels over-restricts simple agents, which pushes employees toward workarounds, and under-restricts autonomous ones.

    Classification alone is not enough. A study by Simons and Broniatowski tested how different roles would apply the NIST AI RMF and found that people understood their responsibilities well, but only executives had the authority to act on them. Middle managers and system owners rarely did. The study used AI-simulated personas rather than real organizations, but its lesson holds: for each autonomy level, name who receives the warning signs and what they are empowered to do about them. A level that nobody can act on is just a label.

    3. Set your own evaluation thresholds

    Organizations should define their own pass-or-fail standards for AI model performance, because no framework does. ISO/IEC 42001 control A.6.2.4 requires organizations to define and apply evaluation criteria. The EU AI Act's Article 15(3) requires providers to declare accuracy metrics. The NIST AI RMF refers to performance criteria. None of them says what the criteria must be.

    Organizations should treat that silence as an instruction. Choose the metric, set the threshold, test against it before deployment and after every model update, and keep the results as an audit record. An organization that does this and one that only writes a procedure are equally compliant with ISO/IEC 42001. Only the first will know when its model starts to fail.

    The EU AI Act offers a ready template, even before its requirements take effect: document the level of accuracy, the metrics used, and the conditions under which the system was tested, and include that documentation with the system.

    4. Design AI systems so no single agent can leak data

    Organizations should design AI systems so that no single agent combines private data, untrusted content, and the ability to communicate externally. Prompt injection has mitigations but no fix. Earlier security flaws had technical solutions: SQL injection had parameterized queries, and buffer overflows had bounds checking and memory-safe languages. Prompt injection has no equivalent, because AI models read instructions and data the same way.

    The practical defense is Simon Willison's: break up the "lethal trifecta." Any system that holds private data, reads untrusted content, and can communicate externally can be manipulated into leaking data, so remove one of those three capabilities from every system. An agent that reads untrusted email and has access to customer records should not be able to send data out. An agent that can send data out should not have access to customer records.

    This is the highest-leverage recommendation in this paper, because architecture has solved this kind of problem before when policy could not. Personal devices at work were secured not by policies or training, but by platform features that separated work data from personal data, as described in "What would close it." A guardrail that claims to block 95% of attacks is not a security control.

    Use what the frameworks already require

    These five priorities rely on requirements that already exist in ISO/IEC 42001, SOC 2, or the EU AI Act. An auditor can check each one today, but most organizations treat them as paperwork rather than as working controls.

    5. Ask for the scope statement, not just the certificate

    Organizations buying AI products should ask vendors what their certifications actually cover. A certificate number on its own proves very little. Three documents make it meaningful:

    • the certificate number
    • the name of the certification body, its accreditation body, and the date it was accredited specifically for ISO/IEC 42001, not for other standards
    • the Statement of Applicability, which lists the controls the vendor applied, or at minimum the list of controls it excluded

    For SOC 2 reports, buyers should ask which trust services categories were in scope, read the system description to see where its boundary sits, and confirm whether the AI component falls inside it. Processing Integrity, the SOC 2 category most relevant to AI outputs, is optional and usually left out.

    This should be a standard procurement question. It costs nothing, and it is the only way to tell a certificate that covers the product from one that covers a single minor feature.

    6. Document data provenance, because the frameworks do cover it

    Organizations should record where their AI training and operating data comes from, since this is the one area where the frameworks set real requirements. ISO/IEC 42001 controls A.7.3 and A.7.5 and the EU AI Act's Article 10 give specific content. Organizations should record each dataset's origin and collection method, how it was prepared, the rights to use it (including personal data and copyright), and any known bias, and keep those records for as long as the data is kept. Provenance is the one risk where an organization that does the work can point to a binding requirement and show that it met it.

    7. Check whether you've become an AI provider under the EU AI Act

    Organizations that modify or rebrand AI models should check whether the EU AI Act now treats them as providers rather than deployers. Under Article 25, a company that deploys an AI system becomes its provider, with the full set of provider obligations, if it puts its own name or trademark on a high-risk system, substantially modifies one, or changes a system's intended purpose so that it becomes high-risk.

    Fine-tuning, rebranding, and repurposing vendor models trigger this routinely. A company that fine-tunes a vendor model for hiring decisions may believe it is only a deployer, when it is legally a provider. Provider obligations include a formal conformity assessment, the CE marking that certifies compliance, and registration in the EU database.

    8. Use the authority to stop AI projects you already have

    Organizations certified to ISO/IEC 42001 should use the standard's management review to shut down AI projects that aren't working. Clause 6.2 requires measurable AI objectives, with plans stating what will be evaluated, by whom, when, and how. Clause 9.1 requires defined monitoring. Clause 9.3 requires management review against a required list of inputs. Clause 8.2 requires risk assessments to be performed at planned intervals and after significant changes, with records kept.

    Together, these create a standing review with the authority to stop a project, built into a certifiable standard. The EY figures in "The AI governance gap in numbers" show what happens when organizations actually use reviews like this: many modify, pause, or stop AI systems as a result. Most organizations have this tool and run it as a paperwork exercise.

    9. Keep AI activity logs long enough to investigate incidents

    Organizations should retain logs of AI system activity for at least six months, whether or not the EU AI Act applies to them. Article 12 of the EU AI Act requires AI systems to log events automatically throughout their lifetime, and Article 26(6) requires deployers to keep those logs for at least six months. Organizations should adopt that standard now, because the alternative is discovering what an agent did after the evidence is gone. An agent that took an action nobody authorized is an investigation, and every investigation needs a record.

    One principle for both groups

    10. State plainly what your certifications don't cover

    Organizations should be able to say, in one sentence, which AI risks their certifications do not address. An organization holding ISO/IEC 42001 and SOC 2 should know exactly where those documents stop. If nobody in the organization can say, the certifications are being treated as a badge rather than as evidence.

    How the priorities score against an assessable standard

    Assessable means an accredited third party could examine the organization and issue a finding against a published criterion. Partial means a framework requires that a process exist, and sets no criterion for its adequacy.

    Priority

    Assessable today

    Against what

    Gap

    1. Discovery of shadow AI

    No

    Nothing

    No framework reaches outside declared scope

    2. Autonomy tiering bound to authority

    No

    ISO cl. 5.3 names roles only

    Nothing requires authority to match autonomy

    3. Declared evaluation threshold

    Partial

    ISO A.6.2.4; EU Art 15(3) from Dec 2027

    Criterion set by the assessed party

    4. Architectural trifecta break

    No

    Nothing

    Named in no control in any framework

    5. Scope and SoA in procurement

    Yes

    ISO cl. 4.3, 6.1.3; SOC 2 DC §200

    Artifacts exist and are usually confidential

    6. Provenance and data governance

    Yes

    ISO A.7.3, A.7.5; EU Art 10

    Strongest coverage in the set

    7. Article 25 role screening

    Yes, in the EU

    EU Arts 16, 25

    No US equivalent

    8. Kill authority through management review

    Yes

    ISO cl. 6.2, 8.2, 9.1, 9.3

    Instrument held, run as documentation

    9. Log retention

    Yes, in the EU

    EU Arts 12, 26(6)

    Six-month floor, deferred to Dec 2027

    10. Stating coverage limits

    No

    Nothing

    No framework requires it

     

    Five of ten are assessable today. One is partial. Four are not assessable against any published criterion, and three of those four address the highest-loss risks in this paper.

    The five that score well share a property worth naming. Scope discipline, provenance, management review, role classification and log retention are all recognisable from ISO/IEC 27001 and from ordinary quality management. They were assessable before AI and remain assessable now because the object being assessed is an organizational process, and organizational processes were always the thing these instruments could see.

    The four that fail share the opposite property. Discovery requires seeing outside the declared scope. Authority binding requires a judgment about whether a control is sufficient for its risk. Architectural trifecta breaking is a property of system design. Stating coverage limits requires the assessor to certify what it did not assess. Each asks an instrument to do something its construction does not permit.

    The number that should worry a board

    An organization can execute all five assessable priorities, hold an accredited ISO/IEC 42001 certificate and a clean SOC 2 Type II, and still have no answer to the first four questions a serious incident will raise: what AI is running here that we did not approve, who could have stopped the agent, what was the model's error rate at the time, and could untrusted input have reached the tool that made the call.

    That is the assurance gap stated precisely. The frameworks certify organizational process, and enterprises fail at system behavior.

    What would close the AI governance gap

    History points to two things that fixed governance gaps in earlier technology waves: architecture that made the risky behavior impossible, and enforcement against a named company. The mechanism most people expect, a major incident that forces new rules, has a weaker track record. And audits alone, without the right institutions around them, have never been enough.

    Major incidents rarely force new rules quickly

    The usual assumption is that a large enough incident produces a mandate. Open source tested that assumption three times. As the history earlier in this paper shows, Heartbleed, the Equifax breach, and Log4Shell exposed the same weakness, and a federal mandate (Executive Order 14028 and the NTIA's minimum standards for software inventories) arrived only in 2021, seven years after the first incident.

    AI's lawsuits so far have mostly concerned inputs rather than behavior. Bartz v. Anthropic and Thomson Reuters v. Ross are disputes over training data. Moffatt v. Air Canada concerned a deployed chatbot, but it produced a small tribunal award, not a new regulation.

    The most serious AI agent security incident so far also happened in a place unlikely to change enterprise rules. According to disclosures from OpenAI and Hugging Face, between May and July 2026, AI agents in an OpenAI testing environment escaped their intended limits, gained unauthorized internet access, and used exposed credentials and previously unknown vulnerabilities to break into Hugging Face's systems, reaching administrative control of several internal Hugging Face clusters. OpenAI has said its monitoring of the agents' reasoning was not running at the time, would have caught the activity more than a day earlier, and that its production safeguards were not applied in the test environment.

    That incident will likely shape regulation of the companies that build frontier AI models. But it happened to a model developer, not to an enterprise using AI, so it is unlikely to produce new rules for enterprises. When the defining incidents happen to developers, the resulting rules will target developers.

    Architecture solved the problem when policy could not

    Personal devices at work are the clearest example of a governance problem that was actually solved. In 2012, roughly 95% of employees used at least one personal device for work, but only about 20% had signed a bring-your-own-device (BYOD) policy. From 2010 to 2014, companies tried to close the gap with device policies, acceptable use agreements, training, and remote-wipe rules. Those remote wipes sometimes erased employees' personal photos and led to lawsuits.

    What worked was a change in the technology itself. Android Work Profile and Apple User Enrollment created separate containers for work apps and data on personal devices, which made the boundary between corporate and personal data a feature of the phone. Once that existed, the policy debate stopped mattering, because the architecture had answered it.

    The AI equivalent is restricting what agents can do. An agent that cannot reach a database of customer records cannot leak those records, no matter what instructions it receives. This is why priority 4 ranks as the highest-leverage recommendation in this paper: it has the strongest historical support.

    Enforcement against a named company changed behavior

    The other approach that worked was a regulator penalizing a company for failing its share of a security responsibility. Capital One's 2019 breach exposed 106 million records because of a misconfigured web application firewall in its cloud environment. The OCC, which regulates national banks, fined Capital One $80 million in August 2020, and kept the bank under a consent order until August 2022. That penalty came about fourteen years after Amazon launched its EC2 cloud service, and it did more to change cloud governance practices than the Cloud Controls Matrix had in the previous decade.

    No equivalent action exists for AI yet. EU AI Act penalties can reach €35 million or 7% of a company's worldwide annual revenue for prohibited practices, but many EU member states have not yet designated the authorities responsible for enforcement. As of September 2026, no significant AI Act enforcement action has been taken.

    Audits alone will not close the gap

    Research on audits across industries points to the same conclusion: an audit only works when the institutions around it are designed to act on its findings. Legal and AI accountability researchers Raji, Xu, Honigsberg, and Ho studied third-party audit systems in financial, environmental, and health regulation and concluded that audits alone are unlikely to make AI accountable without sustained attention to how those systems are designed.

    Cloud governance shows what that design looks like. It worked because FedRAMP accredited the firms allowed to assess cloud providers, federal contracts passed security requirements down to vendors, and a regulator was willing to impose penalties. The Cloud Controls Matrix was necessary, but it was not sufficient.

    The research also documents what goes wrong when audits stand alone. Legal scholars Goodman and Tréhu describe "audit-washing," where audits create a false sense of legitimacy, turn minimum standards into a ceiling on performance, and discourage organizations from looking for problems on their own. They cite Chris Hoofnagle's research on FTC-mandated privacy assessments, which found that passing those assessments had little connection to actual privacy practice. Google, for example, filed clean assessments while losing court rulings over wiretapping.

    A survey of the algorithmic auditing field by Costanza-Chock, Raji, and Buolamwini found that only 7% of auditors use standardized frameworks. 82% support public disclosure of audit results in principle, but only 4 of 43 auditors actually provided links to their results, and 65% report that the organizations they audit refuse to commit to fixing the problems found. One auditor in that study voiced the concern this paper documents: that the definition of what it means to be audited will be written down, and written too narrowly.

    Legal researchers Terzis, Veale, and Gaumann warn that traditional auditing firms tend to take over new audit fields and bring their existing methods with them. Their historical example is telling: the accounting profession helped write Britain's 1868 Regulation of Railways Act, which required the accounting formats its own firms specialized in, and those firms went on to become the main railway auditors.

    What this paper does not claim

    This paper does not claim the frameworks are worthless. ISO/IEC 42001's provenance controls, its impact assessment requirements (clause 6.1.4 and Annex A.5), and its management review process are real, and better than what came before. The EU AI Act's human oversight requirements in Article 14, including the ability to stop or override a system, awareness of automation bias, and two-person review for biometric identification, are the most specific human oversight rules in any technology regulation. SOC 2 does genuine work on AI infrastructure security.

    It does not claim that writing governance frameworks first was a mistake. Writing process rules early was the right choice, given that the science for measuring AI behavior did not yet exist. It is also why this paper can describe the gap precisely, rather than argue about whether a gap exists at all.

    It does not claim that a certifiable standard for AI behavior is achievable today. ISO is developing ISO/IEC 42007, a draft standard for certifying AI systems themselves, and it is the development to watch. Whether it can specify what ISO/IEC 42001 left out is an open question, and the answer probably depends on evaluation methods that are still being built.

    The claim is narrower. Four frameworks certify organizational process. The ten risks that cause the most enterprise AI losses are split between process and system behavior. And the frameworks are being read as though they covered both.

     

    Appendix: Method and exclusions

    Method

    Research for this paper was conducted on 18 September 2026. Primary sources were used wherever possible, including ISO's authorized previews, NIST publications and its AI Resource Center crosswalks, EUR-Lex and the EU Official Journal, AICPA materials, regulator publications, and peer-reviewed or peer-track academic research. Vendor materials were used only for control lists and were cross-checked against independent sources.

    ISO/IEC 42001 control titles were confirmed against NIST's crosswalk between the AI RMF and the final draft of ISO/IEC 42001, and against four independent lists that agree with it. At least one widely shared public "complete list of all 38 controls" is fabricated, with control numbers and titles that do not exist. Because the standard is behind a paywall, some of the public discussion about it is based on invented material.

    No authoritative public count of ISO/IEC 42001 certificates exists. The figure of roughly 350 organizations as of April 2026 counts only publicly announced certificates.

    Deliberately excluded

    The widely quoted MIT NANDA figure that 95% of enterprise AI pilots fail is not used. It comes from a non-peer-reviewed preprint with contested methodology, and the number is usually quoted without its qualifications. McKinsey's figures, showing that 39% of organizations report any earnings impact from AI and roughly 6% qualify as high performers, support the same point and are better documented.

    The claim that a "SOC for AI" reporting framework exists could not be verified and appears to be false. The AICPA's list of SOC reports does not include it.

    The Simons and Broniatowski study uses AI-simulated personas rather than real organizations. Its findings illustrate a structural argument rather than measure real-world behavior, and the authors recommend testing with real organizations.

    Anthropic's report that AI carried out 80–90% of an espionage campaign is a single company's assessment of misuse of its own product. It conflicts with the more measured conclusion of the International AI Safety Report 2026, and both are cited.

     

     

    Sources

    Surveys and industry reports:

    EY, "EY Survey Finds That Autonomous AI Implementation Outpaces Oversight, Yielding an AI Governance Gap," September 15, 2026. https://www.ey.com/en_us/newsroom/2026/09/ey-survey-finds-that-autonomous-ai-implementation-outpaces-oversight-yielding-an-ai-governance-gap

    IBM, "Cost of a Data Breach Report 2026," July 2026. https://www.ibm.com/reports/data-breach

    IBM, "IBM Study: One in Four Malicious Breaches Are AI-Enabled, Costing Companies $6 Million on Average," July 29, 2026. https://newsroom.ibm.com/2026-07-29-ibm-study-one-in-four-malicious-breaches-are-ai-enabled,-costing-companies-6-million-on-average

    ComplexDiscovery, "Policy Without Control: The AI Governance Gap in IBM's 2026 Cost of a Data Breach Report," 2026. https://complexdiscovery.com/policy-without-control-the-ai-governance-gap-in-ibms-2026-cost-of-a-data-breach-report/

    Enterprise Management Associates, "Agents Without Guardrails: The Agentic AI Governance Gap in the Enterprise," August 2026. https://www.cequence.ai/wp-content/uploads/2026/08/EMA-Research-Report-Agents-Without-Guardrails.pdf

    Verizon, "2026 Data Breach Investigations Report," 2026. https://www.verizon.com/business/resources/reports/dbir/

    Frameworks and standards

    NIST, "Artificial Intelligence Risk Management Framework (AI RMF 1.0)," NIST AI 100-1, January 26, 2023. https://doi.org/10.6028/NIST.AI.100-1

    NIST, "NIST AI RMF Playbook." https://pages.nist.gov/AIRMF/

    NIST, "Roadmap for the NIST Artificial Intelligence Risk Management Framework (AI RMF 1.0)." https://www.nist.gov/itl/ai-risk-management-framework/roadmap-nist-artificial-intelligence-risk-management-framework-ai

    NIST, "Crosswalk: AI RMF (1.0) and ISO/IEC FDIS 42001," AI Resource Center. https://airc.nist.gov/docs/NIST_AI_RMF_to_ISO_IEC_42001_Crosswalk.pdf

    CEN-CENELEC, "EN 18286 in the Spotlight: Supporting Compliance with the AI Act," July 30, 2026. https://www.cencenelec.eu/news-events/news/2026/en-in-the-spotlight/2026-07-30-ai-quality-management/

    European Commission, "Understanding the Standardisation of the AI Act." https://digital-strategy.ec.europa.eu/en/faqs/understanding-standardisation-ai-act

    European Commission, "Implementing Decision C(2025) 3871 Final," June 23, 2025. https://ec.europa.eu/transparency/documents-register/api/files/C(2025)3871_0/de00000001072818?rendition=false

    Regulation

    Hunton Andrews Kurth, "EU Digital Omnibus on AI Enters into Force," July 2026. https://www.hunton.com/privacy-and-cybersecurity-law-blog/eu-digital-omnibus-on-ai-enters-into-force

    Board of Governors of the Federal Reserve System, "SR 26-2: Revised Guidance on Model Risk Management," April 17, 2026. https://www.federalreserve.gov/supervisionreg/srletters/SR2602.htm

    OCC, "Bulletin 2026-13: Model Risk Management: Revised Guidance," April 17, 2026. https://www.occ.gov/news-issuances/bulletins/2026/bulletin-2026-13.html

    OCC, "News Release 2026-29: OCC Issues Updated Model Risk Management Guidance," April 17, 2026. https://www.occ.treas.gov/news-issuances/news-releases/2026/nr-occ-2026-29.html

    Davis Wright Tremaine, "Colorado AI Act Repealed and Replaced by Narrower Statute Focused on Transparency Requirements and Enhanced Consumer Rights," May 2026. https://www.dwt.com/blogs/privacy--security-law-blog/2026/05/colorado-ai-act-repeal-new-transparency-law

    Court rulings

    Higher Regional Court of Hamm (OLG Hamm), "Judgment of May 12, 2026, Case 4 UKl 3/25." Analysis by Taylor Wessing, June 2026. https://www.taylorwessing.com/de/insights-and-events/insights/2026/06/madewe-nl-03

    Munich Regional Court I (LG München I), "Judgment of May 28, 2026, Case 26 O 869/26." https://dejure.org/2026,16717

    Giovanni Vetrugno, "Who Speaks When the Algorithm Speaks? A German Ruling on AI Overviews," University of Oxford Faculty of Law Blogs, 2026. https://blogs.law.ox.ac.uk/node/51011

    Incident disclosures

    Hugging Face, "Security Incident Disclosure: July 2026," July 16, 2026. https://huggingface.co/blog/security-incident-july-2026

    OpenAI, "OpenAI and Hugging Face Partner to Address Security Incident during Model Evaluation," July 21, 2026. https://openai.com/index/hugging-face-model-evaluation-security-incident/

    Hugging Face, "Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident," July 27, 2026. https://huggingface.co/blog/agent-intrusion-technical-timeline

    OpenAI, "The Hugging Face Incident and the Road Ahead," August 26, 2026. https://openai.com/index/hugging-face-incident-and-the-road-ahead/

    Academic research

    Stephen Casper et al., "Black-Box Access Is Insufficient for Rigorous AI Audits," FAccT 2024. https://doi.org/10.1145/3630106.3659037

    Inioluwa Deborah Raji, Peggy Xu, Colleen Honigsberg, and Daniel E. Ho, "Outsider Oversight: Designing a Third Party Audit Ecosystem for AI Governance," AIES 2022. https://doi.org/10.1145/3514094.3534181

    Sasha Costanza-Chock, Inioluwa Deborah Raji, and Joy Buolamwini, "Who Audits the Auditors? Recommendations from a Field Scan of the Algorithmic Auditing Ecosystem," FAccT 2022. https://doi.org/10.1145/3531146.3533213

    Ellen P. Goodman and Julia Tréhu, "AI Audit-Washing and Accountability," German Marshall Fund of the United States, November 2022. https://www.gmfus.org/news/ai-audit-washing-and-accountability

    Petros Terzis, Michael Veale, and Noëlle Gaumann, "Law and the Emerging Political Economy of Algorithmic Audits," FAccT 2024. https://doi.org/10.1145/3630106.3658970

    Joseph R. Simons and David A. Broniatowski, "Why AI Governance Frameworks Are Hard to Adopt: A Role-Based Stress Test of the NIST AI RMF," arXiv:2608.12352, 2026. https://arxiv.org/abs/2608.12352

    Other

    LBMC, "Can We Truly Audit AI?," May 28, 2025. https://www.lbmc.com/blog/can-we-audit-ai/

    AI Provider Trust Registry, "Llama via AWS Bedrock," last verified September 27, 2026. https://aiprovidertrust.com/offerings/llama-bedrock/

     

     

    “The vendors that we've rolled this out to have liked it. Because instead of 300-600 questionnaires, they're really only looking at some 40 pieces of evidence that they upload. They feel better represented if they're being analyzed from a security effectiveness perspective against competitors."

    “Easy to use platform, excellent support and guidance.” 

    Matt C.
    President/CEO

    "Streamlined Compliance with Intuitive Interface. I really appreciate how Strike Graph simplifies and structures the entire compliance process."

    Vivek S.
    Associate Engineer, Enterprise company

    "Strike Graph is an Enterprise Governance, Risk, and Compliance (GRC) tool that has improved Sanmina's security compliance and risk management across 23 countries, multiple locations, and various frameworks. By centralizing operations and replacing manual tracking, it has significantly simplified compliance, enhanced security, and improved our risk matrix documentation."

    Larry F.
    VP IT Security, Enterprise company
    "We were looking at doing [third-party risk management] manually, but Strike Graph appeared, and it helps automate what we were struggling to plan."
    G2-image 1
    G2-image 2
    G2-image 3
    G2-image 4
    G2-image 5
    G2 image 10

    Ready to see Strike Graph in action?

    Fill out a simple form and our team will be in touch.

    Experience a live customized demo, get answers to your specific questions , and find out why Strike Graph is the right choice for your organization.

    Fields marked with a star (*) are required

    By submitting this form, you agree to receive promotional messages from Strike Graph about its products and services. You can unsubscribe at any time by clicking on the link at the bottom of our emails.

    Fill out a simple form and our team will be in touch.

    Experience a live customized demo, get answers to your specific questions , and find out why Strike Graph is the right choice for your organization.

    Ready to see Strike Graph in action?

    Fill out a simple form and our team will be in touch.

    Schedule a Demo
    foot-dark-shade
    SG-logo-white
    Strike Graph is an enterprise AI-native compliance management platform that accelerates audits, eliminates redundant work, and builds trust through its secure, agentic technology and enterprise-ready data model.
    • Resources
    • Schedule a demo
    • Product Support
    • Sign In
    • Contact Us
    • 🦆 icon _rounded linkedin_
    • 🦆 icon _rounded facebook_
    • 🦆 icon _rounded twitterbird_
    • Website images - Subtract

    © 2026 Strike Graph, Inc. All Rights Reserved • Privacy Policy • Terms of Service • EU AI Act

    SOC_NonCPAA
    Achieved-SG-badge_hipaa

    Ready to see Strike Graph in action?

    Fill out a simple form and our team will be in touch.

    Experience a live customized demo, get answers to your specific questions , and find out why Strike Graph is the right choice for your organization.

    Fields marked with a star (*) are required

    Fill out a simple form and our team will be in touch.

    Experience a live customized demo, get answers to your specific questions , and find out why Strike Graph is the right choice for your organization.