Quick answer: An AI governance framework is the set of policies, roles, controls and evidence an organization uses to decide which AI systems may be built, how they are tested, who is accountable for them in production, and how they are monitored and retired. A workable enterprise framework has five parts: an AI inventory, risk tiering, lifecycle gates, named ownership, and continuous monitoring, mapped to NIST AI RMF or ISO/IEC 42001.
Most enterprise AI governance frameworks were written in 2023 and 2024 for a world of chatbots and copilots. They were documents. In 2027 the systems they need to govern are different: retrieval pipelines that touch regulated data, fine-tuned models embedded in underwriting and claims, and agents that call tools and move money. The EU AI Act's high-risk obligations are now live, APRA's CPS 230 operational-risk standard is in force in Australia, and US financial regulators expect model-risk discipline to extend to generative systems.
This article is for the people who have to make governance work rather than write it: Chief Data and AI Officers, AI platform owners, heads of risk, and the architects who get asked "is this allowed?" at 4pm on a Friday. It lays out what a framework must contain, who owns which piece, how the two dominant standards differ, and where governance most often becomes gridlock.
What is an AI governance framework, in practice?
Strip away the ethics language and an AI governance framework answers six operational questions:
- What AI do we have? An inventory of every model, prompt-based application and agent in use, including shadow usage and vendor-embedded AI.
- How risky is each one? A tiering scheme that puts a customer-facing credit decision in a different lane from an internal meeting summarizer.
- What must happen before it ships? Lifecycle gates: data approval, evaluation thresholds, security review, human-oversight design, documentation.
- Who is accountable? Named owners for each system, plus a decision body that can say yes, no, or "yes with conditions" within a known timeframe.
- How do we know it is still behaving? Monitoring for drift, quality regression, cost, misuse and incidents, with escalation paths.
- What evidence can we show? Records that would satisfy an internal auditor, an external regulator, or a customer's procurement team.
Frameworks that only cover principles fail at question 3; frameworks that are only a 40-page policy fail at question 5.
What should an enterprise AI governance framework include?
The table below is the minimum component set we see in enterprises that have moved beyond pilots. It aligns to the four NIST AI RMF functions (Govern, Map, Measure, Manage) so you can trace it back to a recognized standard.
- AI inventory and registry — What it contains: Every model, GenAI app, agent, vendor-embedded AI; owner; data used; risk tier · NIST AI RMF function: Map · Common failure mode: Only covers data-science models; misses SaaS copilots and agents
- Risk tiering — What it contains: 3-4 tiers based on decision impact, data sensitivity, autonomy, regulatory exposure · NIST AI RMF function: Map · Common failure mode: One tier for everything, so every project gets the heaviest process
- Lifecycle gates — What it contains: Intake, data approval, evaluation, security and privacy review, pre-production sign-off, change control · NIST AI RMF function: Measure / Manage · Common failure mode: Gates defined but no SLA, so projects queue for months
- Roles and decision rights — What it contains: System owner, model risk, security, legal, data protection, business sponsor; a council with quorum rules · NIST AI RMF function: Govern · Common failure mode: Council exists but no one can approve below it
- Evaluation standards — What it contains: Required tests per tier: accuracy, robustness, bias, red-teaming, hallucination and groundedness for RAG · NIST AI RMF function: Measure · Common failure mode: Pass/fail thresholds never written down
- Monitoring and incident response — What it contains: Drift, quality, cost, misuse signals; AI-specific incident runbook; rollback · NIST AI RMF function: Manage · Common failure mode: Monitoring is a dashboard nobody is paged on
- Third-party and model-supply governance — What it contains: Vendor AI due diligence, model cards, licensing, data-residency terms · NIST AI RMF function: Map / Manage · Common failure mode: Vendor clauses cover software, not models
- Policy and training — What it contains: Acceptable-use policy, AI literacy obligations (EU AI Act Article 4), developer standards · NIST AI RMF function: Govern · Common failure mode: Policy published once, never operationalized
- Evidence and audit trail — What it contains: Decision logs, test results, approvals, change records, retained per tier · NIST AI RMF function: Govern · Common failure mode: Evidence lives in chat threads and slide decks
Two additions matter in 2027 that were optional in 2024. First, generative AI governance needs explicit controls for prompts, retrieval sources and output filtering, because the "model" is now model plus prompt plus knowledge base. Second, agent governance needs controls over actions and tool permissions, not only outputs. Both are covered in more depth in our companion pieces on enterprise AI platform architecture and agentic AI governance.
Who is responsible for AI governance in an enterprise?
Ownership is where most frameworks are quietly broken. A workable split:
- Accountable executive. Increasingly a Chief AI Officer or Chief Data and AI Officer; in regulated firms sometimes the Chief Risk Officer. One person owns the framework, not a committee.
- AI governance council. Cross-functional (risk, legal, security, privacy, data, business lines). Sets policy and tiering, approves tier-1 systems, reviews incidents. Meets on a cadence and has a documented quorum and turnaround time.
- System owners. Every AI system has a named business owner who signs off that it is fit for purpose and accepts residual risk. This is not the data scientist who built it.
- Platform and enablement team. Builds the controls into the platform (gateway policies, evaluation harnesses, logging) so that compliance is the default path, not a side quest.
- Second line (model risk, compliance). Independent challenge, especially in banking where SR 11-7-style model risk management already exists and is being extended to generative systems.
- Third line (internal audit). Periodic assurance that the framework is operating as described.
The practical test: can a team shipping a tier-2 system find out within a week who decides, what evidence is needed, and when they will hear back?
NIST AI RMF vs ISO/IEC 42001: which should you build on?
Both are voluntary and both are being referenced by regulators and procurement teams. They are complementary rather than competing.
- What it is — NIST AI RMF 1.0 (2023) + Generative AI Profile (2024): A risk-management framework: four functions (Govern, Map, Measure, Manage) with subcategories and suggested actions · ISO/IEC 42001:2023: A certifiable management-system standard (like ISO 27001 for security) for an "AI management system"
- Certification — NIST AI RMF 1.0 (2023) + Generative AI Profile (2024): None; self-attestation · ISO/IEC 42001:2023: Third-party certification available
- Strength — NIST AI RMF 1.0 (2023) + Generative AI Profile (2024): Flexible, strong on risk identification and measurement; the GenAI Profile lists concrete risks and actions · ISO/IEC 42001:2023: Forces organizational discipline: policy, roles, objectives, internal audit, management review, continual improvement
- Best fit — NIST AI RMF 1.0 (2023) + Generative AI Profile (2024): US-headquartered firms; organizations that want a risk vocabulary without a certification program · ISO/IEC 42001:2023: Firms that already run ISO 27001 and want AI to plug into the same management system; firms selling into procurement processes that ask for certification
- Regulatory mapping — NIST AI RMF 1.0 (2023) + Generative AI Profile (2024): Referenced by US agencies and state laws (for example Colorado's AI Act allows NIST alignment as a defense) · ISO/IEC 42001:2023: Mapped by many to EU AI Act quality-management requirements for high-risk providers
A common pattern: use NIST AI RMF as the risk taxonomy and control vocabulary, and run the governance machinery (roles, audits, reviews) as an ISO 42001-shaped management system, certifying only if a customer or regulator asks. Add the EU AI Act's obligations as an overlay for any system that is in scope.
How does the EU AI Act change enterprise AI governance?
For an enterprise framework, the Act's main effects are:
- Classification becomes mandatory work. You need to determine for each system whether it is prohibited, high-risk (Annex III use cases such as creditworthiness, employment, essential services, plus safety components), limited-risk (transparency obligations), or minimal. Your inventory and tiering must be able to produce this.
- Deployers have duties, not just providers. Enterprises using high-risk systems must use them per instructions, ensure human oversight by trained people, monitor operation, keep logs, and in some cases run a fundamental-rights impact assessment.
- AI literacy is a legal obligation (Article 4, applicable since February 2025) for staff dealing with AI systems.
- Timeline. Prohibitions applied from February 2025, general-purpose AI model obligations from August 2025, and the main high-risk obligations phased in from August 2026, with further elements following into 2027. The Commission has proposed adjustments to implementation timelines, so check the current state before committing dates in your plan.
- Penalties scale with turnover, up to 7% of global annual turnover for prohibited practices.
UK and Australian enterprises are in scope if their systems or outputs are used in the EU; Australia's Voluntary AI Safety Standard and the UK's sector-regulator approach are lighter but point the same way.
How do you govern generative AI differently from traditional ML?
Classical model risk management assumed a model with a fixed training set, a scored output and a measurable error rate. Generative and retrieval-based systems break several of those assumptions, so the framework needs additions:
- Scope the "system", not the model. The thing you govern is model + system prompt + retrieval corpus + tools + output filters. A change to any part is a change to the system and should go through change control.
- Evaluate for groundedness, not just accuracy. For RAG, test that answers are supported by retrieved sources, that retrieval respects document-level permissions, and that the system refuses when it should.
- Red-team for misuse. Prompt injection, data exfiltration via outputs and jailbreaks are in the OWASP Top 10 for LLM Applications; treat them as required pre-production tests for anything exposed to untrusted input.
- Monitor quality continuously. Foundation models change under you when a vendor updates them. Pin versions where you can, and run regression evaluations on a schedule.
- Govern the corpus. Who approved the documents the system can read? When were they last reviewed? Stale or unauthorized knowledge is the most common root cause of embarrassing GenAI outputs.
In practice: a mid-sized insurer rolled out a claims-handling assistant that drafted responses to policyholders. The model never changed, but the retrieval corpus included an outdated policy wording from a product withdrawn in 2024. The assistant quoted it to several customers before a complaints analyst noticed. The fix was not a model fix: it was a corpus ownership rule (every indexed document has a named owner and an expiry), a groundedness evaluation in the release gate, and a monitoring alert on citations to documents flagged as superseded. That is what "generative AI governance" means operationally.
How much should AI governance cost, and how do you avoid gridlock?
There is no reliable industry benchmark. Spend should be proportional to risk tier, and the fastest-moving enterprises invest in platform-embedded controls (which scale) rather than committee time (which does not). Three design rules keep governance from becoming the reason projects stall:
- Tier aggressively. Most internal productivity use cases should clear a lightweight self-assessment in days. Reserve the full gate sequence for systems that make or materially influence decisions about people, money or safety.
- Set turnaround SLAs for every gate and publish them. A council that takes six weeks to review a tier-3 chatbot has created an incentive for shadow AI.
- Automate the evidence. Evaluation results, approvals, prompts and corpus versions should be captured by the platform, not assembled by hand for each review. Your gateway and MLOps tooling are the control points; see our note on where policy is enforced in the platform.
Governance designed with delivery teams, not for them, is the pattern in the operating models of enterprises that have scaled; see enterprise AI strategy: from pilot to production.
Key takeaways
- A framework that cannot tell you what AI you have, how risky each system is, and who decides is a policy, not governance.
- Map components to NIST AI RMF for a shared risk vocabulary; run the machinery as an ISO/IEC 42001-style management system; overlay EU AI Act obligations where in scope.
- Generative AI changes the unit of governance from "model" to "system": prompts, corpus and tools all need ownership and change control.
- Tier aggressively, publish gate SLAs and automate evidence capture in the platform, or governance becomes the cause of shadow AI.
- Agents extend governance from outputs to actions; plan for that now rather than retrofitting it.
Join your peers at Enterprise AI Global 2027
Enterprise AI Global is the practitioner forum for the people building and leading AI inside the enterprise, with governance, operating-model and production-architecture tracks in every city. Register for Melbourne, 2 March 2027, Sydney, 11 August 2027, New York, 14 October 2027, San Francisco, 21 October 2027 or London, 28 October 2027.
Not ready to register? Sign up for updates on agendas, speakers and new practitioner articles.