AI Governance Has a Grading Problem, Not a Framework Problem
More than 40 AI governance frameworks and ethical guidelines exist across the world today. Almost none of them have been tested to see whether they actually work. That is not my characterization. It is the finding of the Independent International Scientific Panel on AI, the first global scientific body dedicated to AI, established by the United Nations (UN) General Assembly and co-chaired by Yoshua Bengio and Maria Ressa. Its preliminary report, released on July 1, 2026, ahead of the inaugural Global Dialogue on AI Governance in Geneva, adds a second finding that matters even more: many of the safety assessments behind those frameworks are conducted by the companies developing the technology themselves.
That is the thesis of this piece. Governments do not have a framework shortage. They have an evaluation capacity shortage, and writing a forty-first framework is easier than building the capacity to test whether any of the first 40 actually hold up. The same flaw runs through corporate AI governance, where the team that ships a model and the team that grades its safety often report to the same executive. Three weeks before the panel’s report landed, that exact flaw played out in public when a leading AI lab found a flaw in its own safeguards, graded its own fix, and reported the result to the world.
The evidence dilemma the panel describes is not a temporary lag that better funding or faster committees will close. It is the permanent condition of governing a technology that outpaces its own audit trail, and most institutions, public and private, are still built as though the lag will eventually close on its own.
The Evidence Dilemma Is a Condition, Not a Queue
The panel’s central diagnosis is what its own members call an evidence dilemma. Policymakers need reliable data before they can regulate responsibly. By the time enough data exists, the technology has already moved on. Governments today, according to the panel, were not built for a technology evolving this quickly, and the assessment gap between what is deployed and what is understood keeps widening rather than closing.
Most institutions treat this as a temporary problem, a backlog that a bigger budget or a faster committee will eventually clear. That is the wrong model. Consider how central banks operate under a structurally similar condition. A central bank never has complete information about the economy it is steering. Data arrives with a lag, models are always approximations, and by the time a clean picture emerges, the moment for acting on it has often passed. Central banks did not solve this by waiting for better data. They built standing institutional capacity, staff, models, and decision protocols, designed to act competently under permanent uncertainty rather than to eliminate the uncertainty first. AI governance has not made that same shift. Most governments are still organized as though the evidence gap is a queue they are waiting to clear, rather than a permanent operating condition they need standing capacity to manage.
This distinction is not academic. A queue justifies waiting. A permanent condition demands infrastructure. The panel’s own language, that policymakers face an evidence dilemma rather than an evidence lag, is a tell that even the scientists writing the report understand this is closer to the second category. The frameworks being produced in the meantime mostly behave as if they were the first.
Self-Assessment, Twice Over: What the Fable and Mythos Sequence Actually Showed
The panel notes that many safety assessments are conducted by the companies developing the technology, rather than by an independent party. Most readers will nod at that sentence and move on, because it sounds like a structural critique rather than a specific, dated event. It is worth pausing to consider a specific, dated event, because one occurred in the weeks before the report was published.
On June 9, 2026, Anthropic released two frontier models, Fable 5 and Mythos 5. Three days later, on June 12, the United States government issued export control directives requiring Anthropic to restrict foreign access to both, following a report from Amazon researchers describing a technique that bypassed one of Fable 5’s cybersecurity safeguards. Because the restriction took effect immediately and Anthropic had no reliable way to verify user nationality in real time, the company suspended access to both models for everyone, not only for foreign nationals. Anthropic reviewed the Amazon finding, concluded it represented a borderline case for Fable 5’s safeguards rather than a unique Mythos-level capability, retrained its own safety classifier, and reported that the technique was now blocked in more than 99 percent of cases. The controls were lifted on June 30, and Fable 5 began rolling out globally again on July 1, the same day the panel’s report launched. Mythos 5, the more capable of the two models, received only a partial restoration, limited to a specific group of United States (US) organizations.
Walk through who did what in that sequence. An external researcher found the flaw. The company that built the model then led the assessment of how serious it was, built its own fix, and reported its own success rate of more than 99 percent. There was an external check, but it is worth being precise about what kind. Commerce’s Center for AI Standards and Innovation reviewed the safeguards before access was restored. However, that review came through an emergency export control order issued under national security authorities, not through any routine evaluation process. A voluntary pre-release framework, created by executive order on June 2, was not yet operational when Fable 5 launched a week later, and by design it would have depended on Anthropic choosing to submit the model in the first place. The effectiveness figure itself, the more than 99 percent, was the company’s own measurement. And the fix was narrow: a classifier tuned to block the single technique that had been reported, which does nothing for the techniques not yet found, and detection-based safeguards of exactly that kind were what had been defeated to trigger the ban.
This is not a criticism of how any single company handled a difficult three weeks. It is a near-perfect illustration of the panel’s abstract finding, playing out with a name, a date, and a percentage attached. The only independent scrutiny arrived as an emergency national security intervention, triggered because a competitor happened to find the flaw and had the standing to escalate it to the White House. There was no standing evaluation regime that required a check; there was a crisis that produced one. That is the opposite of the durable, institutionalized capacity the panel is asking governments to build.
The same structure exists inside most corporate AI governance functions, just with less public visibility. A model risk committee that reports to the same executive who owns the product roadmap is being asked to grade work that its own reporting line has an interest in passing. Fable and Mythos made that dynamic visible for three weeks. Most enterprise AI programs run this way permanently, without an export control order to force a public accounting of what happened.
Counterpoint: Isn’t Caution With Immature Evidence Exactly What Good Regulation Looks Like
The strongest objection to this argument is that caution in the face of incomplete evidence is not dysfunction. It is what responsible governance is supposed to look like. The panel itself is careful not to prescribe specific rules, precisely because it does not want to lock in premature judgments about a fast-moving technology. Rushing to build evaluation capacity before anyone agrees on what should be evaluated for, this objection runs, risks manufacturing false confidence faster than it manufactures safety. A poorly designed test is arguably worse than no test, since it gives everyone permission to stop worrying.
This is a fair challenge, and it deserves a real answer rather than a dismissal. The answer is that evidentiary caution and evaluation capacity are not competing priorities. They are sequential requirements for the same goal. A government or a company that has no independent capacity to test a system cannot exercise evidentiary caution in any meaningful way. It can only decide whether to trust what it is told. Building the capacity to test does not commit anyone to premature rules. It commits them to being able to check a claim before acting on it, which is the precondition for caution, not a substitute for it. The Fable and Mythos sequence shows exactly why this distinction matters. The one external check that occurred was an emergency government review, triggered by a competitor’s escalation, and the headline effectiveness figure remained the company’s own measurement. Even the leading AI power’s considered answer to the evaluation gap, the June 2 executive order, was a voluntary framework that a developer can decline to enter, resting on a threshold set through a classified process that developers cannot see. A regime a lab can opt out of is not the standing independent capacity that evidentiary caution requires.
Forty Frameworks, Zero Benchmarks
If the evidence dilemma is permanent and self-assessment is the default, the natural question is why governments keep producing new frameworks instead of the evaluation capacity that would let them test the frameworks they already have. The panel counts more than 40 of these documents worldwide and notes that they remain fragmented, inconsistent, and rarely tested to determine whether they work in practice.
The honest answer is that drafting a framework is achievable within a single budget cycle and a single minister’s tenure. Building independent evaluation capacity is not. A framework is a document a ministry can commission, draft, publish, and claim credit for within 18 months. An evaluation body requires specialized technical staff who are expensive and hard to hire away from industry, a multi-year funding commitment that survives a change of government, and a mandate that gives it actual access to the systems it is supposed to test. Every incentive in a typical government points toward the document and away from the institution. The same asymmetry exists inside companies. A board can approve an AI governance policy in a single meeting. Standing up a technical evaluation function that is structurally independent from the product organization takes years, dedicated headcount, and a willingness to let that function occasionally block a launch.
Forty frameworks without benchmarks are not evidence that AI governance is immature. They are evidence that institutions, public and private, keep choosing the achievable output over the necessary one. Fragmentation is the visible symptom. The capacity shortage underneath it is the actual disease.
What Separating the Grader From the Graded Would Actually Require
None of this means governments should stop writing frameworks or that boards should stop approving policies. It means the frameworks and policies need a second, independent institution behind them that can check whether the underlying claims hold up, and that institution needs three things most current efforts lack.
Standing funding that survives a change in leadership. An evaluation body funded project by project will always be vulnerable to being defunded the year its findings become inconvenient. Funding needs to be structured more like an audit function than a grant program, insulated from the political or commercial cycle that produced the systems it evaluates.
A mandate that guarantees access rather than requesting it. The most credible independent evaluators operating today still depend on the voluntary cooperation of the organizations they evaluate. That dependency is a structural weakness worth naming directly, and it sets up the cases in the next section.
A reporting line that does not terminate inside the organization being graded. This is the simplest requirement, and the one governments and companies both violate most often. If the person receiving the evaluation result also controls the evaluator’s budget, promotion, and continued access, the evaluation is advisory at best.
What Evaluation Capacity Looks Like When It Is Actually Built
Three institutions, at very different scales, show what happens when governments take the second and third requirements above seriously. None of them fully solves the funding problem. All three demonstrate that the model can work.
The United Kingdom’s AI Security Institute, created after the 2023 Bletchley Park summit as the AI Safety Institute and renamed in 2025, sits inside the Department for Science, Innovation and Technology but operates as a dedicated evaluation directorate rather than a policy unit. It conducts continuous, independent testing of frontier systems before and after deployment, and has tested more than 30 frontier models to date, feeding its findings back to developers to strengthen safeguards while publishing its own trend reporting separately. The institute anchors a wider network, with counterpart AI Safety Institutes now established in the US, Japan, France, Canada, Australia, and elsewhere, coordinating shared testing methodology rather than each government inventing its own from scratch. Its limitation is the one named above: the institute’s access to pre-deployment systems still depends on developers agreeing to provide it, not on a legal requirement that they must.
METR, a nonprofit evaluator based in California, offers the clearest example of funding structured for independence. It has conducted pre-deployment evaluations for essentially every major frontier model released since GPT-4, and it does not accept payment from the labs whose models it tests, relying instead on philanthropic funding and compute credits donated by the companies themselves. That structure is deliberate. Accepting a fee from the organization being evaluated recreates the exact conflict the evaluation exists to remove. METR has been candid about the limits this independence does not address. It has no legal entitlement to test any lab’s model, so a lab could simply deny access if a relationship became strained. It has also documented what researchers call the adversarial elicitation problem, the risk that a lab could under-invest in helping a model perform well during testing, thereby presenting a deliberately weaker version of a system than the one it plans to deploy. METR’s response has been to publish its own standards for what counts as a fair test, but it cannot independently verify that every lab meets them. This is the honest state of the field’s most independent evaluator, and it argues for exactly the point above: capacity without a binding mandate is progress, not a solution.
The most useful case for a development finance audience is Singapore’s. Its Digital Trust Centre, established with an initial government commitment of S$50 million and operated through Nanyang Technological University, developed the testing methodology behind A.I. Verify, which is now maintained by an independent foundation with more than 90 member organizations worldwide. Singapore was subsequently designated as a national AI Safety Institute and has since chaired the Association of Southeast Asian Nations (ASEAN) Digital Ministers’ Meeting that launched a regional guide on AI governance and ethics, extending its testing methodology into a shared regional standard rather than keeping it a domestic asset. The lesson is not that every country needs Singapore’s resources. It is that meaningful evaluation capacity has been built for a fraction of what a frontier training run costs, by a government that chose to invest in testing infrastructure rather than in a fifth or sixth national AI strategy document, and then made that capacity exportable to its neighbors.
Three institutions, three very different routes to the same capability: a government directorate, an independent nonprofit, and a state-seeded body that handed its methodology to a foundation. The budgets differ by orders of magnitude and the ownership models share almost nothing, yet all three made the same underlying choice: fund the capacity to check a claim, not only the document that states one. What none of them has fully solved is the mandate problem. Evaluation built on voluntary access is real capacity sitting on a fragile foundation, and closing that gap is the unfinished business the UN panel’s own three-year mandate now inherits.
Implications for Decision Makers
The instinct after reading a diagnosis like this is to ask which framework needs rewriting. That is the wrong question. The right one is which independent capacity your organization can fund this year that did not exist last year, because the frameworks will keep multiplying regardless of what you do next.
- Fund evaluation as a standing line item, not a project. A one-time grant or a pilot program signals that testing is optional. Budget it the way an internal audit function is budgeted, as a permanent cost of operating the technology, immune to being cut when its findings become inconvenient. Ask whether your current AI governance spending would survive a change in minister or chief executive.
- Separate the reporting line before you separate the org chart. The Fable and Mythos sequence shows what happens when the entity that built the system also grades its own fix. Before restructuring anything, identify who currently signs off on your organization’s AI safety claims and whether that person’s incentives are aligned with finding a problem or with shipping on schedule.
- Treat existing frameworks as living documents with a review trigger, not finished products. The panel’s finding that 40 frameworks remain untested is a warning about what happens when a framework is written once and filed. Attach a mandatory review date to any AI governance document your organization has already approved, and assign someone whose job depends on the review actually happening.
- Build capacity even where you cannot yet secure a binding mandate. METR’s experience shows that voluntary-access evaluation is still worth building, because the alternative is no independent check at all. Do not wait for a legal requirement to compel access before investing in the technical capability to use that access well once it exists.
- Decide now what you would do if your evaluation function found something serious. An evaluation body that has never been tested against a real bad finding is untested itself. Run a tabletop exercise this year: If your evaluators reported a genuine safety failure in a system already in production, what happens next, and who has the authority to pull it?
Every one of these steps is achievable without waiting for the next UN report or the next national AI strategy. The organizations that come through the next three weeks, like this one, that have already built the capacity to check their own claims will not be the ones caught explaining, after the fact, why they trusted a number nobody outside their walls could verify.
Conclusion
Forty frameworks and zero benchmarks was the panel’s diagnosis before Geneva even convened. Three weeks earlier, a real company had already supplied the case study, investigating its own flaw, grading its own fix, and asking the world to trust the number. The one outside check came only because a rival raised the alarm and a government reached for emergency powers. The question this report leaves for every government and every board is not whether to write framework 41. It is whether framework 41 will be checked by anyone who was not in the room when it was written, and by anyone who did not have to wait for a crisis to look.
