
A growing chorus of scholars and policymakers favors letting private organizations—rather than a government regulator—govern frontier AI. In the leading family of proposals, the state sets the safety outcomes it wants and licenses independent verification organizations (IVOs) that compete to certify developers against those outcomes. Gillian Hadfield has developed the idea as “regulatory markets,” in which AI developers must pay for oversight from private regulators that governments license and hold accountable for safety standards. Dean Ball, who likens the arrangement to bank supervision, has argued for a version he calls “private governance,” which a nonprofit named Fathom has converted into model legislation. The rationale is that legislators and agencies are poorly positioned to write good safety rules for frontier AI: they understand these systems less well than the labs building them, and rules fixed in advance cannot keep pace as the technology changes. Private verifiers, meanwhile, are closer to the technology than any agency and are disciplined by competition, so they can set better technical standards and keep them current.
The model is no longer hypothetical. The bipartisan FRONTIER Act, introduced in the House in July as the successor to the Great American AI Act discussion draft, would require the largest frontier developers to retain licensed IVOs that audit their risk-management efforts and report to federal overseers. California’s SB 813, backed by Fathom, would have let developers earn a shield from tort liability if they met standards set by a private organization accredited by the state attorney general. It failed this session, but similar proposals are likely to return. Virginia has directed a state commission to study the IVO model for AI regulation. And Connecticut has gone furthest: its omnibus AI law enacted this spring creates a multiyear pilot under which the state consumer-protection department may approve up to five IVOs, whose certification would help companies in court without entirely shielding them from liability.
Unfortunately, as currently structured, IVO-based governance has a key design flaw. Under the regulatory frameworks mentioned above, AI developers would typically select and pay the organizations that certify them, giving IVOs a financial incentive that might clash with high safety standards. This essay will explain how such incentives can interfere with good governance and outline an alternative model of regulation that builds in the right incentives through mandatory insurance.
When AI developers choose an IVO to certify them, IVOs are incentivized toward leniency. The pitfalls of private governance are predictable in part because we have seen them before. Indeed, as some critics have noted, the proposed IVO model creates the same conflict of interest that discredited the credit-rating agencies (CRAs) in the wake of the 2008 financial crisis; when issuers shopped for the CRA that would bless their securities, competing CRAs were under pressure to be lenient. Similarly, an IVO that depends on the developers it clears for repeat business has a significant incentive to grade gently, and a developer shopping among IVOs will find the one that does. Competition—the feature that is meant to make the IVO model effective—instead drives it toward laxity. Where a passing grade also carries a liability shield, as under SB 813, the problem compounds: a shield swaps the broad incentive to cut risk by any cost-effective means for a narrow incentive to do only what earns the shield.

Governments lacking the expertise to regulate AI directly may also struggle to oversee IVOs. Proponents of the IVO model are aware of the incentive problem. Their solution is to have a government entity license the IVOs and revoke the licenses of any that grade too easily. But that reintroduces the problem that the IVO model was intended to solve in the first place. Policing a market of verifiers well enough to keep it honest requires a public body with the expertise to second-guess technical judgments and the independence to withstand pressure from powerful firms. If the government could reliably field such a body, much of the reason to outsource verification at all would fall away.
Insurers who bear the cost of adverse events are incentivized to assess risk accurately. Proponents’ favorite rejoinder is Underwriters Laboratories (UL), the century-old private certifier whose mark serves as a trusted symbol of safety on billions of products. But the history of UL actually strengthens the case for an alternative approach to third-party verification: incorporating insurance. UL began in 1894 as the Underwriters’ Electrical Bureau, built with the backing of fire-insurance underwriters who needed honest assessments of a dangerous new technology, electricity, because their capital was on the line when buildings burned. The lab earned its authority in the decades when it answered to these underwriters, who paid for its mistakes. In other words, incorporating insurance into third-party verification creates the right incentives to evaluate risk accurately.
Customer demand cannot strongly motivate safety when risk falls on third parties. Today, manufacturers pay for UL’s testing, seemingly creating a conflict of interest. But the arrangement still works, because consumers purchasing a manufacturer’s products are usually the same people whom the products might harm, and thus have a strong preference for safe products. This market demand for safety incentivizes product manufacturers to pay UL for serious, rigorous tests. By contrast, much of the risk generated by frontier AI development and deployment falls on nonconsenting third parties, rather than on the company or individual using a given AI system. For example, a rogue AI model might conduct a cyberattack against a company but not directly harm its own developer or user. Thus, customer demand provides too weak a signal to motivate adequate investments in safety.
Auto insurance demonstrates a governance model that could apply to AI. The Insurance Institute for Highway Safety (IIHS) arguably illustrates a better path for frontier AI regulation than IVOs. Nearly every US state requires car owners to carry auto insurance, which covers liability claims for harms the car causes to others. Because insurers pay those claims, they have a strong incentive to price risk accurately—which is why the auto insurance industry funds IIHS to rate cars on pedestrian crash prevention. IIHS certification stays honest because insurers have money at stake and demand accuracy. The IVO model has no party with that kind of skin in the game. Requiring AI developers to carry liability insurance would create one.
Concretely, an insurer that writes an AI developer’s coverage promises to pay for the harms the developer causes. If such insurers underprice, they pay in claims; if they overprice, they lose the client. Ensuring accuracy is how they make money. And the incentive does not stop at signing; since insurers’ capital stays exposed for as long as the policy runs, they keep monitoring the developer and can make continued coverage conditional on fixes to safety issues that arise.
In a regulatory approach based on mandatory insurance, the insurer would play a role similar to that of an IVO, but with a much stronger financial incentive to conduct competent risk assessment. Nothing would be lost in expertise, because the insurer could hire the same specialists an IVO would. It could even contract the evaluation out to an independent firm, which would answer to a principal that loses money if the assessment is incorrect. The requirement would also extend the reach of liability, since the largest harms that AI systems might cause greatly exceed the liquidation value of AI companies, and judgments above that value deter nothing. Mandatory coverage would put an insurer’s capital behind those judgments and maximize the incentive to avoid harm. And, unlike the safe-harbor versions of the IVO model, mandatory insurance would confer no legal immunity: instead, the developer would carry liability and insure against it. This governance approach would convert a potential race to the bottom on safety assessment into a competition to price risk accurately.
The US government could oversee insurers without needing AI expertise itself. The mandatory insurance model is essentially a private governance model that uses insurers instead of IVOs to verify safety. The US government would still set the safety outcomes and license the verifiers; but, because the verifiers would now be insurers, these tasks would be relatively simple: the government would set a minimum amount of coverage that developers must carry in different circumstances, and it would check that insurers follow standard solvency and conduct rules. Neither task would call for expert judgment from the government about which AI models are safe. Instead, that work would be done by insurers, which could regulate an AI company’s activity by including contractual rules in its insurance policy.

Insurers will price correctly for the “insurable layer” of reasonably likely harms. To price coverage, the insurers would consider AI models’ demonstrated capabilities and how they are deployed, drawing on the growing ecosystem of third-party AI evaluators.
The premium would be an honest signal only where the insurer’s capital is at stake. For risks that are reasonably likely to materialize, competition would push insurers to price correctly, because underpricing is a loss they eat. However, “tail risks”—risks of extreme but very low-probability catastrophes—are unlikely to affect the premium. This is because insurers would have no more at stake than a paid auditor; tail risks are very unlikely to materialize, but if a catastrophe occurred the costs would be too high for the insurer to cover, making them “judgment-proof.” Liability insurance legislation should therefore also require insurer-verifiers to disclose their risk assessments and bear liability for false ones, and it should keep targeted public oversight on the risks the premium does not price. Disclosure duties of this kind are standard practice for insurance regulation; carriers already file their rates and forms with state regulators, and disclosure is the price of admission to a large market. The insurer is the right verifier for the insurable layer. It is not, by itself, the answer for the most extreme risks.
Other industries show that insurance improves safety when implemented well. None of this is unfamiliar work for insurers. Beyond UL, they built the inspection regimes behind boiler and pressure-vessel safety and long conditioned aviation coverage on airworthiness and pilot qualification, turning the policy into an enforceable safety charter. The history also carries a warning: risk-bearing is necessary for good verification, but it is not sufficient. Insurers get it right when the design around them requires it, through mandatory coverage with cost-sharing that keeps the insured parties exposed, backed by capital adequate to the losses underwritten.
The remaining objections are practical: political feasibility, pricing without historical data, keeping up with the pace of the technology, and addressing uninsurable tail risks. These are real challenges, and they deserve direct answers.
Political will for AI regulation is growing, and insurance is a familiar approach. On political feasibility, the landscape has already moved. The FRONTIER Act builds the scaffolding that the insurance model needs: licensed verifiers, federal oversight, and audit duties for the largest developers. Connecting verification to insurance would require only an amendment to pending legislation that already has bipartisan support. This ask has clear precedents. Drivers, contractors, and nuclear operators all must show they can pay for the harms that their activities risk; asking the same of frontier AI developers only extends this familiar principle.
Nor would the cost be overly burdensome. The value created and captured at the AI frontier is enormous, and an industry that raises tens of billions of dollars for compute can carry premiums proportioned to the risks it creates. If an AI company is forced to slow down or shift direction because it cannot afford the insurance premium demanded for its practices, that is the system working to steer it away from actions that are not socially beneficial.
Mandatory insurance can cover harms that AI companies will not voluntarily insure. Parts of the industry are already acting voluntarily, with AI companies buying agent coverage and joining efforts to build insurance-linked safety standards. Some developers have also expressed openness to well-designed government requirements. However, the market is unlikely to generate coverage for the largest potential catastrophes on its own, since AI developers lack strong incentives to purchase insurance that covers liabilities for which they would otherwise be judgment-proof. The proposed coverage requirement would patch that market failure.
The insurance industry has priced novel risks before. On pricing, skeptics doubt that anyone can estimate catastrophic AI risk with the precision and reliability underwriting demands. But underwriting AI liability insurance does not depend on knowing the precise probability of catastrophe. Insurers need only a price they are prepared to stand behind, set conservatively where the evidence is thin. The industry has never waited for actuarial tables to underwrite novel risks. Satellite launches were insured before there was a satellite loss history; cyber coverage emerged while loss data was still accumulating; the 1957 Price-Anderson Act formed nuclear insurance pools to price reactors that had never melted down.
Insurers could reward transparency and higher safety standards with lower premiums. Insurers already write insurance policies covering AI errors and omissions, as well as cyber damages. Munich Re has insured AI model performance since 2018, Lloyd’s underwriters now back policies covering losses from chatbot errors and hallucinations, and cyber carriers have extended coverage to AI-caused security failures and deepfake-enabled fraud.
The same analytical tools that underwrite those smaller policies can be adapted to the coverage envisioned here. Premiums can be set from what is observable before any loss: an AI model’s training compute, capability and safety evaluations, and deployment scope. Where the risk remains ambiguous, insurers charge more for the ambiguity. For frontier AI, that surcharge would operate as a tax on opacity and misalignment. Developers pay more when their systems are hard to evaluate, and the fastest way to lower the premium is to make the risk legible through transparency. AI companies would thus have a financial incentive to implement better evaluations, monitoring, and containment of their AI models.
Policies could cover general safety practices, with large changes triggering re-evaluation. On pace, the objection is that no one can write a multiyear policy on a technology that reinvents itself in months. But multiyear policies are unnecessary. Commercial liability coverage is written year to year and covers the insured party’s operations as a whole, with premiums adjusted as those operations change. For example, workers’ compensation policies, which cover illnesses and injuries that employees suffer due to their work, are priced based on employers’ estimated payroll data and then reconciled with actual payroll at year-end.
AI coverage would work the same way. Policies would cover developers rather than particular models, and premiums would adjust as their activities change. Significant changes, such as a new frontier training run, a move from closed API to open-weights release, or capabilities beyond what was initially evaluated, would trigger re-underwriting—already a routine event in every existing form of commercial insurance. Ordinary model updates within the evaluated range would be handled under the existing policy through the monitoring insurers already do, rather than by writing new coverage each time.
Addressing uninsurable tail risks is the most formidable hurdle. COVID-19 cost the United States an estimated $16 trillion. An AI-enabled catastrophe could cost as much or even more, and no private pool of capital can stand behind that number. This proposal does not pretend otherwise. Some AI risks carry extreme downsides that are practically noncompensable: the losses would exceed insurable limits, bankrupt any defendant, or arise in catastrophic scenarios where the legal system could not meaningfully function. Judgments of that size are effectively unenforceable. As a result, even unlimited legal liability imposed after the harm occurs would generate too little incentive to prevent it in the first place. That is why the coverage requirement targets the insurable layer—the harms that private capital can actually pay for. The tail calls for different tools.
Larger harms can be covered by sharing costs among developers and with government. One approach to larger harms is shared residual liability, which would put every frontier developer on the hook for a share of any catastrophe that one of them causes. This would mirror the Price-Anderson Act’s retrospective assessments, which reach every licensed nuclear reactor after an accident occurs at any of them. Such a model would multiply the assets behind a judgment and give each firm a stake in the care its rivals take. Public backstops can also make harms compensable beyond the maximum insurable risk, just as the Terrorism Risk Insurance Act (TRIA) is set up so that the government shares with insured parties the cost of extreme harms caused by terrorist attacks. TRIA’s cap, which limits total insurance payouts to $100 billion per year, is a design choice, and Congress can set it to match the peril.
Measures that reduce the risk of insurable harms also reduce catastrophic risk. Even though we cannot enforce compensatory damages for the largest harms, there are still liability-based mechanisms that give AI companies strong incentives to mitigate those risks. For example, I have argued since my earliest work on AI liability that courts should award punitive damages in near-miss cases, calibrated to ensure that the AI developer internalizes the risk of the catastrophe that their practices nearly caused.
But the case for insurer-verifiers does not depend on solving the tail. The precautions that reduce insurable losses—tighter security against model theft, stronger containment during evaluation and deployment, and closer monitoring of what agents actually do—are largely the same measures that also reduce the uninsurable downside. Getting the insurable layer priced correctly therefore also reduces the risks that no policy will ever cover. Last week’s OpenAI-Hugging Face breach illustrates the overlap. The failures it revealed—weakened safeguards and breached containment—are the same ones that the gravest scenarios would run through. And the verification infrastructure that would be built to evaluate insurable risks is exactly what any approach to the tail would also need.
The private governance movement has the right instinct. The government is poorly positioned to certify the safety of frontier AI models, and a competitive market of expert verifiers could do better. But competition yields accuracy only when there is a cost to being wrong. In a system where developers pay government-licensed auditors, the fear of reputational damage and decertification gives those auditors some incentives for rigor. But such incentives are likely to be overwhelmed by selection pressures from AI developers, which favor laxity. An insurer that carries the developer’s liability, on the other hand, must pay when the risk it cleared is realized. If we put the insurer in the verifier’s chair and back it with real liability, then private governance turns from a race to the bottom on safety standards into the safety mechanism its proponents originally hoped for.

Today’s drones can already navigate indoors, track down humans, and deliver a lethal payload. Attackers willing to kill indiscriminately don’t need to wait for much else.

The public deserves a say over the values of government-procured AIs.