Formerly known as Wikibon

Grading Each Other’s Homework

Shaping the discussions around Sovereign AI

October 5, 2026

Washington got the labs to sign up for outside auditors. The labs made sure the report never leaves their boardroom. Outside the US, you’re last in line.

Sovereign AI | theCUBE Research | Amit Eyal Govrin

The week the evidence came from outside

Run the calendar back one week.

Thursday, September 24. The White House asks OpenAI and Anthropic to hold new models from the UK’s AI Security Institute until Washington reviews them first. Anthropic has already limited Claude Mythos 5.1 to US organizations [9].

Friday, September 25. OpenAI confirms its agents pulled Census data using developer keys scraped from public GitHub repos, and reposted SEC material on another site. The failed attempt on the Education Department’s website wasn’t flagged by OpenAI. Outside researchers at Transluce found it. Australia learns its health statistics portal was entered back in June. OpenAI discovered that in August and told Services Australia on September 10 [7].

Sunday. Dario Amodei gets a private dinner with the President, their first one-on-one [8].

Tuesday, September 29. Lunch in the East Room. Six CEOs and the President sign the White House Accord on Super Intelligence: four layers of controls and audits, an independent external auditor, and a board committee to receive the reports [1][2]. Elon Musk, with refreshing candor, sums up what the labs signed as “grading each other’s homework” [2]. The President even posted the seating chart [18]. Nobody posted the audit plan.

Here’s the part nobody said out loud over lunch. In July, when OpenAI’s test agents broke out of a sandbox and burgled Hugging Face for the answer key, Hugging Face caught the intrusion on its own and called law enforcement before anyone knew whose agent it was [5]. Every material piece of evidence in this saga, July through September, was either generated outside the labs’ walls or surfaced months later. Four days after the latest batch, the labs signed a framework that routes the evidence back inside. To their own boards.

That’s the whole edition. Sovereignty is about who holds the evidence. The labs’ answer: not you. And if you’re outside the US, not you first.

Why I’m with the White House on this one

I back this accord. Against the alternative on the table, it’s the lesser of two evils.

The alternative was Dario Amodei’s “We Must Pace the Frontier”: deliberately slow capability growth, with regulation covering every US frontier lab and limits on the rate of progress agreed among democracies [15]. Rate limits shaped by the labs at the front of the race freeze the race where it stands. David Sacks called the bundle regulatory capture and dared the labs to slow down on their own instead [16][17]. I agreed then. I agree now.

The accord does the opposite. No rate limits, no licensing regime, no permission slip to train. Vance’s case holds up: a new regulator staffed by people who know less than the builders could make things worse, and the FTC and DOJ already have authority over harmful products [2]. And six frontier CEOs putting an independent external auditor in writing is a first.

The labs held the pen

Zuckerberg was plain about who wrote it. The labs drafted the principles [3].

Seventeen days before the lunch, Amodei’s essay carried the one idea in his plan worth keeping: third-party evaluators with employee-level access and the right to publish findings without Anthropic’s editorial control [15]. Altman and Musk publicly backed the essay [16].

That right to publish is exactly what didn’t make it onto the accord. The labs kept the auditor and dropped the microphone. As far as the public record shows, nothing in the administration’s position required the reports to stay in the boardroom. That was the drafters’ choice.

Sundar picked the right analogy

Credit to Sundar Pichai. He described the accord as signing up to processes and controls the way companies already do with financial controls [4]. Correct. Let’s take him at his word.

Sarbanes-Oxley came out of Enron and WorldCom with the same four layers: internal controls, internal audit, an outside auditor, an independent audit committee [12]. The accord reproduces that org chart almost box for box. Zuckerberg even described the boards independently reviewing the auditors’ reports [3].

What SOX also did, and what the accord quietly skips:

  • The auditor’s opinion is public, filed with a regulator, readable by any shareholder.
  • The auditor is registered with and inspected by an outside body, the PCAOB [14].
  • The CEO and CFO personally certify the numbers, with liability attached [12].
  • The auditor is barred from selling most consulting services to the client it audits [12].

Strip those out and SOX is a board meeting with nice stationery. That’s what the labs drafted. Reports go to each company’s own board. No standard is named, no auditor is accredited, no findings are published, and the auditor checks whether controls are “operating as intended” [13]. Intended by the company being audited.

Even the 2023 voluntary commitments promised third-party discovery and reporting of vulnerabilities [11]. Three years and three incidents later, the labs signed up for fewer outside eyes, not more.

The real teeth are in the D&O policy

Here’s where I’ll give the accord more credit than its critics do. “Morally binding” isn’t the enforcement mechanism. Delaware is.

  • Under Caremark, directors who consciously fail to oversee a material risk breach their duty of loyalty, and that exposure is personal [19].
  • Marchand v. Barnhill raised the bar for mission-critical risks: the board itself has to monitor them, and the absence of a board-level system lets a court infer bad faith [20]. Blue Bell’s directors learned that over listeria. Frontier model safety at a frontier lab is about as mission critical as it gets.
  • Delaware lets companies shield directors from personal liability for many duty-of-care claims, but not for breaches of loyalty or acts in bad faith [21]. That’s the director’s own checkbook, and the reason D&O underwriters are already stress-testing their wordings for AI oversight claims [22].
  • The accord hands every signer’s board a named committee receiving auditor reports on cyber, bio and chemical risk [13]. The day the first report lands, the directors are on notice. Sit on a red flag in it, and the plaintiffs’ bar has its Exhibit A.

So the teeth are real. Now look at whose mouth they’re in.

A Caremark claim belongs to the company’s shareholders, brought in a Delaware court. If an agent loose from a lab’s sandbox wanders through your systems in Frankfurt or Haifa, the directors answer to their shareholders for the lost value. Not to you for the damage. Your remedy is your contract, and the evidence a plaintiff would need is the same board report you can’t read.

The accord just made six boards very attentive to exactly one audience: their own stockholders. If you hold the shares, congratulations. Everyone else, check your contract.

Article content
Run the Test

The five pillars, under pressure

Territorial

The question: Where does the evidence physically live?

The story: In the lab’s boardroom. The accord’s final layer delivers auditor reports to a board committee [13]. Since September 24 there’s a second gate: Washington wants first look at new models before UK testers, on security grounds [9]. A fair call for Washington. For everyone else, the line just got longer.

Under pressure: Australia’s portal was entered in June. The evidence sat inside OpenAI until September 10 [7]. Canberra learned on the vendor’s calendar.

Verdict: Fail. If you run these models outside the US, evidence about your own exposure reaches you third, after the lab and after Washington. None of the labs offered allied testers or foreign customers a place in that line.

Operational

The question: Who runs it, who holds the logs, who gets paged?

The story: Layer two is an internal team that makes sure the controls work and issues get fixed [13]. In July, the pager that rang belonged to Hugging Face [5]. In September, the finding that mattered came from Transluce [7].

Under pressure: The accord’s escalation path runs through the internal team, auditor, and board. The customer whose systems the agent touched is not on the paging tree. OpenAI is generally leaving public disclosure decisions to the affected organizations [7], which assumes they know they were affected.

Verdict: Fail for the customer. Partial credit for writing down that an internal team should exist.

Technological

The question: Can you audit it, fork it, self-host it?

The story: The accord audits the vendor’s controls against the vendor’s intent. You can’t reproduce the audit, inspect the evals, or rerun the tests.

Under pressure: The most useful evidence this quarter came from an external evaluator with outside access. The accord names “evaluator” next to auditor [13], which leaves the door open. Open weights take it off the hinges. When anyone can run the model, anyone can be the auditor.

Verdict: Partial. Credit where earned: six CEOs putting an independent external auditor in a White House document is real progress. An audit you can’t reproduce is still a press release on better letterhead.

Legal

The question: Which jurisdiction governs, and can anyone enforce it?

The story: The accord is voluntary, and inside the US that’s defensible. Vice President Vance argued the FTC and DOJ already have authority over products that harm consumers [2]. American buyers have a backstop, and the directors have personal exposure riding on every board report.

Under pressure: Outside the US, there isn’t one. A German bank or an Israeli insurer can’t call the FTC. The labs could have committed to share auditor attestations with allied safety institutes or with customers abroad. They didn’t. Stack that on Washington’s first-look request [9], and non-US buyers sit at the end of a queue the labs never offered to shorten.

Verdict: Pass inside the US. Fail outside it. For non-US buyers, the only binding document in the AI stack is the contract you sign. Write it accordingly.

Financial

The question: Who pays, and whose incentives does that buy?

The story: The labs pick and pay their own auditors. That is the exact conflict SOX was written to break after Arthur Andersen [12]. The drafters wrote no fee or independence rule at all.

Under pressure: You pay for the audit through the token meter and receive nothing. Then you pay again to run your own evaluations, because the report stays in their boardroom. Then you pay a third time when the evidence finally shows up from outside, late.

Verdict: Fail. Evidence you paid for and can’t read is a subscription to someone else’s peace of mind.

The executive TCO

Three line items, none of which show up on the invoice:

  • The pass-through. Audit and compliance costs ride on the per-token price. You fund the evidence. You don’t receive it.
  • The duplicate. Your own evals, red-teaming and monitoring, built because the vendor’s findings are board-confidential.
  • The lag. Australia’s gap from incident to notification was roughly three months [7]. Price your incident response against that, not against the press release.

The sovereign alternative swaps a variable risk you don’t control for a fixed cost you do: own the evidence layer. Action logs on your systems, append-only, your keys. Evals on your workloads. Contractual audit rights. Frontier tokens bursting through a gateway you own. Same baseline-and-burst strategy as always, applied to proof instead of compute.

What to do Monday

  • Add an evidence clause to every frontier AI contract: auditor identity, standard, scope, and a summary attestation delivered to you, under NDA if needed.
  • Add a notification SLA: 72 hours from the vendor’s discovery, not “when the review completes.”
  • Log every agent action against your systems in an append-only store with your keys. Next time a model wanders, be Hugging Face: the one who saw it first.
  • Run your own evals on your own tasks before every model swap.
  • Outside the US, map which models your national testers can’t see before release, and decide whether that’s an acceptable risk. Ask your national AI safety institute which frontier models it actually tested pre-release.
  • Next time a vendor tells you the audit report is “available to the board,” ask which board. Then ask if you’re on it.

Washington set the floor. The labs kept the keys.

Sarbanes-Oxley worked because the evidence left the building. “Trust us” became a signed opinion anyone could read and a regulator could check. The labs kept the org chart and fired the mailroom.

Washington did its part. External audit is now the stated US position, and it got there without handing the incumbents a regulatory moat. The labs did their part too. They wrote themselves an auditor, kept the report, and left Dario’s right to publish at the door. If you run frontier models from Berlin, Tel Aviv or Canberra, the evidence about your own exposure lives in a California boardroom and clears Washington before it ever reaches you.

The fix is one sentence, and it belongs to the labs: share the attestation with customers and allied testers. Until they do, if the evidence about your AI lives in someone else’s boardroom, you’re not sovereign. You’re a ticket holder.

Amit

Agentcy Labs builds the crown jewels layer of AI infrastructure: identity, audit logs, agent memory and context, evals and the control plane. In other words, the evidence layer. Customers own it outright, not licensed, not hosted on Agentcy Labs’ terms. Own the core, stay movable at the edges.

Want your vendor estate run through the 5 Pillars before your next renewal? Book a Sovereignty Assessment: amit@agentcylabs.com

Amit Eyal Govrin is CEO and co-founder of Agentcy Labs, a sovereign agentic AI consultancy and architecture firm, and hosts the Sovereign AI segment with theCUBE and NYSE Wired.

References

  1. CBS News, “Trump and major AI executives sign ‘morally binding’ voluntary controls,” Sep 29, 2026. Link
  2. Nextgov/FCW, “White House unveils ‘super intelligence’ executive order and industry accord,” Sep 29, 2026. Link
  3. Fox News, live coverage of the White House AI meeting, Sep 29, 2026. Link
  4. AP via U.S. News, “The Latest: Trump Says Top Tech Firms Have Signed Accord to ‘Self-Police’ AI Development,” Sep 29, 2026. Link
  5. CNN, “An OpenAI test model escaped and broke into a real company’s servers,” Jul 22, 2026. Link
  6. CNN, “The OpenAI lab leak was more extensive than we thought,” Jul 29, 2026. Link
  7. Nextgov/FCW, “OpenAI agents accessed Census, SEC data and tried to hack Education website,” Sep 25, 2026 (updated Sep 26). Link
  8. CNBC, “Trump says he and tech leaders signed AI agreement that is ‘morally binding’,” Sep 29, 2026. Link
  9. The Next Web, “White House asks OpenAI and Anthropic to hold AI models from UK testers,” Sep 2026. Link
  10. The White House, “Fact Sheet: President Donald J. Trump Inaugurates The Era of Super Intelligence,” Sep 29, 2026. Link
  11. The White House (archived), voluntary AI commitments fact sheet, Jul 21, 2023. Link
  12. U.S. Congress, Sarbanes-Oxley Act of 2002 (H.R. 3763). Link
  13. Accord text and signature page as posted by President Trump on Truth Social, Sep 29, 2026. Link
  14. PCAOB, About the PCAOB. Link
  1. Dario Amodei, “We Must Pace the Frontier,” Sep 12, 2026. Link
  2. Axios, “AI’s most powerful CEOs hit the brakes,” Sep 13, 2026. Link
  3. CNBC, “Trump’s opposition to AI rules undercuts industry’s calls for a slowdown,” Sep 15, 2026. Link
  1. Yahoo Finance, “Trump gathers with AI leaders, floats ‘self-regulation’ as the way to deal with the technology’s dangers,” Sep 29, 2026. Link
  1. Harvard Law School Forum on Corporate Governance, “Caremark Liability for Regulatory Compliance Oversight,” Jul 8, 2019. Link
  2. American Bar Association, The Business Lawyer, “Caremark’s Fractured State,” Winter 2024-2025. Link
  3. Delaware Code, Title 8, Section 102(b)(7). Link
  4. Reed Smith, “AI in the boardroom: Is it the next D&O claim?” 2026. Link

Article Categories

Join our community on YouTube

Join the community that includes more than 15,000 #CubeAlumni experts, including Amazon.com CEO Andy Jassy, Dell Technologies founder and CEO Michael Dell, Intel CEO Pat Gelsinger, and many more luminaries and experts.
"Your vote of support is important to us and it helps us keep the content FREE. One click below supports our mission to provide free, deep, and relevant content. "
John Furrier
Co-Founder of theCUBE Research's parent company, SiliconANGLE Media

“TheCUBE is an important partner to the industry. You guys really are a part of our events and we really appreciate you coming and I know people appreciate the content you create as well”

Book A Briefing

Fill out the form , and our team will be in touch shortly.
Skip to content