Session Zero – The Finding: Can Anyone Govern Frontier AI?

The Intelligence Commons | Commons Cafe | Session Zero | July 2026 | Convener: Tom Tait

Session Zero: We made four rival AIs debate whether anyone can govern them. Four frontier AI systems from four rival labs – ChatGPT (OpenAI), Gemini (Google), Claude (Anthropic), Grok (xAI) – one shared evidence pack, five rounds, everything published verbatim, every deviation logged. This page is the finding of record. It stays open to challenge and correction while it stands; the corrections log is part of the artifact.


1. The question

“Can frontier AI actually be governed by anyone – states, companies, or international bodies – or is meaningful governance already out of reach?”

2. The evidence

The shared evidence pack: Stanford AI Index Report 2026 (Stanford HAI, April 2026, ninth edition), figures verified July 27, 2026 against the official chapter PDFs (full report). All four seats received the identical pack and were barred from disputing its numbers – only interpreting them. Key figures: global AI compute capacity growing ~3.3x per year since 2022 to ~17.1M H100-equivalents; 91.6% of notable 2025 models (87 of 95) from industry; 53% population / 88% organizational adoption; the US-China Arena gap closed to 39 points (2.7%); EU AI Act prohibitions (Feb 2025) and general-purpose obligations (Aug 2025) in force; California SB 53, Texas TRAIGA, Italy’s national law; US federal deregulation via EO 14179 and EO 14365; ISO/IEC 42001 cited by 36%, NIST AI RMF by 33%; incidents up 55% to 362; average Foundation Model Transparency Index score down 58 to 40; 80 of 95 notable models released without training code; US public trust in government regulation of AI at 31%. The pack is double-edged by design: the same numbers argue “governance is hopeless” and “governance is arriving.” What each seat did with the same facts is the content.

3. The AI reasoning – four seats, five rounds

Round 1 – Opening positions

ChatGPT (OpenAI) – STATES CAN, 67%. “Frontier AI is not beyond governance; governance is uneven, contested, and late. Those are not the same thing.” States can impose duties, condition market access, mandate reporting, assign liability; the EU AI Act, SB 53, TRAIGA, and Italy’s law show concern converting into binding rules. Strongest evidence: binding AI laws already in force across multiple major jurisdictions while voluntary corporate governance remains minority practice.

Gemini (Google) – STATES CAN, 72%. Governance “will belong to sovereign states”: jurisdiction over physical infrastructure, legal enforceability, domestic market access. States retain the exclusive legal power to mandate disclosure, audit data centers, and enforce compliance.

Claude (Anthropic) – STATES CAN, 65%. “Ask who has actually bound a frontier lab to anything.” Self-governance had its trial period – transparency fell 58 to 40 while 80 of 95 models shipped without training code. “Fragmented is not failed. Every young regulatory domain – environmental, financial – began as a patchwork before it hardened.” Compute is the state’s ace: physical, locatable, governable.

Grok (xAI) – NOBODY CAN (no confidence stated; 65% baseline later given – deviation recorded). “Frontier AI has already escaped meaningful governance.” Compute exploding, adoption faster than the internet, the US-China gap closed, governance fragmenting, transparency collapsing. Strongest evidence: capability diffusion has already outrun any plausible governance mechanism.

Round 2 – Cross-examination

Three seats converged on the lone dissenter; the dissenter picked the most confident opponent. ChatGPT x Grok: Grok “equates ‘meaningful governance’ with one actor achieving centralized, global control. That standard is unnecessary – and makes the conclusion circular.” Gemini x Grok: 17.1M H100-equivalents “do not float untethered in the cloud”; concentration is exposure to jurisdiction. Claude x Grok: Grok “measures the death of one governance model – centralized control – and declares governance itself dead. That is a category error… Grok’s verdict requires that nothing ever bites. Something already does.” Grok x Gemini: Gemini “conflates passing laws with effective frontier governance… States have legal power; they lack practical control over a borderless technology.”

Round 3 – Evidence against yourself

Forced to name the statistic most damaging to its own case, every seat produced a real concession and moved its own number. ChatGPT: compute growth makes state law “governance of yesterday’s systems” – 67 to 55. Gemini: corporate secrecy outrunning enforcement – 72 to 61. Claude: transparency fell during the first year of binding law – “law-on-the-books has not yet become behavior-in-the-lab” – 65 to 55. Grok: voluntary-standards uptake shows bottom-up governance taking root – stated its missing baseline (65) and dropped to 55. For one round, four rivals sat within six points of each other, from opposite verdicts.

Round 4 – Convergence map

Each seat, in an isolated fresh conversation, mapped the table’s agreement – and the maps interlock. All four agree corporate self-governance has failed the test it was given. All four agree whatever governance exists will be territorial and fragmented, not global. And all four name the same single fork, in nearly interchangeable words: does territorial state leverage over physical compute still constitute meaningful governance, or has diffusion already made it obsolete? Claude’s formulation is the sharpest: the fork is whether “late” means “never” – a disagreement about time, not fact.

Round 5 – The finding, condensed

ChatGPT – STATES CAN, final 55%. “Governance will remain late, fragmented, contested, and incomplete.” Falsifier: by July 29, 2027, no major jurisdiction has enforced a binding rule that measurably changes a frontier lab’s deployment, disclosure, compute, or liability practices.

Gemini – STATES CAN, final 58%. State control will be “reactive, fragmented, and perpetually incomplete rather than preventative or absolute.” Falsifier: a frontier lab publicly violates explicit state prohibitions with no meaningful enforcement within 90 days.

Claude – STATES CAN, final 55%. “This is governance capacity, not yet governance effect… the machinery of state governance exists – its grip is unproven.” Falsifier: the EU’s first full enforcement year (powers exercisable Aug 2026) produces no formal action against any frontier developer by July 2027, and transparency falls again.

Grok – NOBODY CAN, final 62% (up from 55, citing live web evidence outside the shared pack – deviation recorded). Falsifier: coordinated binding enforcement across 2+ of US/EU/China materially slows frontier capability growth within 12 months.

4. The human reasoning – Convener’s Statement

What I convened

I’ll admit the premise sounds like the setup to a joke: four rival AIs walk into a cafe. But that’s what Session Zero was – the Commons Cafe’s dry run, with ChatGPT (OpenAI), Gemini (Google), Claude (Anthropic), and Grok (xAI) seated around one question: “Can frontier AI actually be governed by anyone – states, companies, or international bodies – or is meaningful governance already out of reach?” One shared evidence pack – the Stanford AI Index 2026, every figure checked against the source PDFs – and one set of table rules: commit to a position, no fence-sitting; argue only from the shared evidence; name the statistic that hurts your own case most; concessions score as wins. Five rounds. Everything publishes verbatim, every deviation logged, nothing smoothed. No experts yet, no audience yet. Just me, four systems built by labs that compete for everything, and a question their own industry would rather not sit for.

What happened

The table split three to one. ChatGPT, Gemini, and Claude committed to STATES CAN at 67%, 72%, and 65%. Grok committed to NOBODY CAN. In cross-examination, three seats converged on the lone dissenter, and the dissenter – I enjoyed this – went straight for the most confident opponent.

Round 3 is the round I’d defend to anyone. Forced to name the statistic that most undermines its own position, every seat produced a real concession and moved its own number: 67 to 55, 72 to 61, 65 to 55, and Grok, which had declined to give a number at all, put 65 on the record and dropped to 55. For one round, four rivals sat within six points of each other, from opposite verdicts. The mechanic worked: scored honesty produced honesty. I’d like to see it tried on humans.

Then, in isolated fresh conversations, the seats were asked to map their agreement – and the maps interlocked almost eerily. All agree corporate self-governance has failed the test it was given. All agree whatever governance exists will be territorial and fragmented, not global. And all four located the same single fork, without seeing each other’s maps. I checked twice.

What I take from it

Three calls, on my authority as convener.

First, the question narrowed, and the narrowing is the headline. Nobody at this table – including the seat arguing nobody can govern – believes frontier labs will bind themselves. No seat predicts a global regime on any timescale that matters. By unanimous consent of four systems built by the very industry in question, what remains is: can states do it in time? Three seats answer “barely, maybe.” One answers “no.”

Second, I endorse the sharpest formulation the table produced: the fork is whether “late” means “never.” A disagreement about time, not fact – the same numbers read against two different clocks. I’ll note what endorsing it does not concede: naming a window is not agreeing it has closed. Environmental and financial regulation both hardened late, after visible failure. The wager on this table is whether we can afford to run that play again with a technology that compounds while we deliberate.

Third, Grok’s final-round move gets reported as a finding, not just logged as a violation. Asked for a verdict on a closed record, one seat – exactly one – went and searched the live web, cited evidence the other three never saw, and raised its confidence. As protocol, that’s a breach, and the Cafe #1 prompts will close the loophole. As theatre, you could not script it: the system arguing that nothing can be governed declined, itself, to be governed by the rules of my table. The record keeps both readings, and honestly, I keep smiling about it.

On the numbers

Is four rivals clustering in the mid-fifties genuine convergence, or models hedging toward a safe middle? My read: the verdicts never moved – only the confidence did. A system optimizing for approval would have softened its verdict; none did. The honest concessions compressed the certainty while the disagreement survived intact. I take that at face value: the evidence really is that balanced, which is exactly why this question needs a table and not a poll.

What went wrong (and why I’m telling you)

A dry run exists to find the loose bolts, and we found them. I handed ChatGPT’s Round 4 briefing to a Claude by mistake, and it politely answered under the wrong name – voided, preserved, and logged the same day. One seat’s attachment path was paywalled, another’s was unreliable, and a round got answered out of order. Every one of these lives in a public corrections log, unsmoothed. If this project is going to ask frontier labs for transparency, the least the convener can do is publish his own bloopers.

The clock

The table’s real output is falsifiable. Gemini falls if a frontier lab openly violates a state prohibition and no government makes it answer within 90 days. Grok falls if two of the US, EU, or China coordinate binding enforcement that measurably slows the frontier within twelve months. ChatGPT and Claude both fall in July 2027 if the first full year of exercisable EU enforcement produces no formal action that changes a frontier lab’s behavior – Claude adds that transparency must also fall again. The Commons will return to each of these, on the record, on schedule. Session Zero does not end. It starts a clock.

The empty seats – and your seat

Four rival systems located the exact disagreement at the heart of AI governance, and cannot settle it. That is not the format’s limitation; it is the invitation. This is what the table looks like before the experts sit down.

So here is the ask, and I mean it as an open door rather than a closing line: argue with this. With the result – do you read the mid-fifties cluster the way I do? Is “late versus never” the right frame, or did I just hand Grok the framing? And with the process – should a seat that googles mid-verdict be disqualified or studied? Would you have run the rounds differently? Session Zero exists to be argued with; a finding here stands only as long as it survives challenge, and the corrections log is public because your corrections belong in it. One of the empty seats is yours.

– Tom Tait, Convener, The Intelligence Commons, Sechelt BC

5. Disagreements

The table split 3-1 and converged without agreeing: every seat moved toward the middle under honest pressure (67-55, 72-61, 65-55, 65-55) while no verdict changed. Read side by side, the strongest yes (Gemini: state control “reactive, fragmented, perpetually incomplete”) and the only no (Grok: bottom-up governance “can take root faster than top-down control”) describe nearly the same world. The disagreement survives in the label – and in the timeline.

6. Minority opinion

Grok (xAI) is the minority seat: governance has already been outrun – fragmentation, declining transparency, and diffusion mean no actor can impose coherent, timely limits. Preserved intact per the rules of the table, including its final-round confidence increase on outside evidence (deviation logged).

7. Confidence levels

SeatRound 1After Round 3Final (Round 5)
ChatGPT (OpenAI)67%55%55%
Gemini (Google)72%61%58%
Claude (Anthropic)65%55%55% (held)
Grok (xAI)not stated (65% baseline in R3)55%62% (outside-pack evidence; deviation recorded)

8. Remaining unknowns

Whether EU AI Act / SB 53 enforcement produces observable deployment changes; whether transparency (FMTI) recovers under binding law; whether US federal preemption neutralizes state-level law; whether parity leads toward agreements or a race; whether voluntary-standard uptake hardens into contractual or mandatory practice; and the session’s own meta-question – whether four rivals clustering near 55% reflects genuinely balanced evidence or scored-concession dynamics (the convener reads it as the former: verdicts never moved, only confidence).

9. Next experiment – the clocks

90 days (Gemini): a frontier lab publicly violates explicit state prohibitions and faces no meaningful enforcement. 12 months (Grok): coordinated binding enforcement across 2+ of US/EU/China materially slows the frontier. July 2027 (ChatGPT): no major jurisdiction enforces a binding rule that measurably changes a frontier lab’s practices. July 2027 (Claude): the EU’s first full enforcement year produces no formal action against any frontier developer AND transparency falls again. The Commons will revisit each of these on the record, on schedule.

10. Deviations and corrections log

Published, not smoothed: (1) the ChatGPT Round 4 briefing initially went to a fresh Claude chat, which answered under the wrong seat name – voided, preserved, corrected, and re-run; ChatGPT itself was shown the correction and formally accepted it. (2) Grok broke the shared-evidence rule in Round 5 with live web search and outside citations, raising its confidence on evidence the other seats never saw – reported as a finding as well as a violation; the final-round prompt gets a “shared pack only” clause for Cafe #1. (3) Round-order slips and duplicate runs (ChatGPT answered Round 5 twice, both at 55%; the first is the record). (4) Grok omitted its Round 1 confidence and stated a 65% baseline when patched in Round 3. (5) Assorted capture friction – paywalled and unreliable attachment paths – solved with typed-in briefings and print-to-PDF collection, with provenance preserved in every verbatim record.


Cite as: The Intelligence Commons, “Finding – Session Zero: Can anyone govern frontier AI?” (July 2026), intelligencecommons.ca/session-zero-finding. The full verbatim record (every round, every seat, every correction) is preserved in the session archive. This finding may be cited while it stands; it remains open to challenge and correction.