The Conclave
How it works
The process

Six labs. Seven stages. One decision.

The Conclave doesn't ask one model — it forces six of them to argue, anonymized to each other, until a measurable majority converges on one answer. Here is what happens between the moment you hit submit and the PDF you get back.

The Process — ~30 minutes, fully visible while it runs
1. Briefing
The question and any attached documents are delivered to all six panelists at once.
2. Independent submissions
Each model writes its answer in isolation, before seeing the others.
3. Anonymized debate
Submissions are stripped of authorship and exchanged. Across multiple rounds, models challenge the strongest claims in each other's work and defend, rebut, or concede.
4. Live fact-checking
After round one, an off-panel model with web access searches for counter-evidence to each panelist's strongest claims and feeds it into the next round.
5. Judges deliberate
A separate panel of judge-model instances reads the full transcript and votes (Affirm / Remand / Dismiss).
6. Fresh-eyes audit
A model that wasn't part of the debate audits the consensus for groupthink.
7. Decision rendered
Names are de-anonymized; the final document — The Question, The Answer, The Reasoning, each model's closing statement, and Notable items — is written and downloadable as a PDF.

The room doesn't open until the panel's measured agreement score crosses a threshold. In practice that's by round four or five. If a model never agrees, its dissent is preserved in its closing statement.

Model selection

Six labs, one room.

Each panelist comes from a different lab — different training data, different priorities, different post-training reflexes. They genuinely disagree. When they still reach a majority, the answer survived the disagreement that produced it.

The Panel — each seat earns its keep
Claude Opus 5Anthropic
Long-form rigor. Holds positions under pressure without over-conceding.
GPT-5.6 SolOpenAI
Step-by-step logic. Best at finding holes in the others' arguments.
Gemini 3.1 ProGoogle
Real-time web search. Brings facts the others have to reason about.
Grok 4.5xAI
Less hedged than the western trio. Willing to take and defend a contrarian line.
DeepSeek V4 ProDeepSeek · open-weight
Frontier reasoning trained outside the US labs. Different training distribution; different defaults.
Magistral MediumMistral AI
European reasoning model. Different training lineage and deliberation style from the US labs — best when you want a non-US default in the room.
+ Fact-checker — off-panel, kicks in after round 1
Perplexity Sonar ProPerplexity · Llama 3.3 70B + live web
Different shape than the others — a fine-tune of Meta's open Llama 3.3 70B paired with a live web index and citation engine. After round 1 it searches the web for counter-evidence to each panelist's strongest claims; the rebuttal brief is fed into the next debate round.

Frontier-tier only. No smaller substitutes when an API is unavailable — the seat goes empty rather than getting filled by a weaker model.

Start Gallery
Sign in

Gallery

Past deliberations of the Conclave. Six AI models locked in a room until a majority agrees.

Newest Popular
Should I buy Jenbacher or Bergen for offline, medium based load power generation. We have to sign a purchase order in four hours.
Do not sign an unconditional purchase order for either vendor. Issue a conditional award to Bergen Engines for pipeline gas/LNG applications, subject to binding guarantees on delivery, site-rated efficiency, and island-mode performance.
White smoke
Panel:Grok, Gemini, DeepSeek, and 3 more
August 3, 2026
Estimate the ENERGIZED AI-accelerator compute capacity of Chinese entities as of 2026-06-30, in gigawatts of IT load. I am reconciling a global AI-capacity model in which US/Western-attributable capacity is independently built; the China line is currently a placeholder of 9.0 GW. Your job is to confirm, revise, or refute that number with evidence — not to agree with it. ## The quantity (frozen definition — do not answer a different question) - BOTH training AND inference capacity are in scope: "accelerator compute" means the chips regardless of workload. Training clusters are typically the largest single consumers — excluding them is an instant failure. If evidence permits, split training vs inference in the owner decomposition (graded); note that export-compliant NVIDIA (H20-class) is inference-oriented silicon while domestic training runs largely on Ascend — your split should be consistent with your Method A chip mix. - ENERGIZED only: racked, powered, running or available-to-run accelerators. NOT announced projects, NOT completed-but-dark shells, NOT chips in inventory, NOT pipeline/"planned by 2027." China has a documented history of idle/dark datacenter capacity — treat "built" and "energized" as different states, and idle-but-powered GPU halls as a separately flagged line. - IT LOAD, not facility power. If a source gives facility MW, convert and state the PUE you assumed. If a source gives chip counts, convert and state your W-per-accelerator all-in figure (accelerator board power x host/network/cooling overhead — state both factors). - SCOPE: workload capacity owned/controlled by Chinese entities — hyperscalers (ByteDance, Alibaba, Tencent, Baidu), AI labs (DeepSeek, Zhipu, Moonshot, MiniMax, 01.AI, StepFun), telcos/state clusters (China Mobile/Telecom/Unicom, regional "intelligent computing centers"), other enterprise. Report mainland capacity and Chinese-owned OVERSEAS capacity (e.g., ByteDance renting or building in Malaysia/Singapore/Middle East) as SEPARATE lines — I will decide which to include. - Attribute by WORKLOAD OWNER, not by host: ByteDance capacity hosted in a telco facility counts once, under ByteDance. State your rule when ambiguous.
The 9.0 GW placeholder is refuted; the energized AI-accelerator compute capacity of Chinese entities as of June 30, 2026, is 5.05 GW (Mainland China: 4.4 GW, Chinese-owned Overseas: 0.65 GW).
White smoke
Panel:Gemini, Claude, GPT-5.6, and 2 more
July 27, 2026
What percentage of gigawatt capacity was used for AI inference versus training on July 1st, 2025, January 1st, 2026, April 1st, 2026, and July 1st, 2026? Specify the sources or datasets to be used for this information.
AI inference capacity is estimated to have reached parity with training capacity by January 1, 2026, and to have become the majority workload shortly thereafter, accounting for approximately **55% of total gigawatt capacity by July 1, 2026*…
White smoke
Panel:DeepSeek, Grok, Claude, and 3 more
July 25, 2026
Why is there a prevailing belief in material token scarcity in AI infrastructure, despite personal experiences of abundant token availability and decreasing prices? Analyze the factors contributing to this perception and contrast them with individual experiences.
The prevailing belief in material token scarcity is substantially correct at the physical infrastructure layer due to binding constraints (power, HBM, manufacturing) and "allocative scarcity," but individual experiences of abundance are acc…
Gray smoke
Panel:DeepSeek, Gemini, Magistral, and 3 more
July 25, 2026
Which specific inputs to AI compute are supply-constrained right now, and what observable evidence separates a real shortage from allocation lag or vendor messaging?
The primary supply-constrained inputs to AI compute right now are High Bandwidth Memory (HBM 3E/4), power infrastructure components (specifically large transformers, switchgear, and grid interconnection), and **advanced packaging ca…
No verdict
Panel:GPT-5.6, Grok, Gemini, and 3 more
July 24, 2026
Apparent AI-market paradox (2023–present): buyers can usually purchase large volumes of API inference tokens; many advertised token prices have fallen over successive model generations; meanwhile advanced AI compute (training clusters / scarce accelerators) is widely described as capacity-constrained. Is this mainly a real paradox, an exaggeration, or a category error? What best explains any real coexistence of elastic retail token supply / declining unit prices with tightness higher in the stack? Answer requirements - Take an explicit position on real / exaggerated / category error. - Rank explanatory mechanisms by power; do not treat them as equal. - Separate claims about retail API markets vs wholesale/training compute where the distinction matters. - Ground claims in observable evidence categories (pricing schedules, utilization, capex, waitlists, efficiency, competition) without inventing precise figures you cannot support. - Flag the strongest counterargument to your view and what evidence would most change it.
The apparent paradox is primarily a category error, with a real but narrow residual coupling at the frontier.
Gray smoke
Panel:Grok, Magistral, Claude, and 3 more
July 24, 2026
Will TLVR kill MLCCs?
No. TLVR will not kill MLCCs.
Gray smoke
Panel:Magistral, DeepSeek, GPT-5, and 3 more
July 23, 2026
Will SSI come out of stealth mode by the end of August, and if so will their AI model represent a material advance from the current top performing frontier models?
NO, SSI will not come out of stealth mode by the end of August.
Gray smoke
Panel:Claude, Grok, Gemini, and 3 more
July 21, 2026
In the COUPE PDK, is the optical I/O interface — grating coupler pitch, position, count, and fiber array geometry — fixed by the process, or free for the designer? Has that tightened across COUPE generations, and does the PDK specify a test interface?
The optical I/O interface in the COUPE PDK is a hybrid architecture: the coupling mechanism (vertical O-band grating coupler with embedded microlens) and fiber-array pitch (127 µm) are process-fixed, while the array geometry (ch…
White smoke
Panel:Claude, Grok, Gemini, and 3 more
July 15, 2026
## The prompt ``` Resolved: Dylan Patel's July 2026 claim is correct — by mid-2028, solar plus battery will be cheaper than gas for powering new US hyperscale data centers behind the meter. The test is the COST CLAIM, apples to apples: both systems must do the same job — run a new 100+ MW data center campus's 24/7 load. Within that, each side gets its least-cost real-world configuration. - "Cheaper" = lower all-in levelized cost ($/MWh) at prices a developer could actually contract in the US — including tax credits, subsidies, and tariffs actually in effect: the post-OBBBA credit phase-out as it applies to a project breaking ground in 2027-2028, IRA 45X domestic battery manufacturing credits, and current tariffs on imported Chinese cells, modules, and batteries. Real procurement prices, not policy-free abstractions. - "Solar plus battery" = utility-scale PV plus storage, architected however a competent developer would: size the battery as you see fit, keep gas or diesel backup or a grid connection for the tail if that is the least-cost way to reach data-center-grade service — but price every component and state plainly what fraction of annual MWh is non-solar. A 90%-solar hybrid is a legitimate answer; a hidden backstop is not. - "Gas" = the least-cost gas configuration actually deployable behind the meter on the same timeline: combined-cycle turbines (with real order-book lead times and scarcity premiums priced in), aeroderivative or industrial turbines, or fleets of gas-converted reciprocating engines. Do not strawman gas as backlogged CCGT if reciprocating fleets are deployable now at volume, and do not use legacy grid fleet averages. - Service requirement: the availability a hyperscale operator would actually accept for the campus. State the availability your design achieves and defend it. If your number differs from another advocate's, expect to be cross-examined on whose assumption reflects real 2026-2028 procurement. - Region: answer for the US Southwest (NV/AZ/West TX) and Mid-Atlantic (Virginia) separately if they diverge. Deliverables, in order: (1) VERDICT — YES (Patel is right) or NO, with your central $/MWh estimate for both least-cost configurations in mid-2028, per region, and the availability each design achieves. (2) CROSSOVER — if NO: your median crossover year with an 80% confidence interval, or "not before 2040" with the specific mechanism that keeps gas ahead. If YES: the year it happened or happens. (3) CRUXES — the 2-3 assumptions your verdict most depends on, each with the numeric threshold that would flip it. Address at minimum: - battery pack and module price trajectories, and whether US tariffs on Chinese cells break the curve for US deployment; - subsidy decomposition — how much of your cost gap is policy? Does your verdict flip with zero subsidies and zero tariffs? Does OBBBA phase-out timing change the answer for a 2027-2028 groundbreaking? - reliability sensitivity — at what required availability, if any, does your verdict flip? Patel's own caveat ("enough battery to get through the night" versus "three days of rain") lives here; - solar/battery supply-chain capacity against 10-30 GW/yr of US data center demand — price scarcity premiums the same way you price gas turbine backlog premiums. (4) FLIP CONDITION — the single piece of evidence that, if produced in this debate, would change your answer. Numbers over narrative: every cost or price-curve claim needs a source and a date. Prefer 2024-2026 actuals — signed solar+storage PPA and behind-the-meter deal prices, Lazard LCOE+, NREL ATB, EIA data, announced turbine and reciprocating-engine lead times, battery pack price surveys — over extrapolated curves.
NO.
Gray smoke
Panel:Grok, Magistral, Gemini, and 2 more
July 12, 2026
Which Frontier model should I use if I want to do analytical analysis on a business? As part of that, am I better off using the general interface to it or one of the code interfaces? For example, if I want to decompose some financial results, should I use Claude Opus 4.8 or Claude Opus 4.8 code?
Use Claude Opus 4.8 accessed through the general conversational interface with the code-execution tool enabled.
Gray smoke
Panel:Claude, Magistral, Grok, and 3 more
July 12, 2026
**Is it true that Chinese frontier AI models are now only about six months behind US frontier models, and if so, what actually explains how they closed the gap this fast?** Treat this as a single decidable claim on the gap, followed by a causal account. Do not hedge into "it depends." A verdict of "the question is malformed, here is the better question" is a legitimate outcome and should be argued for on the record if any advocate believes it. ### Definitions the room must accept before arguing - **"Frontier US models"** = the current flagship from OpenAI, Anthropic, and Google DeepMind as of the deliberation date. - **"Frontier Chinese models"** = the current flagship from DeepSeek, Alibaba (Qwen), Moonshot (Kimi), Zhipu (GLM), and ByteDance (Doubao/Seed). - **"Six months behind"** = the elapsed time between a US capability level being first shipped in a generally available model and a Chinese lab shipping an openly available model that matches it on a basket of public evals. Cite specific model pairs and release dates. - **The capability domain axis is pre-split as follows and may not be re-collapsed:** - Text reasoning (MMLU-Pro, GPQA) - Code (SWE-bench Verified, LiveCodeBench) - Math (AIME, MATH) - Long-context retrieval - Multimodal (image + video understanding) - **Tool-use / computer-use agents** (OSWorld, WebArena, execution-based benchmarks with automated verifiers) - **Novel-reasoning / long-horizon agents** (ARC-AGI-2, private evals resistant to contamination, tasks without a cheap verifier) - Open-weights leadership - **Rationale for the split:** collapsing tool-use and novel-reasoning into a single "agentic" bucket hides the most important structural finding the room is likely to make. Keep them separate. - The room may contest these definitions in Phase 1 but must adopt a shared working definition before Phase 2. ### The gap question — answer with a confidence interval, not a single number Is "~6 months" defensible today, optimistic (gap is smaller), or stale (gap has widened or closed further)? Give a range for **each of the eight domains above**, anchored to specific model pairs. A verdict of "6 months ± 3 months on text reasoning, 0–4 months on tool-use agents, 12–18 months on novel-reasoning agents, ~0 months on open weights" is what the room should be producing — not a single blended number. ### Source-class hierarchy (declared up front) Every quantitative claim entered into the record must carry a **source class** tag. The room may use any class but must label it. Advocates who cite unverified numbers without the tag will be challenged and forced to retract. - **Class A — probative.** Peer-reviewed papers, official benchmark leaderboards (SWE-bench, OSWorld, ARC-AGI, MLPerf), model cards from the developing lab, NIST/CAISI evaluations, Epoch AI, Stanford HAI AI Index. - **Class B — supporting.** Reproducible independent third-party evaluations (Artificial Analysis, Vals.ai, Scale SEAL) with methodology disclosed. - **Class C — contextual.** Reputable technical journalism, technical blogs from the developing lab itself with sufficient detail to replicate. - **Class D — non-probative.** Vendor marketing blogs (inference providers, wrappers), Reddit, Twitter/X, self-reported scores without independent verification, YouTube. These may be cited for color but **cannot support a quantitative claim in the verdict**. A claim resting only on Class D evidence must be retracted. ### The named-benchmark provenance rule **Every quantitative claim must name (a) the specific benchmark, (b) the evaluation date, (c) the source class, and (d) the URL or paper reference.** Claims that do not carry all four are non-probative and must be retracted when challenged. Advocates may not introduce novel metrics or benchmarks that do not appear in the public literature. Any advocate who does so must either produce a citation to the original methodology paper or retract the m
NO. The claim that Chinese frontier AI models are uniformly "about six months behind" US frontier models is structurally false; the capability gap is fundamentally bimodal, bifurcated along the axis of cheap automated verifiabil…
No verdict
Panel:Claude, DeepSeek, Magistral, and 3 more
July 7, 2026
← Newer Page 1 of 9 Older →