# Conclave Prompt — The China–US Frontier AI Gap
*Prepared for Matt Murray · Substrate Capital · July 7, 2026*
---
## Prompt (paste into Conclave)
**Is it true that Chinese frontier AI models are now only about six months behind US frontier models, and if so, what actually explains how they closed the gap this fast?**
Treat this as a single decidable claim on the gap, followed by a causal account. Do not hedge into "it depends." A verdict of "the question is malformed, here is the better question" is a legitimate outcome and should be argued for on the record if any advocate believes it.
### Definitions the room must accept before arguing
- **"Frontier US models"** = the current flagship from OpenAI, Anthropic, and Google DeepMind as of the deliberation date.
- **"Frontier Chinese models"** = the current flagship from DeepSeek, Alibaba (Qwen), Moonshot (Kimi), Zhipu (GLM), and ByteDance (Doubao/Seed).
- **"Six months behind"** = the elapsed time between a US capability level being first shipped in a generally available model and a Chinese lab shipping an openly available model that matches it on a basket of public evals (MMLU-Pro, GPQA, SWE-bench Verified, LiveCodeBench, AIME, Arena-Hard, long-context retrieval). Cite specific model pairs and release dates.
- The room may contest these definitions in Phase 1 but must adopt a shared working definition before Phase 2.
### The gap question — answer with a confidence interval, not a single number
Is "~6 months" defensible today, optimistic (gap is smaller), or stale (gap has widened or closed further)? Give a range, and break it out by capability domain — text reasoning, code, math, long-context, multimodal, agentic/tool-use, open-weights leadership. A verdict of "6 months ± 3 months on text reasoning, 12–18 months on agentic workloads, ~0 months on open weights" is more honest than a single number. Anchor to named model pairs, not vibes.
### Structure the causal deliberation in two phases
**Phase 1 — Independent nomination (Briefing → Submit stages).** Each of the six advocates must, before seeing the others' submissions, nominate the **two causes they believe are most load-bearing** for the gap-closure. Advocates are explicitly instructed not to converge on a canonical Western-analyst list. Constraints on the nomination round:
- At least one advocate must argue the premise itself is wrong — either the gap is not ~6 months, or "gap" is the wrong frame entirely.
- At least one advocate must nominate a cause that would not appear in a standard US think-tank readout. Steelman a Chinese-industry, hardware-supply-chain, or benchmarks-skeptic view.
- One advocate is designated the **red-team seat** and must argue the strongest version of "the gap is a mirage — Chinese models look close on public benchmarks because those benchmarks are saturated, contaminated, or gameable via post-training, and the real frontier (long-horizon agents, tool use, novel scientific reasoning) still has a 12–24 month gap."
**Phase 2 — Collate, contest, rank (Collate → Cross-examine → Debate stages).** Merge the nominated causes into a working set. Before producing any ranking, the room must:
- Produce a **one-paragraph causal sketch** identifying which causes are independent drivers vs. downstream effects of other causes. Export controls *cause* algorithmic efficiency; open weights *enable* distillation. A flat ranking hides this — a sketch surfaces it.
- Flag any cause that appears in fewer than two advocates' nominations as a **minority hypothesis** and give it a dedicated defender before the room is allowed to dismiss it.
- Explicitly name **what is NOT in the working set and why** — one sentence per omitted candidate. At minimum, the room must consider and either adopt or reject on the record: benchmark contamination and eval overfitting; Chinese-language and industrial/government data advantages; weaker IP and copyright friction on training data; US labs d
No.
Gray smoke