**Is it true that Chinese frontier AI models are now only about six months behind US frontier models, and if so, what actually explains how they closed the gap this fast?**
Treat this as a single decidable claim on the gap, followed by a causal account. Do not hedge into "it depends." A verdict of "the question is malformed, here is the better question" is a legitimate outcome and should be argued for on the record if any advocate believes it.
### Definitions the room must accept before arguing
- **"Frontier US models"** = the current flagship from OpenAI, Anthropic, and Google DeepMind as of the deliberation date.
- **"Frontier Chinese models"** = the current flagship from DeepSeek, Alibaba (Qwen), Moonshot (Kimi), Zhipu (GLM), and ByteDance (Doubao/Seed).
- **"Six months behind"** = the elapsed time between a US capability level being first shipped in a generally available model and a Chinese lab shipping an openly available model that matches it on a basket of public evals. Cite specific model pairs and release dates.
- **The capability domain axis is pre-split as follows and may not be re-collapsed:**
- Text reasoning (MMLU-Pro, GPQA)
- Code (SWE-bench Verified, LiveCodeBench)
- Math (AIME, MATH)
- Long-context retrieval
- Multimodal (image + video understanding)
- **Tool-use / computer-use agents** (OSWorld, WebArena, execution-based benchmarks with automated verifiers)
- **Novel-reasoning / long-horizon agents** (ARC-AGI-2, private evals resistant to contamination, tasks without a cheap verifier)
- Open-weights leadership
- **Rationale for the split:** collapsing tool-use and novel-reasoning into a single "agentic" bucket hides the most important structural finding the room is likely to make. Keep them separate.
- The room may contest these definitions in Phase 1 but must adopt a shared working definition before Phase 2.
### The gap question — answer with a confidence interval, not a single number
Is "~6 months" defensible today, optimistic (gap is smaller), or stale (gap has widened or closed further)? Give a range for **each of the eight domains above**, anchored to specific model pairs. A verdict of "6 months ± 3 months on text reasoning, 0–4 months on tool-use agents, 12–18 months on novel-reasoning agents, ~0 months on open weights" is what the room should be producing — not a single blended number.
### Source-class hierarchy (declared up front)
Every quantitative claim entered into the record must carry a **source class** tag. The room may use any class but must label it. Advocates who cite unverified numbers without the tag will be challenged and forced to retract.
- **Class A — probative.** Peer-reviewed papers, official benchmark leaderboards (SWE-bench, OSWorld, ARC-AGI, MLPerf), model cards from the developing lab, NIST/CAISI evaluations, Epoch AI, Stanford HAI AI Index.
- **Class B — supporting.** Reproducible independent third-party evaluations (Artificial Analysis, Vals.ai, Scale SEAL) with methodology disclosed.
- **Class C — contextual.** Reputable technical journalism, technical blogs from the developing lab itself with sufficient detail to replicate.
- **Class D — non-probative.** Vendor marketing blogs (inference providers, wrappers), Reddit, Twitter/X, self-reported scores without independent verification, YouTube. These may be cited for color but **cannot support a quantitative claim in the verdict**. A claim resting only on Class D evidence must be retracted.
### The named-benchmark provenance rule
**Every quantitative claim must name (a) the specific benchmark, (b) the evaluation date, (c) the source class, and (d) the URL or paper reference.** Claims that do not carry all four are non-probative and must be retracted when challenged. Advocates may not introduce novel metrics or benchmarks that do not appear in the public literature. Any advocate who does so must either produce a citation to the original methodology paper or retract the m
NO. The claim that Chinese frontier AI models are uniformly "about six months behind" US frontier models is structurally false; the capability gap is fundamentally bimodal, bifurcated along the axis of cheap automated verifiabil…
No verdict