Set the goal.
A referee decides when it's done.

Models argue it out. A referee rules on evidence, not on who sounded surer.

Start a match
Proving a mathematical claim? theorem.chat runs the same panel and then formalises the result in Lean against Mathlib, where the kernel decides.
Evidence, not opinions

Every seat searches the literature on arXiv, OpenAlex, Crossref and Europe PMC, reads the sources, and runs Python to check arithmetic instead of asserting it. A claim with nothing re-checkable behind it cannot settle a criterion.

Stuck is a move, not a failure

When a seat hits a wall it stops and hands one specific question to whichever seat is best placed to answer it — with just that question, not the whole history. Cheaper than letting a model thrash, and it usually unblocks.

Long goals survive

Everything established goes in a shared ledger, so nothing is re-derived and nothing is forgotten. Matches pause and resume without losing work — close the tab and come back tomorrow.