ReViSQL reaches a human proxy on corrected mini-dev—not the live BIRD test board
The 92.97% result is real and reproducible in a released artifact, but its split is Arcwise-Plat-SQL; the 92.96% human reference comes from BIRD Test.
Thinking Machines Lab’s ReViSQL result deserves attention—and a split label in the headline. The team reports 91.37% accuracy with greedy decoding on Arcwise-Plat-SQL at $0.035 per task, rising to 92.97% with 16-sample self-consistency at $0.56 per task. Arcwise-Plat-SQL contains 498 questions whose gold SQL was corrected while the original questions and external knowledge were retained (Thinking Machines Lab; ReViSQL artifact).
The important qualifier is that the 92.96% human reference comes from BIRD Test, not Arcwise-Plat-SQL. The live BIRD page labels that number “Human Performance,” while the ReViSQL paper explicitly describes it as a proxy transferred to its Arcwise evaluation (BIRD leaderboard; technical report v4). The 92.97% result therefore clears the cited human proxy by 0.01 percentage point, but it is not a same-split human-versus-model comparison and is not a win on the live BIRD Test leaderboard. As of our rendered-page check on August 28, 2026, ReViSQL was not listed there.
That distinction does not erase the result. It explains what moved. The project’s second split, Arcwise-Plat-Full, also has 498 questions, but corrects the questions and external knowledge as well as the SQL. ReViSQL-BIRD-K2.6 scores 93.8% greedily and 94.2% with 16-sample self-consistency on that fully corrected split; on SQL-only-corrected Arcwise-Plat-SQL, the corresponding rounded figures are 91.4% and 93.0% (technical report v4). The gap shows why “BIRD” is no longer a sufficient split description.
The released artifact makes the comparison inspectable. It includes 2,462 expert-verified training instances, the two 498-question Arcwise files, training code, inference code, and commands for greedy and 16-candidate evaluation. The authors report correcting SQL errors in 52.1% of audited training instances, question flaws in 26.2%, and external-knowledge errors in 18.2%; categories overlap (ReViSQL artifact; Thinking Machines Lab).
The practical reading is narrower and stronger than “BIRD is solved.” Verified data plus reward shaping produced a released model recipe that reaches the transferred human proxy on a corrected mini-dev variant. Its greedy score remains 1.59 points below that proxy, while 16 generations and majority voting cross it. Until the same system is evaluated on BIRD Test—or humans are measured on the Arcwise split—those are two useful milestones, not one interchangeable leaderboard result.
sources
- Thinking Machines Lab — Putting Task Expertise into RL Achieves State-of-the-Art Performance on Text-to-SQLthinkingmachines.ai
- ReViSQL technical report, arXiv:2603.20004arxiv.org
- ReViSQL SIGMOD 2027 artifactgithub.com
- BIRD leaderboardbird-bench.github.io
comments · 0