SiriusAI-SQL is now BIRD’s official test leader at 82.28%
BIRD merged the Tencent system’s row on Sept. 1, moving the held-out test record 0.33 points above AskData + GPT-4o.
Update, Sept. 1, 2026: BIRD maintainers merged pull request #219 at 01:13 UTC. The live leaderboard now lists SiriusAI-SQL at 82.28% test execution accuracy (EX) and 77.77% dev EX, making the Aug. 22 submission the official test-set leader.
The merge moves AskData + GPT-4o’s 81.95% test EX out of first place, a margin of 0.33 percentage points. This article originally reported the SiriusAI-SQL result while it was still an unmerged contributor proposal; the maintainer merge and rendered BIRD page now resolve that status question.
What the merged patch changes
The merged diff inserts SiriusAI-SQL into the main BIRD table with 77.77 dev EX and 82.28 test EX, removes bold formatting from AskData’s 81.95, and adds 74.76 R-VES in the efficiency table. It also removes the superseded SiriusAI-Text2SQL-Agent row, which reported 75.35 dev EX and 77.03 test EX.
On those numbers, the update is larger than the headline margin over the previous leader suggests: the same team’s new row improves its previous test score by 5.25 points and its dev score by 2.42 points. The pull request also standardizes the method name and Tencent Data Computing Platform Department attribution across Sirius entries.
Official does not mean reproducible
The overall-leaderboard row lists model size as unknown and does not link code. The merged change does not add a paper, model card, frozen predictions or evaluation artifact for the 82.28 result. A held-out server score can establish rank without those materials, but outsiders still cannot inspect the system or reproduce its pipeline from the leaderboard entry.
That distinction is now the useful one. Before Sept. 1, the question was whether maintainers would accept the contributor’s proposed row. After the merge, 82.28 is the published BIRD test record; the unresolved question is what architecture, model and inference procedure produced it.
Two other BIRD submissions merged during the same queue cycle fill very different scale regimes. A 2B Qwen3.5 + StructCoT system reports 68.47 test EX, while a 394M model trained from scratch reports 33.15 test EX with seven-sample execution-guided voting. Neither challenges the leader, but both make model scale explicit in a way the new record row does not.
The leaderboard has therefore moved, but its provenance gap remains. SiriusAI-SQL is officially first on BIRD’s held-out test split as of Sept. 1, 2026; readers should not infer from that rank that the underlying system is publicly reproducible.
sources
comments · 0