BIRD’s new top two are 0.06 points apart—but neither links code
SiriusAI-SQL and DataGallery-Text2SQL both clear 82% test execution accuracy; the leaderboard records unknown model size, oracle knowledge, and shrinking public provenance.
A new high, then an immediate near-tie
BIRD’s execution-accuracy table has a new leading pair. A Sept. 1 repository commit added Tencent’s SiriusAI-SQL at 77.77% on the dev split and 82.28% on the test split. The commit labels the result a new state of the art, while the leaderboard row itself is dated Aug. 22, 2026.
One day later, a second commit added Huawei 2012 Labs’ DataGallery-Text2SQL at 77.71% dev execution accuracy, 82.22% test execution accuracy, and 77.92% test R-VES. On the same test-EX split, Sirius leads DataGallery by only 0.06 percentage points. Sirius is 0.33 points above the previous 81.95% test-EX row for AskData + GPT-4o; DataGallery is 0.27 points above it.
The ranking is precise; the systems are not yet inspectable
Both new overall rows mark model size as UNK, indicate use of oracle knowledge, and provide no code link. Sirius’ commit also removes the earlier SiriusAI-Text2SQL-Agent row—75.35% dev EX and 77.03% test EX—and inserts the renamed SiriusAI-SQL result, a 5.25-point difference on test EX between the removed and replacement rows. In the single-trained-model table, the same commit renames the Sirius entry and removes its prior arXiv paper link.
DataGallery arrived with a Notion technical-report link, but an open pull request filed Sept. 3 asks the benchmark maintainers to “temporarily delete” that link. The patch removes only the linked author/report line; it does not change DataGallery’s scores. As of this check, the pull request remains open and has no discussion explaining the temporary removal.
That does not invalidate either test result: BIRD’s public table is the primary record of the evaluated scores. It does limit what outsiders can audit about the systems behind a margin of six hundredths of a point. Until model configuration, prompts, retrieval components, and code are available, the defensible conclusion is narrow: SiriusAI-SQL is the current BIRD test-EX leader at 82.28%, with DataGallery-Text2SQL at 82.22% on the same split. The leaderboard can order those submissions; it cannot yet show whether the tiny difference is reproducible.
sources
- BIRD commit: Update SiriusAI-SQL leaderboard resultsgithub.com
- BIRD commit: Add new DataGallery resultsgithub.com
- Open PR #225: Update DataGallery entry infogithub.com
- BIRD-SQL leaderboardbird-bench.github.io
comments · 0