BIRD has a new leader—but no method card yet
DataGallery-Text2SQL leads both BIRD execution-accuracy splits by narrow margins, while its public submission still promises technical details later.
BIRD’s execution-accuracy leaderboard has a new first-place system. DataGallery-Text2SQL’s submission page reports 78.10% execution accuracy on the 1,534-question development split and 82.39% on the 1,789-question test split, with a submission date of September 7, 2026. The page identifies the team as DataGallery-TextSQL at Huawei 2012 Labs.
Those numbers are now reflected on the live BIRD leaderboard, where DataGallery sits above SiriusAI-SQL. Sirius reports 77.77% dev EX and 82.28% test EX, so the new leader’s margins are only 0.33 percentage points on dev and 0.11 points on test. This is a real leaderboard move, but not evidence of a broad step-change in text-to-SQL accuracy.
The score arrived before the method
The public submission is unusually thin for a leading result. It lists the two EX scores and 77.64% test R-VES, leaves dev R-VES blank, and says only that more details will be released in a future arXiv preview. The BIRD row likewise has no code link and labels model size as unknown. That means readers can verify the leaderboard position, split sizes and reported scores, but cannot yet inspect the model stack, prompting strategy, retrieval setup, inference budget or reproducibility artifacts.
The repository history also shows why benchmark records need versioned snapshots. A September 9 BIRD commit replaced an earlier DataGallery entry dated September 2 at 77.71% dev / 82.22% test with the current September 7 result at 78.10% / 82.39%. The same commit changed the test R-VES value from 77.92% to 77.64% and removed an older DataGallery row at 74.64% dev / 77.53% test.
None of those edits is inherently suspicious: resubmissions and corrected rows are normal. But a top-line score without a method card is not yet enough to explain why the system won, and a mutable HTML table is not enough to reconstruct what was compared at a given moment.
For now, the defensible conclusion is narrow. DataGallery-Text2SQL is BIRD’s current execution-accuracy leader on both named splits, by less than half a point on each. The next meaningful evidence will not be another leaderboard decimal; it will be the promised paper, a code release, or an evaluation card that fixes the model, context, oracle-knowledge and inference settings behind the result.
sources
- DataGallery-Text2SQL — BIRD Benchmark Submissiondatagallery.cn
- BIRD-SQL leaderboardbird-bench.github.io
- BIRD commit: Update DataGallery Entry (New SOTA)github.com
comments · 0