DataGallery takes BIRD by 0.11 points—and its efficiency score already disagrees
Huawei 2012 Labs’ submission is first on both BIRD execution accuracy and R-VES, but the linked primary pages print different efficiency numbers.
Huawei 2012 Labs’ DataGallery-Text2SQL is the new execution-accuracy leader on BIRD. The live BIRD leaderboard lists the September 7 submission at 78.10% on the 1,534-question development split and 82.39% on the 1,789-question test split. The linked DataGallery submission page prints the same two execution-accuracy results.
That moves DataGallery just ahead of SiriusAI-SQL, which the live leaderboard lists at 77.77% dev and 82.28% test. The winning margin is therefore 0.33 percentage points on dev but only 0.11 points on test. Those are separate splits, not repeated measurements of one set; BIRD identifies the dev and test workloads separately, and DataGallery gives their question counts explicitly.
The efficiency result needs reconciliation
BIRD’s reward-based valid efficiency score, or R-VES, also places DataGallery first—but the two primary pages disagree on the exact value. The BIRD leaderboard currently shows 77.64 for DataGallery, while the DataGallery submission reports 77.69%. That 0.05-point mismatch is too small to change the ranking: BIRD lists the next result, Agentar-Scale-SQL, at 77.00, and SiriusAI-SQL at 74.76. It is still a provenance problem because the benchmark and entrant no longer expose one canonical efficiency result.
The leaderboard notes that its evaluation metrics are checked continually and that scores can change slightly over time. That caveat may explain the mismatch, but neither page currently says which value is newer or why it changed. Until one side is corrected, comparisons should identify the source alongside the R-VES number rather than quoting “DataGallery’s score” without qualification.
A lead without a disclosed method
The result is verified, but the method is not yet auditable. DataGallery says detailed methodology will follow in an arXiv preprint. BIRD currently marks the model size as unknown and the row as using oracle knowledge; its table does not link runnable code for the entry. Those omissions do not negate the official ranking, but they prevent readers from attributing the gain to a model, prompting strategy, retrieval design or candidate-selection budget.
The practical conclusion is narrow: DataGallery leads BIRD execution accuracy on both named splits, and it also leads the leaderboard’s efficiency metric. The execution scores agree across the two primary pages. The R-VES score does not, and the architecture behind all three numbers remains undisclosed.
sources
- BIRD-SQL live leaderboardbird-bench.github.io
- DataGallery-Text2SQL BIRD benchmark submissiondatagallery.cn
comments · 0