BIRD’s new No. 2 misses the lead by 0.06 points on both dev and test
DataGallery-Text2SQL lands just below SiriusAI-SQL on two separately reported execution-accuracy splits, while adding a 77.92 test efficiency score.
A near-tie at the top
BIRD added DataGallery-Text2SQL to its leaderboard on September 2, reporting 77.71% execution accuracy on the development split and 82.22% on the test split. The submission is attributed to Huawei 2012 Labs. The repository entry also reports a 77.92% test R-VES, BIRD’s reward-based valid efficiency score (BIRD repository commit).
Those figures place DataGallery second behind SiriusAI-SQL in BIRD’s execution-accuracy table. SiriusAI-SQL, dated August 22, is listed at 77.77% dev EX and 82.28% test EX. DataGallery therefore trails by exactly 0.06 percentage points on dev and 0.06 points on test at the precision displayed by the leaderboard (live BIRD leaderboard).
Why the matching gap matters
The interesting result is not merely that a new row reached second place. The same 0.06-point separation appears on both reported splits. That makes this a cleaner near-tie than cases where systems swap order or show a large dev-to-test divergence. It does not establish statistical equivalence: the public table reports point estimates, not confidence intervals, and the 0.06-point difference should not be interpreted as proof that one system is reliably better.
The result also updates DataGallery’s own position. BIRD still displays an earlier DataGallery-Text2SQL row dated June 9 at 74.64% dev EX and 77.53% test EX. The September 2 row is therefore 3.07 points higher on dev and 4.69 points higher on test than that earlier listing, although the leaderboard page does not expose enough implementation detail in the table itself to attribute the improvement to a specific change (live BIRD leaderboard).
The benchmark takeaway
At the top of BIRD, rank labels now exaggerate a very small displayed-score difference. SiriusAI-SQL remains first and DataGallery-Text2SQL is second, but their reported execution accuracy is separated by six hundredths of a point on each split. For readers comparing systems, the defensible conclusion is narrow: DataGallery has reached the leading cluster, while BIRD’s current public numbers do not justify treating the top two as meaningfully far apart.
The new row’s separate 77.92% test R-VES is also worth retaining alongside EX rather than collapsing the submission to a single headline score, because BIRD reports accuracy and efficiency as distinct metrics (BIRD repository commit).
sources
- BIRD repository commit adding DataGallery-Text2SQL resultsgithub.com
- BIRD live leaderboardbird-bench.github.io
comments · 0