NL2SQL leaderboards are still moving—just not at the top
Fresh entries continue to arrive on BIRD and Spider 2.0, but neither benchmark has crowned a new leader for months.
The two most-watched public NL2SQL boards are still accepting and posting systems. What they are not doing is changing leaders.
As rendered on August 29, BIRD’s test execution-accuracy leader remained AskData + GPT-4o at 81.95%, dated December 16, 2025. The newest visible BIRD entry was RAS, dated August 19, 2026, at 79.82% test EX and 72.49% dev EX. That leaves the latest submission 2.13 percentage points below a leader that has held for more than eight months. These are BIRD’s own test and dev fields; they should not be merged or treated as interchangeable. BIRD leaderboard
Spider 2.0-Snow tells the same story over a shorter interval. Genloop’s Sentinel Agent v2 Pro remained first at 96.70, dated March 1, 2026. A newer visible system, Omni 2.0, was posted July 8 at 82.81; several May entries also sit below the March leader. The board therefore continues to receive results without changing first place for nearly six months. Spider 2.0 leaderboard
Those percentages must not be compared against each other. BIRD reports execution accuracy on its own development and held-out test splits, while Spider 2.0-Snow describes a 547-example Snowflake text-to-SQL task and ranks methods with its own score. Spider also warns that its evaluation metrics are checked continuously and scores can move slightly over time. The defensible comparison is within each board over time—not 81.95 versus 96.70. BIRD benchmark description Spider 2.0 settings
The negative result matters because a busy submission queue can create the impression that capability is advancing at the frontier every week. On these boards, recent activity is instead filling positions below established leaders. That may reflect durable leaders, weaker new systems, or evaluation choices that reward already-optimized approaches; the public tables alone cannot distinguish among those explanations.
For practitioners, the useful signal is stability rather than a new state of the art. BIRD’s newest visible submission did not erase the gap to its December leader, and Spider 2.0-Snow’s newer entries did not displace its March leader. Until a submission changes those rows—or an evaluation revision changes the ordering—“new result” and “new frontier” remain different claims.
sources
- BIRD-SQL official leaderboardbird-bench.github.io
- Spider 2.0 official leaderboardspider2-sql.github.io
comments · 0