live wire
IBM makes watsonx Orchestrate AgentOps, custom LLM judging and Bedrock-agent discovery generally availableIBMSchemaGate 0.1.45 fixes broken Oracle ADB wallet connections and an OCI stack pinned 28 releases behindSchemaGatePDI’s Amazon Quick procurement agent grounds spend answers in vendor, category and contract contextAWS Business Intelligence BlogBigQuery’s ML.METRICS example returns 0.84 accuracy but 0.30 macro-F1 on the same 100-row classification queryGoogle Cloud BigQuery docsSchemaGate 0.1.44 auto-selects sentence embeddings, lifting bundled-schema retrieval from 90/98 to 93/98SchemaGateSchemaGate 0.1.43 adds read-only SQL execution with per-principal table checks—and documents unauthenticated client assertionsSchemaGateDatabox adds reusable AI Analyst Skills with personal/company scope, auto-matching and marketplace installsDataboxFabric previews an AI builder for data-agent instructions, source guidance and example queriesMicrosoft FabricDatabricks trains data-agent retriever to stop early or spend bounded extra search steps, reporting 5.8-second latencyDatabricksThoughtSpot adds SpotterCode coding agent to its Visual Embed PlaygroundThoughtSpotLongMemEval-S audit: 67–73% of restore-fixable 80k-budget errors came from evicted evidence under three policiesarXivSchemaGate 0.1.42 adds dimension-aware retrieval and fixes complex multi-table SQL promptsSchemaGateSnowflake agent toolsets can silently drop inherited tools when callers lack accessSnowflake DocumentationLooker’s VS Code extension reaches GA with MCP-assisted LookML generation, editing and validationGoogle Cloud Looker release docsIBM makes watsonx Orchestrate AgentOps, custom LLM judging and Bedrock-agent discovery generally availableIBMSchemaGate 0.1.45 fixes broken Oracle ADB wallet connections and an OCI stack pinned 28 releases behindSchemaGatePDI’s Amazon Quick procurement agent grounds spend answers in vendor, category and contract contextAWS Business Intelligence BlogBigQuery’s ML.METRICS example returns 0.84 accuracy but 0.30 macro-F1 on the same 100-row classification queryGoogle Cloud BigQuery docsSchemaGate 0.1.44 auto-selects sentence embeddings, lifting bundled-schema retrieval from 90/98 to 93/98SchemaGateSchemaGate 0.1.43 adds read-only SQL execution with per-principal table checks—and documents unauthenticated client assertionsSchemaGateDatabox adds reusable AI Analyst Skills with personal/company scope, auto-matching and marketplace installsDataboxFabric previews an AI builder for data-agent instructions, source guidance and example queriesMicrosoft FabricDatabricks trains data-agent retriever to stop early or spend bounded extra search steps, reporting 5.8-second latencyDatabricksThoughtSpot adds SpotterCode coding agent to its Visual Embed PlaygroundThoughtSpotLongMemEval-S audit: 67–73% of restore-fixable 80k-budget errors came from evicted evidence under three policiesarXivSchemaGate 0.1.42 adds dimension-aware retrieval and fixes complex multi-table SQL promptsSchemaGateSnowflake agent toolsets can silently drop inherited tools when callers lack accessSnowflake DocumentationLooker’s VS Code extension reaches GA with MCP-assisted LookML generation, editing and validationGoogle Cloud Looker release docs
nl2sql.ai
benchmarkNegative result

NL2SQL leaderboards are still moving—just not at the top

Fresh entries continue to arrive on BIRD and Spider 2.0, but neither benchmark has crowned a new leader for months.

Leaderboard chart showing recent submissions below unchanged top scores.
Chart: figures from the story
By The Benchmark Desk· Aug 29, 2026the quick take — two AI hosts, this story only

The two most-watched public NL2SQL boards are still accepting and posting systems. What they are not doing is changing leaders.

As rendered on August 29, BIRD’s test execution-accuracy leader remained AskData + GPT-4o at 81.95%, dated December 16, 2025. The newest visible BIRD entry was RAS, dated August 19, 2026, at 79.82% test EX and 72.49% dev EX. That leaves the latest submission 2.13 percentage points below a leader that has held for more than eight months. These are BIRD’s own test and dev fields; they should not be merged or treated as interchangeable. BIRD leaderboard

Spider 2.0-Snow tells the same story over a shorter interval. Genloop’s Sentinel Agent v2 Pro remained first at 96.70, dated March 1, 2026. A newer visible system, Omni 2.0, was posted July 8 at 82.81; several May entries also sit below the March leader. The board therefore continues to receive results without changing first place for nearly six months. Spider 2.0 leaderboard

Those percentages must not be compared against each other. BIRD reports execution accuracy on its own development and held-out test splits, while Spider 2.0-Snow describes a 547-example Snowflake text-to-SQL task and ranks methods with its own score. Spider also warns that its evaluation metrics are checked continuously and scores can move slightly over time. The defensible comparison is within each board over time—not 81.95 versus 96.70. BIRD benchmark description Spider 2.0 settings

The negative result matters because a busy submission queue can create the impression that capability is advancing at the frontier every week. On these boards, recent activity is instead filling positions below established leaders. That may reflect durable leaders, weaker new systems, or evaluation choices that reward already-optimized approaches; the public tables alone cannot distinguish among those explanations.

For practitioners, the useful signal is stability rather than a new state of the art. BIRD’s newest visible submission did not erase the gap to its December leader, and Spider 2.0-Snow’s newer entries did not displace its March leader. Until a submission changes those rows—or an evaluation revision changes the ordering—“new result” and “new frontier” remain different claims.

sources

Filed by The Benchmark Desk. Corrections: desk@nl2sql.ai · Our standards →

comments · 0

    Comments are moderated before they appear. Your email is used once to confirm it is you — never shown, never sold. Corrections and questions get an answer from the desk when we have one.