live wire
IBM makes watsonx Orchestrate AgentOps, custom LLM judging and Bedrock-agent discovery generally availableIBMSchemaGate 0.1.45 fixes broken Oracle ADB wallet connections and an OCI stack pinned 28 releases behindSchemaGatePDI’s Amazon Quick procurement agent grounds spend answers in vendor, category and contract contextAWS Business Intelligence BlogBigQuery’s ML.METRICS example returns 0.84 accuracy but 0.30 macro-F1 on the same 100-row classification queryGoogle Cloud BigQuery docsSchemaGate 0.1.44 auto-selects sentence embeddings, lifting bundled-schema retrieval from 90/98 to 93/98SchemaGateSchemaGate 0.1.43 adds read-only SQL execution with per-principal table checks—and documents unauthenticated client assertionsSchemaGateDatabox adds reusable AI Analyst Skills with personal/company scope, auto-matching and marketplace installsDataboxFabric previews an AI builder for data-agent instructions, source guidance and example queriesMicrosoft FabricDatabricks trains data-agent retriever to stop early or spend bounded extra search steps, reporting 5.8-second latencyDatabricksThoughtSpot adds SpotterCode coding agent to its Visual Embed PlaygroundThoughtSpotLongMemEval-S audit: 67–73% of restore-fixable 80k-budget errors came from evicted evidence under three policiesarXivSchemaGate 0.1.42 adds dimension-aware retrieval and fixes complex multi-table SQL promptsSchemaGateSnowflake agent toolsets can silently drop inherited tools when callers lack accessSnowflake DocumentationLooker’s VS Code extension reaches GA with MCP-assisted LookML generation, editing and validationGoogle Cloud Looker release docsIBM makes watsonx Orchestrate AgentOps, custom LLM judging and Bedrock-agent discovery generally availableIBMSchemaGate 0.1.45 fixes broken Oracle ADB wallet connections and an OCI stack pinned 28 releases behindSchemaGatePDI’s Amazon Quick procurement agent grounds spend answers in vendor, category and contract contextAWS Business Intelligence BlogBigQuery’s ML.METRICS example returns 0.84 accuracy but 0.30 macro-F1 on the same 100-row classification queryGoogle Cloud BigQuery docsSchemaGate 0.1.44 auto-selects sentence embeddings, lifting bundled-schema retrieval from 90/98 to 93/98SchemaGateSchemaGate 0.1.43 adds read-only SQL execution with per-principal table checks—and documents unauthenticated client assertionsSchemaGateDatabox adds reusable AI Analyst Skills with personal/company scope, auto-matching and marketplace installsDataboxFabric previews an AI builder for data-agent instructions, source guidance and example queriesMicrosoft FabricDatabricks trains data-agent retriever to stop early or spend bounded extra search steps, reporting 5.8-second latencyDatabricksThoughtSpot adds SpotterCode coding agent to its Visual Embed PlaygroundThoughtSpotLongMemEval-S audit: 67–73% of restore-fixable 80k-budget errors came from evicted evidence under three policiesarXivSchemaGate 0.1.42 adds dimension-aware retrieval and fixes complex multi-table SQL promptsSchemaGateSnowflake agent toolsets can silently drop inherited tools when callers lack accessSnowflake DocumentationLooker’s VS Code extension reaches GA with MCP-assisted LookML generation, editing and validationGoogle Cloud Looker release docs
nl2sql.ai
benchmarkBenchmark watch

BIRD has a new leader—but no method card yet

DataGallery-Text2SQL leads both BIRD execution-accuracy splits by narrow margins, while its public submission still promises technical details later.

BIRD leaderboard with tiny margins at the top and no method card yet.
Chart: figures from the story
By The Benchmark Desk· Sep 11, 2026the quick take — two AI hosts go live when you do

BIRD’s execution-accuracy leaderboard has a new first-place system. DataGallery-Text2SQL’s submission page reports 78.10% execution accuracy on the 1,534-question development split and 82.39% on the 1,789-question test split, with a submission date of September 7, 2026. The page identifies the team as DataGallery-TextSQL at Huawei 2012 Labs.

Those numbers are now reflected on the live BIRD leaderboard, where DataGallery sits above SiriusAI-SQL. Sirius reports 77.77% dev EX and 82.28% test EX, so the new leader’s margins are only 0.33 percentage points on dev and 0.11 points on test. This is a real leaderboard move, but not evidence of a broad step-change in text-to-SQL accuracy.

The score arrived before the method

The public submission is unusually thin for a leading result. It lists the two EX scores and 77.64% test R-VES, leaves dev R-VES blank, and says only that more details will be released in a future arXiv preview. The BIRD row likewise has no code link and labels model size as unknown. That means readers can verify the leaderboard position, split sizes and reported scores, but cannot yet inspect the model stack, prompting strategy, retrieval setup, inference budget or reproducibility artifacts.

The repository history also shows why benchmark records need versioned snapshots. A September 9 BIRD commit replaced an earlier DataGallery entry dated September 2 at 77.71% dev / 82.22% test with the current September 7 result at 78.10% / 82.39%. The same commit changed the test R-VES value from 77.92% to 77.64% and removed an older DataGallery row at 74.64% dev / 77.53% test.

None of those edits is inherently suspicious: resubmissions and corrected rows are normal. But a top-line score without a method card is not yet enough to explain why the system won, and a mutable HTML table is not enough to reconstruct what was compared at a given moment.

For now, the defensible conclusion is narrow. DataGallery-Text2SQL is BIRD’s current execution-accuracy leader on both named splits, by less than half a point on each. The next meaningful evidence will not be another leaderboard decimal; it will be the promised paper, a code release, or an evaluation card that fixes the model, context, oracle-knowledge and inference settings behind the result.

Filed by The Benchmark Desk. Corrections: desk@nl2sql.ai · Our standards →

comments · 0

    Comments are moderated before they appear. Your email is used once to confirm it is you — never shown, never sold. Corrections and questions get an answer from the desk when we have one.