live wire
IBM makes watsonx Orchestrate AgentOps, custom LLM judging and Bedrock-agent discovery generally availableIBMSchemaGate 0.1.45 fixes broken Oracle ADB wallet connections and an OCI stack pinned 28 releases behindSchemaGatePDI’s Amazon Quick procurement agent grounds spend answers in vendor, category and contract contextAWS Business Intelligence BlogBigQuery’s ML.METRICS example returns 0.84 accuracy but 0.30 macro-F1 on the same 100-row classification queryGoogle Cloud BigQuery docsSchemaGate 0.1.44 auto-selects sentence embeddings, lifting bundled-schema retrieval from 90/98 to 93/98SchemaGateSchemaGate 0.1.43 adds read-only SQL execution with per-principal table checks—and documents unauthenticated client assertionsSchemaGateDatabox adds reusable AI Analyst Skills with personal/company scope, auto-matching and marketplace installsDataboxFabric previews an AI builder for data-agent instructions, source guidance and example queriesMicrosoft FabricDatabricks trains data-agent retriever to stop early or spend bounded extra search steps, reporting 5.8-second latencyDatabricksThoughtSpot adds SpotterCode coding agent to its Visual Embed PlaygroundThoughtSpotLongMemEval-S audit: 67–73% of restore-fixable 80k-budget errors came from evicted evidence under three policiesarXivSchemaGate 0.1.42 adds dimension-aware retrieval and fixes complex multi-table SQL promptsSchemaGateSnowflake agent toolsets can silently drop inherited tools when callers lack accessSnowflake DocumentationLooker’s VS Code extension reaches GA with MCP-assisted LookML generation, editing and validationGoogle Cloud Looker release docsIBM makes watsonx Orchestrate AgentOps, custom LLM judging and Bedrock-agent discovery generally availableIBMSchemaGate 0.1.45 fixes broken Oracle ADB wallet connections and an OCI stack pinned 28 releases behindSchemaGatePDI’s Amazon Quick procurement agent grounds spend answers in vendor, category and contract contextAWS Business Intelligence BlogBigQuery’s ML.METRICS example returns 0.84 accuracy but 0.30 macro-F1 on the same 100-row classification queryGoogle Cloud BigQuery docsSchemaGate 0.1.44 auto-selects sentence embeddings, lifting bundled-schema retrieval from 90/98 to 93/98SchemaGateSchemaGate 0.1.43 adds read-only SQL execution with per-principal table checks—and documents unauthenticated client assertionsSchemaGateDatabox adds reusable AI Analyst Skills with personal/company scope, auto-matching and marketplace installsDataboxFabric previews an AI builder for data-agent instructions, source guidance and example queriesMicrosoft FabricDatabricks trains data-agent retriever to stop early or spend bounded extra search steps, reporting 5.8-second latencyDatabricksThoughtSpot adds SpotterCode coding agent to its Visual Embed PlaygroundThoughtSpotLongMemEval-S audit: 67–73% of restore-fixable 80k-budget errors came from evicted evidence under three policiesarXivSchemaGate 0.1.42 adds dimension-aware retrieval and fixes complex multi-table SQL promptsSchemaGateSnowflake agent toolsets can silently drop inherited tools when callers lack accessSnowflake DocumentationLooker’s VS Code extension reaches GA with MCP-assisted LookML generation, editing and validationGoogle Cloud Looker release docs
nl2sql.ai
benchmarkBENCHMARK AUDIT

DataGallery takes BIRD by 0.11 points—and its efficiency score already disagrees

Huawei 2012 Labs’ submission is first on both BIRD execution accuracy and R-VES, but the linked primary pages print different efficiency numbers.

Leaderboard chart highlighting tiny benchmark leads and a score discrepancy.
Chart: figures from the story
By The Benchmark Desk· Sep 9, 2026the quick take — two AI hosts go live when you do

Huawei 2012 Labs’ DataGallery-Text2SQL is the new execution-accuracy leader on BIRD. The live BIRD leaderboard lists the September 7 submission at 78.10% on the 1,534-question development split and 82.39% on the 1,789-question test split. The linked DataGallery submission page prints the same two execution-accuracy results.

That moves DataGallery just ahead of SiriusAI-SQL, which the live leaderboard lists at 77.77% dev and 82.28% test. The winning margin is therefore 0.33 percentage points on dev but only 0.11 points on test. Those are separate splits, not repeated measurements of one set; BIRD identifies the dev and test workloads separately, and DataGallery gives their question counts explicitly.

The efficiency result needs reconciliation

BIRD’s reward-based valid efficiency score, or R-VES, also places DataGallery first—but the two primary pages disagree on the exact value. The BIRD leaderboard currently shows 77.64 for DataGallery, while the DataGallery submission reports 77.69%. That 0.05-point mismatch is too small to change the ranking: BIRD lists the next result, Agentar-Scale-SQL, at 77.00, and SiriusAI-SQL at 74.76. It is still a provenance problem because the benchmark and entrant no longer expose one canonical efficiency result.

The leaderboard notes that its evaluation metrics are checked continually and that scores can change slightly over time. That caveat may explain the mismatch, but neither page currently says which value is newer or why it changed. Until one side is corrected, comparisons should identify the source alongside the R-VES number rather than quoting “DataGallery’s score” without qualification.

A lead without a disclosed method

The result is verified, but the method is not yet auditable. DataGallery says detailed methodology will follow in an arXiv preprint. BIRD currently marks the model size as unknown and the row as using oracle knowledge; its table does not link runnable code for the entry. Those omissions do not negate the official ranking, but they prevent readers from attributing the gain to a model, prompting strategy, retrieval design or candidate-selection budget.

The practical conclusion is narrow: DataGallery leads BIRD execution accuracy on both named splits, and it also leads the leaderboard’s efficiency metric. The execution scores agree across the two primary pages. The R-VES score does not, and the architecture behind all three numbers remains undisclosed.

Filed by The Benchmark Desk. Corrections: desk@nl2sql.ai · Our standards →

comments · 0

    Comments are moderated before they appear. Your email is used once to confirm it is you — never shown, never sold. Corrections and questions get an answer from the desk when we have one.