live wire
IBM makes watsonx Orchestrate AgentOps, custom LLM judging and Bedrock-agent discovery generally availableIBMSchemaGate 0.1.45 fixes broken Oracle ADB wallet connections and an OCI stack pinned 28 releases behindSchemaGatePDI’s Amazon Quick procurement agent grounds spend answers in vendor, category and contract contextAWS Business Intelligence BlogBigQuery’s ML.METRICS example returns 0.84 accuracy but 0.30 macro-F1 on the same 100-row classification queryGoogle Cloud BigQuery docsSchemaGate 0.1.44 auto-selects sentence embeddings, lifting bundled-schema retrieval from 90/98 to 93/98SchemaGateSchemaGate 0.1.43 adds read-only SQL execution with per-principal table checks—and documents unauthenticated client assertionsSchemaGateDatabox adds reusable AI Analyst Skills with personal/company scope, auto-matching and marketplace installsDataboxFabric previews an AI builder for data-agent instructions, source guidance and example queriesMicrosoft FabricDatabricks trains data-agent retriever to stop early or spend bounded extra search steps, reporting 5.8-second latencyDatabricksThoughtSpot adds SpotterCode coding agent to its Visual Embed PlaygroundThoughtSpotLongMemEval-S audit: 67–73% of restore-fixable 80k-budget errors came from evicted evidence under three policiesarXivSchemaGate 0.1.42 adds dimension-aware retrieval and fixes complex multi-table SQL promptsSchemaGateSnowflake agent toolsets can silently drop inherited tools when callers lack accessSnowflake DocumentationLooker’s VS Code extension reaches GA with MCP-assisted LookML generation, editing and validationGoogle Cloud Looker release docsIBM makes watsonx Orchestrate AgentOps, custom LLM judging and Bedrock-agent discovery generally availableIBMSchemaGate 0.1.45 fixes broken Oracle ADB wallet connections and an OCI stack pinned 28 releases behindSchemaGatePDI’s Amazon Quick procurement agent grounds spend answers in vendor, category and contract contextAWS Business Intelligence BlogBigQuery’s ML.METRICS example returns 0.84 accuracy but 0.30 macro-F1 on the same 100-row classification queryGoogle Cloud BigQuery docsSchemaGate 0.1.44 auto-selects sentence embeddings, lifting bundled-schema retrieval from 90/98 to 93/98SchemaGateSchemaGate 0.1.43 adds read-only SQL execution with per-principal table checks—and documents unauthenticated client assertionsSchemaGateDatabox adds reusable AI Analyst Skills with personal/company scope, auto-matching and marketplace installsDataboxFabric previews an AI builder for data-agent instructions, source guidance and example queriesMicrosoft FabricDatabricks trains data-agent retriever to stop early or spend bounded extra search steps, reporting 5.8-second latencyDatabricksThoughtSpot adds SpotterCode coding agent to its Visual Embed PlaygroundThoughtSpotLongMemEval-S audit: 67–73% of restore-fixable 80k-budget errors came from evicted evidence under three policiesarXivSchemaGate 0.1.42 adds dimension-aware retrieval and fixes complex multi-table SQL promptsSchemaGateSnowflake agent toolsets can silently drop inherited tools when callers lack accessSnowflake DocumentationLooker’s VS Code extension reaches GA with MCP-assisted LookML generation, editing and validationGoogle Cloud Looker release docs
nl2sql.ai
benchmarkBENCHMARK WATCH

BIRD’s new top two are 0.06 points apart—but neither links code

SiriusAI-SQL and DataGallery-Text2SQL both clear 82% test execution accuracy; the leaderboard records unknown model size, oracle knowledge, and shrinking public provenance.

Two benchmark leaders differ by 0.06 points.
Chart: figures from the story
By The Benchmark Desk· Sep 5, 2026the quick take — two AI hosts go live when you do

A new high, then an immediate near-tie

BIRD’s execution-accuracy table has a new leading pair. A Sept. 1 repository commit added Tencent’s SiriusAI-SQL at 77.77% on the dev split and 82.28% on the test split. The commit labels the result a new state of the art, while the leaderboard row itself is dated Aug. 22, 2026.

One day later, a second commit added Huawei 2012 Labs’ DataGallery-Text2SQL at 77.71% dev execution accuracy, 82.22% test execution accuracy, and 77.92% test R-VES. On the same test-EX split, Sirius leads DataGallery by only 0.06 percentage points. Sirius is 0.33 points above the previous 81.95% test-EX row for AskData + GPT-4o; DataGallery is 0.27 points above it.

The ranking is precise; the systems are not yet inspectable

Both new overall rows mark model size as UNK, indicate use of oracle knowledge, and provide no code link. Sirius’ commit also removes the earlier SiriusAI-Text2SQL-Agent row—75.35% dev EX and 77.03% test EX—and inserts the renamed SiriusAI-SQL result, a 5.25-point difference on test EX between the removed and replacement rows. In the single-trained-model table, the same commit renames the Sirius entry and removes its prior arXiv paper link.

DataGallery arrived with a Notion technical-report link, but an open pull request filed Sept. 3 asks the benchmark maintainers to “temporarily delete” that link. The patch removes only the linked author/report line; it does not change DataGallery’s scores. As of this check, the pull request remains open and has no discussion explaining the temporary removal.

That does not invalidate either test result: BIRD’s public table is the primary record of the evaluated scores. It does limit what outsiders can audit about the systems behind a margin of six hundredths of a point. Until model configuration, prompts, retrieval components, and code are available, the defensible conclusion is narrow: SiriusAI-SQL is the current BIRD test-EX leader at 82.28%, with DataGallery-Text2SQL at 82.22% on the same split. The leaderboard can order those submissions; it cannot yet show whether the tiny difference is reproducible.

Filed by The Benchmark Desk. Corrections: desk@nl2sql.ai · Our standards →

comments · 0

    Comments are moderated before they appear. Your email is used once to confirm it is you — never shown, never sold. Corrections and questions get an answer from the desk when we have one.