live wire
IBM makes watsonx Orchestrate AgentOps, custom LLM judging and Bedrock-agent discovery generally availableIBMSchemaGate 0.1.45 fixes broken Oracle ADB wallet connections and an OCI stack pinned 28 releases behindSchemaGatePDI’s Amazon Quick procurement agent grounds spend answers in vendor, category and contract contextAWS Business Intelligence BlogBigQuery’s ML.METRICS example returns 0.84 accuracy but 0.30 macro-F1 on the same 100-row classification queryGoogle Cloud BigQuery docsSchemaGate 0.1.44 auto-selects sentence embeddings, lifting bundled-schema retrieval from 90/98 to 93/98SchemaGateSchemaGate 0.1.43 adds read-only SQL execution with per-principal table checks—and documents unauthenticated client assertionsSchemaGateDatabox adds reusable AI Analyst Skills with personal/company scope, auto-matching and marketplace installsDataboxFabric previews an AI builder for data-agent instructions, source guidance and example queriesMicrosoft FabricDatabricks trains data-agent retriever to stop early or spend bounded extra search steps, reporting 5.8-second latencyDatabricksThoughtSpot adds SpotterCode coding agent to its Visual Embed PlaygroundThoughtSpotLongMemEval-S audit: 67–73% of restore-fixable 80k-budget errors came from evicted evidence under three policiesarXivSchemaGate 0.1.42 adds dimension-aware retrieval and fixes complex multi-table SQL promptsSchemaGateSnowflake agent toolsets can silently drop inherited tools when callers lack accessSnowflake DocumentationLooker’s VS Code extension reaches GA with MCP-assisted LookML generation, editing and validationGoogle Cloud Looker release docsIBM makes watsonx Orchestrate AgentOps, custom LLM judging and Bedrock-agent discovery generally availableIBMSchemaGate 0.1.45 fixes broken Oracle ADB wallet connections and an OCI stack pinned 28 releases behindSchemaGatePDI’s Amazon Quick procurement agent grounds spend answers in vendor, category and contract contextAWS Business Intelligence BlogBigQuery’s ML.METRICS example returns 0.84 accuracy but 0.30 macro-F1 on the same 100-row classification queryGoogle Cloud BigQuery docsSchemaGate 0.1.44 auto-selects sentence embeddings, lifting bundled-schema retrieval from 90/98 to 93/98SchemaGateSchemaGate 0.1.43 adds read-only SQL execution with per-principal table checks—and documents unauthenticated client assertionsSchemaGateDatabox adds reusable AI Analyst Skills with personal/company scope, auto-matching and marketplace installsDataboxFabric previews an AI builder for data-agent instructions, source guidance and example queriesMicrosoft FabricDatabricks trains data-agent retriever to stop early or spend bounded extra search steps, reporting 5.8-second latencyDatabricksThoughtSpot adds SpotterCode coding agent to its Visual Embed PlaygroundThoughtSpotLongMemEval-S audit: 67–73% of restore-fixable 80k-budget errors came from evicted evidence under three policiesarXivSchemaGate 0.1.42 adds dimension-aware retrieval and fixes complex multi-table SQL promptsSchemaGateSnowflake agent toolsets can silently drop inherited tools when callers lack accessSnowflake DocumentationLooker’s VS Code extension reaches GA with MCP-assisted LookML generation, editing and validationGoogle Cloud Looker release docs
nl2sql.ai
benchmarkLEADERBOARD UPDATE

SiriusAI-SQL is now BIRD’s official test leader at 82.28%

BIRD merged the Tencent system’s row on Sept. 1, moving the held-out test record 0.33 points above AskData + GPT-4o.

Leaderboard numbers showing a pending new top score just above the published leader.
Chart: figures from the story
By The Benchmark Desk· Aug 29, 2026the quick take — two AI hosts, this story only

Update, Sept. 1, 2026: BIRD maintainers merged pull request #219 at 01:13 UTC. The live leaderboard now lists SiriusAI-SQL at 82.28% test execution accuracy (EX) and 77.77% dev EX, making the Aug. 22 submission the official test-set leader.

The merge moves AskData + GPT-4o’s 81.95% test EX out of first place, a margin of 0.33 percentage points. This article originally reported the SiriusAI-SQL result while it was still an unmerged contributor proposal; the maintainer merge and rendered BIRD page now resolve that status question.

What the merged patch changes

The merged diff inserts SiriusAI-SQL into the main BIRD table with 77.77 dev EX and 82.28 test EX, removes bold formatting from AskData’s 81.95, and adds 74.76 R-VES in the efficiency table. It also removes the superseded SiriusAI-Text2SQL-Agent row, which reported 75.35 dev EX and 77.03 test EX.

On those numbers, the update is larger than the headline margin over the previous leader suggests: the same team’s new row improves its previous test score by 5.25 points and its dev score by 2.42 points. The pull request also standardizes the method name and Tencent Data Computing Platform Department attribution across Sirius entries.

Official does not mean reproducible

The overall-leaderboard row lists model size as unknown and does not link code. The merged change does not add a paper, model card, frozen predictions or evaluation artifact for the 82.28 result. A held-out server score can establish rank without those materials, but outsiders still cannot inspect the system or reproduce its pipeline from the leaderboard entry.

That distinction is now the useful one. Before Sept. 1, the question was whether maintainers would accept the contributor’s proposed row. After the merge, 82.28 is the published BIRD test record; the unresolved question is what architecture, model and inference procedure produced it.

Two other BIRD submissions merged during the same queue cycle fill very different scale regimes. A 2B Qwen3.5 + StructCoT system reports 68.47 test EX, while a 394M model trained from scratch reports 33.15 test EX with seven-sample execution-guided voting. Neither challenges the leader, but both make model scale explicit in a way the new record row does not.

The leaderboard has therefore moved, but its provenance gap remains. SiriusAI-SQL is officially first on BIRD’s held-out test split as of Sept. 1, 2026; readers should not infer from that rank that the underlying system is publicly reproducible.

Filed by The Benchmark Desk. Corrections: desk@nl2sql.ai · Our standards →

comments · 0

    Comments are moderated before they appear. Your email is used once to confirm it is you — never shown, never sold. Corrections and questions get an answer from the desk when we have one.