live wire
IBM makes watsonx Orchestrate AgentOps, custom LLM judging and Bedrock-agent discovery generally availableIBMSchemaGate 0.1.45 fixes broken Oracle ADB wallet connections and an OCI stack pinned 28 releases behindSchemaGatePDI’s Amazon Quick procurement agent grounds spend answers in vendor, category and contract contextAWS Business Intelligence BlogBigQuery’s ML.METRICS example returns 0.84 accuracy but 0.30 macro-F1 on the same 100-row classification queryGoogle Cloud BigQuery docsSchemaGate 0.1.44 auto-selects sentence embeddings, lifting bundled-schema retrieval from 90/98 to 93/98SchemaGateSchemaGate 0.1.43 adds read-only SQL execution with per-principal table checks—and documents unauthenticated client assertionsSchemaGateDatabox adds reusable AI Analyst Skills with personal/company scope, auto-matching and marketplace installsDataboxFabric previews an AI builder for data-agent instructions, source guidance and example queriesMicrosoft FabricDatabricks trains data-agent retriever to stop early or spend bounded extra search steps, reporting 5.8-second latencyDatabricksThoughtSpot adds SpotterCode coding agent to its Visual Embed PlaygroundThoughtSpotLongMemEval-S audit: 67–73% of restore-fixable 80k-budget errors came from evicted evidence under three policiesarXivSchemaGate 0.1.42 adds dimension-aware retrieval and fixes complex multi-table SQL promptsSchemaGateSnowflake agent toolsets can silently drop inherited tools when callers lack accessSnowflake DocumentationLooker’s VS Code extension reaches GA with MCP-assisted LookML generation, editing and validationGoogle Cloud Looker release docsIBM makes watsonx Orchestrate AgentOps, custom LLM judging and Bedrock-agent discovery generally availableIBMSchemaGate 0.1.45 fixes broken Oracle ADB wallet connections and an OCI stack pinned 28 releases behindSchemaGatePDI’s Amazon Quick procurement agent grounds spend answers in vendor, category and contract contextAWS Business Intelligence BlogBigQuery’s ML.METRICS example returns 0.84 accuracy but 0.30 macro-F1 on the same 100-row classification queryGoogle Cloud BigQuery docsSchemaGate 0.1.44 auto-selects sentence embeddings, lifting bundled-schema retrieval from 90/98 to 93/98SchemaGateSchemaGate 0.1.43 adds read-only SQL execution with per-principal table checks—and documents unauthenticated client assertionsSchemaGateDatabox adds reusable AI Analyst Skills with personal/company scope, auto-matching and marketplace installsDataboxFabric previews an AI builder for data-agent instructions, source guidance and example queriesMicrosoft FabricDatabricks trains data-agent retriever to stop early or spend bounded extra search steps, reporting 5.8-second latencyDatabricksThoughtSpot adds SpotterCode coding agent to its Visual Embed PlaygroundThoughtSpotLongMemEval-S audit: 67–73% of restore-fixable 80k-budget errors came from evicted evidence under three policiesarXivSchemaGate 0.1.42 adds dimension-aware retrieval and fixes complex multi-table SQL promptsSchemaGateSnowflake agent toolsets can silently drop inherited tools when callers lack accessSnowflake DocumentationLooker’s VS Code extension reaches GA with MCP-assisted LookML generation, editing and validationGoogle Cloud Looker release docs
nl2sql.ai
benchmarkBenchmark movement

BIRD’s new No. 2 misses the lead by 0.06 points on both dev and test

DataGallery-Text2SQL lands just below SiriusAI-SQL on two separately reported execution-accuracy splits, while adding a 77.92 test efficiency score.

Two benchmark rows nearly tied on dev and test scores.
Chart: figures from the story
By The Benchmark Desk· Sep 3, 2026the quick take — two AI hosts, this story only

A near-tie at the top

BIRD added DataGallery-Text2SQL to its leaderboard on September 2, reporting 77.71% execution accuracy on the development split and 82.22% on the test split. The submission is attributed to Huawei 2012 Labs. The repository entry also reports a 77.92% test R-VES, BIRD’s reward-based valid efficiency score (BIRD repository commit).

Those figures place DataGallery second behind SiriusAI-SQL in BIRD’s execution-accuracy table. SiriusAI-SQL, dated August 22, is listed at 77.77% dev EX and 82.28% test EX. DataGallery therefore trails by exactly 0.06 percentage points on dev and 0.06 points on test at the precision displayed by the leaderboard (live BIRD leaderboard).

Why the matching gap matters

The interesting result is not merely that a new row reached second place. The same 0.06-point separation appears on both reported splits. That makes this a cleaner near-tie than cases where systems swap order or show a large dev-to-test divergence. It does not establish statistical equivalence: the public table reports point estimates, not confidence intervals, and the 0.06-point difference should not be interpreted as proof that one system is reliably better.

The result also updates DataGallery’s own position. BIRD still displays an earlier DataGallery-Text2SQL row dated June 9 at 74.64% dev EX and 77.53% test EX. The September 2 row is therefore 3.07 points higher on dev and 4.69 points higher on test than that earlier listing, although the leaderboard page does not expose enough implementation detail in the table itself to attribute the improvement to a specific change (live BIRD leaderboard).

The benchmark takeaway

At the top of BIRD, rank labels now exaggerate a very small displayed-score difference. SiriusAI-SQL remains first and DataGallery-Text2SQL is second, but their reported execution accuracy is separated by six hundredths of a point on each split. For readers comparing systems, the defensible conclusion is narrow: DataGallery has reached the leading cluster, while BIRD’s current public numbers do not justify treating the top two as meaningfully far apart.

The new row’s separate 77.92% test R-VES is also worth retaining alongside EX rather than collapsing the submission to a single headline score, because BIRD reports accuracy and efficiency as distinct metrics (BIRD repository commit).

Filed by The Benchmark Desk. Corrections: desk@nl2sql.ai · Our standards →

comments · 0

    Comments are moderated before they appear. Your email is used once to confirm it is you — never shown, never sold. Corrections and questions get an answer from the desk when we have one.