live wire
IBM makes watsonx Orchestrate AgentOps, custom LLM judging and Bedrock-agent discovery generally availableIBMSchemaGate 0.1.45 fixes broken Oracle ADB wallet connections and an OCI stack pinned 28 releases behindSchemaGatePDI’s Amazon Quick procurement agent grounds spend answers in vendor, category and contract contextAWS Business Intelligence BlogBigQuery’s ML.METRICS example returns 0.84 accuracy but 0.30 macro-F1 on the same 100-row classification queryGoogle Cloud BigQuery docsSchemaGate 0.1.44 auto-selects sentence embeddings, lifting bundled-schema retrieval from 90/98 to 93/98SchemaGateSchemaGate 0.1.43 adds read-only SQL execution with per-principal table checks—and documents unauthenticated client assertionsSchemaGateDatabox adds reusable AI Analyst Skills with personal/company scope, auto-matching and marketplace installsDataboxFabric previews an AI builder for data-agent instructions, source guidance and example queriesMicrosoft FabricDatabricks trains data-agent retriever to stop early or spend bounded extra search steps, reporting 5.8-second latencyDatabricksThoughtSpot adds SpotterCode coding agent to its Visual Embed PlaygroundThoughtSpotLongMemEval-S audit: 67–73% of restore-fixable 80k-budget errors came from evicted evidence under three policiesarXivSchemaGate 0.1.42 adds dimension-aware retrieval and fixes complex multi-table SQL promptsSchemaGateSnowflake agent toolsets can silently drop inherited tools when callers lack accessSnowflake DocumentationLooker’s VS Code extension reaches GA with MCP-assisted LookML generation, editing and validationGoogle Cloud Looker release docsIBM makes watsonx Orchestrate AgentOps, custom LLM judging and Bedrock-agent discovery generally availableIBMSchemaGate 0.1.45 fixes broken Oracle ADB wallet connections and an OCI stack pinned 28 releases behindSchemaGatePDI’s Amazon Quick procurement agent grounds spend answers in vendor, category and contract contextAWS Business Intelligence BlogBigQuery’s ML.METRICS example returns 0.84 accuracy but 0.30 macro-F1 on the same 100-row classification queryGoogle Cloud BigQuery docsSchemaGate 0.1.44 auto-selects sentence embeddings, lifting bundled-schema retrieval from 90/98 to 93/98SchemaGateSchemaGate 0.1.43 adds read-only SQL execution with per-principal table checks—and documents unauthenticated client assertionsSchemaGateDatabox adds reusable AI Analyst Skills with personal/company scope, auto-matching and marketplace installsDataboxFabric previews an AI builder for data-agent instructions, source guidance and example queriesMicrosoft FabricDatabricks trains data-agent retriever to stop early or spend bounded extra search steps, reporting 5.8-second latencyDatabricksThoughtSpot adds SpotterCode coding agent to its Visual Embed PlaygroundThoughtSpotLongMemEval-S audit: 67–73% of restore-fixable 80k-budget errors came from evicted evidence under three policiesarXivSchemaGate 0.1.42 adds dimension-aware retrieval and fixes complex multi-table SQL promptsSchemaGateSnowflake agent toolsets can silently drop inherited tools when callers lack accessSnowflake DocumentationLooker’s VS Code extension reaches GA with MCP-assisted LookML generation, editing and validationGoogle Cloud Looker release docs
nl2sql.ai
benchmarkSplit check

ReViSQL reaches a human proxy on corrected mini-dev—not the live BIRD test board

The 92.97% result is real and reproducible in a released artifact, but its split is Arcwise-Plat-SQL; the 92.96% human reference comes from BIRD Test.

Accuracy and cost chart comparing ReViSQL scores to the human proxy.
Chart: figures from the story
By The Benchmark Desk· Aug 28, 2026the quick take — two AI hosts go live when you do

Thinking Machines Lab’s ReViSQL result deserves attention—and a split label in the headline. The team reports 91.37% accuracy with greedy decoding on Arcwise-Plat-SQL at $0.035 per task, rising to 92.97% with 16-sample self-consistency at $0.56 per task. Arcwise-Plat-SQL contains 498 questions whose gold SQL was corrected while the original questions and external knowledge were retained (Thinking Machines Lab; ReViSQL artifact).

The important qualifier is that the 92.96% human reference comes from BIRD Test, not Arcwise-Plat-SQL. The live BIRD page labels that number “Human Performance,” while the ReViSQL paper explicitly describes it as a proxy transferred to its Arcwise evaluation (BIRD leaderboard; technical report v4). The 92.97% result therefore clears the cited human proxy by 0.01 percentage point, but it is not a same-split human-versus-model comparison and is not a win on the live BIRD Test leaderboard. As of our rendered-page check on August 28, 2026, ReViSQL was not listed there.

That distinction does not erase the result. It explains what moved. The project’s second split, Arcwise-Plat-Full, also has 498 questions, but corrects the questions and external knowledge as well as the SQL. ReViSQL-BIRD-K2.6 scores 93.8% greedily and 94.2% with 16-sample self-consistency on that fully corrected split; on SQL-only-corrected Arcwise-Plat-SQL, the corresponding rounded figures are 91.4% and 93.0% (technical report v4). The gap shows why “BIRD” is no longer a sufficient split description.

The released artifact makes the comparison inspectable. It includes 2,462 expert-verified training instances, the two 498-question Arcwise files, training code, inference code, and commands for greedy and 16-candidate evaluation. The authors report correcting SQL errors in 52.1% of audited training instances, question flaws in 26.2%, and external-knowledge errors in 18.2%; categories overlap (ReViSQL artifact; Thinking Machines Lab).

The practical reading is narrower and stronger than “BIRD is solved.” Verified data plus reward shaping produced a released model recipe that reaches the transferred human proxy on a corrected mini-dev variant. Its greedy score remains 1.59 points below that proxy, while 16 generations and majority voting cross it. Until the same system is evaluated on BIRD Test—or humans are measured on the Arcwise split—those are two useful milestones, not one interchangeable leaderboard result.

Filed by The Benchmark Desk. Corrections: desk@nl2sql.ai · Our standards →

comments · 0

    Comments are moderated before they appear. Your email is used once to confirm it is you — never shown, never sold. Corrections and questions get an answer from the desk when we have one.