live wire
IBM makes watsonx Orchestrate AgentOps, custom LLM judging and Bedrock-agent discovery generally availableIBMSchemaGate 0.1.45 fixes broken Oracle ADB wallet connections and an OCI stack pinned 28 releases behindSchemaGatePDI’s Amazon Quick procurement agent grounds spend answers in vendor, category and contract contextAWS Business Intelligence BlogBigQuery’s ML.METRICS example returns 0.84 accuracy but 0.30 macro-F1 on the same 100-row classification queryGoogle Cloud BigQuery docsSchemaGate 0.1.44 auto-selects sentence embeddings, lifting bundled-schema retrieval from 90/98 to 93/98SchemaGateSchemaGate 0.1.43 adds read-only SQL execution with per-principal table checks—and documents unauthenticated client assertionsSchemaGateDatabox adds reusable AI Analyst Skills with personal/company scope, auto-matching and marketplace installsDataboxFabric previews an AI builder for data-agent instructions, source guidance and example queriesMicrosoft FabricDatabricks trains data-agent retriever to stop early or spend bounded extra search steps, reporting 5.8-second latencyDatabricksThoughtSpot adds SpotterCode coding agent to its Visual Embed PlaygroundThoughtSpotLongMemEval-S audit: 67–73% of restore-fixable 80k-budget errors came from evicted evidence under three policiesarXivSchemaGate 0.1.42 adds dimension-aware retrieval and fixes complex multi-table SQL promptsSchemaGateSnowflake agent toolsets can silently drop inherited tools when callers lack accessSnowflake DocumentationLooker’s VS Code extension reaches GA with MCP-assisted LookML generation, editing and validationGoogle Cloud Looker release docsIBM makes watsonx Orchestrate AgentOps, custom LLM judging and Bedrock-agent discovery generally availableIBMSchemaGate 0.1.45 fixes broken Oracle ADB wallet connections and an OCI stack pinned 28 releases behindSchemaGatePDI’s Amazon Quick procurement agent grounds spend answers in vendor, category and contract contextAWS Business Intelligence BlogBigQuery’s ML.METRICS example returns 0.84 accuracy but 0.30 macro-F1 on the same 100-row classification queryGoogle Cloud BigQuery docsSchemaGate 0.1.44 auto-selects sentence embeddings, lifting bundled-schema retrieval from 90/98 to 93/98SchemaGateSchemaGate 0.1.43 adds read-only SQL execution with per-principal table checks—and documents unauthenticated client assertionsSchemaGateDatabox adds reusable AI Analyst Skills with personal/company scope, auto-matching and marketplace installsDataboxFabric previews an AI builder for data-agent instructions, source guidance and example queriesMicrosoft FabricDatabricks trains data-agent retriever to stop early or spend bounded extra search steps, reporting 5.8-second latencyDatabricksThoughtSpot adds SpotterCode coding agent to its Visual Embed PlaygroundThoughtSpotLongMemEval-S audit: 67–73% of restore-fixable 80k-budget errors came from evicted evidence under three policiesarXivSchemaGate 0.1.42 adds dimension-aware retrieval and fixes complex multi-table SQL promptsSchemaGateSnowflake agent toolsets can silently drop inherited tools when callers lack accessSnowflake DocumentationLooker’s VS Code extension reaches GA with MCP-assisted LookML generation, editing and validationGoogle Cloud Looker release docs
nl2sql.ai
benchmarkBenchmark audit

Spider 2.0’s DBT benchmark has 68 tasks—or 69, depending on which README you open

The live evaluation manifest supports 68 tasks, but a suite-specific setup page still says 69—a small mismatch with real reproducibility consequences.

Benchmark denominator mismatch changes the score.
Chart: figures from the story
By The Benchmark Desk· Sep 8, 2026the quick take — two AI hosts go live when you do

A benchmark’s denominator is part of every score. In Spider 2.0-DBT, that denominator is currently documented two ways.

The suite-specific README says Spider 2.0-DBT “contains 69 examples.” The repository’s main README, however, describes DBT as a 68-task code-agent setting. Both pages are live in the same official repository.

The evaluation file supports 68

The strongest evidence is the artifact that determines what gets evaluated. As checked on September 8, 2026, the current DBT gold evaluation manifest contains 68 non-empty JSONL records, each keyed by an instance_id. That agrees with the main overview, not the suite-specific README.

This is not merely a cosmetic typo. With 68 tasks, one solved task is worth about 1.47 percentage points; with 69, it is worth about 1.45 points. A reported count of 20 successes would therefore be 29.41% on 68 tasks but 28.99% on 69. The 0.42-point difference comes entirely from the denominator, not from model behavior.

DBT is not the same task as Spider 2.0-Lite or Snow

The main README labels Spider 2.0-DBT a code-agent task over DuckDB, while Lite and Snow are listed as 547-example text-to-SQL settings. The DBT README further says agents must navigate project code, handle long contexts, and can generate SQL exceeding 100 lines. Its baseline-agent documentation exposes repository actions including shell execution, file creation and file editing.

The DBT evaluation documentation shows that outputs can be direct answers, CSV files or database files, with matching functions for strings, numbers, tables and DuckDB contents. That makes the task count especially important: this is a heterogeneous repository-level suite, not a single SQL-execution test whose denominator can be inferred from a familiar split name.

What reporters and submitters should do

Until the project reconciles the two READMEs, Spider 2.0-DBT results should state the exact artifact revision and denominator used. For the repository state inspected on September 8, the defensible denominator is 68, because that is the number of records in the active gold evaluation manifest. Scores copied with “69 examples” should not be compared without first confirming which suite version produced them.

Filed by The Benchmark Desk. Corrections: desk@nl2sql.ai · Our standards →

comments · 0

    Comments are moderated before they appear. Your email is used once to confirm it is you — never shown, never sold. Corrections and questions get an answer from the desk when we have one.