BigQuery’s own example shows why accuracy alone is a weak evaluation
ML.METRICS puts four classification measures in one SQL row; its sample output makes the case for reading all four.
ML.METRICS puts four classification measures in one SQL row; its sample output makes the case for reading all four.
DataGallery-Text2SQL leads both BIRD execution-accuracy splits by narrow margins, while its public submission still promises technical details later.
The paper’s 5.1%–32.5% result is promising, but it measures a warmed system whose evidence checker still skips needed escalation on more than 5% of queries in many datasets.
ML.CORRELATION expands conversational analytics into statistical SQL, but Google publishes a capability contract, not evidence that agents choose the right target, method and dimensions.
The paper’s strongest result comes from 900 private queries, while its Spider 2.0 number uses only a 120-query subset. Both constraints matter more than the headline margin.
Two open benchmark harnesses report large token savings from graph-based schema retrieval—but their small, vendor-authored tests are a starting point, not a verdict.
The open-source engine reports a large Spider 2.0-lite lead over Genie and Cortex Analyst—but its more actionable finding is that dialect fixes, refusal policy and repeatability changed outcomes more than brute-force inference.
Join expansion cut DIN-SQL by 27.24 points on a 58-query stress set, while fine-grained metrics separated over-prediction from under-prediction. But 12.1% of generated prompts lost intent.
Huawei 2012 Labs’ submission is first on both BIRD execution accuracy and R-VES, but the linked primary pages print different efficiency numbers.
CostBench’s continuous-load design is useful, but its composite score, matched pairwise windows and unavailable source data demand a careful reading.
Adding AI.PREDICT to conversational analytics means an answer can be syntactically valid yet statistically poor. Teams should grade the agent and TabFM separately.
The live evaluation manifest supports 68 tasks, but a suite-specific setup page still says 69—a small mismatch with real reproducibility consequences.
A new public entity-resolution benchmark finds that aggregate scores hide an almost total collapse on subsidiaries whose names reveal no shared token with their parent.
The preprint isolates SQL candidate selection on fixed pools, pairing retrieved memories with permutation aggregation to improve accuracy while reducing inference cost.
BlockTabBench spans more than 100 production datasets and changes the winner when data volume and difficulty change. The result is not about NL2SQL—but the evaluation design is directly useful to teams deploying it.
A new 1,393-task benchmark ships with code and clause-level annotations. Its strongest result is not just a gain: gold retrieval leaves another 6.39 points on the table.
SiriusAI-SQL and DataGallery-Text2SQL both clear 82% test execution accuracy; the leaderboard records unknown model size, oracle knowledge, and shrinking public provenance.
A reproducible 25-workload comparison quantifies the price of routing rules through SPARQL, while showing that evaluation strategy can matter more than architecture.
The PVLDB paper’s density-aware vector-result cache beats fixed-threshold baselines, but its synthetic workload and single-thread setup define the limits of the claim.
A merged SiEval port exposes a denominator trap: 207 Snowflake cases are unsupported by the referenced evaluator, but its headline score still divides by all 547 questions.
The released code constructs 100 fixed, benchmark-style cases locally. It contains no sampler, split IDs or seed—and the paper’s 91% figure cannot be a raw score on 50 cases.
The release replaces approximate “rule of three” targets with Wilson-bound minima, and keeps held-out customer content out of the repository.
An 18.5× warm-run lead on data cleaning falls to 8.4× on failure reproduction, and the scaling test explains why.
DataGallery-Text2SQL lands just below SiriusAI-SQL on two separately reported execution-accuracy splits, while adding a 77.92 test efficiency score.
A reproducible oracle separates exact, bounded, indeterminate and anomalous floating-point results—and shows that the engine’s algorithm sets the boundary.
A runnable declarative compiler exposed temporal errors in hand-written benchmark SQL—and shows why corrected labels need a new score baseline.
Data Agent MNIST evaluates model correctness, cost and speed on 201 production-derived questions—and ships a harness for companies to build private benchmarks against their own schemas.
A production benchmark found larger structural gains from optimizing reusable SQL reference cards than from tuning retrieval tools and prompts—but its public BEAVER check remains directional.
Ontology2SQL’s published artifacts turn SQL-dialect portability into a measurable benchmark result: 70.20% EX on SQLite, 65.80% after deterministic PostgreSQL adaptation.
MetaSieve’s strongest result is a measured 11.8× training-time cut with a higher test AUC—but it is a relational-learning benchmark, not an NL2SQL score.
An open patch adds eight synthetic rows so execution evaluation can distinguish `=` from `<>` in one IPL workflow; the fix is not yet official.
The proposed patch would cut 23 of 547 tasks after live-data drift and revoked Snowflake shares; it remains unmerged, so current scores are unchanged.
At 10 Mbps, self-sizing IBLT was 1.55× faster on a large, scattered-difference workload—but not on the tiny clustered case.
An agent repaired most sampled inference pipelines, but the paper’s own hallucination filter removed nearly nine points of Type II coverage—and still needs human review.
BIRD merged the Tencent system’s row on Sept. 1, moving the held-out test record 0.33 points above AskData + GPT-4o.
Fresh entries continue to arrive on BIRD and Spider 2.0, but neither benchmark has crowned a new leader for months.
The paper combines measured component results into a 7.8× realistic estimate, but it has not yet run the composed optimizer end to end at that scale.
A 1,000-question enterprise evaluation shows why SQL post-processing belongs inside the benchmark’s audited correctness boundary.
The 92.97% result is real and reproducible in a released artifact, but its split is Arcwise-Plat-SQL; the 92.96% human reference comes from BIRD Test.
A 24-hour PostgreSQL/MySQL experiment rewards path-aware test generation, but its 80% plateau is a boundary, not a correctness score.
On 22 TPC-H SF10 queries running on one H100, GPT-5.5’s best CUDA configuration delivered 2.112× suite-level speedup with every query passing.
The OOPSLA 2026 artifact does not report an accuracy score. It offers something benchmark suites often lack: machine-checked rules for deciding whether supported GQL queries and result schemas are well formed.
A new eight-dataset BEIR study finds binary-first retrieval can preserve top-10 quality while cutting mean latency—but only inside a clearly bounded, single-machine regime.
The vendor’s new benchmark gives teams useful sizing numbers—but ordered queries, uploads and small agent responses need separate acceptance tests.
A new Hacker News workload exposes how badly nested loops handle SQL LIKE joins, but the headline peak comes from one query, one machine and an Umbra prototype.
Execution accuracy is not enough: a new seven-system study shows generated SQL can consume minutes before the query even starts.
A new reproducibility study finds that carefully designed database tools helped 3B–8B local models outperform both raw SQL access and a thin generic tool layer—but its task-aligned setup limits the claim.
Usage, projected tokens and hourly latency turn production conversations into an evaluation surface—if teams preserve correctness as the gate.
A 499-question MediaSum study finds a large token-efficiency gain, but its strongest accuracy comparison is a statistical tie—and the evaluation remains author-run.
A controlled ChromaDB study separates visible deletion from semantic erasure—but its strongest result is narrower than a cross-vendor privacy verdict.