live wire
IBM makes watsonx Orchestrate AgentOps, custom LLM judging and Bedrock-agent discovery generally availableIBMSchemaGate 0.1.45 fixes broken Oracle ADB wallet connections and an OCI stack pinned 28 releases behindSchemaGatePDI’s Amazon Quick procurement agent grounds spend answers in vendor, category and contract contextAWS Business Intelligence BlogBigQuery’s ML.METRICS example returns 0.84 accuracy but 0.30 macro-F1 on the same 100-row classification queryGoogle Cloud BigQuery docsSchemaGate 0.1.44 auto-selects sentence embeddings, lifting bundled-schema retrieval from 90/98 to 93/98SchemaGateSchemaGate 0.1.43 adds read-only SQL execution with per-principal table checks—and documents unauthenticated client assertionsSchemaGateDatabox adds reusable AI Analyst Skills with personal/company scope, auto-matching and marketplace installsDataboxFabric previews an AI builder for data-agent instructions, source guidance and example queriesMicrosoft FabricDatabricks trains data-agent retriever to stop early or spend bounded extra search steps, reporting 5.8-second latencyDatabricksThoughtSpot adds SpotterCode coding agent to its Visual Embed PlaygroundThoughtSpotLongMemEval-S audit: 67–73% of restore-fixable 80k-budget errors came from evicted evidence under three policiesarXivSchemaGate 0.1.42 adds dimension-aware retrieval and fixes complex multi-table SQL promptsSchemaGateSnowflake agent toolsets can silently drop inherited tools when callers lack accessSnowflake DocumentationLooker’s VS Code extension reaches GA with MCP-assisted LookML generation, editing and validationGoogle Cloud Looker release docsIBM makes watsonx Orchestrate AgentOps, custom LLM judging and Bedrock-agent discovery generally availableIBMSchemaGate 0.1.45 fixes broken Oracle ADB wallet connections and an OCI stack pinned 28 releases behindSchemaGatePDI’s Amazon Quick procurement agent grounds spend answers in vendor, category and contract contextAWS Business Intelligence BlogBigQuery’s ML.METRICS example returns 0.84 accuracy but 0.30 macro-F1 on the same 100-row classification queryGoogle Cloud BigQuery docsSchemaGate 0.1.44 auto-selects sentence embeddings, lifting bundled-schema retrieval from 90/98 to 93/98SchemaGateSchemaGate 0.1.43 adds read-only SQL execution with per-principal table checks—and documents unauthenticated client assertionsSchemaGateDatabox adds reusable AI Analyst Skills with personal/company scope, auto-matching and marketplace installsDataboxFabric previews an AI builder for data-agent instructions, source guidance and example queriesMicrosoft FabricDatabricks trains data-agent retriever to stop early or spend bounded extra search steps, reporting 5.8-second latencyDatabricksThoughtSpot adds SpotterCode coding agent to its Visual Embed PlaygroundThoughtSpotLongMemEval-S audit: 67–73% of restore-fixable 80k-budget errors came from evicted evidence under three policiesarXivSchemaGate 0.1.42 adds dimension-aware retrieval and fixes complex multi-table SQL promptsSchemaGateSnowflake agent toolsets can silently drop inherited tools when callers lack accessSnowflake DocumentationLooker’s VS Code extension reaches GA with MCP-assisted LookML generation, editing and validationGoogle Cloud Looker release docs
nl2sql.ai
benchmark tracker

Reference points and reported results on the beat's two load-bearing evaluations. Every number links to its submission, paper, or leaderboard. Maintained continuously by the benchmark desk.

Arcwise-Plat-SQL

official leaderboard ↗
systemmetricvaluereportedsource
ReViSQL-BIRD-K2.6Accuracy (498-question eval, SQL-only corrections), SC-16 T=192.97Aug 27, 2026link ↗
ReViSQL-BIRD-K2.6Accuracy (498-question eval, SQL-only corrections), greedy T=091.37Aug 27, 2026link ↗

BEAVER

official leaderboard ↗
systemmetricvaluereportedsource
Claude Sonnet 4.5 + optimized cards and raw SQLExecution accuracy, dw dev held-out subset (n=300)9.00Aug 24, 2026link ↗
Claude Sonnet 4.5 few-shot reproductionExecution accuracy, dw dev held-out subset (n=300)6.33Aug 24, 2026link ↗

BEIR

official leaderboard ↗
systemmetricvaluereportedsource
scrydb Hamming + cosint8 (Qwen3-Embedding-8B)FiQA test nDCG@100.65Aug 26, 2026link ↗

BFSI business-query QA

official leaderboard ↗
systemmetricvaluereportedsource
CypherAgent-[ME]-[FL]-[CE] + GPT-4oBFSI all-query numeric-answer accuracy (%)85.57Jun 17, 2026link ↗
SQLAgent (ReAct)BFSI all-query numeric-answer accuracy (%)73.45Jun 17, 2026link ↗

BIRD

Big Bench for Large-Scale Database Grounded Text-to-SQL Evaluation: 12,751 questions over 95 databases (33 GB) across 37 domains, emphasizing dirty values, external knowledge, and SQL efficiency. The de facto standard leaderboard for single-database text-to-SQL.

official leaderboard ↗
systemmetricvaluereportedsource
CHESS on SQLMorph JQE expansion setEX, dev-derived expansion set (n=58)44.83Sep 8, 2026link ↗
MAC-SQL on SQLMorph JQE expansion setEX, dev-derived expansion set (n=58)39.66Sep 8, 2026link ↗
DIN-SQL on SQLMorph JQE expansion setEX, dev-derived expansion set (n=58)32.76Sep 8, 2026link ↗
DataGallery-Text2SQLTest execution accuracy (EX), 1,789 questions82.39Sep 7, 2026link ↗
DataGallery-Text2SQLExecution Accuracy (test split)82.39Sep 7, 2026link ↗
DataGallery-Text2SQLExecution Accuracy (dev split)78.10Sep 7, 2026link ↗
DataGallery-Text2SQLDev execution accuracy (EX), 1,534 questions78.10Sep 7, 2026link ↗
MaP-SQL + Agentar-Scale-SQL-Generation-32B (32 candidates)Execution accuracy (dev)73.08Sep 7, 2026link ↗
SQL-Zero-7B iter3Execution accuracy (dev, greedy@1)58.40Sep 4, 2026link ↗
SQL-Zero-3B iter2Execution accuracy (dev, greedy@1)45.30Sep 4, 2026link ↗
DataGallery-Text2SQLExecution Accuracy — test split (%)82.22Sep 2, 2026link ↗
DataGallery-Text2SQL (Huawei 2012 Labs, Sep. 2 submission)Execution Accuracy (EX), test split82.22Sep 2, 2026link ↗
DataGallery-Text2SQLTest execution accuracy (EX)82.22Sep 2, 2026link ↗
DataGallery-Text2SQLTest reward-based valid efficiency score (R-VES)77.92Sep 2, 2026link ↗
DataGallery-Text2SQLR-VES — test split (%)77.92Sep 2, 2026link ↗
DataGallery-Text2SQLExecution Accuracy — dev split (%)77.71Sep 2, 2026link ↗
DataGallery-Text2SQL (Huawei 2012 Labs, Sep. 2 submission)Execution Accuracy (EX), dev split77.71Sep 2, 2026link ↗
DataGallery-Text2SQLDev execution accuracy (EX)77.71Sep 2, 2026link ↗
SiriusAI-SQLTest execution accuracy (official held-out test)82.28Sep 1, 2026link ↗
SiriusAI-SQLDev execution accuracy77.77Sep 1, 2026link ↗
Reflect-SQL + Claude Sonnet 4.5Execution accuracy (dev)72.03Sep 1, 2026link ↗
SQL-Trail-7B (majority vote)BIRD dev execution accuracy (%)64.20Aug 31, 2026link ↗
SQL-Trail-7BExecution accuracy (dev, majority vote)64.20Aug 31, 2026link ↗
SQL-Trail-7BExecution accuracy (dev, greedy)60.10Aug 31, 2026link ↗
ReToolSQL SFT→RFT + SC@16Dev execution accuracy (EX)74.77Aug 28, 2026link ↗

BIRD dev generated-SQL verification

official leaderboard ↗
systemmetricvaluereportedsource
TraceSQLF1 on 1,521 generated candidates / 11 held-out dev DBs66.47Aug 18, 2026link ↗
TraceSQLROC-AUC on 1,521 generated candidates / 11 held-out dev DBs64.48Aug 18, 2026link ↗

BIRD-History

official leaderboard ↗
systemmetricvaluereportedsource
OpenSearch-SQL with gold history retrievalExecution Accuracy (full 1,393-task set)67.34Aug 29, 2026link ↗
OpenSearch-SQL + BIRD-History retriever (Q+C)Execution Accuracy (full 1,393-task set)60.95Aug 29, 2026link ↗
OpenSearch-SQL without history retrievalExecution Accuracy (full 1,393-task set)54.27Aug 29, 2026link ↗
DAIL-SQL + BIRD-History retrieverExecution Accuracy (full 1,393-task set)51.33Aug 29, 2026link ↗

BIRD Mini-Dev PostgreSQL

official leaderboard ↗
systemmetricvaluereportedsource
Ontology2SQL + DeepSeek V4 Flash, deterministic dialect adaptationExecution accuracy (public Mini-Dev, PostgreSQL replay)65.80Aug 21, 2026link ↗

BIRD Mini-Dev SQLite

official leaderboard ↗
systemmetricvaluereportedsource
Ontology2SQL + DeepSeek V4 FlashExecution accuracy (public Mini-Dev, SQLite)70.20Aug 21, 2026link ↗

BIRD Schema Representation Study

official leaderboard ↗
systemmetricvaluereportedsource
Gemini 2.5 FlashL1·S3 minus L6·S1 EX gap (pp), corrected 397-question subset18.20Aug 21, 2026link ↗
Phi-4L1·S3 minus L6·S1 EX gap (pp), corrected 397-question subset11.30Aug 21, 2026link ↗
Qwen2.5-Coder-14BL1·S3 minus L6·S1 EX gap (pp), corrected 397-question subset10.60Aug 21, 2026link ↗
OLMo-2-13B-InstructL1·S3 minus L6·S1 EX gap (pp), corrected 397-question subset2.70Aug 21, 2026link ↗

BIRD-SQL

official leaderboard ↗
systemmetricvaluereportedsource
RAS (Adya AI)Execution Accuracy (EX), test split79.82Aug 19, 2026link ↗
RAS (Adya AI)Execution accuracy — test split (%)79.82Aug 19, 2026link ↗
RAS (Adya AI)Reward-based Valid Efficiency Score (R-VES), test split74.95Aug 19, 2026link ↗
RAS (Adya AI)Execution accuracy — dev split (%)72.49Aug 19, 2026link ↗
RAS (Adya AI)Execution Accuracy (EX), original dev split72.49Aug 19, 2026link ↗
EGV-SQL + Qwen3.6-27B (Tampere University)Execution accuracy — test split (%)72.16Aug 17, 2026link ↗
EGV-SQL + Qwen3.6-27B (Tampere University)Execution accuracy — dev split (%)71.64Aug 17, 2026link ↗

BIRD-SQL Execution Accuracy

official leaderboard ↗
systemmetricvaluereportedsource
AskData + GPT-4oTest execution accuracy (%)81.95Dec 16, 2025link ↗
Agentar-Scale-SQLTest execution accuracy (%)81.67Sep 25, 2025link ↗

BIRD-SQL Reward-based Valid Efficiency Score

official leaderboard ↗
systemmetricvaluereportedsource
AskData + GPT-4oTest R-VES76.31Dec 16, 2025link ↗
Agentar-Scale-SQLTest R-VES77.00Sep 25, 2025link ↗

Bolo Model Remediation

official leaderboard ↗
systemmetricvaluereportedsource
Bolo agentic remediationType II runtime-error-free pipeline coverage97.27Aug 29, 2026link ↗
Bolo after HalluVer filteringType II retained coverage after hallucination filter88.50Aug 29, 2026link ↗
Bolo agentic remediationType III runtime-error-free pipeline coverage86.08Aug 29, 2026link ↗

BranchBench

official leaderboard ↗
systemmetricvaluereportedsource
Git4DataSF100 data_cleaning warm speedup vs DoltDB18.50Sep 3, 2026link ↗
Git4DataSF100 failure_repro warm speedup vs DoltDB8.40Sep 3, 2026link ↗

CLEVER semantic-cache evaluation

official leaderboard ↗
systemmetricvaluereportedsource
LFU (MiniLM, LMSYS, 10% cache)Raw cache hit rate (%)57.00Aug 20, 2026link ↗

CorpFam

official leaderboard ↗
systemmetricvaluereportedsource
MiniLM-L6-v2 sentence embedding matcherInvisible-name stratum recall (test, %)4.70Sep 3, 2026link ↗

Cortex AISQL compositional online-learning

official leaderboard ↗
systemmetricvaluereportedsource
Larch-Sel + GAMCALAnalytical independence upper bound; N=1M, 5-predicate workload11.40Aug 27, 2026link ↗
Larch-Sel + GAMCALAnalytical realistic cost reduction; N=1M, 5-predicate workload7.80Aug 27, 2026link ↗

D² Data Investigations (50-case benchmark)

official leaderboard ↗
systemmetricvaluereportedsource
Correctness F1 (%) — all 50 cases95.26Sep 3, 2026link ↗
OpenClaw + Claude Sonnet 5Correctness F1 (%) — all 50 cases94.64Sep 3, 2026link ↗
Claude Code + Claude Sonnet 5Correctness F1 (%) — all 50 cases93.37Sep 3, 2026link ↗
Completeness F1 (%) — all 50 cases76.02Sep 3, 2026link ↗
Verifiability ease (%) — all 50 cases68.88Sep 3, 2026link ↗
OpenClaw + Claude Sonnet 5Completeness F1 (%) — all 50 cases32.35Sep 3, 2026link ↗
Claude Code + Claude Sonnet 5Completeness F1 (%) — all 50 cases32.22Sep 3, 2026link ↗
Claude Code + Claude Sonnet 5Completeness F1 (%) — all 50 cases32.00Sep 3, 2026link ↗
OpenClaw + Claude Sonnet 5Completeness F1 (%) — all 50 cases32.00Sep 3, 2026link ↗
Claude Code + Claude Sonnet 5Verifiability ease (%) — all 50 cases2.91Sep 3, 2026link ↗
OpenClaw + Claude Sonnet 5Verifiability ease (%) — all 50 cases2.22Sep 3, 2026link ↗

DAGSmith dbt pipeline optimization

official leaderboard ↗
systemmetricvaluereportedsource
DAGSmithTuva 255-model, frequency-weighted BigQuery slot-time reduction (%)67.70Aug 23, 2026link ↗
DAGSmithStripe 65-model project, BigQuery frequency-weighted slot-time reduction (%)54.80Aug 23, 2026link ↗
DAGSmithTuva 255-model claims module, BigQuery single-run slot-time reduction (%)51.10Aug 23, 2026link ↗
DAGSmithStripe 65-model project, BigQuery single-run slot-time reduction (%)47.40Aug 23, 2026link ↗
DAGSmithTuva 255-model, frequency-weighted elapsed-time reduction (%)42.60Aug 23, 2026link ↗

Data-Agent Benchmark

official leaderboard ↗
systemmetricvaluereportedsource
GPT-5.6 Sol Codex agent — no persistent contextAverage turns/query, 42-task held-out split5.90Sep 2, 2026link ↗
GPT-5.6 Sol Codex agent — self-curated schema contextAverage turns/query, 42-task held-out split4.60Sep 2, 2026link ↗
GPT-5.6 Sol Codex agent — self-curated latency contextAverage turns/query, 42-task held-out split4.50Sep 2, 2026link ↗

DataBench v1

official leaderboard ↗
systemmetricvaluereportedsource
Opus 5 HighBlended score — 100 tasks88.00Aug 13, 2026link ↗
Fable 5 MaxBlended score — 100 tasks85.00Aug 13, 2026link ↗
GPT-5.6 Sol XHighBlended score — 100 tasks75.00Aug 13, 2026link ↗
GPT-5.6 Luna XHighBlended score — 100 tasks73.00Aug 13, 2026link ↗
Opus 5 MaxBlended score — 100 tasks70.00Aug 13, 2026link ↗

DataKernelBench

official leaderboard ↗
systemmetricvaluereportedsource
GPT-5.5 — CUDA, full-queryOverall speedup vs compiled TorchPlan — TPC-H SF10, 22 queries, 1 H1002.11Aug 25, 2026link ↗
Claude Sonnet 4.6 — Triton, full-queryOverall speedup vs compiled TorchPlan — TPC-H SF10, 22 queries, 1 H1001.54Aug 25, 2026link ↗
Claude Opus 4.7 — CUDA, full-queryOverall speedup vs compiled TorchPlan — TPC-H SF10, 22 queries, 1 H1001.51Aug 25, 2026link ↗
Gemini 3.1 Pro Preview — CUDA, full-queryOverall speedup vs compiled TorchPlan — TPC-H SF10, 22 queries, 1 H1001.44Aug 25, 2026link ↗
Claude Haiku 4.5 — Triton, full-queryOverall speedup vs compiled TorchPlan — TPC-H SF10, 22 queries, 1 H1001.30Aug 25, 2026link ↗
Qwen3.5-397B-A17B-FP8 — Triton, full-queryOverall speedup vs compiled TorchPlan — TPC-H SF10, 22 queries, 1 H1001.26Aug 25, 2026link ↗
GPT-OSS-120B — CUDA, full-queryOverall speedup vs compiled TorchPlan — TPC-H SF10, 22 queries, 1 H1001.26Aug 25, 2026link ↗
DeepSeek-V4-Flash — Triton, coreOverall speedup vs compiled TorchPlan — TPC-H SF10, 22 queries, 1 H1001.23Aug 25, 2026link ↗
MiniMax-M2.5 — Triton, full-queryOverall speedup vs compiled TorchPlan — TPC-H SF10, 22 queries, 1 H1001.19Aug 25, 2026link ↗
Devstral-2-123B-Instruct-2512 — Triton, full-queryOverall speedup vs compiled TorchPlan — TPC-H SF10, 22 queries, 1 H1001.08Aug 25, 2026link ↗

DBcover SQL test generation

official leaderboard ↗
systemmetricvaluereportedsource
DBcoverMySQL 8.0.33 line coverage after 24 hours82.30Aug 27, 2026link ↗
DBcoverPostgreSQL 17.0 line coverage after 24 hours80.10Aug 27, 2026link ↗

DevRev NL2SQL Benchmark

official leaderboard ↗
systemmetricvaluereportedsource
DevRev cost-aware NL-to-SQL architectureAnswer correctness (private 900-query set)91.70Sep 4, 2026link ↗

DRL Enterprise NL2SQL Verification Suite

official leaderboard ↗
systemmetricvaluereportedsource
GPT-4oPostgreSQL execution match, corrected harness, schema-linked, 1,000-pair suite52.90Aug 28, 2026link ↗
Claude Sonnet 4.5PostgreSQL execution match, corrected harness, schema-linked, 1,000-pair suite52.80Aug 28, 2026link ↗
Gemini 2.5 FlashPostgreSQL execution match, corrected harness, schema-linked, 1,000-pair suite52.10Aug 28, 2026link ↗

DS-NL2SQL

official leaderboard ↗
systemmetricvaluereportedsource
Dial (Qwen-3-Max)strict all-six-engine executability, test (%)97.33Mar 8, 2026link ↗
Dial (Qwen-3-Max)strict all-six-engine execution accuracy, test (%)48.39Mar 8, 2026link ↗

EHRSQL

official leaderboard ↗
systemmetricvaluereportedsource
MaP-SQL + Arctic-Text2SQL-R1-7B (32 candidates)Execution accuracy44.71Sep 7, 2026link ↗

Enterprise Context Artifacts

official leaderboard ↗
systemmetricvaluereportedsource
Claude Sonnet 4.6 + optimized SQL reference cardsAST similarity, internal production holdout (n=517)0.55Aug 24, 2026link ↗
Qwen Coder 3-30B + optimized SQL reference cardsAST similarity, internal production holdout (n=517)0.51Aug 24, 2026link ↗

EXPLAIN Yourself planner-stall search

official leaderboard ↗
systemmetricvaluereportedsource
GPT-5.5 agentic SQL search across seven DBMSesDBMSes with ≥1 EXPLAIN planning time >3 min; TPC-H SF100; no split7.00Aug 24, 2026link ↗

Ghost Echoes retrieval deletion audit

official leaderboard ↗
systemmetricvaluereportedsource
ChromaDB 0.4.24 / all-MiniLM-L6-v2median Top-5 retrieval-centroid drift after target deletion0.15Aug 24, 2026link ↗

Hacker News wildcard benchmark

official leaderboard ↗
systemmetricvaluereportedsource
UmbraOursWildcard Join Query 3 throughput, 32 threads (queries/s)4.59Aug 25, 2026link ↗
DuckDB 1.4.4Wildcard Join Query 3 throughput, 32 threads (queries/s)0.15Aug 25, 2026link ↗
UmbraBaseWildcard Join Query 3 throughput, 32 threads (queries/s)0.04Aug 25, 2026link ↗

IDS sliding-window attribution runtime

official leaderboard ↗
systemmetricvaluereportedsource

KnowFeat public tabular suite

official leaderboard ↗
systemmetricvaluereportedsource
KnowFeatAverage rank, frozen 20% test across 12 datasets (5 model seeds)2.30Sep 4, 2026link ↗
KnowFeatChurn AUC, frozen 20% test (5 model seeds)0.85Sep 4, 2026link ↗

KnowFeat Public Tabular Suite

official leaderboard ↗
systemmetricvaluereportedsource
KnowFeatAverage rank on frozen 20% test across 12 datasets (lower is better)2.30Sep 3, 2026link ↗
KnowFeatChurn AUC on frozen 20% test split0.85Sep 3, 2026link ↗

LiveSQLBench Base Lite

official leaderboard ↗
systemmetricvaluereportedsource
o3-miniExecution accuracy — PostgreSQL split47.78Jul 22, 2025link ↗
o3-miniExecution accuracy — SQLite split42.59Jul 22, 2025link ↗
Claude 3.7 SonnetExecution accuracy — SQLite split41.11Jul 22, 2025link ↗
Claude 3.7 SonnetExecution accuracy — PostgreSQL split39.26Jul 22, 2025link ↗
DeepSeek R1-0528Execution accuracy — PostgreSQL split38.14Jul 22, 2025link ↗
DeepSeek R1-0528Execution accuracy — SQLite split32.96Jul 22, 2025link ↗
Mixtral 8x7B InstructExecution accuracy — SQLite split8.89Jul 22, 2025link ↗
Mixtral 8x7B InstructExecution accuracy — PostgreSQL split2.59Jul 22, 2025link ↗

LongMemEval-S

official leaderboard ↗
systemmetricvaluereportedsource
Random eviction, GPT-4o-mini reader, 80k budget, top-k=60 retrievalIrreversible share among restore-corrected errors0.73Sep 8, 2026link ↗
FIFO eviction, GPT-4o-mini reader, 80k budget, top-k=60 retrievalIrreversible share among restore-corrected errors0.71Sep 8, 2026link ↗
Redundancy-aware eviction, GPT-4o-mini reader, 80k budget, top-k=60 retrievalIrreversible share among restore-corrected errors0.67Sep 8, 2026link ↗
LLM-importance eviction, GPT-4o-mini reader, 80k budget, top-k=60 retrievalIrreversible share among restore-corrected errors0.60Sep 8, 2026link ↗

MCP Blueprint Sakila reproducibility benchmark

official leaderboard ↗
systemmetricvaluereportedsource
Verticalized domain tool pack (Approach B)Pooled mean score (17 tasks, four 3B-8B models, 3 reps)0.94Aug 22, 2026link ↗
Raw SQL execute_sql interface (Approach A)Pooled mean score (17 tasks, four 3B-8B models, 3 reps)0.67Aug 22, 2026link ↗
Generic thin-tool pack (Approach C)Pooled mean score (17 tasks, four 3B-8B models, 3 reps)0.61Aug 22, 2026link ↗

MediaSum ISC held-out QA

official leaderboard ↗
systemmetricvaluereportedsource
Contextualized chunks + hybrid retrieval + rerankerAnswer correctness (%) — held-out 499-question split, 2,048-token budget88.00Aug 21, 2026link ↗
Ingest-time semantic compilation (compiled claims)Answer correctness (%) — held-out 499-question split, 2,048-token budget85.20Aug 21, 2026link ↗

MemTrapBench

official leaderboard ↗
systemmetricvaluereportedsource
Gemini-3-Flash-Preview without memoryAverage score across four memory-trap scenarios (%)85.16Aug 20, 2026link ↗
Gemini-3-Flash-Preview + EverMemOSAverage score across four memory-trap scenarios (%)71.17Aug 20, 2026link ↗

MMQA

official leaderboard ↗
systemmetricvaluereportedsource
MDB-Link (Qwen2.5-14B)Exact match (%)51.41Aug 10, 2026link ↗
MDB-Link (Qwen2.5-14B)EM (%)51.41Aug 10, 2026link ↗

MM-quecat multi-model query evaluation

official leaderboard ↗
systemmetricvaluereportedsource
Post-hoc oracle physical-representation selectorMean-latency reduction vs best fixed single-DBMS environment (%)20.00Sep 7, 2026link ↗

PLSQLBench

official leaderboard ↗
systemmetricvaluereportedsource
GPT-5.6-Sol + Codex database agentSpider2-MT-test Mean Test Pass@1 (%)81.29Aug 16, 2026link ↗
GPT-5.4Spider2-MT-test Mean Test Pass@1 (%)73.19Aug 16, 2026link ↗
GPT-5.6-Sol + Codex database agentSpider2-MT-test Episode Pass@1 (%)41.27Aug 16, 2026link ↗

ProcArena Direct

official leaderboard ↗
systemmetricvaluereportedsource
Gemini-3.1 ProExecution accuracy — Direct, pooled PostgreSQL + Oracle62.20Sep 6, 2026link ↗

ProcArena Interactive

official leaderboard ↗
systemmetricvaluereportedsource
Gemini-3.1 ProExecution accuracy — Interactive, pooled PostgreSQL + Oracle57.80Sep 6, 2026link ↗

RelBench beer-churn

official leaderboard ↗
systemmetricvaluereportedsource
MetaSieve + RelGT (3-hop, k=50)Seconds per training epoch1494.00Aug 26, 2026link ↗
MetaSieve + RelGT (3-hop, k=50)Test AUCROC0.80Aug 26, 2026link ↗
Random sampling + RelGT (3-hop, k=50)Test AUCROC0.78Aug 26, 2026link ↗

RTGL RelBench driver-dnf

official leaderboard ↗
systemmetricvaluereportedsource
GraphSAGEtest accuracy (fraction, corrected RTGL task table)0.72Sep 1, 2026link ↗
HGTtest accuracy (fraction, corrected RTGL task table)0.69Sep 1, 2026link ↗

ScienceBenchmark

official leaderboard ↗
systemmetricvaluereportedsource
MIRAExecution Accuracy Improvement (%)8.78Aug 7, 2026link ↗

Self-sizing IBLT production-shape replay

official leaderboard ↗
systemmetricvaluereportedsource
Self-sizing IBLTEnd-to-end speedup vs tuned Merkle-style localization, P1, 10 Mbps1.55Aug 27, 2026link ↗

SmallBank on ChainMaker execution layer

official leaderboard ↗
systemmetricvaluereportedsource
LanternThroughput ktxn/s (Zipf skew 0.9, CFBS enabled)14.78Sep 4, 2026link ↗

Spider

official leaderboard ↗
systemmetricvaluereportedsource
MaP-SQL + Arctic-Text2SQL-R1-7B (32 candidates)Execution accuracy (test)87.59Sep 7, 2026link ↗
SQL-Zero-7B iter1Execution accuracy (test, greedy@1)82.40Sep 4, 2026link ↗
SQL-Zero-3B iter2Execution accuracy (test, greedy@1)74.70Sep 4, 2026link ↗
GPT-5.4 + BIRD-derived E2/S2/G2/R1 pipelineExecution accuracy, dev split87.04Aug 28, 2026link ↗
Gemini-2.5-Flash + BIRD-derived E2/S2/G2/R1 pipelineExecution accuracy, dev split84.72Aug 28, 2026link ↗
Qwen2.5-Instruct-7B, 4-bit, grammar-constrained beam width 4Development execution accuracy (%)66.20Aug 26, 2026link ↗
Qwen2.5-Instruct-3B, 4-bit, grammar-constrained beam width 8Development execution accuracy (%)56.20Aug 26, 2026link ↗
Qwen2.5-Instruct-1.5B, 4-bit, grammar-constrained beam width 8Development execution accuracy (%)50.50Aug 26, 2026link ↗
Qwen2.5-Instruct-0.5B, 4-bit, grammar-constrained beam width 8Development execution accuracy (%)24.50Aug 26, 2026link ↗
Qwen3-8B + LoRA (CoT SFT)test execution accuracy (EX, 2,147 examples)82.24Aug 15, 2026link ↗
Qwen3-8B + LoRA (No-CoT SFT)test execution accuracy (EX, 2,147 examples)77.04Aug 15, 2026link ↗
LLaMA-3.1-8B + LoRA (CoT SFT)test execution accuracy (EX, 2,147 examples)76.01Aug 15, 2026link ↗
LLaMA-3.1-8B + LoRA (No-CoT SFT)test execution accuracy (EX, 2,147 examples)76.01Aug 15, 2026link ↗
GLM-4 (3-shot)test execution accuracy (EX, 2,147 examples)66.28Aug 15, 2026link ↗
DeepSeek V3 (3-shot)test execution accuracy (EX, 2,147 examples)51.47Aug 15, 2026link ↗
SafeQL (hybrid) + DAIL-SQL + GPT-OSS-120BEX (%)91.70Aug 10, 2026link ↗
SafeQL hybrid + DAIL-SQL (GPT-OSS-120B)Execution accuracy (EX, %)91.70Aug 10, 2026link ↗
SafeQL search + DAIL-SQL (GPT-OSS-120B)Execution accuracy (EX, %)91.30Aug 10, 2026link ↗
MERIT + Qwen2.5-7B-InstructExecution Accuracy69.79Aug 6, 2026link ↗
SERL-SQLSpider-Test Execution Accuracy89.92Aug 4, 2026link ↗
AttnLink-SSchema-linking mAP99.22Aug 1, 2026link ↗
SIRIUS-SQL + Gemini-3.1 ProExecution accuracy (test split)91.20May 31, 2026link ↗

Spider 2.0

Successor to Yale's Spider, built on real enterprise workflows: databases on Snowflake/BigQuery with 1,000+ column schemas, multiple dialects, and multi-step agentic tasks. Frontier models scored under 20% at launch — the current reality check for the field.

official leaderboard ↗
systemmetricvaluereportedsource
Genloop's Sentinel Agent v2 ProScore96.70Aug 11, 2026link ↗
Native miniScore96.53Aug 11, 2026link ↗
QUVI-3 + Gemini-3-pro-previewScore94.15Aug 11, 2026link ↗
MDB-Link (Qwen2.5-14B)Spider2-Snow exact match (%)9.17Aug 10, 2026link ↗
MDB-Link (Qwen2.5-14B)EM (%)9.17Aug 10, 2026link ↗
APEX-SQLOfficial leaderboard score (current evaluator)73.13Feb 28, 2026link ↗
o1-preview agentic baseline (paper)task success rate17.00Nov 1, 2024link ↗

Spider 2.0-AIFunc

official leaderboard ↗
systemmetricvaluereportedsource
Claude Opus 4.6 + minimal Spider-AgentExecution accuracy, full 465-instance benchmark70.30Jul 7, 2026link ↗
Claude Sonnet 4.6 + minimal Spider-AgentExecution accuracy, full 465-instance benchmark69.00Jul 7, 2026link ↗
Kimi K2.5 + minimal Spider-AgentExecution accuracy, full 465-instance benchmark58.10Jul 7, 2026link ↗

Spider 2.0-DBT

official leaderboard ↗
systemmetricvaluereportedsource
SignalPilot AgentOfficial leaderboard score (68-example DBT split)65.60Aug 16, 2026link ↗
SignalPilot AgentOfficial leaderboard score (68-example DBT code-agent track)65.60May 29, 2026link ↗
SignalPilot AgentScore65.60May 29, 2026link ↗

Spider 2.0-Lite

official leaderboard ↗
systemmetricvaluereportedsource
Tianqiong Data Agent + GLM 5.2Official leaderboard score (%) — 547-task Lite split76.23Aug 21, 2026link ↗
Tianqiong Data Agent + GLM 5.2Official score (547-example Lite split)76.23Aug 20, 2026link ↗
Tianqiong Data Agent + GLM 5.2Execution accuracy (official leaderboard, 547-example Lite split)76.23Aug 16, 2026link ↗
Tianqiong Data Agent + GLM 5.2Score76.23Jul 28, 2026link ↗
Tianqiong Data Agent + GLM 5.2Official leaderboard score (547-example Lite track)76.23Jul 28, 2026link ↗
Tianqiong Data Agent + GLM 5.2execution success rate (%)76.23Jul 28, 2026link ↗
Tianqiong Data Agent + GLM 5.2official score, 547-example Spider 2.0-Lite set76.23Jul 28, 2026link ↗
Tianqiong Data Agent + GLM 5.2Official score (547-example Lite set)76.23Jul 28, 2026link ↗
DecisionX Agentofficial score, 547-example Spider 2.0-Lite set74.95Jul 28, 2026link ↗
DecisionX AgentOfficial score (547-example Lite set)74.95Jul 28, 2026link ↗
DecisionX AgentScore74.95Jul 28, 2026link ↗

Spider 2.0-Snow

official leaderboard ↗
systemmetricvaluereportedsource
Genloop's Sentinel Agent v2 ProOfficial leaderboard score (%) — 547-task Snow split96.70Aug 21, 2026link ↗
Genloop's Sentinel Agent v2 ProOfficial score (547-example Snow split)96.70Aug 20, 2026link ↗
Genloop's Sentinel Agent v2 ProExecution accuracy (official leaderboard, 547-example Snow split)96.70Aug 16, 2026link ↗
Genloop's Sentinel Agent v2 ProOfficial leaderboard score (547-example Snow track)96.70Mar 1, 2026link ↗
Genloop's Sentinel Agent v2 Proexecution success rate (%)96.70Mar 1, 2026link ↗
Genloop's Sentinel Agent v2 Proofficial score, 547-example Spider 2.0-Snow set96.70Mar 1, 2026link ↗

Spider2-SQLite

official leaderboard ↗
systemmetricvaluereportedsource
AttnLink-SSchema-linking mAP83.29Aug 1, 2026link ↗

Spider Dev

official leaderboard ↗
systemmetricvaluereportedsource
SPOC-SQL (DeepSeek-V3, full simulated intervention)execution accuracy (dev split)95.60Aug 24, 2026link ↗

Spider-Realistic

official leaderboard ↗
systemmetricvaluereportedsource
SPOC-SQL (DeepSeek-V3, full simulated intervention)execution accuracy93.10Aug 24, 2026link ↗

Spider-Syn

official leaderboard ↗
systemmetricvaluereportedsource
SQL-Zero-7B iter1Execution accuracy (test, greedy@1; fixed no-retrieval)73.70Sep 4, 2026link ↗
SQL-Zero-3B iter3Execution accuracy (test, greedy@1; fixed no-retrieval)67.00Sep 4, 2026link ↗

SRL Engine Performance

official leaderboard ↗
systemmetricvaluereportedsource
Eyeleng 1.2.2 vs Comunica query-shacl-rule 1.0.0Speedup, 3,000-rule firing-count workload (median ratio; 25 runs)6.44Sep 3, 2026link ↗
Eyeleng 1.2.2 vs Comunica query-shacl-rule 1.0.0Speedup, 25-node recursive chain workload (median ratio; 25 runs)2.15Sep 3, 2026link ↗

Thinkingbox-bench

official leaderboard ↗
systemmetricvaluereportedsource
GPT-5.4pass@20, at least one success in 20 trials (%)91.12Aug 20, 2026link ↗
GPT-5.4overall pass@1, 507-task test set, 20 trials/task (%)65.36Aug 20, 2026link ↗
GPT-5.4pass^20, success in all 20 trials (%)25.25Aug 20, 2026link ↗

TIDE-Bench

official leaderboard ↗
systemmetricvaluereportedsource
Claude-Sonnet-4.6Dialogue Execution Accuracy — Drift-only split (%)45.10Aug 30, 2026link ↗
Claude-Sonnet-4.6Dialogue Execution Accuracy — Joint chain+drift split (%)27.80Aug 30, 2026link ↗

TPC-H

official leaderboard ↗
systemmetricvaluereportedsource
GaussDB on 32 Kunpeng bare-metal serversComposite QphH@30TB, millions (unaudited)39.51Aug 28, 2026link ↗

YCSB on ChainMaker execution layer

official leaderboard ↗
systemmetricvaluereportedsource
LanternThroughput txn/s (Zipf skew 0.9, 32-core server)6131.00Sep 4, 2026link ↗