benchmark tracker
Reference points and reported results on the beat's two load-bearing evaluations. Every number links to its submission, paper, or leaderboard. Maintained continuously by the benchmark desk.
Arcwise-Plat-SQL
official leaderboard ↗BEAVER
official leaderboard ↗BEIR
official leaderboard ↗| system | metric | value | reported | source |
|---|---|---|---|---|
| scrydb Hamming + cosint8 (Qwen3-Embedding-8B) | FiQA test nDCG@10 | 0.65 | Aug 26, 2026 | link ↗ |
BFSI business-query QA
official leaderboard ↗BIRD
Big Bench for Large-Scale Database Grounded Text-to-SQL Evaluation: 12,751 questions over 95 databases (33 GB) across 37 domains, emphasizing dirty values, external knowledge, and SQL efficiency. The de facto standard leaderboard for single-database text-to-SQL.
official leaderboard ↗| system | metric | value | reported | source |
|---|---|---|---|---|
| CHESS on SQLMorph JQE expansion set | EX, dev-derived expansion set (n=58) | 44.83 | Sep 8, 2026 | link ↗ |
| MAC-SQL on SQLMorph JQE expansion set | EX, dev-derived expansion set (n=58) | 39.66 | Sep 8, 2026 | link ↗ |
| DIN-SQL on SQLMorph JQE expansion set | EX, dev-derived expansion set (n=58) | 32.76 | Sep 8, 2026 | link ↗ |
| DataGallery-Text2SQL | Test execution accuracy (EX), 1,789 questions | 82.39 | Sep 7, 2026 | link ↗ |
| DataGallery-Text2SQL | Execution Accuracy (test split) | 82.39 | Sep 7, 2026 | link ↗ |
| DataGallery-Text2SQL | Execution Accuracy (dev split) | 78.10 | Sep 7, 2026 | link ↗ |
| DataGallery-Text2SQL | Dev execution accuracy (EX), 1,534 questions | 78.10 | Sep 7, 2026 | link ↗ |
| MaP-SQL + Agentar-Scale-SQL-Generation-32B (32 candidates) | Execution accuracy (dev) | 73.08 | Sep 7, 2026 | link ↗ |
| SQL-Zero-7B iter3 | Execution accuracy (dev, greedy@1) | 58.40 | Sep 4, 2026 | link ↗ |
| SQL-Zero-3B iter2 | Execution accuracy (dev, greedy@1) | 45.30 | Sep 4, 2026 | link ↗ |
| DataGallery-Text2SQL | Execution Accuracy — test split (%) | 82.22 | Sep 2, 2026 | link ↗ |
| DataGallery-Text2SQL (Huawei 2012 Labs, Sep. 2 submission) | Execution Accuracy (EX), test split | 82.22 | Sep 2, 2026 | link ↗ |
| DataGallery-Text2SQL | Test execution accuracy (EX) | 82.22 | Sep 2, 2026 | link ↗ |
| DataGallery-Text2SQL | Test reward-based valid efficiency score (R-VES) | 77.92 | Sep 2, 2026 | link ↗ |
| DataGallery-Text2SQL | R-VES — test split (%) | 77.92 | Sep 2, 2026 | link ↗ |
| DataGallery-Text2SQL | Execution Accuracy — dev split (%) | 77.71 | Sep 2, 2026 | link ↗ |
| DataGallery-Text2SQL (Huawei 2012 Labs, Sep. 2 submission) | Execution Accuracy (EX), dev split | 77.71 | Sep 2, 2026 | link ↗ |
| DataGallery-Text2SQL | Dev execution accuracy (EX) | 77.71 | Sep 2, 2026 | link ↗ |
| SiriusAI-SQL | Test execution accuracy (official held-out test) | 82.28 | Sep 1, 2026 | link ↗ |
| SiriusAI-SQL | Dev execution accuracy | 77.77 | Sep 1, 2026 | link ↗ |
| Reflect-SQL + Claude Sonnet 4.5 | Execution accuracy (dev) | 72.03 | Sep 1, 2026 | link ↗ |
| SQL-Trail-7B (majority vote) | BIRD dev execution accuracy (%) | 64.20 | Aug 31, 2026 | link ↗ |
| SQL-Trail-7B | Execution accuracy (dev, majority vote) | 64.20 | Aug 31, 2026 | link ↗ |
| SQL-Trail-7B | Execution accuracy (dev, greedy) | 60.10 | Aug 31, 2026 | link ↗ |
| ReToolSQL SFT→RFT + SC@16 | Dev execution accuracy (EX) | 74.77 | Aug 28, 2026 | link ↗ |
BIRD dev generated-SQL verification
official leaderboard ↗BIRD-History
official leaderboard ↗| system | metric | value | reported | source |
|---|---|---|---|---|
| OpenSearch-SQL with gold history retrieval | Execution Accuracy (full 1,393-task set) | 67.34 | Aug 29, 2026 | link ↗ |
| OpenSearch-SQL + BIRD-History retriever (Q+C) | Execution Accuracy (full 1,393-task set) | 60.95 | Aug 29, 2026 | link ↗ |
| OpenSearch-SQL without history retrieval | Execution Accuracy (full 1,393-task set) | 54.27 | Aug 29, 2026 | link ↗ |
| DAIL-SQL + BIRD-History retriever | Execution Accuracy (full 1,393-task set) | 51.33 | Aug 29, 2026 | link ↗ |
BIRD Mini-Dev PostgreSQL
official leaderboard ↗| system | metric | value | reported | source |
|---|---|---|---|---|
| Ontology2SQL + DeepSeek V4 Flash, deterministic dialect adaptation | Execution accuracy (public Mini-Dev, PostgreSQL replay) | 65.80 | Aug 21, 2026 | link ↗ |
BIRD Mini-Dev SQLite
official leaderboard ↗| system | metric | value | reported | source |
|---|---|---|---|---|
| Ontology2SQL + DeepSeek V4 Flash | Execution accuracy (public Mini-Dev, SQLite) | 70.20 | Aug 21, 2026 | link ↗ |
BIRD Schema Representation Study
official leaderboard ↗| system | metric | value | reported | source |
|---|---|---|---|---|
| Gemini 2.5 Flash | L1·S3 minus L6·S1 EX gap (pp), corrected 397-question subset | 18.20 | Aug 21, 2026 | link ↗ |
| Phi-4 | L1·S3 minus L6·S1 EX gap (pp), corrected 397-question subset | 11.30 | Aug 21, 2026 | link ↗ |
| Qwen2.5-Coder-14B | L1·S3 minus L6·S1 EX gap (pp), corrected 397-question subset | 10.60 | Aug 21, 2026 | link ↗ |
| OLMo-2-13B-Instruct | L1·S3 minus L6·S1 EX gap (pp), corrected 397-question subset | 2.70 | Aug 21, 2026 | link ↗ |
BIRD-SQL
official leaderboard ↗| system | metric | value | reported | source |
|---|---|---|---|---|
| RAS (Adya AI) | Execution Accuracy (EX), test split | 79.82 | Aug 19, 2026 | link ↗ |
| RAS (Adya AI) | Execution accuracy — test split (%) | 79.82 | Aug 19, 2026 | link ↗ |
| RAS (Adya AI) | Reward-based Valid Efficiency Score (R-VES), test split | 74.95 | Aug 19, 2026 | link ↗ |
| RAS (Adya AI) | Execution accuracy — dev split (%) | 72.49 | Aug 19, 2026 | link ↗ |
| RAS (Adya AI) | Execution Accuracy (EX), original dev split | 72.49 | Aug 19, 2026 | link ↗ |
| EGV-SQL + Qwen3.6-27B (Tampere University) | Execution accuracy — test split (%) | 72.16 | Aug 17, 2026 | link ↗ |
| EGV-SQL + Qwen3.6-27B (Tampere University) | Execution accuracy — dev split (%) | 71.64 | Aug 17, 2026 | link ↗ |
BIRD-SQL Execution Accuracy
official leaderboard ↗BIRD-SQL Reward-based Valid Efficiency Score
official leaderboard ↗Bolo Model Remediation
official leaderboard ↗| system | metric | value | reported | source |
|---|---|---|---|---|
| Bolo agentic remediation | Type II runtime-error-free pipeline coverage | 97.27 | Aug 29, 2026 | link ↗ |
| Bolo after HalluVer filtering | Type II retained coverage after hallucination filter | 88.50 | Aug 29, 2026 | link ↗ |
| Bolo agentic remediation | Type III runtime-error-free pipeline coverage | 86.08 | Aug 29, 2026 | link ↗ |
BranchBench
official leaderboard ↗CLEVER semantic-cache evaluation
official leaderboard ↗| system | metric | value | reported | source |
|---|---|---|---|---|
| LFU (MiniLM, LMSYS, 10% cache) | Raw cache hit rate (%) | 57.00 | Aug 20, 2026 | link ↗ |
CorpFam
official leaderboard ↗| system | metric | value | reported | source |
|---|---|---|---|---|
| MiniLM-L6-v2 sentence embedding matcher | Invisible-name stratum recall (test, %) | 4.70 | Sep 3, 2026 | link ↗ |
Cortex AISQL compositional online-learning
official leaderboard ↗D² Data Investigations (50-case benchmark)
official leaderboard ↗| system | metric | value | reported | source |
|---|---|---|---|---|
| D² | Correctness F1 (%) — all 50 cases | 95.26 | Sep 3, 2026 | link ↗ |
| OpenClaw + Claude Sonnet 5 | Correctness F1 (%) — all 50 cases | 94.64 | Sep 3, 2026 | link ↗ |
| Claude Code + Claude Sonnet 5 | Correctness F1 (%) — all 50 cases | 93.37 | Sep 3, 2026 | link ↗ |
| D² | Completeness F1 (%) — all 50 cases | 76.02 | Sep 3, 2026 | link ↗ |
| D² | Verifiability ease (%) — all 50 cases | 68.88 | Sep 3, 2026 | link ↗ |
| OpenClaw + Claude Sonnet 5 | Completeness F1 (%) — all 50 cases | 32.35 | Sep 3, 2026 | link ↗ |
| Claude Code + Claude Sonnet 5 | Completeness F1 (%) — all 50 cases | 32.22 | Sep 3, 2026 | link ↗ |
| Claude Code + Claude Sonnet 5 | Completeness F1 (%) — all 50 cases | 32.00 | Sep 3, 2026 | link ↗ |
| OpenClaw + Claude Sonnet 5 | Completeness F1 (%) — all 50 cases | 32.00 | Sep 3, 2026 | link ↗ |
| Claude Code + Claude Sonnet 5 | Verifiability ease (%) — all 50 cases | 2.91 | Sep 3, 2026 | link ↗ |
| OpenClaw + Claude Sonnet 5 | Verifiability ease (%) — all 50 cases | 2.22 | Sep 3, 2026 | link ↗ |
DAGSmith dbt pipeline optimization
official leaderboard ↗| system | metric | value | reported | source |
|---|---|---|---|---|
| DAGSmith | Tuva 255-model, frequency-weighted BigQuery slot-time reduction (%) | 67.70 | Aug 23, 2026 | link ↗ |
| DAGSmith | Stripe 65-model project, BigQuery frequency-weighted slot-time reduction (%) | 54.80 | Aug 23, 2026 | link ↗ |
| DAGSmith | Tuva 255-model claims module, BigQuery single-run slot-time reduction (%) | 51.10 | Aug 23, 2026 | link ↗ |
| DAGSmith | Stripe 65-model project, BigQuery single-run slot-time reduction (%) | 47.40 | Aug 23, 2026 | link ↗ |
| DAGSmith | Tuva 255-model, frequency-weighted elapsed-time reduction (%) | 42.60 | Aug 23, 2026 | link ↗ |
Data-Agent Benchmark
official leaderboard ↗| system | metric | value | reported | source |
|---|---|---|---|---|
| GPT-5.6 Sol Codex agent — no persistent context | Average turns/query, 42-task held-out split | 5.90 | Sep 2, 2026 | link ↗ |
| GPT-5.6 Sol Codex agent — self-curated schema context | Average turns/query, 42-task held-out split | 4.60 | Sep 2, 2026 | link ↗ |
| GPT-5.6 Sol Codex agent — self-curated latency context | Average turns/query, 42-task held-out split | 4.50 | Sep 2, 2026 | link ↗ |
DataBench v1
official leaderboard ↗| system | metric | value | reported | source |
|---|---|---|---|---|
| Opus 5 High | Blended score — 100 tasks | 88.00 | Aug 13, 2026 | link ↗ |
| Fable 5 Max | Blended score — 100 tasks | 85.00 | Aug 13, 2026 | link ↗ |
| GPT-5.6 Sol XHigh | Blended score — 100 tasks | 75.00 | Aug 13, 2026 | link ↗ |
| GPT-5.6 Luna XHigh | Blended score — 100 tasks | 73.00 | Aug 13, 2026 | link ↗ |
| Opus 5 Max | Blended score — 100 tasks | 70.00 | Aug 13, 2026 | link ↗ |
DataKernelBench
official leaderboard ↗| system | metric | value | reported | source |
|---|---|---|---|---|
| GPT-5.5 — CUDA, full-query | Overall speedup vs compiled TorchPlan — TPC-H SF10, 22 queries, 1 H100 | 2.11 | Aug 25, 2026 | link ↗ |
| Claude Sonnet 4.6 — Triton, full-query | Overall speedup vs compiled TorchPlan — TPC-H SF10, 22 queries, 1 H100 | 1.54 | Aug 25, 2026 | link ↗ |
| Claude Opus 4.7 — CUDA, full-query | Overall speedup vs compiled TorchPlan — TPC-H SF10, 22 queries, 1 H100 | 1.51 | Aug 25, 2026 | link ↗ |
| Gemini 3.1 Pro Preview — CUDA, full-query | Overall speedup vs compiled TorchPlan — TPC-H SF10, 22 queries, 1 H100 | 1.44 | Aug 25, 2026 | link ↗ |
| Claude Haiku 4.5 — Triton, full-query | Overall speedup vs compiled TorchPlan — TPC-H SF10, 22 queries, 1 H100 | 1.30 | Aug 25, 2026 | link ↗ |
| Qwen3.5-397B-A17B-FP8 — Triton, full-query | Overall speedup vs compiled TorchPlan — TPC-H SF10, 22 queries, 1 H100 | 1.26 | Aug 25, 2026 | link ↗ |
| GPT-OSS-120B — CUDA, full-query | Overall speedup vs compiled TorchPlan — TPC-H SF10, 22 queries, 1 H100 | 1.26 | Aug 25, 2026 | link ↗ |
| DeepSeek-V4-Flash — Triton, core | Overall speedup vs compiled TorchPlan — TPC-H SF10, 22 queries, 1 H100 | 1.23 | Aug 25, 2026 | link ↗ |
| MiniMax-M2.5 — Triton, full-query | Overall speedup vs compiled TorchPlan — TPC-H SF10, 22 queries, 1 H100 | 1.19 | Aug 25, 2026 | link ↗ |
| Devstral-2-123B-Instruct-2512 — Triton, full-query | Overall speedup vs compiled TorchPlan — TPC-H SF10, 22 queries, 1 H100 | 1.08 | Aug 25, 2026 | link ↗ |
DBcover SQL test generation
official leaderboard ↗DevRev NL2SQL Benchmark
official leaderboard ↗| system | metric | value | reported | source |
|---|---|---|---|---|
| DevRev cost-aware NL-to-SQL architecture | Answer correctness (private 900-query set) | 91.70 | Sep 4, 2026 | link ↗ |
DRL Enterprise NL2SQL Verification Suite
official leaderboard ↗| system | metric | value | reported | source |
|---|---|---|---|---|
| GPT-4o | PostgreSQL execution match, corrected harness, schema-linked, 1,000-pair suite | 52.90 | Aug 28, 2026 | link ↗ |
| Claude Sonnet 4.5 | PostgreSQL execution match, corrected harness, schema-linked, 1,000-pair suite | 52.80 | Aug 28, 2026 | link ↗ |
| Gemini 2.5 Flash | PostgreSQL execution match, corrected harness, schema-linked, 1,000-pair suite | 52.10 | Aug 28, 2026 | link ↗ |
DS-NL2SQL
official leaderboard ↗EHRSQL
official leaderboard ↗| system | metric | value | reported | source |
|---|---|---|---|---|
| MaP-SQL + Arctic-Text2SQL-R1-7B (32 candidates) | Execution accuracy | 44.71 | Sep 7, 2026 | link ↗ |
Enterprise Context Artifacts
official leaderboard ↗EXPLAIN Yourself planner-stall search
official leaderboard ↗| system | metric | value | reported | source |
|---|---|---|---|---|
| GPT-5.5 agentic SQL search across seven DBMSes | DBMSes with ≥1 EXPLAIN planning time >3 min; TPC-H SF100; no split | 7.00 | Aug 24, 2026 | link ↗ |
Ghost Echoes retrieval deletion audit
official leaderboard ↗| system | metric | value | reported | source |
|---|---|---|---|---|
| ChromaDB 0.4.24 / all-MiniLM-L6-v2 | median Top-5 retrieval-centroid drift after target deletion | 0.15 | Aug 24, 2026 | link ↗ |
Hacker News wildcard benchmark
official leaderboard ↗IDS sliding-window attribution runtime
official leaderboard ↗| system | metric | value | reported | source |
|---|
KnowFeat public tabular suite
official leaderboard ↗KnowFeat Public Tabular Suite
official leaderboard ↗LiveSQLBench Base Lite
official leaderboard ↗| system | metric | value | reported | source |
|---|---|---|---|---|
| o3-mini | Execution accuracy — PostgreSQL split | 47.78 | Jul 22, 2025 | link ↗ |
| o3-mini | Execution accuracy — SQLite split | 42.59 | Jul 22, 2025 | link ↗ |
| Claude 3.7 Sonnet | Execution accuracy — SQLite split | 41.11 | Jul 22, 2025 | link ↗ |
| Claude 3.7 Sonnet | Execution accuracy — PostgreSQL split | 39.26 | Jul 22, 2025 | link ↗ |
| DeepSeek R1-0528 | Execution accuracy — PostgreSQL split | 38.14 | Jul 22, 2025 | link ↗ |
| DeepSeek R1-0528 | Execution accuracy — SQLite split | 32.96 | Jul 22, 2025 | link ↗ |
| Mixtral 8x7B Instruct | Execution accuracy — SQLite split | 8.89 | Jul 22, 2025 | link ↗ |
| Mixtral 8x7B Instruct | Execution accuracy — PostgreSQL split | 2.59 | Jul 22, 2025 | link ↗ |
LongMemEval-S
official leaderboard ↗| system | metric | value | reported | source |
|---|---|---|---|---|
| Random eviction, GPT-4o-mini reader, 80k budget, top-k=60 retrieval | Irreversible share among restore-corrected errors | 0.73 | Sep 8, 2026 | link ↗ |
| FIFO eviction, GPT-4o-mini reader, 80k budget, top-k=60 retrieval | Irreversible share among restore-corrected errors | 0.71 | Sep 8, 2026 | link ↗ |
| Redundancy-aware eviction, GPT-4o-mini reader, 80k budget, top-k=60 retrieval | Irreversible share among restore-corrected errors | 0.67 | Sep 8, 2026 | link ↗ |
| LLM-importance eviction, GPT-4o-mini reader, 80k budget, top-k=60 retrieval | Irreversible share among restore-corrected errors | 0.60 | Sep 8, 2026 | link ↗ |
MCP Blueprint Sakila reproducibility benchmark
official leaderboard ↗| system | metric | value | reported | source |
|---|---|---|---|---|
| Verticalized domain tool pack (Approach B) | Pooled mean score (17 tasks, four 3B-8B models, 3 reps) | 0.94 | Aug 22, 2026 | link ↗ |
| Raw SQL execute_sql interface (Approach A) | Pooled mean score (17 tasks, four 3B-8B models, 3 reps) | 0.67 | Aug 22, 2026 | link ↗ |
| Generic thin-tool pack (Approach C) | Pooled mean score (17 tasks, four 3B-8B models, 3 reps) | 0.61 | Aug 22, 2026 | link ↗ |
MediaSum ISC held-out QA
official leaderboard ↗| system | metric | value | reported | source |
|---|---|---|---|---|
| Contextualized chunks + hybrid retrieval + reranker | Answer correctness (%) — held-out 499-question split, 2,048-token budget | 88.00 | Aug 21, 2026 | link ↗ |
| Ingest-time semantic compilation (compiled claims) | Answer correctness (%) — held-out 499-question split, 2,048-token budget | 85.20 | Aug 21, 2026 | link ↗ |
MemTrapBench
official leaderboard ↗MMQA
official leaderboard ↗MM-quecat multi-model query evaluation
official leaderboard ↗| system | metric | value | reported | source |
|---|---|---|---|---|
| Post-hoc oracle physical-representation selector | Mean-latency reduction vs best fixed single-DBMS environment (%) | 20.00 | Sep 7, 2026 | link ↗ |
PLSQLBench
official leaderboard ↗ProcArena Direct
official leaderboard ↗| system | metric | value | reported | source |
|---|---|---|---|---|
| Gemini-3.1 Pro | Execution accuracy — Direct, pooled PostgreSQL + Oracle | 62.20 | Sep 6, 2026 | link ↗ |
ProcArena Interactive
official leaderboard ↗| system | metric | value | reported | source |
|---|---|---|---|---|
| Gemini-3.1 Pro | Execution accuracy — Interactive, pooled PostgreSQL + Oracle | 57.80 | Sep 6, 2026 | link ↗ |
RelBench beer-churn
official leaderboard ↗RTGL RelBench driver-dnf
official leaderboard ↗ScienceBenchmark
official leaderboard ↗| system | metric | value | reported | source |
|---|---|---|---|---|
| MIRA | Execution Accuracy Improvement (%) | 8.78 | Aug 7, 2026 | link ↗ |
Self-sizing IBLT production-shape replay
official leaderboard ↗| system | metric | value | reported | source |
|---|---|---|---|---|
| Self-sizing IBLT | End-to-end speedup vs tuned Merkle-style localization, P1, 10 Mbps | 1.55 | Aug 27, 2026 | link ↗ |
SmallBank on ChainMaker execution layer
official leaderboard ↗| system | metric | value | reported | source |
|---|---|---|---|---|
| Lantern | Throughput ktxn/s (Zipf skew 0.9, CFBS enabled) | 14.78 | Sep 4, 2026 | link ↗ |
Spider
official leaderboard ↗| system | metric | value | reported | source |
|---|---|---|---|---|
| MaP-SQL + Arctic-Text2SQL-R1-7B (32 candidates) | Execution accuracy (test) | 87.59 | Sep 7, 2026 | link ↗ |
| SQL-Zero-7B iter1 | Execution accuracy (test, greedy@1) | 82.40 | Sep 4, 2026 | link ↗ |
| SQL-Zero-3B iter2 | Execution accuracy (test, greedy@1) | 74.70 | Sep 4, 2026 | link ↗ |
| GPT-5.4 + BIRD-derived E2/S2/G2/R1 pipeline | Execution accuracy, dev split | 87.04 | Aug 28, 2026 | link ↗ |
| Gemini-2.5-Flash + BIRD-derived E2/S2/G2/R1 pipeline | Execution accuracy, dev split | 84.72 | Aug 28, 2026 | link ↗ |
| Qwen2.5-Instruct-7B, 4-bit, grammar-constrained beam width 4 | Development execution accuracy (%) | 66.20 | Aug 26, 2026 | link ↗ |
| Qwen2.5-Instruct-3B, 4-bit, grammar-constrained beam width 8 | Development execution accuracy (%) | 56.20 | Aug 26, 2026 | link ↗ |
| Qwen2.5-Instruct-1.5B, 4-bit, grammar-constrained beam width 8 | Development execution accuracy (%) | 50.50 | Aug 26, 2026 | link ↗ |
| Qwen2.5-Instruct-0.5B, 4-bit, grammar-constrained beam width 8 | Development execution accuracy (%) | 24.50 | Aug 26, 2026 | link ↗ |
| Qwen3-8B + LoRA (CoT SFT) | test execution accuracy (EX, 2,147 examples) | 82.24 | Aug 15, 2026 | link ↗ |
| Qwen3-8B + LoRA (No-CoT SFT) | test execution accuracy (EX, 2,147 examples) | 77.04 | Aug 15, 2026 | link ↗ |
| LLaMA-3.1-8B + LoRA (CoT SFT) | test execution accuracy (EX, 2,147 examples) | 76.01 | Aug 15, 2026 | link ↗ |
| LLaMA-3.1-8B + LoRA (No-CoT SFT) | test execution accuracy (EX, 2,147 examples) | 76.01 | Aug 15, 2026 | link ↗ |
| GLM-4 (3-shot) | test execution accuracy (EX, 2,147 examples) | 66.28 | Aug 15, 2026 | link ↗ |
| DeepSeek V3 (3-shot) | test execution accuracy (EX, 2,147 examples) | 51.47 | Aug 15, 2026 | link ↗ |
| SafeQL (hybrid) + DAIL-SQL + GPT-OSS-120B | EX (%) | 91.70 | Aug 10, 2026 | link ↗ |
| SafeQL hybrid + DAIL-SQL (GPT-OSS-120B) | Execution accuracy (EX, %) | 91.70 | Aug 10, 2026 | link ↗ |
| SafeQL search + DAIL-SQL (GPT-OSS-120B) | Execution accuracy (EX, %) | 91.30 | Aug 10, 2026 | link ↗ |
| MERIT + Qwen2.5-7B-Instruct | Execution Accuracy | 69.79 | Aug 6, 2026 | link ↗ |
| SERL-SQL | Spider-Test Execution Accuracy | 89.92 | Aug 4, 2026 | link ↗ |
| AttnLink-S | Schema-linking mAP | 99.22 | Aug 1, 2026 | link ↗ |
| SIRIUS-SQL + Gemini-3.1 Pro | Execution accuracy (test split) | 91.20 | May 31, 2026 | link ↗ |
Spider 2.0
Successor to Yale's Spider, built on real enterprise workflows: databases on Snowflake/BigQuery with 1,000+ column schemas, multiple dialects, and multi-step agentic tasks. Frontier models scored under 20% at launch — the current reality check for the field.
official leaderboard ↗| system | metric | value | reported | source |
|---|---|---|---|---|
| Genloop's Sentinel Agent v2 Pro | Score | 96.70 | Aug 11, 2026 | link ↗ |
| Native mini | Score | 96.53 | Aug 11, 2026 | link ↗ |
| QUVI-3 + Gemini-3-pro-preview | Score | 94.15 | Aug 11, 2026 | link ↗ |
| MDB-Link (Qwen2.5-14B) | Spider2-Snow exact match (%) | 9.17 | Aug 10, 2026 | link ↗ |
| MDB-Link (Qwen2.5-14B) | EM (%) | 9.17 | Aug 10, 2026 | link ↗ |
| APEX-SQL | Official leaderboard score (current evaluator) | 73.13 | Feb 28, 2026 | link ↗ |
| o1-preview agentic baseline (paper) | task success rate | 17.00 | Nov 1, 2024 | link ↗ |
Spider 2.0-AIFunc
official leaderboard ↗| system | metric | value | reported | source |
|---|---|---|---|---|
| Claude Opus 4.6 + minimal Spider-Agent | Execution accuracy, full 465-instance benchmark | 70.30 | Jul 7, 2026 | link ↗ |
| Claude Sonnet 4.6 + minimal Spider-Agent | Execution accuracy, full 465-instance benchmark | 69.00 | Jul 7, 2026 | link ↗ |
| Kimi K2.5 + minimal Spider-Agent | Execution accuracy, full 465-instance benchmark | 58.10 | Jul 7, 2026 | link ↗ |
Spider 2.0-DBT
official leaderboard ↗Spider 2.0-Lite
official leaderboard ↗| system | metric | value | reported | source |
|---|---|---|---|---|
| Tianqiong Data Agent + GLM 5.2 | Official leaderboard score (%) — 547-task Lite split | 76.23 | Aug 21, 2026 | link ↗ |
| Tianqiong Data Agent + GLM 5.2 | Official score (547-example Lite split) | 76.23 | Aug 20, 2026 | link ↗ |
| Tianqiong Data Agent + GLM 5.2 | Execution accuracy (official leaderboard, 547-example Lite split) | 76.23 | Aug 16, 2026 | link ↗ |
| Tianqiong Data Agent + GLM 5.2 | Score | 76.23 | Jul 28, 2026 | link ↗ |
| Tianqiong Data Agent + GLM 5.2 | Official leaderboard score (547-example Lite track) | 76.23 | Jul 28, 2026 | link ↗ |
| Tianqiong Data Agent + GLM 5.2 | execution success rate (%) | 76.23 | Jul 28, 2026 | link ↗ |
| Tianqiong Data Agent + GLM 5.2 | official score, 547-example Spider 2.0-Lite set | 76.23 | Jul 28, 2026 | link ↗ |
| Tianqiong Data Agent + GLM 5.2 | Official score (547-example Lite set) | 76.23 | Jul 28, 2026 | link ↗ |
| DecisionX Agent | official score, 547-example Spider 2.0-Lite set | 74.95 | Jul 28, 2026 | link ↗ |
| DecisionX Agent | Official score (547-example Lite set) | 74.95 | Jul 28, 2026 | link ↗ |
| DecisionX Agent | Score | 74.95 | Jul 28, 2026 | link ↗ |
Spider 2.0-Snow
official leaderboard ↗| system | metric | value | reported | source |
|---|---|---|---|---|
| Genloop's Sentinel Agent v2 Pro | Official leaderboard score (%) — 547-task Snow split | 96.70 | Aug 21, 2026 | link ↗ |
| Genloop's Sentinel Agent v2 Pro | Official score (547-example Snow split) | 96.70 | Aug 20, 2026 | link ↗ |
| Genloop's Sentinel Agent v2 Pro | Execution accuracy (official leaderboard, 547-example Snow split) | 96.70 | Aug 16, 2026 | link ↗ |
| Genloop's Sentinel Agent v2 Pro | Official leaderboard score (547-example Snow track) | 96.70 | Mar 1, 2026 | link ↗ |
| Genloop's Sentinel Agent v2 Pro | execution success rate (%) | 96.70 | Mar 1, 2026 | link ↗ |
| Genloop's Sentinel Agent v2 Pro | official score, 547-example Spider 2.0-Snow set | 96.70 | Mar 1, 2026 | link ↗ |
Spider2-SQLite
official leaderboard ↗| system | metric | value | reported | source |
|---|---|---|---|---|
| AttnLink-S | Schema-linking mAP | 83.29 | Aug 1, 2026 | link ↗ |
Spider Dev
official leaderboard ↗| system | metric | value | reported | source |
|---|---|---|---|---|
| SPOC-SQL (DeepSeek-V3, full simulated intervention) | execution accuracy (dev split) | 95.60 | Aug 24, 2026 | link ↗ |
Spider-Realistic
official leaderboard ↗| system | metric | value | reported | source |
|---|---|---|---|---|
| SPOC-SQL (DeepSeek-V3, full simulated intervention) | execution accuracy | 93.10 | Aug 24, 2026 | link ↗ |
Spider-Syn
official leaderboard ↗SRL Engine Performance
official leaderboard ↗Thinkingbox-bench
official leaderboard ↗TIDE-Bench
official leaderboard ↗TPC-H
official leaderboard ↗| system | metric | value | reported | source |
|---|---|---|---|---|
| GaussDB on 32 Kunpeng bare-metal servers | Composite QphH@30TB, millions (unaudited) | 39.51 | Aug 28, 2026 | link ↗ |
YCSB on ChainMaker execution layer
official leaderboard ↗| system | metric | value | reported | source |
|---|---|---|---|---|
| Lantern | Throughput txn/s (Zipf skew 0.9, 32-core server) | 6131.00 | Sep 4, 2026 | link ↗ |