Run Your Own Benchmarks

Everyone says you should. BenchBox makes it easy to do. Use as CLI, Python library, or MCP for AI assistants. No Docker. No compilers. Just pip install.

CLI, Library, or MCP
# CLI - Quick benchmarking
benchbox run --platform duckdb --benchmark tpch

# MCP - AI assistant integration
"Run TPC-H on DuckDB" # Claude executes via MCP

# Library - Deep Python integration
from benchbox import TPCH
BenchBox - Database Benchmarking Toolkit

What Makes It Simple

From measurements to answers

Explore and Compare Benchmark Results

The Results Explorer turns BenchBox output into a browsable view of performance. Use the public corpus to understand how platforms behave on the same workload, or open your own result file in the browser before you share it.

Open the Results Explorer →

TPC-H · Scale factor 10

Four DuckDB versions, one comparable cohort

View comparison ↗
Platforms
4
Queries
22
Best power
281,041
Phase
Power
VersionGeomeanPower@Size
DuckDB v2.0.0…128 ms 281,041
DuckDB v1.5.5152 ms 236,191
DuckDB v1.4.4168 ms 213,021
DuckDB v1.3.2189 ms 198,175
Latency distributionRange and middle 50% for all 22 queries Explore ↗
v2.0.0…v1.5.5 v1.4.4v1.3.2 70100 150200 300500 ms Latency (log scale)
Cumulative latencyShare of queries completed by each latency Explore ↗
0%25% 50%75% 100% 70100 150200 300500 ms Latency (log scale)
  • v2.0.0…
  • v1.5.5
  • v1.4.4
  • v1.3.2
1

Find relevant runs

Browse by benchmark or platform, then narrow the public corpus to results that match the workload and scale you care about.

2

Compare like with like

Compare compatible runs, inspect query-level timings, and see where one platform gains or loses time instead of relying on a single headline number.

3

Check your own result

Open a BenchBox result JSON file locally in the browser. Review its charts and run receipt without uploading it or adding it to the public rankings.

Benchmarks by Category

TPC Standards

Official industry standards for comparing databases

Academic Benchmarks

Research benchmarks from academia

Industry Benchmarks

Real-world benchmarks from practitioners

Real-World Data

Benchmarks built on public real-world datasets

Time-Series Benchmarks

Workloads for time-series databases and columnar monitoring engines

BenchBox Primitives

Fundamental database operation testing

AI & ML Benchmarks

Vector similarity and other AI-shaped analytical workloads

BenchBox Experimental

Experimental benchmarks for specialized testing

Supported Platforms

Single-Node Analytics Engines

In-process and local columnar databases

Row-Based & Postgres-Compatible

Traditional relational databases

Cloud Data Platforms

Enterprise cloud data warehouses and lakehouses

Distributed Query Engines

Open source MPP and federated SQL engines

Managed Spark Services

Cloud-managed Spark for lakehouse and data lake analytics

Time Series Databases

Optimized for time-stamped data

DataFrame Platforms

Native DataFrame APIs instead of SQL

Open Table Formats

Convert benchmark data to modern columnar formats for optimized storage, ACID transactions, and time travel capabilities.

File Formats

Columnar storage for fast analytics

Table Formats

ACID transactions, time travel, and schema evolution

Convert benchmark data to any format
# Convert TPC-H data to different formats
benchbox convert --input ./tpch_sf1 --format parquet
benchbox convert --input ./tpch_sf1 --format delta
benchbox convert --input ./tpch_sf1 --format iceberg
benchbox convert --input ./tpch_sf1 --format ducklake

AI Assistant Integration

BenchBox includes an MCP (Model Context Protocol) server for Claude Code and other AI assistants. Run benchmarks with natural language instead of memorizing CLI flags.

Instruct a coding agent

Pick a surface, platform, benchmark, and scale. Copy the prompt into your coding agent. Defaults to a safe local DuckDB + TPC-H quickstart.

Quick Setup

Add to Claude Code
claude mcp add benchbox -- uv run python -m benchbox.mcp

What You Can Do

Natural language commands for benchmarking workflows

Discover

"What benchmarks are available?"

"Which platforms support TPC-DS?"

Explore 22 benchmarks and 50 platforms without reading documentation.

Execute

"Run TPC-H on DuckDB at scale 0.1"

"Compare Polars and Pandas on SSB"

Run benchmarks without memorizing CLI syntax or options.

Analyze

"Which queries were slowest?"

"Compare results from my last two runs"

Get AI-powered analysis of performance patterns and regressions.

Get Started

1. Install

uv add benchbox

2. Run as CLI

# Quick TPC-H benchmark on DuckDB
benchbox run --platform duckdb --benchmark tpch --scale 0.1

# Preview on cloud before spending credits
benchbox run --platform databricks --benchmark tpch --dry-run ./preview

3. Or Use as Library

from benchbox import TPCH

# Initialize and generate benchmark data
tpch = TPCH(scale_factor=0.1)
data_files = tpch.generate_data()

# Get schema and queries
create_sql = tpch.get_create_tables_sql()
query = tpch.get_query(1)  # Q1: Pricing Summary Report

Requirements

  • Python 3.11 or higher
  • No external dependencies for data generation
  • Optional: database drivers (duckdb, sqlite3, etc.)