BenchLM / Data
Query reference

Query reference

What you can query

Free access includes current final BenchAlign scores and ranks, licensed benchmark results, model identities, and benchmark catalog metadata. Pro includes three rolling calendar months by UTC date. Research includes all retained eligible history with no age cap. Both require available coverage for the requested fields.

REST endpoints and MCP tools
DataREST endpointMCP tool
Account usage and resetsGET /v1/usagebenchlm_get_usage
Current BenchAlign rankingsGET /v1/rankings/currentbenchlm_list_current_rankings
One model’s current rankingsGET /v1/models/{identifier}/rankings/currentbenchlm_get_current_model_rankings
Model catalogGET /v1/modelsbenchlm_list_models
One modelGET /v1/models/{identifier}benchlm_get_model
Benchmark catalogGET /v1/benchmarksbenchlm_list_benchmarks
Licensed benchmark result coverageGET /v1/benchmarks/results/coveragebenchlm_benchmark_result_coverage
Current licensed benchmark resultsGET /v1/benchmarks/results/currentbenchlm_list_current_benchmark_results
Published evaluation runs (Pro / Research)GET /v1/history/benchmark_resultsbenchlm_benchmark_result_history
Publisher result snapshots (Pro / Research)GET /v1/history/benchmark_result_snapshotsbenchlm_benchmark_result_snapshot_history
Available historical coverageGET /v1/coveragebenchlm_history_coverage
Stored BenchAlign rank historyGET /v1/history/rank_historybenchlm_rank_history
Benchmark measurements over timeGET /v1/history/benchmark_historybenchlm_benchmark_history
Benchmark additions and changesGET /v1/history/benchmark_eventsbenchlm_benchmark_events
Stored ranking at a dateGET /v1/history/ranking_snapshotbenchlm_ranking_snapshot

Sign in to check available historical methods and dates.

Current rankings take surface, limit and offset. Use catalog IDs to choose a model and benchmark. Ranking history takes modelKey; benchmark measurements take modelKey, category and benchmarkKey. Add since, through and limit to choose a historical page. A ranking snapshot takes at.

Historical queries return up to 500 rows per page. Saved-state history covers up to 90 days per request and uses committed record dates. Source-run history preserves the publisher’s reported run start; use date filters of up to 90 days. ECI release dates do not establish evaluation run times. BullshitBench v2 contains one pinned snapshot with native judge ratings, rates, counts, and published ranks; its artifact generation time is not a per-model run date. Arena overall snapshots preserve the publisher’s rating, rank, confidence interval, variance, and vote count, with three scoring methods kept separate. Publication dates do not establish model run times. A ranking snapshot contains the covered subset and identifies its coverage.

Current ranking responses contain final BenchAlign outputs and their scoring version. Licensed benchmark result responses contain source scores with units, source model profiles, and attribution; filter them by category or boardId. For source-run and publisher snapshot history, use boardId and sourceModelId from the result coverage and current-result tools. Snapshot history accepts since, through, limit, and cursor; optional artifactId selects an exact retained artifact. Historical rankings are stored BenchAlign outputs. Benchmark measurement history contains scores. Published Arena snapshot ranks come from the source’s dated leaderboard; run records do not establish ranks.

Each successful data page uses one read. MCP discovery, coverage and usage checks, and unsuccessful queries do not use your monthly allowance. Free allows 10 requests per minute; Pro and Research allow 60. Use GET /v1/usage or benchlm_get_usage to check both reset times. API key rotation does not reset usage.