Run-centric benchmark evidence

Choose evidence,
not a leaderboard.

Compare models only when the benchmark contract and execution conditions support it. Every score links to task outcomes, trajectories, raw responses, provenance, and explicit caveats.

30
immutable run bundles
13
models with retained evidence
3
distinct benchmark suites

Start with the workflow

A model can be strong in one workflow and weak in another. Suite scores are never blended.

Selected evidence

Select up to four model runs to compare.

What this proves

Task-level performance under the recorded prompt, tool schema, fixture, seed, runtime, and lifecycle controls.

What it does not prove

General intelligence, unseen-task reliability, another quantization or host, or provider-wide service quality.

Choose models · grouped by provider

Exactness funnel

Core requirements → full contract → strict pass. Bars use both labels and values; the table below is the accessible equivalent.

Quality–speed frontier

Upper-left is better: higher task score with less recorded task time. The dashed line marks non-dominated selected evidence. Provider-reported monetary cost is shown separately below; local infrastructure cost is excluded.

Quality, reliability, speed, and cost

Replicate distributions are median, interquartile range, and min–max. Single runs are labeled as such. Cost is provider-reported workload cost; $0 local does not mean zero infrastructure cost.

ModelProviderStatusScoreStrict passTotal timeProvider costReplicates / spreadEvidence

Task matrix

Mean task score across selected replicates. The text value is authoritative; color is only a secondary cue.

Failure taxonomy