Choose evidence,
not a leaderboard.
Compare models only when the benchmark contract and execution conditions support it. Every score links to task outcomes, trajectories, raw responses, provenance, and explicit caveats.
Start with the workflow
A model can be strong in one workflow and weak in another. Suite scores are never blended.
Selected evidence
What this proves
Task-level performance under the recorded prompt, tool schema, fixture, seed, runtime, and lifecycle controls.
What it does not prove
General intelligence, unseen-task reliability, another quantization or host, or provider-wide service quality.
Choose models · grouped by provider
Exactness funnel
Core requirements → full contract → strict pass. Bars use both labels and values; the table below is the accessible equivalent.
Quality–speed frontier
Upper-left is better: higher task score with less recorded task time. The dashed line marks non-dominated selected evidence. Provider-reported monetary cost is shown separately below; local infrastructure cost is excluded.
Quality, reliability, speed, and cost
Replicate distributions are median, interquartile range, and min–max. Single runs are labeled as such. Cost is provider-reported workload cost; $0 local does not mean zero infrastructure cost.
| Model | Provider | Status | Score | Strict pass | Total time | Provider cost | Replicates / spread | Evidence |
|---|
Task matrix
Mean task score across selected replicates. The text value is authoritative; color is only a secondary cue.