Research methodology
This page explains the test inputs, every reported measure, and how to use the results when cost and usable capacity must be weighed together.
Evidence pipeline
Every new suite result is stored as one immutable bundle per tested model and replicate. A bundle includes its manifest, normalized task outcomes, ordered trajectory, raw response data, configuration, lineage, and artifact hashes. The registry is derived from validated bundles; charts are never the source of truth.
- Comparability requires matching suite version, fixture, task, numeric seed, seed-document, prompt, and tool-schema hashes.
- Local-versus-hosted or different runtime comparisons are labeled directional even when the benchmark contract matches.
- Infrastructure failures remain in the scheduled denominator and are separated from model and tool-protocol failures.
- Repeated runs report median, interquartile range, and full range. A single run is explicitly labeled.
Inspect the task-level record in the run explorer or download the machine-readable run registry.
What this research is for
The goal is a specific, repeatable comparison. It does not ask which model is best at everything. It asks what a model can do for the same real work, and what that work costs. A frontier model, a third-party API model, and a self-hosted model can then be compared on the same job.
Objective measurements matter because model names, parameter counts, and price cards do not tell us whether a model completes the work. The same seed, prompts, fixtures, files, rubric, and calculation rules give an apples-to-apples starting point. Human judgment still exists in the rubric, so the page also names the evaluator limits.
Use the results to choose the lowest-cost model that meets the needed standard for a named use case. Do not use a high score as proof that a model is safe, current, or right for every job.
Two separate test tracks
This project has two tracks. They answer different questions, so their scores must not be added or treated as one ranking.
- Assistant track: 32 records: 9 self-hosted local models and 23 OpenRouter models. It tests tool-using assistant work.
- Editorial track: one client content bundle per model. It tests the quality, reading level, and recorded cost of that bundle.
A model may be strong at one track and weak at the other. That is useful information, not a conflict.
Assistant test: same work, fixed conditions
The fixed suite is labeled 2026-08-12. The current data combines a local round and a later OpenRouter extension. Every model received the same 21 tasks: project work, research, structured tool calls, state tracking, English communication, and safe handling of requests.
- Temperature 0: removes random sampling as much as the provider allows. This makes a result easier to repeat.
- Seed 42: fixes the test starting point. It helps make task conditions equal.
- 768-token cap: gives each model the same maximum reply budget. This prevents a model from winning only by writing much more.
- Fresh fixture and chat: each task begins with new task data and no earlier conversation. This prevents hidden memory from helping a later task.
- One request at a time: local models are unloaded, then the runner waits 10 seconds before the next model. This reduces cross-model server state and load effects.
Assistant score dimensions
The weighted assistant score is a 0–100 summary. The weights show which kinds of failure matter most for this use case. A high score is more valuable than a low price only when the model can do the parts of the job you need.
- Outcome — 30%
- Whether the task reaches the requested result. It matters most because an inexpensive answer has little value if the work is not finished.
- Tool use — 25%
- Whether the model calls the right tool with usable inputs and uses the result. It matters when the assistant must act on calendars, files, data, or other systems.
- Grounding — 15%
- Whether the answer stays tied to the supplied facts and tool results. It matters because unsupported claims can create costly follow-up work.
- State — 10%
- Whether the model keeps track of task facts across turns. It matters for multi-step work where lost details cause rework.
- English — 10%
- Whether the answer is clear, direct, and usable in English. It matters because a correct plan must still be understood and acted on.
- Safety — 5%
- Whether the model handles risky or disallowed requests safely. It matters because a low-cost model can have a high operational cost if it creates unsafe output.
- Efficiency — 5%
- Whether the model completes work without wasteful extra turns. It matters because extra turns add time, cost, and failure chances.
Assistant data fields
- Tasks passed and task pass rate
- The number and share of the 21 tasks that passed. This is the clearest view of how often the stated workflow worked.
- Tool-call success rate
- The share of detected tool calls that succeeded. It helps separate good writing from reliable tool execution.
- Run status
okmeans the run completed cleanly.partialmeans a tool-loop limit or API error affected it. There are 5okrecords and 27partialrecords. Partial rows retain their scheduled task slots so failure risk stays visible; do not read them as clean runs.- Cold load, first response, warm request, and total time
- Cold load is startup time. First response is time to the first token after a cold start. The warm OpenClaw-style request measures a later request after the model is ready. Total time combines the recorded parts. These are speed and capacity-planning measures, not quality points. Self-hosted inference is expected to be slower in this environment.
Sources: JSON, CSV, and notes. Fixture hash: ea2601bcb637a9c66563e91015f15a007a0a06f7d88532ec218c8c178903efb9. Task hash: 907017aa0d29639967cd0c0702764b73e7cc8ebbf254625fcc07b63e13b4428e.
Editorial test: one weekly content bundle
Each model receives the same client content-bundle job, pipeline structure, seed, and optional-track settings. A bundle is broader than one article. It shows whether a model can make connected pieces that are useful together.
- Context
- The brief with the audience, topic, and constraints. It sets the same starting facts for each model.
- Lesson
- The teaching piece. It tests whether the model can explain the central idea.
- Long-form letter (
index.md) - The main reader-facing letter. It is scored for argument, grounding, action, reading control, and closing voice.
- Instructions (
INSTRUCTIONS.md) - The practical guide that turns the idea into steps. It is scored for argument, framework, reading control, and voice.
- Worksheet copy, worksheet, and masked worksheet
- Working prompts and their usable forms. They show whether the idea can become an action tool, including a version that hides answers or guidance when needed.
- Newsletter email and community post
- Short distribution pieces that bring the weekly idea to readers in different channels.
- Cover prompt and rendered hero image
- The visual brief and its output. They test whether the model can direct a useful, on-brand visual asset.
- Pipeline configuration
- The recorded run settings. It makes the case study traceable and helps explain differences.
- OpenRouter usage record
- For paid-provider runs, the provider-reported token usage and charge record. This is the source for recorded bundle cost.
The content-quality scorer directly evaluates only index.md and INSTRUCTIONS.md. Readability is calculated for LESSON.md, index.md, INSTRUCTIONS.md, and the combined text. Other bundle pieces are retained as evidence and for practical review; they are not silently turned into rubric points.
Editorial quality rubric
The editorial rubric is versioned and fixed for this round. It keeps rule-based measurements and model-based scores in separate fields. A required criterion can lower the result when it is missed. Category weights show the parts of quality that have more influence on each artifact score.
Letter categories
- Argument — 40%
- The letter has a clear idea that develops. It matters because the main piece must give the reader a reason to change or decide.
- Grounding — 15%
- The letter uses real, concrete life details. It matters because abstract advice is hard to trust or use.
- Action — 15%
- The letter leads to a real choice or next step. It matters because content capacity includes helping a reader do something.
- Readability — 15%
- The letter stays clear without becoming empty. It matters because a strong idea fails if the target reader cannot follow it.
- Closure — 15%
- The ending and tone stay firm without sounding like a guru, preacher, therapist, or hype message. It matters because voice affects trust and brand fit.
Instruction categories
- Argument — 40%
- The guide explains why the work matters. It stops steps from becoming empty commands.
- Framework — 30%
- The guide makes the named method usable. It matters because the reader needs a repeatable process, not only good prose.
- Readability — 15%
- The guide is easy to scan and follow. It matters because instructions must work during action.
- Voice — 15%
- The guide stays direct without hype, therapy language, preaching, or guru language. It matters because trust is part of usefulness.
Every scored criterion
The items below define the hard values in the quality-standards table and radar. A shared item is checked in both scored files unless the text says otherwise.
- Perspective shift
- The piece helps the reader see the problem in a new way. This matters when paying more for insight instead of familiar advice.
- Argument progression
- The reasoning moves in a clear order. This matters because a reader can follow and test the claim.
- Concrete opening — letter only
- The letter begins with a real situation or detail. This matters because it makes the subject immediate and believable.
- Psychological interpretation
- The piece explains the thought, habit, or pressure behind the problem. This matters because useful change needs more than surface tips.
- Behavioral grounding
- The claims connect to actions people can observe. This matters because it makes advice testable in daily life.
- Practical philosophy
- The piece turns a bigger idea into a useful rule. This matters because capacity includes sound judgment, not only instructions.
- Intellectual depth
- The piece goes beyond a simple slogan. This matters because premium model cost should buy more than polished filler.
- Qualification and nuance — letter only
- The letter states limits or exceptions where needed. This matters because overconfident writing can mislead readers.
- Personal responsibility
- The reader has a clear part to play. This matters because the content should support agency, not passive consumption.
- Perspective progression — letter only
- The letter carries the reader from the old view to the new view. This matters because a strong ending depends on a complete change path.
- Decision — letter only
- The letter names the choice the reader should make. This matters because vague inspiration is hard to act on.
- Prescribed action — letter only
- The letter gives a specific next action. This matters because a content bundle must lead to use, not only reflection.
- Concrete and abstract balance — letter only
- The letter balances examples with ideas. This matters because either extreme can make a piece shallow or unclear.
- Everyday realism — letter only
- The advice fits normal time, limits, and pressures. This matters because impractical advice creates no real capacity.
- Intellectual simplicity control
- The language is simple without making the idea childish. This matters because readers need clear, credible writing.
- Sentence rhythm
- Sentence length and shape vary in a useful way. This matters because readable pacing keeps a long piece usable.
- Semantic economy
- Each sentence does useful work with few wasted words. This matters because concise output lowers edit time and delivery cost.
- Anti-motivational tone
- The piece avoids empty cheering or pressure. This matters because hype can damage trust.
- Anti-therapy tone
- The piece avoids acting like treatment or diagnosis. This matters because content should stay in its intended role.
- Anti-preacher tone
- The piece avoids moralizing at the reader. This matters because readers need respect, not a lecture.
- Anti-guru tone
- The piece avoids false certainty or special-authority claims. This matters because reliable content shows limits.
- Closing strength — letter only
- The ending lands the idea with force and clarity. This matters because the final message shapes recall and action.
- Model or framework lands — instructions only
- The named method is explained well enough to use. This matters because the guide must deliver the promised framework.
- Unique letter-specific action — instructions only
- The guide gives an action that fits this letter, not a generic productivity step. This matters because specificity shows real understanding.
- Argument to framework to action — instructions only
- The guide connects the reason, the method, and the next step. This matters because a reader needs a complete path from insight to use.
Readability and text measurements
Markdown presentation syntax is removed first. One deterministic English syllable rule is then applied to every model output. These numbers describe the text; they do not decide whether the advice is true or good. They show likely reader effort and editing effort, which both affect the real cost of a model choice.
- Characters
- The count of letters and numbers in the normalized text. It shows total text size and helps explain processing or editing volume.
- Letters
- Alphabetic characters. It is an input for some reading formulas and helps compare writing density.
- Words
- Detected English word forms. It shows length and is a basic editing-work measure.
- Sentences
- Detected sentences. It helps show how the model breaks ideas into units a reader can follow.
- Paragraphs
- Detected blocks of text. It helps show page structure and scanability.
- Syllables
- Estimated spoken parts of words. More syllables often mean harder words and more reading effort.
- Polysyllables
- Words with three or more syllables. It is a direct signal used by some grade-level formulas.
- Long words
- Words with seven or more letters. It is a simple sign of dense language that may need editing.
- Long sentences
- Sentences with 20 or more words. It shows where readers may need to hold too much in mind.
- Average words per sentence
- Words divided by sentences. Higher values often raise reading effort, so it helps compare pacing.
- Average syllables per word
- Syllables divided by words. Higher values often mean more complex vocabulary.
- Average characters per word
- Letters divided by words. It is another simple view of word length and density.
- Flesch reading ease
- A 0–100 style formula based on sentence and word length; higher is easier. It gives a quick direction for general ease.
- Flesch-Kincaid grade
- An estimated U.S. school grade needed to read the text; lower is easier. Grade 6 is a reference target, not a pass rule.
- Gunning fog
- An estimated school grade based on sentence length and complex words; lower is easier. It checks whether long words are making prose heavy.
- Coleman-Liau
- An estimated school grade based on letters and sentences; lower is easier. It gives a second view that does not use syllable estimates.
- Automated readability index
- An estimated school grade based on letters and words; lower is easier. It helps catch dense writing with long words.
- SMOG
- An estimated school grade based mainly on polysyllables; lower is easier. It highlights complex vocabulary.
- Lexical diversity
- The share of distinct words. It shows word variety, but very high variety is not always better if it makes text needlessly hard.
The readability radar converts each raw measure to a within-measure rank so unlike units can share one shape. It is for quick comparison only. Use the raw tables and reading report for decisions.
Cost, capacity, and chart rules
Cost is the recorded price of one completed editorial bundle when provider usage records exist. It is not a price-per-token estimate and it is not a full business cost. Capacity means the measured ability to complete this specific work at the needed quality and reliability.
- Frontier and paid third-party models
- These models may offer high measured quality or speed, but their recorded provider charge can be higher. The value question is whether better output saves enough editing, review, or failure cost to justify that charge.
- Self-hosted local models
- These models run on the research machine. The editorial chart marks them as $0 direct provider cost because no paid API bill is recorded. That is not a claim that local inference is free: hardware, power, setup, maintenance, and slower local time are outside this field.
- Cost burden
- The radar reverses recorded cost into a rank where lower cost gives a better value. It is a visual aid, not a quality score. A cheap model with weak quality is not automatically the better choice.
- Quality-and-cost frontier
- The decision chart highlights observed models that offer the best known quality for a cost level in this sample. A point near the frontier is a candidate for the use case; it is not a universal winner.
- Missing cost or score
- Missing data stays missing. A model without
CONTENT_SCORE.jsonis marked as awaiting quality scoring. It keeps readability and cost values where available, but the report does not invent an editorial score. - Radar rank and raw table value
- Radar shapes use normalized ranks so measures with different units can be displayed together. Tables keep the actual score, seconds, dollars, word counts, and grade estimates. Use tables for a purchasing decision.
The editorial sources are the comparison JSON, readability JSON, and the saved case-study folders. Provider usage records are the source of paid bundle cost.
Judge and evidence limits
The local editorial round used GPT-5.6 Luna to find evidence, score the files, and break ties. Older rows used DeepSeek, GPT-5.4-mini, or both. More than the final judge changed between these historical groups.
- No saved content has a complete score from both DeepSeek and Luna.
- We cannot claim that a score difference is Luna’s effect alone.
- A fair judge comparison saves one output and scores it twice with the same rubric.
See the DeepSeek and Luna note for the evidence and paired-test plan.
What the results do and do not prove
These results are evidence for the listed prompts, bundle, tasks, model versions, settings, hardware, and provider runs. They do not prove general intelligence, future price, all-purpose reliability, or a model’s value on a different machine. Use them to make a specific model choice, then run a small validation on your own high-risk work before scaling it.