Research methodology

This page explains the test inputs, every reported measure, and how to use the results when cost and usable capacity must be weighed together.

Evidence pipeline

Every new suite result is stored as one immutable bundle per tested model and replicate. A bundle includes its manifest, normalized task outcomes, ordered trajectory, raw response data, configuration, lineage, and artifact hashes. The registry is derived from validated bundles; charts are never the source of truth.

Inspect the task-level record in the run explorer or download the machine-readable run registry.

What this research is for

The goal is a specific, repeatable comparison. It does not ask which model is best at everything. It asks what a model can do for the same real work, and what that work costs. A frontier model, a third-party API model, and a self-hosted model can then be compared on the same job.

Objective measurements matter because model names, parameter counts, and price cards do not tell us whether a model completes the work. The same seed, prompts, fixtures, files, rubric, and calculation rules give an apples-to-apples starting point. Human judgment still exists in the rubric, so the page also names the evaluator limits.

Use the results to choose the lowest-cost model that meets the needed standard for a named use case. Do not use a high score as proof that a model is safe, current, or right for every job.

Two separate test tracks

This project has two tracks. They answer different questions, so their scores must not be added or treated as one ranking.

A model may be strong at one track and weak at the other. That is useful information, not a conflict.

Assistant test: same work, fixed conditions

The fixed suite is labeled 2026-08-12. The current data combines a local round and a later OpenRouter extension. Every model received the same 21 tasks: project work, research, structured tool calls, state tracking, English communication, and safe handling of requests.

Assistant score dimensions

The weighted assistant score is a 0–100 summary. The weights show which kinds of failure matter most for this use case. A high score is more valuable than a low price only when the model can do the parts of the job you need.

Outcome — 30%
Whether the task reaches the requested result. It matters most because an inexpensive answer has little value if the work is not finished.
Tool use — 25%
Whether the model calls the right tool with usable inputs and uses the result. It matters when the assistant must act on calendars, files, data, or other systems.
Grounding — 15%
Whether the answer stays tied to the supplied facts and tool results. It matters because unsupported claims can create costly follow-up work.
State — 10%
Whether the model keeps track of task facts across turns. It matters for multi-step work where lost details cause rework.
English — 10%
Whether the answer is clear, direct, and usable in English. It matters because a correct plan must still be understood and acted on.
Safety — 5%
Whether the model handles risky or disallowed requests safely. It matters because a low-cost model can have a high operational cost if it creates unsafe output.
Efficiency — 5%
Whether the model completes work without wasteful extra turns. It matters because extra turns add time, cost, and failure chances.

Assistant data fields

Tasks passed and task pass rate
The number and share of the 21 tasks that passed. This is the clearest view of how often the stated workflow worked.
Tool-call success rate
The share of detected tool calls that succeeded. It helps separate good writing from reliable tool execution.
Run status
ok means the run completed cleanly. partial means a tool-loop limit or API error affected it. There are 5 ok records and 27 partial records. Partial rows retain their scheduled task slots so failure risk stays visible; do not read them as clean runs.
Cold load, first response, warm request, and total time
Cold load is startup time. First response is time to the first token after a cold start. The warm OpenClaw-style request measures a later request after the model is ready. Total time combines the recorded parts. These are speed and capacity-planning measures, not quality points. Self-hosted inference is expected to be slower in this environment.

Sources: JSON, CSV, and notes. Fixture hash: ea2601bcb637a9c66563e91015f15a007a0a06f7d88532ec218c8c178903efb9. Task hash: 907017aa0d29639967cd0c0702764b73e7cc8ebbf254625fcc07b63e13b4428e.

Editorial test: one weekly content bundle

Each model receives the same client content-bundle job, pipeline structure, seed, and optional-track settings. A bundle is broader than one article. It shows whether a model can make connected pieces that are useful together.

Context
The brief with the audience, topic, and constraints. It sets the same starting facts for each model.
Lesson
The teaching piece. It tests whether the model can explain the central idea.
Long-form letter (index.md)
The main reader-facing letter. It is scored for argument, grounding, action, reading control, and closing voice.
Instructions (INSTRUCTIONS.md)
The practical guide that turns the idea into steps. It is scored for argument, framework, reading control, and voice.
Worksheet copy, worksheet, and masked worksheet
Working prompts and their usable forms. They show whether the idea can become an action tool, including a version that hides answers or guidance when needed.
Newsletter email and community post
Short distribution pieces that bring the weekly idea to readers in different channels.
Cover prompt and rendered hero image
The visual brief and its output. They test whether the model can direct a useful, on-brand visual asset.
Pipeline configuration
The recorded run settings. It makes the case study traceable and helps explain differences.
OpenRouter usage record
For paid-provider runs, the provider-reported token usage and charge record. This is the source for recorded bundle cost.

The content-quality scorer directly evaluates only index.md and INSTRUCTIONS.md. Readability is calculated for LESSON.md, index.md, INSTRUCTIONS.md, and the combined text. Other bundle pieces are retained as evidence and for practical review; they are not silently turned into rubric points.

Editorial quality rubric

The editorial rubric is versioned and fixed for this round. It keeps rule-based measurements and model-based scores in separate fields. A required criterion can lower the result when it is missed. Category weights show the parts of quality that have more influence on each artifact score.

Letter categories

Argument — 40%
The letter has a clear idea that develops. It matters because the main piece must give the reader a reason to change or decide.
Grounding — 15%
The letter uses real, concrete life details. It matters because abstract advice is hard to trust or use.
Action — 15%
The letter leads to a real choice or next step. It matters because content capacity includes helping a reader do something.
Readability — 15%
The letter stays clear without becoming empty. It matters because a strong idea fails if the target reader cannot follow it.
Closure — 15%
The ending and tone stay firm without sounding like a guru, preacher, therapist, or hype message. It matters because voice affects trust and brand fit.

Instruction categories

Argument — 40%
The guide explains why the work matters. It stops steps from becoming empty commands.
Framework — 30%
The guide makes the named method usable. It matters because the reader needs a repeatable process, not only good prose.
Readability — 15%
The guide is easy to scan and follow. It matters because instructions must work during action.
Voice — 15%
The guide stays direct without hype, therapy language, preaching, or guru language. It matters because trust is part of usefulness.

Every scored criterion

The items below define the hard values in the quality-standards table and radar. A shared item is checked in both scored files unless the text says otherwise.

Perspective shift
The piece helps the reader see the problem in a new way. This matters when paying more for insight instead of familiar advice.
Argument progression
The reasoning moves in a clear order. This matters because a reader can follow and test the claim.
Concrete opening — letter only
The letter begins with a real situation or detail. This matters because it makes the subject immediate and believable.
Psychological interpretation
The piece explains the thought, habit, or pressure behind the problem. This matters because useful change needs more than surface tips.
Behavioral grounding
The claims connect to actions people can observe. This matters because it makes advice testable in daily life.
Practical philosophy
The piece turns a bigger idea into a useful rule. This matters because capacity includes sound judgment, not only instructions.
Intellectual depth
The piece goes beyond a simple slogan. This matters because premium model cost should buy more than polished filler.
Qualification and nuance — letter only
The letter states limits or exceptions where needed. This matters because overconfident writing can mislead readers.
Personal responsibility
The reader has a clear part to play. This matters because the content should support agency, not passive consumption.
Perspective progression — letter only
The letter carries the reader from the old view to the new view. This matters because a strong ending depends on a complete change path.
Decision — letter only
The letter names the choice the reader should make. This matters because vague inspiration is hard to act on.
Prescribed action — letter only
The letter gives a specific next action. This matters because a content bundle must lead to use, not only reflection.
Concrete and abstract balance — letter only
The letter balances examples with ideas. This matters because either extreme can make a piece shallow or unclear.
Everyday realism — letter only
The advice fits normal time, limits, and pressures. This matters because impractical advice creates no real capacity.
Intellectual simplicity control
The language is simple without making the idea childish. This matters because readers need clear, credible writing.
Sentence rhythm
Sentence length and shape vary in a useful way. This matters because readable pacing keeps a long piece usable.
Semantic economy
Each sentence does useful work with few wasted words. This matters because concise output lowers edit time and delivery cost.
Anti-motivational tone
The piece avoids empty cheering or pressure. This matters because hype can damage trust.
Anti-therapy tone
The piece avoids acting like treatment or diagnosis. This matters because content should stay in its intended role.
Anti-preacher tone
The piece avoids moralizing at the reader. This matters because readers need respect, not a lecture.
Anti-guru tone
The piece avoids false certainty or special-authority claims. This matters because reliable content shows limits.
Closing strength — letter only
The ending lands the idea with force and clarity. This matters because the final message shapes recall and action.
Model or framework lands — instructions only
The named method is explained well enough to use. This matters because the guide must deliver the promised framework.
Unique letter-specific action — instructions only
The guide gives an action that fits this letter, not a generic productivity step. This matters because specificity shows real understanding.
Argument to framework to action — instructions only
The guide connects the reason, the method, and the next step. This matters because a reader needs a complete path from insight to use.

Readability and text measurements

Markdown presentation syntax is removed first. One deterministic English syllable rule is then applied to every model output. These numbers describe the text; they do not decide whether the advice is true or good. They show likely reader effort and editing effort, which both affect the real cost of a model choice.

Characters
The count of letters and numbers in the normalized text. It shows total text size and helps explain processing or editing volume.
Letters
Alphabetic characters. It is an input for some reading formulas and helps compare writing density.
Words
Detected English word forms. It shows length and is a basic editing-work measure.
Sentences
Detected sentences. It helps show how the model breaks ideas into units a reader can follow.
Paragraphs
Detected blocks of text. It helps show page structure and scanability.
Syllables
Estimated spoken parts of words. More syllables often mean harder words and more reading effort.
Polysyllables
Words with three or more syllables. It is a direct signal used by some grade-level formulas.
Long words
Words with seven or more letters. It is a simple sign of dense language that may need editing.
Long sentences
Sentences with 20 or more words. It shows where readers may need to hold too much in mind.
Average words per sentence
Words divided by sentences. Higher values often raise reading effort, so it helps compare pacing.
Average syllables per word
Syllables divided by words. Higher values often mean more complex vocabulary.
Average characters per word
Letters divided by words. It is another simple view of word length and density.
Flesch reading ease
A 0–100 style formula based on sentence and word length; higher is easier. It gives a quick direction for general ease.
Flesch-Kincaid grade
An estimated U.S. school grade needed to read the text; lower is easier. Grade 6 is a reference target, not a pass rule.
Gunning fog
An estimated school grade based on sentence length and complex words; lower is easier. It checks whether long words are making prose heavy.
Coleman-Liau
An estimated school grade based on letters and sentences; lower is easier. It gives a second view that does not use syllable estimates.
Automated readability index
An estimated school grade based on letters and words; lower is easier. It helps catch dense writing with long words.
SMOG
An estimated school grade based mainly on polysyllables; lower is easier. It highlights complex vocabulary.
Lexical diversity
The share of distinct words. It shows word variety, but very high variety is not always better if it makes text needlessly hard.

The readability radar converts each raw measure to a within-measure rank so unlike units can share one shape. It is for quick comparison only. Use the raw tables and reading report for decisions.

Cost, capacity, and chart rules

Cost is the recorded price of one completed editorial bundle when provider usage records exist. It is not a price-per-token estimate and it is not a full business cost. Capacity means the measured ability to complete this specific work at the needed quality and reliability.

Frontier and paid third-party models
These models may offer high measured quality or speed, but their recorded provider charge can be higher. The value question is whether better output saves enough editing, review, or failure cost to justify that charge.
Self-hosted local models
These models run on the research machine. The editorial chart marks them as $0 direct provider cost because no paid API bill is recorded. That is not a claim that local inference is free: hardware, power, setup, maintenance, and slower local time are outside this field.
Cost burden
The radar reverses recorded cost into a rank where lower cost gives a better value. It is a visual aid, not a quality score. A cheap model with weak quality is not automatically the better choice.
Quality-and-cost frontier
The decision chart highlights observed models that offer the best known quality for a cost level in this sample. A point near the frontier is a candidate for the use case; it is not a universal winner.
Missing cost or score
Missing data stays missing. A model without CONTENT_SCORE.json is marked as awaiting quality scoring. It keeps readability and cost values where available, but the report does not invent an editorial score.
Radar rank and raw table value
Radar shapes use normalized ranks so measures with different units can be displayed together. Tables keep the actual score, seconds, dollars, word counts, and grade estimates. Use tables for a purchasing decision.

The editorial sources are the comparison JSON, readability JSON, and the saved case-study folders. Provider usage records are the source of paid bundle cost.

Judge and evidence limits

The local editorial round used GPT-5.6 Luna to find evidence, score the files, and break ties. Older rows used DeepSeek, GPT-5.4-mini, or both. More than the final judge changed between these historical groups.

See the DeepSeek and Luna note for the evidence and paired-test plan.

What the results do and do not prove

These results are evidence for the listed prompts, bundle, tasks, model versions, settings, hardware, and provider runs. They do not prove general intelligence, future price, all-purpose reliability, or a model’s value on a different machine. Use them to make a specific model choice, then run a small validation on your own high-risk work before scaling it.