A model comparison starts with a specific task: what did we ask for, and did the output meet the requirements? We test AI models on GlobalGPT using the GlobalGPT CLI. Exceptions are identified in the relevant article.
AI checks outputs for errors and prompt adherence. When outputs pass those checks, we compare speed, price, and token use where applicable. Authors also assess subjective qualities, such as writing style and visual appeal. Audio evaluation is performed entirely by AI using Gemini 3.6 Flash.
We are the GlobalGPT Review Team, part of GlobalGPT. This methodology covers the tests we publish in the Blog when new models launch, along with model comparisons. Research reports based on GlobalGPT user data explain their own datasets and analysis methods.
Our Test Environment: GlobalGPT CLI
The GlobalGPT CLI is our default test environment. A test that uses a different interface, API route, or native coding tool needs to be read in that context. Model names, available settings, tools, and access can vary between environments.
The model that generates an output and the AI that evaluates it have different roles. An evaluator’s judgment is evidence about the returned output; it does not establish how the tested model works internally.
Task Design: Prompts, Inputs, and Requirements
We give models concrete tasks with identifiable requirements. A writing task might require a summary supported only by supplied notes. A coding task might require a function to handle an empty input. An image task might require exact wording on a poster.
For a comparison, the prompt and input materials establish what each model is being asked to do. Differences in supported formats, settings, or available tools matter when interpreting the results. A fixed prompt comparison and a tutorial that improves a prompt over several attempts answer different questions.
Output Evaluation: Errors, Instructions, and Efficiency
Errors and Prompt Adherence
AI evaluates whether the output follows the task and contains identifiable errors. Relevant checks might cover incorrect calculations, unsupported claims, missing fields, or a failure to follow a required format.
Passing means no error was identified by the applied checks. It does not guarantee that the output is free of every possible error. AI evaluators can overlook mistakes or make incorrect judgments.
Execution evidence is a separate part of the assessment. In our Kimi coding plan test, generated code was followed by five local checks. Reading code, asking a model to inspect its own answer, and running tests provide different evidence.
Speed, Price, and Token Use
Once an output meets the task requirements, efficiency helps distinguish useful options. A quick answer that misses a required calculation has not completed the same job as a correct answer.
| 度量 | How to Read It |
|---|---|
| 時間 | Check the measurement window. Request-to-completion time includes the route and processing involved; it is not necessarily provider-side model latency. |
| Token use | Use the categories returned by the service. Missing input, output, cached, or reasoning-token data does not mean zero usage. |
| 成本 | Distinguish recorded GlobalGPT charges or credits from estimates based on provider API prices. Subscription fees are another measure. |
| 媒體世代 | Consider the output settings and applicable credits or per-output charges. Token counts may not be available or meaningful. |
A price estimate needs a stated calculation basis. Credits cannot be converted into dollars without a verified conversion rule, and a zero-valued response field alone does not establish that a request was free.
Author Judgment for Subjective Tasks
Authors also compare outputs when preference affects usefulness. For writing or images, that can mean explaining which result has clearer language, a more readable layout, or a composition that better suits the brief.
我們的 Seedream 5.0 Pro 與 GPT Image 2 比較 shows paired images across specific tasks. Visible wording and layout give readers something to inspect alongside the author’s preference. An aesthetic preference remains distinct from an objectively checked requirement.
AI Audio Evaluation With Gemini 3.6 Flash
Our audio listening evaluations are performed entirely by AI. We use Gemini 3.6 Flash to interpret the actual audio. These are AI assessments, not human listening-panel results.
The evaluation depends on the task: requested speech, language, sound events, timing, or musical structure. Our Seed Audio 1.0 review includes generated audio and AI analysis of content, events, timing, and artifacts.
File properties such as format and duration are separate from content judgments. An audio track’s presence does not prove that the requested words or sounds are correct. Any published score must be read with the evaluation criteria and the audio it describes.
How to Read Test Dates, Settings, and Output Evidence
The individual article is where the detailed test evidence belongs. Useful context includes:
- The test date, model identifier, and environment.
- The task, prompt, input materials, and required outcome.
- Available settings, number of attempts, and relevant changes.
- The evaluator, applied checks, and reasons for any author preference.
- Actual outputs or excerpts, screenshots or media, and available timing, usage, and cost records.
Records differ across articles and older tests. Missing measurements cannot be reconstructed as observed facts. The Blog & Guides section of a model collection page links to the relevant reviews and this methodology; full prompts and output evidence remain in the Blog articles.
Interpreting Retries, Failed Requests, and Revisions
A first output, a selected result, and a revised attempt tell readers different things. The article’s attempt history matters, including technical failures, empty responses, truncated answers, and changes to prompts or settings.
A completed request may still return an unusable answer. A failed request may reveal a problem with the tested route without establishing a general model weakness. Our articles use different test scopes; there is no universal number of runs or blind-testing procedure behind every result.
What Task Results Can and Cannot Tell You
A result supports a conclusion about the tested task, model, date, and environment. Small task sets and individual timing measurements cannot establish a universal performance ranking. If only one side of a comparison was tested, that evidence cannot support a tested winner between both models.
Official specifications, external benchmarks, previous tests, and an article’s own outputs provide different kinds of evidence. Any score depends on its stated scale, weighting, and exclusions. Readers should be able to connect the verdict to the evidence, including partial or unsuccessful results.
How Article Updates Differ From Model Retesting
We generally review articles every two to four weeks, as described in our Editorial Standards. A content check does not mean every model test has been rerun.
Updating a price, correcting an evaluation of an existing output, and generating a new output are separate actions. Historical findings remain tied to their original test context; a newer article update date alone does not make them new test results.



