Skip to content
AI

Step 5 Preview Puts AI Cost per Accepted Result in Focus

StepFun’s updated model report and independent measurements show why AI pricing needs to be assessed against accepted work, retries and review time.

Share Email
Colorful software code displayed on a computer monitor
Source code displayed on a computer monitor. Photo: Markus Spiske via Wikimedia Commons. CC0 1.0 Universal Public Domain Dedication. Center-cropped from 5,760 × 3,840 pixels to 16:9 and resized to 2,400 × 1,350 pixels; no generative or content-altering edits.

StepFun’s Step 5 Preview arrives with a proposition that matters well beyond another model launch: useful AI work at a lower price. Its benchmark report, updated with results dated September 20, gives software teams a fresh reason to examine that proposition. The consequential question is how much a completed, accepted job costs after failed attempts and review are counted.

StepFun describes the model as built for agentic work and offers a million-token context window with image input. Artificial Analysis independently lists its release as September 18 and gives it an Intelligence Index score of 44. Those facts establish a new candidate for evaluation. They do not establish the economics of replacing an existing workflow.

The price is only one part of the bill

Artificial Analysis lists prices of $1 per million input tokens and $2.70 per million output tokens. Its evaluation also records 160 million output tokens, against a comparison median of 92 million. That combination is significant: a low unit price can coexist with substantial token consumption. The evaluator’s measurements describe its test workload, not every customer’s application.

TENS Magazine’s analysis is that buyers should define the unit of value before comparing models. For a coding assistant, that might be a change accepted after tests and human review. For a document workflow, it might be a report whose claims survive checking. A response that looks finished but requires extensive repair should remain part of the cost calculation.

A transparent illustration shows why. Assume a hypothetical request consumes 100,000 uncached input tokens and 10,000 billable output tokens at the listed rates. The input charge is ten cents and the output charge is 2.7 cents, making 12.7 cents for one attempt. If three equally sized attempts are needed before an acceptable result, the token charge becomes 38.1 cents. This is arithmetic, not a measured Step 5 success rate.

The example excludes tool charges, computing infrastructure and staff time. It also assumes the same token volume on each attempt; a real retry may include a longer history. Its purpose is to expose the denominator: comparing prices without counting accepted outputs can reward a system for producing inexpensive work that nobody can use.

Benchmarks should guide the trial

StepFun’s report publishes results across several coding and agent evaluations, including company-developed tests. It also explicitly warns that some results use different subsets and are not directly comparable. That disclosure matters because a table can look uniform while its rows represent different conditions. The report’s September 20 update is a timely reminder to preserve benchmark versions alongside scores.

A practical comparison would therefore start with the team’s own recurring work and a fixed acceptance rule. Run the incumbent and the candidate against the same task set, with the same available tools and a declared spending limit. Record each attempt, including abandoned runs. Keep the reviewer unaware of the model identity where practical, so presentation style does not quietly become the scoring rule.

The useful output would be a small set of paired measurements: accepted jobs, total model charges, elapsed time and review minutes. A candidate could win on one and lose on another. That is actionable information. Compressing all four into a single claim that a model is cheaper would conceal the operational choice the team actually faces.

A larger workspace needs a selection rule

The million-token context window raises a related question about what teams send to the model. A large capacity makes room for extensive material, but the capacity figure alone says nothing about which documents are necessary for a particular task. Including everything available should be a testable decision, not an automatic default.

Consider an illustrative software maintenance job. One trial could provide the relevant module, its tests and the reported failure; another could add a much broader repository history. If both produce an equally acceptable fix, the extra context needs another justification. If broader context prevents a regression, its value belongs in the outcome record. This comparison tests the workflow rather than treating a maximum input size as a benefit already realized.

These are proposed evaluation methods, not findings from a TENS production trial. TENS has not measured Step 5 Preview’s reliability on a private customer workload. The available evidence supports taking the model seriously as a candidate while leaving deployment outcomes open.

The broader significance of this release is the pressure it puts on how AI value is described. Lower token prices can widen experimentation, but the next useful comparison is work that passes a defined standard for a known total cost. Step 5 Preview’s arrival gives teams another opportunity to make that comparison explicit.

Image: Illustrative software code on a monitor, not a StepFun interface. Photo: Markus Spiske via Wikimedia Commons. CC0 1.0 Universal Public Domain Dedication. Center-cropped to 16:9 and resized from 5,760 × 3,840 to 2,400 × 1,350 pixels; no generative or content-altering edits.