A coding assistant that works well with an expert may still leave a beginner unable to tell whether the job is finished. New research makes that gap measurable, and puts a practical question ahead of another model ranking: how much technical work must the user supply before an AI-built feature can be trusted?
SWE-Journey, a preprint submitted October 8 by Hexuan Deng and colleagues, evaluates extended coding work with simulated users. Its authors report average functionality-test pass rates above 75% with a software-architect persona and below 25% with a non-coder persona. Those are benchmark outcomes under simulated interactions, not success rates measured across ordinary customers.
TENS Magazine’s analysis is that accessible coding tools need a connected record of intent and evidence. A user should be able to see what the assistant understood, which behavior it tested and what remains unresolved. Comparing SWE-Journey with the separate SWE-Game research shows why improving conversation and improving verification belong in the same product decision.
Who supplies the missing specification?
SWE-Journey constructs extended tasks and uses four user personas informed by interaction data. The authors acknowledge that simulation cannot capture the full complexity of people and that generated tasks can inherit biases. The findings therefore identify a weakness worth investigating; they do not establish a universal penalty for lacking programming experience.
The editorial implication is a change in what counts as assistance. If a person must diagnose the code, identify the relevant module and explain the repair before an assistant succeeds, part of the apparent capability belongs to that person. A product intended for beginners should account for that contribution when presenting its performance.
Consider a hypothetical request to make an event sign-up form work. An expert might specify duplicate-registration rules, confirmation states and storage behavior. A beginner might describe only the screen. Treating those as equivalent starting points would obscure the work needed to turn an intention into a testable requirement.
The useful response is not an endless questionnaire. The assistant could show a short example of the proposed behavior: what happens after a valid submission, what happens after a repeated submission and which information is retained. The user can then correct a concrete interpretation without first learning the application’s internal architecture.
A convincing demonstration needs an independent check
SWE-Game, a separate preprint by Xiaoyu Chen and colleagues, provides a complementary perspective. Its 247 tasks draw on 41 reference games. Evaluation uses running-game checks and input replay, with visual assessment considered separately. In a comparison against human-labeled behaviors from 100 agent-built games, executable checks achieved 92.59% balanced accuracy, versus 78.41% for a video-based AI judge.
That result concerns the reliability of evaluators in this particular benchmark. It does not mean that 92.59% of generated games were correct. Nor can it be combined numerically with SWE-Journey’s persona results: one measures judgments of behavior, while the other reports completion of requested functionality.
Read together, the studies suggest two distinct places to look for failure. The assistant may build a faithful version of the wrong requirement, or it may understand the requirement and deliver software that fails it. A polished demonstration cannot resolve both questions unless the intended behavior was agreed and the demonstration actually exercises it.
For the hypothetical sign-up form, a success message is only one observation. A useful check would also inspect whether a registration was stored and whether a repeated submission followed the agreed rule. The evidence should be available to the user in ordinary language, with a clear distinction between observed results and assumptions.
Measure how much expertise the workflow still needs
TENS proposes a practical comparison for teams evaluating these tools: keep the requested outcome fixed, vary the technical detail supplied by the user and record the human interventions needed to reach acceptance. Separate corrections to the requirement from corrections to the implementation. That would reveal whether a tool reduces technical dependence or merely moves it into the conversation.
This is an evaluation proposal, not an experiment TENS has conducted. It also needs safeguards against rewarding verbosity. More messages, more tests or longer explanations do not automatically mean a better result. A concise clarification that prevents a wrong feature can be more useful than a large test suite built around an unconfirmed assumption.
A second useful measure is whether the user can challenge a completion claim. An assistant should expose the check that supports the claim and identify behavior it did not examine. If the user reports a conflicting result, the next step should investigate that discrepancy against the agreed requirement rather than simply repeat the earlier assurance.
Both studies remain research benchmarks, with particular tasks, models and evaluators. They do not settle how any individual product will perform in a workplace. Their combined value is a more precise question for future coding systems: can the software help a less technical person define success and inspect the evidence for it, while keeping the responsibility for unresolved work visible?
Illustrative photograph: source code on a computer monitor. Photo: Markus Spiske via Wikimedia Commons. License: CC0 1.0 Universal Public Domain Dedication. Modifications: center-cropped from 5,760 × 3,840 pixels to 16:9 and resized to 2,400 × 1,350 pixels; no generative edits.


