The news for teams shipping agent-written code is that Shopify plans for rejection. In the words of its Helix post, "An attempt is allowed to be wrong." Code simply cannot ship until the gates pass. The post gives no retry rates or checkpoint counts, so the cost of those gates is not public.

  • The backstory: on September 10, 2026, Shopify said it was moving its mobile apps back to native code. @fnthawar (https://x.com/fnthawar/status/2098048371172823215) wrote that LLMs changed one of the core assumptions behind the company's 2020 React Native decision, and that the Shop app had already been migrated.
  • How the work is checked: the Helix post (https://shopify.engineering/helix) says each migration is cut into checkpoints, small ordered slices that start with screen skeletons. Each one passes behaviour tests against the reference app, a Gemini screenshot review and two independent adversarial AI reviewers. Engineer approval, the fourth gate, is mandatory by default; in autonomous mode Helix can skip it, but the three automated gates still apply.
  • The scale claimed: according to the same post, Shopify rebuilt and published the Shop app in 12 weeks, and the Shopify app, with more than 300 screens, is next.
  • Why developers are sharing it: @norvex1029 (https://x.com/norvex1029/status/2104474504609055178) argues that the winning teams may be the ones with the best verification system, whatever model they use. @benawad (https://x.com/benawad/status/2098531589802348774) thinks this may become how most teams build apps within a couple of model releases.
  • What an agency can try this week: this is our reading, not Shopify's advice. As a starting point, split one client job into small checkpoints and attach a check to each before anyone reviews it. @kloss_xyz (https://x.com/kloss_xyz/status/2104088147739242660) shares a lighter step: have the agent label each assumption verified or guessed, then stop for manual review.

What does a Helix checkpoint have to pass?

Four gates, in order: command-line behaviour tests against the existing app, a visual review where Gemini lists spacing, alignment and sizing differences, two independent AI agents that check the code against documented architecture, and, by default, an engineer. Per the Helix post (https://shopify.engineering/helix), the engineer's feedback goes back to the agent, which re-runs the gates, and into memory for later checkpoints. When approvals are skipped in autonomous mode, every checkpoint still has to clear the first three gates before the next one starts.

Can a small team copy this without Shopify's resources?

Only the principles transfer directly: small slices, a check the agent cannot mark as passed on its own, and a person at the end. OpenAI's account of building with Codex (https://openai.com/index/harness-engineering/) points the same way, reporting about a million lines and roughly 1,500 merged pull requests from a team that started with three engineers, backed by review and automated checks.

How is this different from asking the agent to check its own work?

In Helix, the reviewers are separate from the agent that wrote the code, and by default a human signs off last. A single agent reviewing its own output lacks that separation; the @kloss_xyz prompt (https://x.com/kloss_xyz/status/2104088147739242660) narrows the gap by making the agent say which assumptions it actually verified in the code.

Sources