Two runs, same model, same prompt, opposite results. As relayed by @JafarNajafov (https://x.com/JafarNajafov/status/2103764227404312730), Anthropic asked Opus 4.5 to build a 2D retro game editor twice. Without a harness the run cost about $9, took 20 minutes and did not work. With a full harness it cost about $200, took six hours and produced a game you could play. The account is secondhand, so treat the exact figures as reported, but the shape of the result is the point of this piece.

If you run a store and have written careful prompts only to get a generic product page back, that experiment describes your situation. The difference between your result and the polished one someone posted is usually the system around the model. This article explains what that system is in plain terms, what a store-sized version looks like, and the three places where even a good one still lets errors through.

The prompt is about a third of the system

The clearest framing this month came from a recap of Andrej Karpathy's Stanford lecture on AI engineering, shared by @dkare1009 (https://x.com/dkare1009/status/2103764008801280248). The progression runs from the raw model at roughly 10%, to a prompt at 30%, an agent at 50%, a loop at 70%, and a full graph of connected steps at 100%. The recap's summary: AI engineering is about building systems around models, giving them context, memory, tools, feedback loops and data flows.

An Anthropic engineer put it more bluntly in a talk summarised by @0xnicc0 (https://x.com/0xnicc0/status/2095609874197413983): "You're not supposed to prompt Claude. You're supposed to build a system that prompts itself." @ghumare64 (https://x.com/ghumare64/status/2104112175425950177) opened a long list of skills for AI engineers with the same line: harness engineering, not just prompt engineering, and context engineering, not just long prompts.

For a merchant, "harness" translates to four ordinary things. What the model is told about your business before it starts. The rules it must follow every time. The check that runs before anything goes live. And the record of what was decided so the next run does not repeat a mistake. A polished result posted on X almost always has those four, and the post rarely shows them.

What Shopify does before it trusts an agent's answer

Shopify's own engineering story shows the harness at full scale. As described by @norvex1029 (https://x.com/norvex1029/status/2104474504609055178), when Shopify pointed a model at its React Native app and asked for a native rewrite in one shot, it got a large pile of unshippable code. The team then built a system it calls Helix. Each screen is cut into small checkpoints. Each checkpoint has to prove its behaviour with tests, match the running app in a visual review, and pass two adversarial AI code reviewers. Only then does an engineer look at it.

That is how Shopify rebuilt and published its Shop app natively in 12 weeks, according to the same post, with the 300-screen Shopify app as the next target. The lesson @norvex1029 draws is the one that matters for a store owner: the teams that win may have the best verification system rather than the best model.

A smaller example makes the cost visible. @jamiegrove (https://x.com/jamiegrove/status/2102810220648976526) had AI proofread 682 product pages on a Shopify store. The first run burned about $185 of usage and never finished. The second run, same prompt and same model, cost $12.60. Their conclusion: the model was not the big lever, the harness was.

The store-sized version of the same idea

You do not need Helix to get most of the benefit. Each of the four parts has a version a two-person store can build without code.

Context. @shannholmberg (https://x.com/shannholmberg/status/2098850243760808268) describes a folder of plain files: what the business does, who the customer is in their own words, the offer, the positioning, the voice, the proof. Pull it from reviews and support emails you already have. Without this, every tool is guessing at your customer.

Rules that accumulate. A CLAUDE.md template attributed to Claude Code's creator and shared by @Divyyanshishrma (https://x.com/Divyyanshishrma/status/2104169439289651427) keeps past mistakes, conventions and rules in one file the agent reads every session. The habit to copy is simple: every error you catch becomes a permanent line in that file.

A check before publish. @Voxyz_ai (https://x.com/Voxyz_ai/status/2103977414711767244) shared a 20-point list to run before shipping a vibe-coded site, covering consistent design tokens, mobile layout, and proper loading, empty and error states. For a product page, the equivalent list is shorter: prices match the catalog, claims match your facts file, required policy text is present, and the page renders on a phone.

A staging copy and a measurement. @dashboardlim (https://x.com/dashboardlim/status/2101390400292290676) rewrote a Shopify product page with AI using exactly this loop: research competing pages, generate drafts, review them on a hidden store copy, then split live traffic. They report 23.75% higher revenue per visitor for the rewrite, a figure from their own test.

Build those four and you have a harness. The rest of this piece is about why that is still not enough.

Where a good harness still lets errors through

The same posts that make the case for harnesses also document where they fail. Three patterns repeat, and each has a store equivalent.

The agent reports done when it is not. @mr_kozh (https://x.com/mr_kozh/status/2103895408795689397) summarised how OpenAI produced about a million lines of code with three engineers and 1,500 pull requests, and noted that even inside that system the agents still forgot decisions, chose wrong tools, duplicated code and marked broken work as complete. The fix was not a longer prompt: a spec that defines done, isolated environments, behaviour tests, checkpoints and reviewers who inspect evidence. For a store, "done" must mean a page that passed your checklist, not a page the tool says it finished.

The context fills up and the model starts inventing state. @dr_cintas (https://x.com/dr_cintas/status/2104274873476337815) reports that letting Opus 5.5 run freely across a 200,000-token context breaks down on complex features: the model forgets earlier tool outputs and produces hallucinated file states or circular edits. Their fix is an external SESSION_STATE.md file the agent must update before touching code, with the objective, active changes, decided architecture and verification status written down outside the conversation.

The store version of this failure is a long session where the tool rewrites forty product descriptions and, by the thirtieth, has lost track of which facts belong to which SKU. The fix is the same: short sessions, one product at a time, and a written state the agent reads back before each step.

Memory keeps the wrong decision too. @beamnxw (https://x.com/beamnxw/status/2103890773972308479) described 64 agent sessions sharing one append-only log of decisions and reasons. Over three days there were 178 lookups; 68% confirmed a prior decision and about 4.5% changed the agent's action. The case that matters: an agent decided a database index looked unused and removed it in an isolated commit. A human retested under realistic conditions, found the index useful and reversed it. The log had faithfully preserved a mistake until a person checked.

That is the risk nobody mentions when they praise memory files. Your rules file will hold whatever you put in it, including the wrong lesson from a bad week. @DamiDefi (https://x.com/DamiDefi/status/2104496839064178804) summarised Daniel Miessler's answer: turn every agent mistake into a permanent system upgrade, and "garden" the codebase so agents do not amplify existing debt. Gardening means someone rereads the rules file and prunes it.

Why this lands on the customer first

At Shopify, Helix puts two AI reviewers and an engineer between the agent and production. At a two-person store, there is often nobody between the AI and the live page. Shopify's CEO named the failure mode for his own staff this month, as reported by @interesting_aIl (https://x.com/interesting_aIl/status/2102007936809841117): unreviewed AI output passed to colleagues, which he called slop grenades, creates more work for everyone who receives it. On a storefront, the person who receives it is a shopper, and the cost is a lost sale or a return rather than a colleague's afternoon.

No public data yet shows how often AI-written product pages ship with factual errors, so the size of the problem for stores is unknown. The mechanism, however, is documented above by the teams with the most resources to prevent it.

What to check on your own workflow this week

  1. Find the four parts. Context folder, rules file, pre-publish checklist, staging copy. Whichever is missing is where your generic output comes from.
  2. Define done in writing. One sentence per job: "A product description is done when every number matches the catalog and the checklist passes." Give that sentence to the tool and to the person reviewing.
  3. Cap session length. One product, one page or one email per session. If a tool insists on batch mode, review the last item in the batch as carefully as the first; that is where drift shows.
  4. Reread the rules file monthly. Delete lines that came from a one-off incident. A rule that says "never mention shipping times" because one page once had a wrong date will quietly strip useful information from every page after.
  5. Keep one human step before live. Approve changes individually, not as a batch of forty. The tool that lets you accept seven and reject three is the one built with a harness in mind.

Start with the rules file. Open it, or create it, and add the last three mistakes you caught. That single habit is the part of Shopify's system a store of any size can run tonight.

Sources