A small team rewrote a Shopify product page with AI and recorded 23.75% higher revenue per visitor in a live split test. That is the headline from @dashboardlim (https://x.com/dashboardlim/status/2101390400292290676), and it is worth more attention than its 519 views suggest, because the post describes the whole workflow rather than just the result. The client had hundreds of descriptions to update across two stores, so the team built a process the client's own staff could keep running: research competing pages, generate drafts, review them on a hidden store copy, split live traffic, and send completion reports to Slack with errors visible.

If you run a store with 50 to 500 products, no developer, and a bundle page that carries a lot of your revenue, that is the loop you want. Not an agent that publishes on its own. An agent that proposes, and a person who approves. Here is how to build it in a week, step by step, using what teams have actually shown working.

Day 1: decide what the agent is allowed to touch

Before any AI writes a word, write down the boundary. The store that @chesny (https://x.com/chesny/status/2102778825591689255) described, where an agent ran a dropshipping business for eight weeks, worked because the human kept exactly two jobs: approving spend and deciding what looked good. Everything else was delegated, and "nothing goes live until the agent has checked" and the human signed off.

For a bundle page, the boundary looks like this:

  • The agent may propose copy, section order, tier labels, FAQ content, image alt text and comparison tables.
  • The agent may not change prices, discount math, subscription defaults, guarantee terms or anything the shopper sees at checkout.
  • Every proposal goes to a staging copy, never to the live page.

That last rule is the one @dashboardlim built their whole process around: "review drafts on a hidden store copy before testing." It is also how Shopify's own Sidekick behaves. @sh_sakamoto (https://x.com/sh_sakamoto/status/2101446357173289307) showed Sidekick proposing SEO changes, marking each one, and waiting for approval before applying anything.

Day 2: feed it the structure of a bundle page that already converts

An agent asked to "improve this page" will produce generic output. An agent given a proven structure and told to fill it produces something you can evaluate. The clearest public template for a skincare bundle page this month came from @johntech778 (https://x.com/johntech778/status/2103228002385441194), breaking down a night repair serum page:

  1. A tier ladder where the "Most Popular" two-bottle option drops the per-bottle price, with a visible Save badge.
  2. An authority claim on the hero image and a volume claim next to it.
  3. Proof expressed as numbers, in this case a clinical result percentage, not adjectives.
  4. Risk reversal at the top: money-back guarantee and free shipping before the shopper scrolls.

A second post from the same designer (https://x.com/johntech778/status/2101787604719190198) is the more useful one for briefing an agent, because the product, price and traffic stayed identical and only the order of information changed: proof stacked above the title, four benefit tiles beside the image, a scannable benefit grid, and the discount framed as a reason rather than a number.

Give the agent that skeleton and your own product facts. Then add the rule @ridark_eth (https://x.com/ridark_eth/status/2098891678480310680) highlighted from a solo operator's prompt set: when the agent cannot find enough evidence for a claim, it must return UNKNOWN rather than guess. A bundle page that says "clinically proven" because the model assumed it is not a variant. It is a liability.

Day 3: generate to staging, and budget for the second run

@dashboardlim's production step is simple to describe: "generate descriptions, comparison tables, faqs and image alt text." The trap is cost and reliability at volume, and @jamiegrove (https://x.com/jamiegrove/status/2102810220648976526) documented it precisely. They had AI proofread 682 product pages built from JSON templates. The first attempt burned about $185 of usage and never finished. The second run, same prompt and same model, cost $12.60. Their conclusion: "The model wasn't the big lever. The harness was."

For a first week, that means three things. Start with one bundle page, not the catalog. Have the agent output one section at a time, so a bad generation costs you a paragraph and not a page. And log every run, so when something goes wrong you can see which step produced it. @dashboardlim's version of that log was completion reports sent to Slack "with errors visible." A shared channel or even a spreadsheet works. The point is that failures are seen, not silently retried.

Day 4: run a QA pass before a human ever looks

The review step gets skipped when it is tedious, so make the agent do the tedious part first. @jay_neyer (https://x.com/jay_neyer/status/2102178562375610443) described moving more than 1,000 products between platforms in under eight hours, then having AI run QA against the original catalog and flag mismatches for a human to check. "The hours that disappeared were the repetitive ones, and the checking stayed."

For a bundle page variant, the QA pass compares the draft against your source of truth and flags:

  • Any number that does not match the catalog: price, size, count, percentage.
  • Any claim not present in your approved facts list.
  • Any change to the tier that is pre-selected, the subscription default, or the guarantee wording.
  • Missing required text. @sepe2x (https://x.com/sepe2x/status/2102727984922665037) learned the hard way that Shopify can penalise a store for removing recurring-charge disclosure in its first 30 days. An agent optimising for a cleaner buy box will happily delete that text unless a rule stops it.

Only drafts that pass this check reach the approval queue. The human then reviews a short list of flagged differences, not an entire page from scratch.

Day 5: approve on staging, then split test with real traffic

The approval itself should feel like Sidekick's flow: see the diff, accept or reject each change, publish only what you accepted. Then measure. @dashboardlim split live traffic between the original page and the rewrite and reported revenue per visitor, which is the right metric for a bundle page because it captures both conversion and order value.

Two more examples show what a clean test looks like. @navuud (https://x.com/navuud/status/2102478479073907010) tested a presell page against sending traffic straight to the product page for two weeks and reported a 23% lift in landing page conversion rate. @Frontend_Prince (https://x.com/Frontend_Prince/status/2102036061366939763) claimed a 628% increase in average order value from a wellness page redesign built around a visual bundle selector and subscription architecture. Treat both numbers as claims from the people who did the work, not as benchmarks. The lesson is the method: one variant, one metric, one clear window, then a decision.

What nobody has shown yet: the agent finding the problem on its own

You may have seen pitches for agents that watch your analytics, spot a drop-off, build a fix and queue it for approval. It is a good idea. As of this writing there is no public case, with numbers, of that full loop running on a Shopify bundle page. @dashboardlim's team chose which pages to rewrite. @chesny's agent left a scorecard each morning, but the human decided what mattered. @jay_neyer's AI found mismatches, not opportunities.

So do not build week one around detection. Build it around the part that is proven: you pick the page, the agent proposes, QA filters, you approve, the test decides. If a tool later adds reliable detection on top, the approval gate you built this week is exactly where its output should land.

The one-week plan on a single page

DayYou doThe agent doesGate
1Write the may / may not list and pick one bundle pageNothing yetBoundary agreed
2Supply product facts and the buy box structureDrafts sections one at a time, returns UNKNOWN where facts are missingFacts list
3Set up a hidden copy of the page and a log channelGenerates to staging, logs every runStaging only
4Read flagged differences onlyCompares draft to catalog, flags numbers, claims, defaults, required textQA pass
5Accept or reject each change, publish approved version as variant BNothing goes live on its ownApproval
6 to 7Start the split test, choose revenue per visitor as the metricReports results to the same logTest window

At the end of the week you will have one tested variant, a log of every generation, and a repeatable loop. That is more than most stores running "AI optimisation" can show. Start with the bundle page that carries the most revenue, and keep the publish button on your side of the table.

Sources