Proofreading 682 product pages with Claude cost $12.60 on the second try, down from about $185 on a first attempt that never finished, according to @jamiegrove, who runs the Shopify store and posted the breakdown on X in September 2026.

The store is not named in the post, and all figures are the merchant's own.

Before

The store's product pages are built from JSON theme templates: stacked sections of story, specs, history and background. Each new page starts as a copy of a similar product's template, so one missed section leaves a page describing the wrong product. The first attempt at checking them:

  • Setup: one Claude Code session per page, looped 682 times
  • Cost: about $185, and the run did not finish
  • Billed tokens per page: about 22,000, of which the page itself was about 3,500

The difference was overhead: each session reloaded the agent's instructions, tool definitions, skills and memory, uncached, for every page.

What they changed

  1. Did the cheap work in code. They pulled every template with the Shopify CLI, mapped each template to its product, and used a small script to render the page a shopper reads, skipping hidden sections. No AI was used in that step.
  2. Made one model call per page. Each page went to Claude Opus 5.5 through the API with the product it should describe and one task: find what is wrong.
  3. Checked two layers against each other. The review compares the product description with every template section, which catches sections about another product, contradictions in dates, weights and counts, factual slips, placeholder text and doubled sections.
  4. Sorted every finding. An "error" is a defect a shopper would see; a "heads-up" needs a human look. Each page gets a one-line verdict and a direct link to where the fix goes: the admin description or the theme editor.
  5. Batched the calls and used structured outputs. Batch processing halved the per-page cost. Opus 5.5 rejects forced tool calls, so they switched to structured outputs with a JSON schema.

After

AI proofreading of 682 product pages: first try vs second. Cost per page: First try $0.33, Second try $0.018. Total cost: First try about $185, unfinished, Second try $12.60
First trySecond try
SetupClaude Code session per page (Opus 5)Direct API call with Batch API (Opus 5.5)
Cost per page$0.33$0.018
Totalabout $185, unfinished$12.60

A direct API call without batching cost $0.036 a page. The review found 1,264 errors across 558 pages and 1,236 heads-ups, and only 24 pages came back clean. A random spot check of 30 errors found all 30 were real. From now on, only templates changed since the last pull get reviewed.

How can you test this on your store?

  1. Pull your theme templates with the Shopify CLI and list which template each product uses.
  2. Turn each template into the text a shopper actually sees, in code, skipping hidden sections.
  3. Send one model call per page with the page text and the product name, and ask for errors and heads-ups in a fixed JSON format.
  4. Send the calls as a batch if the results can wait an hour, and log tokens per page so you know what a full run costs.
  5. Spot-check 30 findings at random before fixing at scale, then rerun only on templates that change.

Frequently asked questions

Why did the first try cost so much?

Each Claude Code session reloaded the full agent setup for every page, uncached. That made each call about 22,000 tokens for about 3,500 tokens of actual page content, 682 times.

Why use a language model instead of a spell checker?

A spell checker only checks words, so it passes a well-written section about the wrong product. The model compares the whole page with the product it should describe, and it also skips things that are intentional, such as hidden sections or related background.

How accurate were the findings?

@jamiegrove says a random check of 30 errors found all 30 were real. The post does not report how many real problems the review missed.

Sources