Shopify's CEO gave the industry a phrase this month, and it is going to stick. In a conversation on The Knowledge Project podcast, shared by @shaneparrish (https://x.com/shaneparrish/status/2099823553726070973), Tobi Lütke described staff who send unreviewed AI-generated emails and code to colleagues. He called that output "slop grenades." A summary of the remark from @interesting_aIl (https://x.com/interesting_aIl/status/2102007936809841117) pulled 713 likes and more than 76,000 views, and the line is now on the agenda for his Vercel Ship interview, per @rauchg (https://x.com/rauchg/status/2102436383923339629).
Shopify CEO warns employees are passing unreviewed AI-generated emails and code to colleagues without checking them He is calling the lazy output "slop grenades" that create more work for everyone https://t.co/HojVPlM8Q7
— Interesting AF (@interesting_aIl) · View on X
I build software for Shopify merchants, and the warning landed differently for me than it did for most people reposting it. Tobi was talking about employees. But the pattern he describes, generate fast and push the checking onto someone else, is exactly how a lot of AI features inside Shopify apps work today. The "someone else" is the merchant.
The grenade is not the AI output, it is the transferred checking
The sharpest reading of Tobi's point came from @tmiyatake1 (https://x.com/tmiyatake1/status/2101824478212669521), who noted that the definition of lazy work has flipped. It used to mean producing too little. In the AI era it means producing too much and making other people verify it.
Translate that to an app. A feature that rewrites twenty product descriptions in one click and publishes them has produced a lot of output. Whether it has produced value depends entirely on what happens next. If the merchant has to open each page, compare old against new, and hunt for the sentence where the model invented a material or a size, the app has handed them a grenade with a friendly button on it.
The reverse is also true. Migration work described by @jay_neyer (https://x.com/jay_neyer/status/2102178562375610443) moved more than a thousand products in under eight hours, then had AI run QA against the original catalog and flag mismatches for a human. The line that stuck with me: the repetitive hours disappeared, and the checking stayed. That is the design target. Remove the repetition, keep the checking, and make the checking cheap.
Sidekick just set the bar for what "reviewable" looks like
Shopify shipped a redesigned admin on September 24 with Sidekick placed front and center. The announcement from @katarinabatina (https://x.com/katarinabatina/status/2103161312901816394) drew 1,451 likes and over 238,000 views. That placement matters to every app builder, because merchants will judge our AI features against the one Shopify put next to their navigation.
What does Sidekick do when it touches store content? @sh_sakamoto (https://x.com/sh_sakamoto/status/2101446357173289307) walked through a product page SEO review: Sidekick proposes changes to the title, meta description, image alt text and URL handle, marks each change, and waits for the merchant to approve before anything is applied.
Shopify Sidekickにお願いすると、商品説明文だけでなく、SEOタイトル・メタディスクリプション・画像alt・URLハンドルなど、SEOに関わる情報をまとめて確認し、最適化案を提示してくれます。 変更をかけた部分がマーキングされ、問題なければ承認してそのまま反映できます。商品ページのSEOを一通り見直したいときに、かなり手軽に使えます。 サンプルプロンプトはリプライに載せてあります。 #Shopify
— 坂本@StoreHero (@sh_sakamoto) · View on X
Three things are happening in that flow. The merchant can see the change. The merchant decides. Nothing goes live until they do. That is the whole anti-grenade pattern in one screen, and it is now the default experience inside the admin.
App developer @junaidkbr (https://x.com/junaidkbr/status/2102778882617463095) asked the question that follows: with Sidekick in that spot, will merchants expect it to work with every app? I think the expectation goes further. Merchants will expect every AI feature, from any app, to behave the way Sidekick behaves. Show me, let me approve, let me undo.
Considering the Sidekick placement, are merchants going to expect that Sidekick works with every app? I hope @Shopify is already committed and working on improving the poor UI + UX this new dashboard introduced. https://t.co/VLK5zl92uA
— Junaid Ahmed (@junaidkbr) · View on X
Even Shopify's own assistant misses two times out of three
The case for review is not that AI is bad. It is that AI is wrong often enough that the design must assume it. @TRPage_dev (https://x.com/TRPage_dev/status/2099492186857275745) asked Sidekick to recommend an app for a gift card migration. Two of the three suggestions did not meet the ask. The third was good, and even pointed to a useful Matrixify angle.
That ratio is fine when the output is a suggestion the merchant reads and discards. It is a disaster when the output is a change the merchant did not see. The same model quality produces opposite outcomes depending on whether a human sits between generation and publication.
A related lesson came from @jamiegrove (https://x.com/jamiegrove/status/2102810220648976526), who had AI proofread 682 product pages. The first attempt burned about $185 of usage and never finished. The second run, same prompt and same model, cost $12.60. The difference was the harness around the model, not the model. For a builder, the harness is the product. It is where you decide how much the model does before a person looks.
What I look for before an AI feature ships
I am not going to quote internal numbers here, because we have not published any, and I would rather not dress up anecdotes as data. What I can share is the set of questions I now put to any AI feature, ours or anyone else's, after watching this week's conversation.
Can the merchant see the diff? Not the new version alone. The old and the new, side by side or with changes marked, the way @sh_sakamoto (https://x.com/sh_sakamoto/status/2101446357173289307) showed Sidekick doing it. If a feature only shows the result, the merchant has to reconstruct what changed from memory. That is transferred checking.
Is there exactly one approval step, and is it real? A confirmation modal that says "Apply 40 changes?" is not review. Review means the merchant can accept some changes and reject others. The agent-run store described by @chesny (https://x.com/chesny/status/2102778825591689255) worked because the human kept two jobs, approving spend and deciding what looked good, and nothing went live until the agent's own checks passed and the human signed off.
Can it be undone in one action? If the answer is "restore from a theme backup," the feature is not safe to hand to a merchant who runs the store between customer emails.
Does it say when it is unsure? The solo-operator workflow shared by @ridark_eth (https://x.com/ridark_eth/status/2098891678480310680) had a rule that stood out: when the prompt cannot find enough evidence, it must return UNKNOWN rather than guess. A product-page generator that flags "I could not find the material for this SKU" is worth more than one that quietly writes "premium cotton."
How much does it generate per request? This is the one builders resist, because volume demos well. But the more a feature produces in one go, the more the merchant has to read before trusting it. Smaller batches with a fast approve loop beat one enormous output with a single publish button.
When the platform does it better, delete your version
There is a second lesson in @junaidkbr's week. Alfred retired its admin sidebar toggle because Shopify's new admin shipped a built-in one that was better, as they wrote at https://x.com/junaidkbr/status/2103215196387119176 (286 likes). They sounded happy about it.
That is the right instinct for AI features too. If Sidekick already handles product-page SEO suggestions with a mark-and-approve flow, an app that offers a worse version of the same thing is adding review burden, not value. The place for app builders is the work Sidekick does not do, or does not do with enough context: page layouts, section-level copy tied to a campaign, conversion elements that need the merchant's brand judgment. Build the reviewable version of that, and let Shopify own the generic layer.
A checklist for merchants and agencies choosing an AI app
You do not need to read the code to tell whether an app was designed to avoid grenades. Run this in the trial:
- Ask it to change something small, like one product title. Does it show you the before and after, or only the after?
- Ask it to change ten things at once. Can you accept seven and reject three, or is it all or nothing?
- Accept a change, then try to undo it. Count the clicks.
- Give it a product with missing information and see whether it admits the gap or fills it with a plausible guess.
- Check who did the checking. If the app's marketing says "publish in one click," ask what you are expected to verify afterward, and how long that takes for your catalog size.
If an app fails three of these five, it will cost you more time in cleanup than it saves in generation, whatever the demo looked like.
Sources
- @shaneparrish: https://x.com/shaneparrish/status/2099823553726070973
- @interesting_aIl: https://x.com/interesting_aIl/status/2102007936809841117
- @rauchg: https://x.com/rauchg/status/2102436383923339629
- @tmiyatake1: https://x.com/tmiyatake1/status/2101824478212669521
- @jay_neyer: https://x.com/jay_neyer/status/2102178562375610443
- @katarinabatina: https://x.com/katarinabatina/status/2103161312901816394
- @sh_sakamoto: https://x.com/sh_sakamoto/status/2101446357173289307
- @junaidkbr: https://x.com/junaidkbr/status/2102778882617463095
- @junaidkbr: https://x.com/junaidkbr/status/2103215196387119176
- @TRPage_dev: https://x.com/TRPage_dev/status/2099492186857275745
- @jamiegrove: https://x.com/jamiegrove/status/2102810220648976526
- @chesny: https://x.com/chesny/status/2102778825591689255
- @ridark_eth: https://x.com/ridark_eth/status/2098891678480310680