On 6 September, @scheemunai did something most people posting about AI never do. They spent 24 hours running the newest model, GPT-6 Astra, on the jobs a store owner actually needs done: branding, copy, positioning, conversion rate analysis. Their verdict: "it's not much better than Fable 5 or even Opus." The feed, they noted, "is full of 'Told Astra to do this and that' but it's all Blender 3D stuff and threejs games nobody plays."
The post drew over 2,200 engagements and 249 replies. That is a lot of people who had quietly noticed the same thing.
If you run a store doing $20k to $500k a month, you have probably felt the gap. Every morning the timeline says one person with an agent can now run a business. Every afternoon you are still writing product descriptions by hand because the AI version needed more fixing than starting over. This piece is about that gap: what the viral posts are really showing, what separates the tools that work from the demos that do not, and how to find out which side a tool is on before you change how you work.
The feed is a demo reel, not a results report
Look closely at the biggest AI posts in the Shopify conversation this month and notice what they have in common.
@TheMattBerman's post about Jev analysing 724 ads in 40 seconds reached close to 8,000 engagements. The tool is not publicly available yet; the post says it "will be avail in @stealads + mcp." @gregisenberg's post about one non-technical person running a business on Grok Bot agents reached nearly four million views. The example is a newsletter, not a store with inventory, returns and a payment processor.
@PrajwalTomar_ is the most honest of the viral posts, because it shows both halves. Asked to design a landing page, Astra's first attempt was "typical AI slop. Tilted phone, sage green, poetic headlines." It only got good after being fed 600,000 real websites and told to critique its own work. The impressive result and the disappointing one came from the same model on the same day. The difference was what it was given.
That is the pattern to hold onto. The demo shows the output. It rarely shows the input, and the input is where the work is.
The tools that work are sitting on data you do not have
The clearest explanation of why some AI works and some does not came from Shopify's CEO, and it was not about a marketing tool at all.
@tobi posted that Shopify's ML team had fine-tuned a 0.8-billion-parameter model, tiny by current standards, and it beat a frontier model on one narrow task. His reason: "a great self improving recursive flywheel." Shopify had data on that task nobody else had, and the small model trained on it outperformed the big model that had not seen it.
Training tiny models for special purpose use cases works so incredibly well if you have a great self improving recursive flywheel. Shopify ML team is on fire. finetuned 0.8b model beats GPT 5.6-sol xhigh in this very specialized task. https://t.co/w6OCWyWRi5
— tobi lutke (@tobi) · View on X
Apply that to the tools in your feed. Jev works because it has read hundreds of thousands of live ads. Higgsfield works because it has seen which creatives won. The general model that @scheemunai tested for 24 hours does not know your customer, your margin or your last twelve months of ad results. It gives you an average answer because an average is all it has.
@mattepstein relayed the same point from a different direction, quoting Sam Altman on how Tobi used early ChatGPT: "He gave the model full access to company context, trained it on exact workflows, and assigned it multi-layered projects. Most people on this app give ChatGPT four words, zero background, and complain when they get 'ai slop'."
@insomnia_vip described what "full context" looks like for someone in ecommerce: a former Meta employee who fed Claude a 2.34 GB knowledge base, 42 ad strategies, 18 campaign breakdowns, 27 creative insights and 31 testing documents, before letting it produce anything. Her formula is 20% human, 60% AI, 20% human. The first 20% is loading the context. The last 20% is fact-checking every page and cutting the filler.
A FORMER META EMPLOYEE JUST REVEALED WHY MOST AI PRODUCTS SELL FOR $0 Her formula is simple: 20% human → 60% AI → 20% human. After 7 years working with Meta ads and scaling e-commerce brands, she fed Claude a 2.34 GB knowledge base containing 42 ad strategies, 18 campaign breakdowns, 27 creative insights and 31 testing documents before letting AI build the first version The middle 60% is where Claude does the heavy lifting organizing the research, finding gaps, building the chapters, decision trees and frameworks, then turning all of that knowledge into a complete digital product instead of
— Insomnia (@insomnia_vip) · View on X
@shannholmberg gave the lightweight version any store can build: a folder of plain files. What the business does, who the customer is in their own words, the offer, the positioning, the voice, the proof. Built from sales calls and customer emails you already have. Without that folder, every tool is guessing. With it, even the model that disappointed @scheemunai has something to work from.
What following blindly costs
The risk is not that AI produces nothing. It is that it produces something plausible, and plausible is expensive to catch.
Shopify's CEO put a name on it this week. As reported by @interesting_aIl, he warned staff about passing unreviewed AI-generated emails and code to colleagues, calling the lazy output "slop grenades" that create more work for everyone. This is a company with thousands of engineers and a review process. At a two-person store, the slop grenade lands on a live product page and the first person to review it is a customer.
Shopify CEO warns employees are passing unreviewed AI-generated emails and code to colleagues without checking them He is calling the lazy output "slop grenades" that create more work for everyone https://t.co/HojVPlM8Q7
— Interesting AF (@interesting_aIl) · View on X
The market is already pricing this in. @jackorgbuild is hiring a video editor for brands spending over $1M a month on Meta, and the job spec asks for someone who can "build AI video that doesn't look like AI slop." The skill being paid for is not generating. It is judging. @vikingmute pointed to Shopify's own engineering write-up on rebuilding its Shop app natively in 12 weeks with AI assistance, and noted that a good chunk of it is about how to avoid slop and shorten feedback loops. Even the company shipping fastest with AI spends its words on quality control.
And there is one experiment worth watching precisely because it is not a demo. @rtwlz described a research lab that gave AI agents real small businesses, each with a bank account, a debit card, an inbox and a domain, and let them try to sell real things to real customers. The post does not report results. That silence is more informative than most of the feed. As of this week there is no public data showing an unsupervised agent running a store profitably.
The winner that never changed
Against all of this, the best-performing account in the ecommerce conversation this month was not using AI to make ads at all.
@carlmonkft shared their entire TikTok ads setup: one campaign per region, around ten proven creatives each, and an ads library they had left almost untouched for weeks, apart from a single new ad, because making more felt like producing AI slop. The winners were found over time and simply left running.
And the most-shared line of the month, from @madsf88 at close to 16,000 engagements: "a girl recommending something to her friend group has a higher conversion rate than your entire paid ads budget."
Neither post is anti-AI. Both are a reminder of what the tools are for. @kdseifu, who credits a $450k August to a stack of Claude, Higgsfield and Rapid Ads, is not running a different business from @carlmonkft. Both found winners. One uses AI to find them faster. The product still has to be worth a friend's recommendation, and no tool in this piece changes that.
A test you can run this week, on your own store
The question was whether to change how you work because of what the feed says. The answer is: not because of the feed. Because of a test you run yourself. Here is one that takes an afternoon and costs nothing you are not already paying.
- Pick one job, not a workflow. Rewriting the description of your best-selling product. Producing five hook variations of your best ad. Drafting the reply to your most common support email. One job with an output you can compare to what you have now.
- Build the context folder first. Follow @shannholmberg's list: what you sell, who buys it in their own words (pull five real reviews and three real support emails), what you promise, what you never say. This takes an hour and it is the step every disappointing test skipped.
- Run the job with and without the folder. Same model, same prompt, one run with your context loaded and one with four words. If the two outputs look the same, the tool is not using your data and it will never beat what you already have. If they differ sharply, you have found where the value is.
- Judge by one number, after the last 20%. Do @insomnia_vip's final pass yourself: fact-check, cut filler, fix voice. Then compare on one metric you already track. Click-through on the ad. Add-to-cart rate on the page. Reply rate on the email. Not "does it look good." Whether it moved the number.
- Only then change the workflow. If the number moved, make that one job the AI's job, permanently, with your review at the end. If it did not, you have lost an afternoon and gained the right to ignore the next viral post about that tool.
Nobody in the feed this month published a test like this on a real store. Until someone does, your own afternoon is the best data you have.
Sources
- @scheemunai: 24 hours with Astra on real marketing work
- @TheMattBerman: Jev breaks down 724 live ads
- @gregisenberg: Running a business on Grok Bot agents
- @PrajwalTomar_: Astra landing page, before and after 600k examples
- @tobi: Fine-tuned 0.8B model beats a frontier model
- @mattepstein: Four words, zero background, and "ai slop"
- @insomnia_vip: The 20/60/20 formula and a 2.34 GB knowledge base
- @shannholmberg: A knowledge base for marketing agents
- @interesting_aIl: Shopify CEO on "slop grenades"
- @jackorgbuild: Hiring for AI video that "doesn't look like AI slop"
- @vikingmute: Shopify's native mobile rebuild and avoiding slop
- @rtwlz: Agents running real small businesses
- @carlmonkft: Ads library untouched since July
- @madsf88: A friend's recommendation beats your ads budget
- @kdseifu: The subscriptions behind a $450k month