From AI Demo to Trustworthy Shopify Agent: What I’m Learning Building Jam
“The first principle is that you must not fool yourself — and you are the easiest person to fool.” — Richard Feynman
Building an AI demo is surprisingly easy.
Connect an LLM to a few tools, give it a goal, and watch it produce something impressive.
Building an AI product that merchants can safely trust with their store is an entirely different problem.
I’ve been building Jam, an AI-powered operating assistant for Shopify merchants. Jam connects to a merchant’s store, understands business activity, identifies opportunities, and helps execute useful work across business analysis, merchandising, and customer recovery.
In this article from my Building Jam in Public Series, I’ll share:
- What Jam is designed to do
- Why a convincing AI demo isn’t enough
- How I’m approaching approvals, calculations, retries, and safety
- The engineering mistakes and lessons I’ve encountered
- What “trustworthy” means for an AI agent connected to a real business
🛍 What is Jam?
Shopify merchants have access to enormous amounts of useful data: orders, products, inventory, refunds, customers, and abandoned checkouts.
The problem is that turning this data into decisions and actions takes time.
Jam is being built to help merchants answer questions such as:
- What changed in my business today?
- Which opportunities need my attention?
- Which products are frequently purchased together?
- Are there abandoned checkouts worth recovering?
- What action should I take next?
The long-term goal is not to give merchants another analytics dashboard they must constantly study.
The goal is for Jam to proactively notice important events, explain them clearly, recommend an action, and safely carry out approved work.
Today, Jam can already:
- Synchronize Shopify products, orders, inventory, refunds, and locations
- Generate business briefings and identify opportunities
- Surface proactive work through an AI Inbox
- Analyze product relationships and prepare bundle recommendations
- Require approval before publishing a merchandising action
- Prepare abandoned-checkout recovery messages
- Enforce consent, unsubscribe, and suppression rules before marketing execution
- Record what happened and verify external actions
However, getting these features to work once was the easy part.
Getting them to behave correctly when permissions are missing, jobs are retried, Shopify is slow, or users change their minds has been the real engineering work.
🎭 The AI Demo Trap
An AI demo usually demonstrates the happiest possible path.
The model receives clean data, makes the correct decision, calls a tool successfully, and returns a polished answer.
Real software operates in a messier world:
- APIs time out
- Permissions change
- Workers restart
- Users click buttons twice
- Webhooks arrive late or more than once
- A request succeeds externally but the local response is lost
- Store data may be incomplete
- An AI-generated statement may sound correct while being unsupported
One of my biggest lessons has been this:
The model is only one component of an agentic product. Reliability comes from the systems surrounding it.
A trustworthy agent needs more than a good prompt. It needs deterministic calculations, typed tools, policy enforcement, idempotency, durable state, observability, approval records, and external verification.
Without those systems, an agent is mostly a confident interface sitting on top of unreliable automation.
🧮 LLMs Should Not Calculate Business Truth
Jam uses models to interpret, prioritize, draft, and explain.
It does not use an LLM as the source of truth for revenue, margins, inventory, order counts, or attribution.
Those calculations are handled using deterministic SQL and tested Python logic.
This distinction is important.
If Jam says that two products were frequently purchased together, that conclusion must come from actual synchronized order data. The model may explain why the opportunity matters, but it should not invent the underlying evidence.
The general architecture looks like this:
- Shopify provides the source commerce data
- Deterministic services calculate the business facts
- The model interprets those facts and prepares a recommendation
- Server-side policy decides whether an action is permitted
- The merchant approves sensitive actions
- A typed integration executes the action
- Jam verifies the result against the external system
This structure makes the AI useful without asking it to become the database, calculator, security layer, and decision-maker simultaneously.
✋ Approval Must Be a Real Product Event
Suppose Jam recommends creating a product bundle.
The recommendation includes:
- The exact products
- The proposed title
- The proposed price
- The intended Shopify publication
- The supporting evidence
- The expected outcome
- The associated risk
The merchant can review and approve that exact version.
If the title, price, or products change, the previous approval becomes invalid. The merchant must approve the new version.
Jam also does not treat a chat message containing “approved” as authorization.
Approval must be recorded as an explicit product event from an authorized user. The server checks it again immediately before executing the external action.
This might feel stricter than a typical AI demo, but a vague conversation should not be enough to publish products, contact customers, or modify a merchant’s store.
Good agent UX should make powerful actions feel simple without making authorization ambiguous.
🔁 Retries Are More Complicated Than They Look
External actions rarely have only two states: success or failure.
Imagine Jam requests Shopify to create a bundle, but the connection drops before Shopify’s response reaches Jam.
Did the operation fail?
Maybe.
But Shopify may have successfully created the product.
Blindly retrying could create a duplicate. Marking it as failed could also be incorrect.
Jam therefore needs to distinguish between:
- Confirmed success
- Confirmed permanent failure
- Transient failure
- Unknown outcome requiring reconciliation
Every side-effecting operation needs an idempotency key. When an outcome is uncertain, Jam should inspect Shopify before deciding whether to retry.
I learned this lesson through real bugs involving duplicate event sequences and reconciliation attempts. The fixes were not prompt changes. They involved database constraints, transaction boundaries, durable execution state, and better external-state verification.
This has changed how I think about agents:
An agent isn’t reliable because it retries. It is reliable because it knows when retrying is safe.
🔐 Permissions Are Part of the Product Experience
Shopify apps require specific permissions to access products, orders, inventory, publications, and protected customer data.
Initially, it was tempting to expose all this technical state directly:
- Missing scopes
- Partial synchronization
- Provider error codes
- Resource-specific failures
- Historical data coverage
That information is useful for debugging, but it can create a terrible experience for ordinary merchants.
A merchant should not need to understand OAuth scopes to use an AI assistant.
Jam still needs to enforce least-privilege permissions internally, but the product should translate those requirements into clear outcomes:
- Reconnect Shopify
- Jam is checking access
- Recovery is ready
- Some historical analysis is still becoming available
One of my continuing goals is to hide infrastructure complexity without hiding important limitations.
Simple UX does not mean pretending everything succeeded. It means explaining the next useful action in language the merchant understands.
📧 Marketing Automation Needs Stronger Guardrails
Abandoned-checkout recovery sounds like a straightforward agent workflow:
- Find an abandoned checkout
- Draft an email
- Send it
In reality, the difficult part is everything around the draft.
Before sending, Jam must check:
- Whether the customer provided valid marketing consent
- Whether the recipient has unsubscribed
- Whether the address appears on a suppression list
- Whether the merchant has configured a valid sender
- Whether the frequency limit permits another message
- Whether the approved message version is still current
- Whether the checkout has already converted
- Whether the operation has already been submitted to the provider
The LLM can write recovery copy, but it cannot override these controls.
This is one of the clearest examples of where agentic automation must remain bounded. A creative model is useful; an unconstrained marketing sender is dangerous.
🧪 Testing Months of Store Activity Without Waiting Months
Another practical challenge was testing Jam’s proactive behavior.
Some recommendations require meaningful order history. Waiting for a development store to naturally accumulate months of activity is not realistic.
I therefore built a development-only store activity simulator that can:
- Generate deterministic historical orders
- Create refunds
- Produce realistic co-purchase patterns
- Prepare safe abandoned-checkout scenarios
- Stream new orders at visible intervals
The simulator is tenant-scoped, explicitly allowlisted, and rejected outside development. It cannot call Shopify, send emails, approve recommendations, or publish anything.
This allows me to observe how Jam responds to business activity without weakening the real production boundaries.
It does not replace actual Shopify testing. Real storefront visits, checkouts, webhooks, publishing, and email delivery must still be tested through Shopify and the relevant providers.
The simulator validates Jam’s internal behavior; controlled integration testing validates the external boundaries.
🧠 What “Trustworthy” Means for Jam
I don’t think trustworthy AI means the model never makes a mistake.
It means the system anticipates that models, APIs, networks, and humans can all make mistakes.
For Jam, trustworthiness means:
- Evidence: Recommendations are grounded in real store data
- Determinism: Business facts are calculated outside the model
- Transparency: Important limitations are shown clearly
- Approval: Sensitive actions require explicit authorization
- Isolation: Every record and action belongs to the correct store
- Idempotency: Repeated requests do not create duplicate effects
- Verification: Jam checks whether the external action actually happened
- Auditability: Decisions, approvals, attempts, and outcomes are recorded
- Consent: Marketing rules are enforced independently of model output
- Safe failure: Uncertainty pauses execution instead of pretending success
These systems are less exciting to show in a short demo, but they are what turn a prototype into software that can eventually be trusted with a real business.
———
🚀 Building Jam in Public
Jam is still being built.
There are more workflows, integrations, evaluations, deployment work, and UX improvements ahead. I’m deliberately avoiding the claim that it is a fully autonomous business operator.
The direction is clear, though.
Merchants should not have to configure multiple “agents,” learn prompt engineering, or become data analysts. They should connect their store, understand what matters, review important actions, and let Jam handle the repetitive work safely.
Throughout this series, I’ll document both the features and the engineering underneath them:
- Agent architecture and orchestration
- Shopify integration lessons
- Human approval design
- Idempotent external actions
- Model routing and cost control
- Multi-tenant data safety
- Evaluation and regression testing
- Product mistakes and failed assumptions
The most valuable lessons have not come from the moments when the AI looked impressive.
They have come from the moments when something failed and the system had to explain, recover, and avoid making the situation worse.
🔗 Resources to Explore
-
———
Next article: Why Jam Never Lets an AI Agent “Just Do Things”