Newfoundland

From AI Demo to Trustworthy Shopify Agent: What I’m Learning Building Jam

“The first principle is that you must not fool yourself — and you are the easiest person to fool.” — Richard Feynman

Building an AI demo is surprisingly easy.

Connect an LLM to a few tools, give it a goal, and watch it produce something impressive.

Building an AI product that merchants can safely trust with their store is an entirely different problem.

I’ve been building Jam, an AI-powered operating assistant for Shopify merchants. Jam connects to a merchant’s store, understands business activity, identifies opportunities, and helps execute useful work across business analysis, merchandising, and customer recovery.

In this article from my Building Jam in Public Series, I’ll share:

🛍 What is Jam?

Shopify merchants have access to enormous amounts of useful data: orders, products, inventory, refunds, customers, and abandoned checkouts.

The problem is that turning this data into decisions and actions takes time.

Jam is being built to help merchants answer questions such as:

The long-term goal is not to give merchants another analytics dashboard they must constantly study.

The goal is for Jam to proactively notice important events, explain them clearly, recommend an action, and safely carry out approved work.

Today, Jam can already:

However, getting these features to work once was the easy part.

Getting them to behave correctly when permissions are missing, jobs are retried, Shopify is slow, or users change their minds has been the real engineering work.

🎭 The AI Demo Trap

An AI demo usually demonstrates the happiest possible path.

The model receives clean data, makes the correct decision, calls a tool successfully, and returns a polished answer.

Real software operates in a messier world:

One of my biggest lessons has been this:

The model is only one component of an agentic product. Reliability comes from the systems surrounding it.

A trustworthy agent needs more than a good prompt. It needs deterministic calculations, typed tools, policy enforcement, idempotency, durable state, observability, approval records, and external verification.

Without those systems, an agent is mostly a confident interface sitting on top of unreliable automation.

🧮 LLMs Should Not Calculate Business Truth

Jam uses models to interpret, prioritize, draft, and explain.

It does not use an LLM as the source of truth for revenue, margins, inventory, order counts, or attribution.

Those calculations are handled using deterministic SQL and tested Python logic.

This distinction is important.

If Jam says that two products were frequently purchased together, that conclusion must come from actual synchronized order data. The model may explain why the opportunity matters, but it should not invent the underlying evidence.

The general architecture looks like this:

  1. Shopify provides the source commerce data
  2. Deterministic services calculate the business facts
  3. The model interprets those facts and prepares a recommendation
  4. Server-side policy decides whether an action is permitted
  5. The merchant approves sensitive actions
  6. A typed integration executes the action
  7. Jam verifies the result against the external system

This structure makes the AI useful without asking it to become the database, calculator, security layer, and decision-maker simultaneously.

✋ Approval Must Be a Real Product Event

Suppose Jam recommends creating a product bundle.

The recommendation includes:

The merchant can review and approve that exact version.

If the title, price, or products change, the previous approval becomes invalid. The merchant must approve the new version.

Jam also does not treat a chat message containing “approved” as authorization.

Approval must be recorded as an explicit product event from an authorized user. The server checks it again immediately before executing the external action.

This might feel stricter than a typical AI demo, but a vague conversation should not be enough to publish products, contact customers, or modify a merchant’s store.

Good agent UX should make powerful actions feel simple without making authorization ambiguous.

🔁 Retries Are More Complicated Than They Look

External actions rarely have only two states: success or failure.

Imagine Jam requests Shopify to create a bundle, but the connection drops before Shopify’s response reaches Jam.

Did the operation fail?

Maybe.

But Shopify may have successfully created the product.

Blindly retrying could create a duplicate. Marking it as failed could also be incorrect.

Jam therefore needs to distinguish between:

Every side-effecting operation needs an idempotency key. When an outcome is uncertain, Jam should inspect Shopify before deciding whether to retry.

I learned this lesson through real bugs involving duplicate event sequences and reconciliation attempts. The fixes were not prompt changes. They involved database constraints, transaction boundaries, durable execution state, and better external-state verification.

This has changed how I think about agents:

An agent isn’t reliable because it retries. It is reliable because it knows when retrying is safe.

🔐 Permissions Are Part of the Product Experience

Shopify apps require specific permissions to access products, orders, inventory, publications, and protected customer data.

Initially, it was tempting to expose all this technical state directly:

That information is useful for debugging, but it can create a terrible experience for ordinary merchants.

A merchant should not need to understand OAuth scopes to use an AI assistant.

Jam still needs to enforce least-privilege permissions internally, but the product should translate those requirements into clear outcomes:

One of my continuing goals is to hide infrastructure complexity without hiding important limitations.

Simple UX does not mean pretending everything succeeded. It means explaining the next useful action in language the merchant understands.

📧 Marketing Automation Needs Stronger Guardrails

Abandoned-checkout recovery sounds like a straightforward agent workflow:

  1. Find an abandoned checkout
  2. Draft an email
  3. Send it

In reality, the difficult part is everything around the draft.

Before sending, Jam must check:

The LLM can write recovery copy, but it cannot override these controls.

This is one of the clearest examples of where agentic automation must remain bounded. A creative model is useful; an unconstrained marketing sender is dangerous.

🧪 Testing Months of Store Activity Without Waiting Months

Another practical challenge was testing Jam’s proactive behavior.

Some recommendations require meaningful order history. Waiting for a development store to naturally accumulate months of activity is not realistic.

I therefore built a development-only store activity simulator that can:

The simulator is tenant-scoped, explicitly allowlisted, and rejected outside development. It cannot call Shopify, send emails, approve recommendations, or publish anything.

This allows me to observe how Jam responds to business activity without weakening the real production boundaries.

It does not replace actual Shopify testing. Real storefront visits, checkouts, webhooks, publishing, and email delivery must still be tested through Shopify and the relevant providers.

The simulator validates Jam’s internal behavior; controlled integration testing validates the external boundaries.

🧠 What “Trustworthy” Means for Jam

I don’t think trustworthy AI means the model never makes a mistake.

It means the system anticipates that models, APIs, networks, and humans can all make mistakes.

For Jam, trustworthiness means:

  1. Evidence: Recommendations are grounded in real store data
  2. Determinism: Business facts are calculated outside the model
  3. Transparency: Important limitations are shown clearly
  4. Approval: Sensitive actions require explicit authorization
  5. Isolation: Every record and action belongs to the correct store
  6. Idempotency: Repeated requests do not create duplicate effects
  7. Verification: Jam checks whether the external action actually happened
  8. Auditability: Decisions, approvals, attempts, and outcomes are recorded
  9. Consent: Marketing rules are enforced independently of model output
  10. Safe failure: Uncertainty pauses execution instead of pretending success

These systems are less exciting to show in a short demo, but they are what turn a prototype into software that can eventually be trusted with a real business.

———

🚀 Building Jam in Public

Jam is still being built.

There are more workflows, integrations, evaluations, deployment work, and UX improvements ahead. I’m deliberately avoiding the claim that it is a fully autonomous business operator.

The direction is clear, though.

Merchants should not have to configure multiple “agents,” learn prompt engineering, or become data analysts. They should connect their store, understand what matters, review important actions, and let Jam handle the repetitive work safely.

Throughout this series, I’ll document both the features and the engineering underneath them:

The most valuable lessons have not come from the moments when the AI looked impressive.

They have come from the moments when something failed and the system had to explain, recover, and avoid making the situation worse.

🔗 Resources to Explore