Article

Fit Confidence

Suddha Ray | Senior Consultant Consultant | Retail Reply, London, UK
Version 1.1 | July 2026

Fit Confidence

Reducing fashion returns with a guided shopping assistant

 

How We Got Here

Online shops began as digital catalogues. A shopper typed a word; the system matched it against product titles and tags and back came a list. This worked well when people knew exactly what they wanted: a brand, a size, a colour.

It stops working the moment the question gets softer. “A jacket for a smart casual weekend in Edinburgh, not too formal” is not a search term. It describes an occasion, a season and a look. No amount of filtering by category, size and price answers that. The shopper is left to open twenty product pages, guess at sizing, and either give up or buy the wrong thing.

That gap is not new. What has changed is who is stepping into it. Google now uses its Gemini system to answer product questions directly inside search results. Amazon's Rufus assistant already handles fit, comparison and price history questions for its own shoppers. If a retailer's product data and stock levels cannot be read and understood by tools like these, the retailer risks becoming invisible at the exact moment a customer starts looking, and not only on its own website.

At home, the cost of this gap shows up in a number every fashion finance director already watches. UK online fashion returns run somewhere between 23 and 30 per cent, and poor fit is repeatedly named by shoppers as one of the leading reasons for sending an item back. ¹ A tool that helps someone choose the right size before they buy has a direct, measurable effect on a cost that is already on the books.

So, the real question is not whether conversational shopping is coming. It is already here, running on someone else's platform. The question is whether to build the capability in-house, on the retailer's own terms, or to leave that part of the customer relationship to Amazon and Google.

 

What We Are Proposing To Build

We propose a layer that sits between your existing systems and the shopper: a service that reads your product data, stock and pricing through secure interfaces, known as APIs (application programming interfaces - the agreed way that two pieces of software exchange information), and turns a shopper's question into a specific, checkable action.

We are not proposing to replace your ecommerce platform, your product data system (often called a PIM, or product information management system), your order management system, your search engine, your customer records system, or your checkout. This layer sits on top of all of them and translates between the shopper and the systems you already run.

Consumer research (Gartner) shows that shoppers do not trust a retailer's own AI assistant any more than they trust a similar tool built by a third party. ² So the case for building this in-house is not that customers will trust you more for owning it. The case is that you keep control of your product truth, your pricing, your customer data and your policies, rather than handing that control to a platform you do not run. That is an argument about data and governance, not about trust, and it should be presented to your board as one.

The shopper stays in charge throughout. The assistant can inform, shortlist and compare. It cannot add anything to a basket without the shopper's explicit confirmation. This is not a safety feature bolted on afterwards. It is the design, and the evidence for it is strong. Gartner's most recent consumer research found that three in four shoppers feel stressed by the idea of an AI system making a purchase decision on their behalf, and that people's openness to a fully autonomous shopping agent has barely moved in ten years. ³ What shoppers do want is help making a better decision themselves: a shortlist, a price history, a straight answer on whether something is in stock. That is what this is built around.

 

What Has To Be True Before We Start

A drop-in layer is only as easy to drop in as the systems underneath it allows. Three things need to be confirmed before a pilot makes sense.

 

Working Access To The Data

The assistant needs, at minimum, read access to product data, stock levels, pricing and promotions. Adding items to a basket need write access to the cart and the checkout handoff. If these are not already available as clean, documented APIs, the first phase becomes an integration project rather than a pilot, and the timeline has to say so honestly.

 

Content You Can Stand Behind

The product catalogue alone is not enough. To answer a fit question properly, the assistant needs the official sizing guide, not a scrape of customer reviews. For care and material questions it needs the verified specification. For allergen information in grocery it must draw only on sources the retailer has checked itself. The rule is simple: if we cannot point to where a claim comes from, the assistant does not make it. This is enforced at the point the answer is shown to the shopper, checked against an approved list of sources, not simply written into the assistant's instructions and hoped for.

 

Agreed Moments For The Assistant To Speak Up

The assistant should not greet every visitor the moment they land on the site. It should stay quiet and present, becoming active only at agreed moments where there is evidence the shopper needs help: a poor search result, repeated comparison between the same few items, a long pause on one product page, a size or colour that has sold out, hesitation over a review or a sizing chart, or a basket that looks like it is missing something obvious. These moments need to be agreed before building starts, because they decide what data the widget has to capture, which APIs it needs from your systems, and how we will know afterwards whether it actually helped.

 

The Pilot: One journey, One problem

Rather than trying to prove several things at once, we recommend starting with a single journey that has a clear problem, a measurable result, and a real advantage over what shoppers can already get elsewhere. We recommend fashion fit and confidence, with one starting trigger: hesitation over size or fit while browsing, comparing, or reading a product page. Other moments, such as an empty basket at the end of a session or an item that has gone out of stock, are better left for a later phase unless the chosen retailer has a stronger reason to start there.

 

Why Fashion, And Why Now

UK online fashion returns run at roughly 23 to 30 per cent, and poor fit is one of the reasons shoppers most often give for sending things back. ¹ Improving confidence before purchase has a direct effect on that figure, and it is a number retailers already track.

M&S, ASOS, Next and John Lewis all offer size guides and fit tools already. Some are experimenting further with AI styling and visual search, while Amazon has set the pace for conversational shopping through Rufus. The opening for Retail Reply is not to claim that Rufus is missing something. Rufus works, and shoppers already use it. The opening is orchestration the retailer actually owns product truth, stock, policy, customer history and basket working together under the retailer's own control, not a third party's. What follows in the table below is drawn from public information only. Some retailers may run features privately, in apps, or in limited trials that do not show up in a public scan, so this should be treated as a starting point for conversation, not a finished picture.
 

RetailerWhat is visible publiclyWhat that suggests
M&SAI styling and personalisation features reported; no full public assistant covering the whole site.Adjacent capability exists, not full journey orchestration.
ASOSAI stylist and visual search; the closest public signal to a full assistant.Proof that demands exists, not proof of stock or basket orchestration.
WaitroseNo full public assistant found.Open ground, particularly for grocery.
TK MaxxNo full public assistant found.Depends on how ready the underlying data and APIs are.
NextNo full public assistant found; Next's own strategic focus is building and selling its Total Platform infrastructure to other retailers.Next may prefer to build in-house, given its existing platform business.
John LewisNo full public assistant found; John Lewis Partnership has publicly described work on AI foundations and is exploring agentic AI.⁴Opportunity, but likely to land better after foundational data work than instead of it.

The Customer Journey

The table below follows one fashion fit session and shows what the assistant does at each step, and which of the retailer's systems are involved. The proactive moment is the single trigger: one instance of fit hesitation. The steps after that show how the conversation continues once the shopper chooses to engage.

MomentWhat the shopper says or doesWhat the assistant doesSystems involved
Search“A smart casual jacket for a city weekend in Edinburgh, not too formal”Turns a loose description into occasion, weather and fit constraints. Returns a shortlist of two or three items.Product catalogue, search index
Comparison“Which of those two is the safer choice?”Compares using the sizing guide, verified review summaries, fabric details and return patterns.Product content, reviews, sizing guides
Availability“I like the first one, do you have it in a 14?”Checks live stock. If it is not available, suggests the nearest alternative and explains the difference plainly.Inventory system
Decision“I'll take the blue one in a 14”Confirms the addition explicitly. The shopper must approve before anything changes in the basket.Cart system, checkout handoff
Cross-sell“Does it come with a matching scarf?”Suggests items from verified product relationships. Does not invent pairings that do not exist.Product catalogue


A Note On Price, Left Out On Purpose

Gartner's research found that price information, price history and price comparison are the single most requested capability shoppers want from an AI shopping tool, ranked above style advice and above availability checks. ⁵ We have left price comparison out of this first pilot on purpose. Fit is the narrower, more contained problem. Mixing two hard problems into one pilot makes it difficult to tell afterwards which one actually moved the return rate. Price comparison is the clear candidate for a second pilot, once this one has proven the pattern works and the retailer trusts the process.

 

A Note On Trust, Shown Rather Than Told

The same research found that more than half of shoppers who had already used an AI tool while shopping felt they had to double check everything it told them. ⁶ Accuracy enforced quietly behind the scenes does not close that gap by itself. The widget should show its working: where a fit claim comes from, when a stock or price figure was last checked. A shopper who can see where an answer comes from trusts it more than a shopper who is simply told to trust it.

What Retail Reply Builds, And What It Connects To

This distinction matters for how the work is scoped and how it is priced.


What We Build

The layer that reads what the shopper is asking and routes it to the right system. The connectors that talk to your commerce APIs. The set of rules that stops the assistant making a claim it cannot back up. The service that grounds every answer in approved content. The on-site widget the shopper actually talks to.


What We Connect To

Your existing systems: product catalogue, stock feeds, pricing, customer records, cart and checkout. None of these are replaced.


A Note On The AI Model Itself

The large language model, the piece of software that actually generates the assistant's replies, sits inside this layer and is not fixed to one supplier. If a retailer has requirements about where data is processed, or a preferred procurement route, the model underneath can be swapped without rebuilding anything else.

 

A Note On Data Handling

Conversation history that carries across a shopping session may include information that identifies the shopper. How that is stored, for how long, and by whom needs to be agreed at the start of the engagement, not discovered afterwards. Any deployment in the UK or the EU has to meet UK GDPR (General Data Protection Regulation) requirements, and how the EU AI Act applies should be confirmed by the retailer's own legal team before anything goes live.

 

Delivery, Phase By Phase
We cannot assume the APIs are ready or the product data is clean before checking. A short discovery phase comes first, and the timeline for everything after it depends on what that phase finds.
 

PhaseWhat happensLengthHow we know it is done
DiscoveryCheck the APIs, assess data quality, agree the trigger points, map the rules the assistant has to follow.3–5 weeksAPIs documented and tested. Content sources checked. Trigger points agreed.
CrawlA read-only assistant for the one chosen trigger: answers questions, compares products, checks stock. No basket actions.3–4 weeks after discoveryLive with sample products. Questions answered accurately, with nothing unsupported.
WalkA live pilot with basket additions gated on confirmation. Measurement in place from day one.6–8 weeksA measurable difference in return rate and conversion between assisted and unassisted sessions.
RunWider category coverage. Personalisation where the retailer has consent and the profile data to support it.OngoingThe pattern applied to a second retailer or category.

If the discovery phase finds that APIs need to be built rather than simply connected to, the discovery phase itself extends. That should be flagged to the client before any dates are agreed, not after.

 

How We Will Judge Whether It Worked

For the pilot, we recommend tracking one measure: the return rate on assisted sessions compared with similar unassisted sessions. Every fashion retailer already tracks return rate, and with a clean way of separating the two groups, it is easier to attribute than softer engagement numbers. It also has a straightforward pound value attached to it.

A second measure, conversion rate, should sit alongside it, using an A/B test designed to separate the assistant's effect from everything else happening in the same session.

We should not publish a return on investment figure for the pilot before the way of measuring it exists. Saying “12 to 18 per cent higher conversion” without showing how that number was worked out will not survive contact with a finance director. If we cannot explain how we arrived at a number, it will be waved away. It is better to propose a pilot that produces evidence someone can check than to promise a headline figure that cannot be defended later.

 

Why Retail Reply

Calling an AI model through an API is not the hard part of this. Most competent engineering teams can do that in an afternoon. The hard part is connecting it safely to live commerce systems, making sure it only says things it can support, and governing what it is and is not allowed to do in front of real customers spending real money.


What Tends To Go Wrong Without That Experience

The data feeding the assistant turns out to be stale or inconsistent. The APIs behave unpredictably once real traffic hits them. The rules meant to keep the assistant honest are written down as policy but not actually checked at the point an answer is shown. The pilot produces engagement numbers nobody can cleanly attribute to anything. The legal team finds a problem with how session data was handled, after it was already live.

Retail Reply's value is knowing where those failures happen, from having connected commerce platforms before, and having delivery patterns that move at a reasonable pace without creating a liability for the client along the way.

The parts of this that can be reused across clients, the on-site widget, the retrieval service, the rule set, the API connectors, and the way we test and measure it all, add up to something closer to a product than a one-off build. Each engagement improves the shared toolkit for the next one.
 

Three Decisions Before This Becomes A Pilot Proposal

  1. One client, One Category
    Which account has the clearest API readiness and the strongest case for cutting returns? Fashion is our recommendation, but the actual choice should be driven by where the door is most open and the starting point most workable, not by which category has the biggest headline market.
  2. One Trigger Point For The Pilot
    Search friction, unavailable stock, or an incomplete basket. We need to pick one. Running all three at once makes it hard to tell what worked and hard to manage the build. The other triggers do not go away; they simply move to a later phase.
  3. One Measure Of Success
    Return rate is the most defensible measure for a fashion fit pilot. If the retailer can cleanly separate assisted and unassisted sessions, attribution is reasonably solid. Conversion rate sits alongside it as the second measure. Agreeing how success will be measured before building starts avoids an argument about the results after the fact.