Most store managers can recite last week’s sales figure from memory and almost none can explain it. Revenue is an outcome, not a cause. It moves when traffic moves, when the conversion rate moves, when baskets get bigger or smaller, and when payroll hours land in the wrong part of the week. A weekly number on its own gives a manager nothing to act on, which is why so many Monday morning reviews end with a vague instruction to “push harder” and no change in behavior on the floor.
The fix is not a bigger dashboard. Most retail reporting systems already produce more metrics than any store team can absorb, and the volume is part of the problem. What works is a short, fixed list of retail store KPIs reviewed on the same day every week, calculated the same way every time, and tied to a specific decision. Five metrics cover almost every question a store manager needs to answer: conversion rate, units per transaction, average transaction value, sales per labor hour, and traffic itself.
In short
- Revenue is a result, not a diagnosis. Two stores can post identical sales weeks with completely different traffic, conversion and basket profiles, and the fixes for each are opposites.
- Conversion rate is the highest-leverage store metric, but it is only as trustworthy as the traffic count underneath it, and most stores count traffic badly.
- Units per transaction (UPT) reads the floor: it reflects attachment selling, adjacency, stock availability and how confident staff are with the range.
- Sales per labor hour (SPLH) is the only metric on the list that connects payroll to output, which makes it the one most likely to change a schedule.
- Benchmarks from your own history beat published industry averages. Compare the same store to the same week last year and to its own trailing median, then act on the gap.
Why revenue alone tells a store manager nothing useful
Consider two stores that both finish the week at $84,000. Store A served 1,200 transactions at an average of $70. Store B served 2,100 transactions at an average of $40. Those are different businesses operating under the same banner, and a single instruction issued to both will help one and hurt the other.
Now add traffic. Store A saw 6,000 visitors, so it converted 20 percent. Store B saw 21,000 visitors and converted 10 percent. Store A has a traffic problem and a healthy floor. Store B has the opposite: plenty of people walking in and a conversion failure that costs it more revenue than any marketing campaign could recover. Revenue alone hid both diagnoses completely.
This is the practical case for a fixed KPI set rather than a sales report. Each metric isolates one stage of the funnel, so a movement in the outcome can be traced to a stage. The wider operating discipline that surrounds these numbers, covering staffing, stock and floor standards, is set out in the retail store operations playbook, and the KPIs below are the measurement layer that sits on top of it.
The five metrics and what each one answers
| Metric | Formula | Question it answers | Primary lever |
|---|---|---|---|
| Traffic | Counted visitors, staff excluded | Are enough people coming in? | Marketing, location, windows, hours |
| Conversion rate | Transactions ÷ traffic × 100 | Are we turning visitors into buyers? | Coverage, service, stock, queue |
| Units per transaction | Units sold ÷ transactions | Are we selling more than one thing? | Attachment, adjacency, training |
| Average transaction value | Net sales ÷ transactions | What is each sale worth? | Mix, price architecture, trade-up |
| Sales per labor hour | Net sales ÷ hours worked | Is payroll producing output? | Schedule shape, task load |
Two rules make this table usable. First, calculate every metric on net sales, after returns and discounts, so that a heavy return week does not flatter the numbers. Second, use the same transaction definition everywhere: a single receipt is one transaction whether it contains one item or fourteen. Mixing definitions across reports is the fastest way to lose the team’s trust in the data.
Conversion rate: counting traffic properly before you trust it
Conversion rate is the most useful number on the list and the most frequently miscalculated. The formula is trivial. The denominator is not. If the traffic count is wrong, conversion is wrong by exactly the same proportion, and a manager can spend a quarter chasing a problem that exists only in the sensor.
A 12 percent conversion rate means nothing in isolation. Convenience and grocery formats routinely run above 90 percent because almost everyone who enters intends to buy. Furniture, jewelry and automotive showrooms can operate profitably in the low single digits because a single conversion is worth thousands. The only comparison that matters is your store against itself, and your store against a store of the same format and size in the same banner.
What actually counts as a visitor
A visitor is a person who enters the shopping area with the potential to buy. That definition excludes several groups that door sensors happily count: staff arriving and leaving, delivery drivers, contractors, mall walkers cutting through a corner unit, and the same customer re-entering after stepping outside to take a call.
Staff exclusion alone is worth doing properly. A store with 8 associates, each crossing the threshold 6 times a day across breaks and stockroom runs, generates roughly 48 phantom visits daily. In a small store counting 300 visitors a day, that is a 16 percent inflation of the denominator and a conversion rate understated by the same amount. Most modern counters support a staff-exclusion zone or a badge-based filter, and turning it on is usually a configuration change rather than a purchase.
Choosing a counting method
| Method | Typical accuracy | Strengths | Limitations |
|---|---|---|---|
| Horizontal beam | Lowest | Cheap, simple to install | Cannot separate groups or direction; miscounts wide doorways |
| Overhead infrared | Moderate | Directional, low cost, no images captured | Struggles with dense groups and strong sunlight |
| Stereo video / depth camera | High | Separates groups, handles children and carts, supports zones | Higher cost; needs a privacy review before deployment |
| Wi-Fi or Bluetooth sensing | Variable | Captures dwell and repeat visits | Only counts people with a device and radio enabled; consent rules apply |
Whichever method a store runs, validate it before trusting the output. Have a supervisor manually count entries for two separate 60-minute windows, once during a quiet weekday morning and once during a Saturday peak, then compare against the sensor for the same clock period. A gap under roughly 5 percent is workable. A gap above 10 percent means the conversion trend is noise, and the sensor needs recalibration or repositioning before any conclusion is drawn from it.
Reading a conversion movement
When conversion falls and traffic holds steady, the cause is almost always inside the store: a coverage gap at a peak hour, a stockout in a headline category, a queue at the till, or a fitting room left unstaffed. When conversion rises while traffic falls sharply, the store is usually not improving. It is simply seeing a higher-intent mix because the casual browsers stopped coming, which is a demand warning dressed up as good news.
Traffic quality is itself a lever rather than a fixed input. A window that communicates the current offer pulls in people already primed to buy, which lifts conversion without any change in service, and the mechanics of that are covered in this guide to window displays that pull foot traffic off the street.
Units per transaction and what it says about the floor
UPT is units sold divided by transactions. It answers one question: when someone decides to buy, do they leave with one item or several? Because it strips out price entirely, UPT is the cleanest read on selling behavior available in a standard point-of-sale report.
A UPT of 1.0 means every customer bought exactly one thing, which for most apparel, beauty, hardware and specialty formats indicates that nobody is asking a second question at the counter. Movement in UPT tends to have four causes, and separating them is the whole exercise.
- Attachment behavior: whether associates suggest a complementary item, and whether they do it at the fitting room or only at the till, where it is far less effective.
- Adjacency and layout: whether the complementary product is physically near the anchor product, or three aisles away in a different department.
- Stock availability: a broken size run or a missing accessory caps UPT no matter how good the service is.
- Promotional structure: multi-buy offers lift UPT mechanically while often reducing average selling price, so the two must be read together.
That last point deserves care. A “buy two, get one free” mechanic will raise UPT and lower average transaction value at the same time, and a manager reading only one of the two will draw the wrong conclusion about whether the promotion worked. Read UPT and ATV as a pair, always, and check gross margin dollars per transaction before declaring a win.
Where UPT is won
The practical lever is the point of the conversation rather than its content. An attachment suggestion made while the customer is still deciding, in the aisle or at the fitting room, converts far more often than the same suggestion made once the card is already out. Stores that shift attachment coaching from the counter to the floor typically see UPT move within a fortnight, and the associate-led version of this discipline, where staff maintain relationships and prepared recommendations, is developed further in this piece on clienteling and turning store associates into a channel.
Average basket versus average transaction value
These two terms are used interchangeably in most stores, and the ambiguity causes real reporting errors. Fix the definitions locally and write them down.
Average transaction value (ATV) is net sales divided by transactions. It is what one receipt is worth. Average basket is used two different ways in practice: sometimes as a synonym for ATV, and sometimes to mean the average value of a customer’s basket including items abandoned before checkout, which only stores with basket-level tracking can measure. If your reporting cannot distinguish the two, use ATV exclusively and retire the word basket from internal reports.
ATV decomposes cleanly into two drivers: ATV equals UPT multiplied by average selling price (ASP). That identity is the fastest diagnostic in retail reporting, because it converts a vague “baskets are down” observation into a specific question about which of the two factors moved.
| What moved | UPT | ASP | Most likely cause | First thing to check |
|---|---|---|---|---|
| ATV down | Down | Flat | Attachment stopped, or an accessory line went out of stock | Availability report on top accessory SKUs |
| ATV down | Flat | Down | Markdown depth or a shift to opening price points | Discount rate and mix by price band |
| ATV up | Down | Up | Fewer, higher-value sales; possible loss of entry-level stock | Sell-through on opening price point |
| ATV flat | Up | Down | Multi-buy promotion working as designed | Gross margin dollars per transaction |
Run this decomposition before any conversation about pricing. A large share of perceived pricing problems turn out to be availability problems in a single category, and they resolve with a replenishment fix rather than a markdown.
Sales per labor hour: the metric that links payroll to results
Sales per labor hour is net sales divided by total hours worked in the same period. It is the only metric on this list that touches the single largest controllable cost in a store, which is why it tends to be the one that actually changes a decision.
Two definitional choices determine whether SPLH is useful or misleading. First, decide whether the denominator includes all scheduled hours or only selling-floor hours. Including stockroom, delivery and administrative time gives a truer productivity picture but depresses the number, so consistency matters more than the choice itself. Second, use hours actually worked rather than hours scheduled, or the metric will quietly reward poor attendance management.
Reading SPLH by hour, not by week
A weekly SPLH figure tells a manager whether the store is broadly over or under-hourly. It does not tell them what to change. The actionable version is SPLH by hour of day, laid against traffic by hour of day, because the mismatch between the two is where money is lost.
The classic pattern: a store schedules heavily from 9am to 5pm because that is when deliveries and administration happen, while traffic peaks between 4pm and 7pm. Conversion sags in the evening because coverage thins exactly when the most people are in the building. Nothing about the total hour count is wrong; the shape is wrong. Rebuilding the schedule around the traffic curve rather than the task curve is usually a same-cost change, and the mechanics of doing that under peak-season pressure are laid out in this guide to building a store labor schedule that survives peak season.
The trap in chasing a higher number
SPLH improves whenever hours are cut, right up to the point where it collapses. Strip coverage far enough and conversion falls, queues form, replenishment stops, and sales drop faster than the payroll saving. Because the damage lands in a different metric from the one being optimized, this failure is easy to miss for several weeks.
Guard against it by never reviewing SPLH alone. Pair it with conversion rate in the same view. If SPLH is rising while conversion is falling, the store is not becoming more productive; it is becoming understaffed, and the trend will reverse itself expensively. Both metrics have to move in the same direction for a coverage change to count as a genuine improvement.
Setting benchmarks when you only have your own history
Published industry averages are close to useless at store level. A national conversion benchmark blends convenience formats with jewelry showrooms, mall units with high-street units, and stores with accurate counters with stores whose sensors have been broken since installation. Aggregate US retail context is genuinely useful for reading demand, and the US Census Bureau monthly retail trade reports are the standard reference for that, but the number that governs a single store’s targets has to come from its own history.
Build the baseline in three steps. Pull at least 13 weeks of data for each metric, ideally 52 so that seasonality is visible. Take the median rather than the mean, because a single blowout week or a storm closure will drag an average somewhere unhelpful. Then record the interquartile range, meaning the spread between the 25th and 75th percentile weeks, which defines what normal variation looks like for that store.
What counts as a real movement
Once the spread is known, a movement inside it is noise and needs no response. A movement outside it, sustained for two consecutive weeks, is a signal worth investigating. This single rule prevents the most common failure in weekly reviews, which is reacting to random variation and exhausting the team with instructions that reverse themselves seven days later.
Comparison sets matter too. Compare each week against the same week last year rather than the week before, so that seasonality does not masquerade as performance. Compare against the store’s own trailing median for the last 13 weeks to catch drift. Compare against sibling stores of the same format and size band, never against the chain average, which is dominated by whichever format has the most doors. For temporary formats where no history exists, published cost and revenue reference points are the only starting point available, and this set of costs and revenue benchmarks for a 30-day pop-up shows the shape those figures usually take.
Labor context for SPLH targets
SPLH targets have to be revisited whenever wage rates move, because the same dollar output covers a different share of payroll cost. Tracking the relevant wage series, published by the US Bureau of Labor Statistics, is enough to know when a target set eighteen months ago has quietly stopped being achievable at the old staffing level. A store held to a stale SPLH number under higher wages is being asked to cut hours it cannot afford to cut.
Turning the weekly numbers into two concrete actions
A weekly KPI review that ends without a decision is a meeting, not a management process. The output of every review should be at most two actions, each with a named owner and a check date. Two is not an arbitrary limit: a store team can genuinely change two things at once and cannot change six.
A workable sequence takes about twenty minutes. Start with traffic, because it sets the context for everything below it. Then conversion, then UPT and ATV together, then SPLH against conversion. At each step ask only whether the metric moved outside its normal range for two weeks running. Stop at the first metric that did, and diagnose that one rather than continuing down the list.
- Traffic outside range: the cause is usually external (weather, local events, a competitor opening, a marketing change). Note it and move on, because the store cannot fix demand this week.
- Conversion outside range: check coverage at peak hours, availability in the top three categories, queue length and fitting room staffing, in that order.
- UPT outside range: check accessory availability first, then where attachment conversations are happening on the floor.
- ATV outside range: decompose into UPT and ASP before touching anything, then check discount rate and mix by price band.
- SPLH outside range: compare the schedule shape against the traffic curve by hour, and confirm conversion is not falling at the same time.
Write the two chosen actions somewhere the team sees them daily, in specific language: “Second associate on the floor 4pm to 7pm Thursday through Saturday, owner Dana, review next Monday” rather than “improve evening coverage”. Then, and this is the part most stores skip, actually check the result on the stated date. An action that is never reviewed teaches the team that the KPI review has no consequences, and attention drains away within a month.
These five metrics are a measurement layer, not a management system. They tell a manager where to look; they do not tell anyone how to run a shift, set standards, or handle stock flow. Pairing the weekly numbers with the broader operating routines in the retail store operations playbook is what turns a report into a change on the floor.
FAQ on store KPIs
What is a good conversion rate for a retail store?
There is no universal figure, because conversion depends almost entirely on format. Convenience and grocery stores often exceed 90 percent, specialty apparel commonly sits somewhere in the 15–30 percent band, and high-consideration categories such as furniture or jewelry can be profitable in the low single digits. The useful benchmark is your own store’s trailing 13-week median, compared against the same week last year.
How often should store KPIs be reviewed?
Weekly, on a fixed day, is the right cadence for a store team. Daily review invites reaction to random variation, since a single day contains too few transactions to be statistically meaningful in most stores. Monthly review is too slow to correct a coverage or availability problem before it costs a full period of sales.
What is the difference between UPT and average basket size?
UPT counts items and ignores price: it is units divided by transactions. Average transaction value counts dollars: it is net sales divided by transactions. The two are linked by average selling price, since ATV equals UPT multiplied by ASP. A promotion can raise UPT while lowering ATV, so both should always be read together.
Should staff be excluded from the traffic count?
Yes, and the effect is larger than most managers expect. In a store counting a few hundred visitors a day, staff crossing the threshold for breaks, deliveries and stockroom runs can inflate the count by 10 percent or more, which understates conversion by the same proportion. Most counters support a staff-exclusion setting; enabling it is usually a configuration change rather than new hardware.
How do I calculate sales per labor hour correctly?
Divide net sales by the hours actually worked in the same period, not the hours scheduled. Decide once whether the denominator includes non-selling time such as stockroom and administration, document that choice, and never change it mid-year. Consistency matters more than which definition you pick, because the metric is only meaningful as a trend.
Which KPI should a store fix first?
Conversion rate, in almost every case. It usually carries the largest addressable gap, it responds to changes the store controls directly (coverage, availability, queue and service), and improvements there flow straight through to revenue without any additional traffic or marketing spend.
Is a rising sales per labor hour always good news?
No. SPLH rises automatically whenever hours are cut, including cuts that damage the business. If SPLH is climbing while conversion falls, the store is understaffed rather than more productive, and sales will follow conversion down. Treat the two metrics as a pair and require both to improve before calling a coverage change successful.
How much data do I need before setting a target?
Thirteen weeks is the practical minimum for a stable median, and 52 weeks is needed to see seasonality clearly. Use the median rather than the mean so one exceptional week does not distort the baseline, and record the spread between the 25th and 75th percentile weeks so the team knows what normal variation looks like before reacting to it.
Can these KPIs work for a store without a traffic counter?
Partly. UPT, average transaction value and sales per labor hour all come from point-of-sale and payroll data and need no sensor at all, so a store can run four fifths of this framework immediately. Conversion rate is the exception, since it requires a traffic denominator, and manual counts during sample hours are a workable interim substitute while a counter is specified.