Grading and authenticating used goods at scale without a lab

Resale looks like a margin business until the first full quarter of data arrives. Then the returns rate on used inventory comes in at two or three times the rate on new goods, the customer service queue fills with “not as described” claims, and the marketplace account health dashboard starts flashing. In almost every case the root cause is not fraud or bad luck. It is grading: two people looking at the same jacket, applying the same written scale, and reaching different answers.

Grading and authentication are the two operational disciplines that decide whether a resale channel earns money or quietly bleeds it. Both are usually treated as judgment calls made by whoever is standing at the intake table. Both can be turned into a documented, repeatable process that a seasonal hire can run on day two, without a testing lab, a spectrometer or a six-figure software contract. This guide sets out how to build that process, what to measure, and where the honest limits of in-house authentication sit.

In short

  • Inconsistent grading is the single largest driver of resale returns. The same physical item graded differently by two staff members produces a dispute rate that looks like a quality problem but is actually a definitions problem.
  • A four or five point grade scale beats a ten point scale because graders can hold the definitions in working memory, and buyers can understand the difference at a glance before they click.
  • Photo standards do more dispute prevention than description text. A fixed shot list, consistent lighting and a mandatory flaw close-up remove most of the ambiguity a buyer would otherwise resolve in their own favor.
  • Authentication is category specific. Apparel, sneakers and electronics fail in different ways, so a single checklist applied across all three catches very little in any of them.
  • Audit grades against returns data monthly, not annually. Return rate by grade, by grader and by category tells you within one cycle whether your scale is calibrated or whether one person is quietly inflating every item they touch.

Why grading consistency drives resale returns and disputes

A new product has one condition: new. A used product has a condition somewhere on a continuum, and the buyer’s expectation is set entirely by the label you attach to it. When that label means different things depending on who applied it, the gap between expectation and reality shows up as a return, a refund request, or a negative review that suppresses the listing for months. This is the mechanism that separates a profitable resale channel from an expensive one, and it is covered in more depth in our guide to recommerce and how resale became a real retail channel.

The financial asymmetry is worth stating plainly. A return on a new item costs you shipping, handling and a restock. A return on a used item costs you all of that plus a re-grade, a re-photograph, a re-list, and the near certainty that the item now sells at a lower price than it would have first time. One mis-graded item can absorb the margin of three correctly graded ones.

The three failure modes of an uncalibrated scale

Grade inflation is the most common. Under pressure to clear an intake backlog, graders round up: a “good” becomes an “excellent” because the flaw is small and the shift ends in twenty minutes. Inflation feels harmless per item and is expensive in aggregate.

Grade drift is the second. Definitions written in January are interpreted differently by August, because nobody re-read them and the original team has turned over. Drift is invisible in daily operations and obvious in a year-on-year returns chart.

Category blindness is the third. A scale written for apparel gets applied to handbags, then to small electronics, and the words stop describing anything meaningful. “Light wear” means something for a cotton t-shirt and almost nothing for a laptop.

What consistency is actually worth

Consistency compounds in three places at once. Returns fall because expectations are met. Customer service time falls because there is a photograph that settles the argument. Repeat purchase rates rise because a buyer who received exactly what the grade promised will trust the next grade you publish.

There is a reputational layer too. Buyers who feel misled by a condition label rarely blame the individual grader; they blame the seller, and increasingly they blame the sustainability claim attached to the channel. That is where the operational problem becomes a brand problem, a pattern we cover in what sustainable retail actually means beyond the marketing.

Building a grade scale customers understand at a glance

The temptation is to build a granular scale, because granularity feels precise. In practice, precision the grader cannot apply consistently is worse than a coarser scale they can. Four or five grades is the range where most teams land after a year of iteration.

Each grade needs three things: a short customer-facing label, a plain-language promise of what the buyer will receive, and an internal rule that a grader can test against the item in front of them. The internal rule is the part most teams skip, and it is the part that does all the work.

Grade Customer-facing promise Internal test rule Typical price index vs. new
Like new Shows no visible signs of use at normal viewing distance. Original packaging or tags may be present. No flaw visible at arm’s length under standard lighting. Zero functional deviation. Fewer than 5 estimated uses. 70 to 85 percent
Excellent Very light signs of use, nothing that draws the eye. Fully functional. At most one minor flaw, smaller than a fingernail, not on a prominent surface. Full function. 55 to 70 percent
Good Visible wear consistent with regular use, photographed and described. Works as intended. Up to three visible flaws, each photographed. No structural damage. Full core function. 35 to 55 percent
Fair Significant wear or a named defect. Priced accordingly and sold as described. Four or more flaws, or one major flaw, or a partial function limitation that is explicitly stated. 15 to 35 percent
For parts or repair Not fully functional. Sold for components, repair or recycling. Any failure of core function. Never listed as usable. 5 to 15 percent

Price indices in that table are illustrative anchors for building a pricing ladder, not market rates. Actual recovery varies enormously by category, brand strength and season, and the ladder needs to be calibrated against your own realized sale prices. The related discipline of setting those numbers without undercutting your own new-goods business is covered in our piece on pricing used items without eating new sales.

Write rules that can be falsified

A good internal rule can be proven wrong by looking at the item. “Minor wear” cannot. “At most one flaw smaller than a fingernail, not on a prominent surface” can. The test is whether two graders who disagree can resolve the disagreement by re-reading the rule rather than by escalating to a manager.

Anchor every subjective term to something physical. Viewing distance, a coin, a fingernail, a business card and a standard sheet of paper are all better reference objects than adjectives. Distance rules work particularly well: “visible at arm’s length” and “visible only on close inspection” split most flaws cleanly.

Never let a grade carry a hidden functional claim

The most expensive disputes come from items that look right and do not work. Separate cosmetic grade from functional status in your data model, even if you display them as one label. An “excellent” phone with a battery at 71 percent health is not the same product as an “excellent” phone at 94 percent, and the buyer will treat the difference as misrepresentation rather than as a detail.

Publish the scale where the buyer will actually read it

A grading scale buried in a footer help page does not set expectations. Put the specific grade definition on the listing itself, next to the grade label, in the words the buyer will use if they later complain. This single change removes a meaningful share of “not as described” claims because the claim now has to argue with text the buyer already saw.

Photo and description standards that prevent disputes

Photographs are the evidence layer of a resale operation. When a dispute reaches a marketplace mediator, the seller who has a date-stamped, consistent, flaw-forward photo set wins, and the seller who has three flattering angles does not. Treat the shot list as a compliance artifact rather than a marketing one.

The economics favor speed plus consistency over artistry. A fixed setup that produces eight adequate photographs in ninety seconds beats a beautiful setup that produces four in six minutes, because the marginal dispute prevented by photograph number seven is worth more than the marginal click won by better styling.

The mandatory shot list

Every item, regardless of category or grade, should carry the same base set: front, back, both sides, a label or serial detail, and one close-up per declared flaw. Items graded below “excellent” should carry a wide shot that shows the overall condition honestly, not just a crop of the best surface.

Flaw photographs need context. A macro shot of a scuff with nothing else in frame is ambiguous about scale, so include a reference object or shoot the flaw within the wider panel so the buyer can see how small it actually is. Sellers frequently under-photograph flaws in the belief that it hurts conversion. It hurts conversion slightly and it cuts returns considerably.

Fix the lighting and the background once

Two variables cause most of the visual inconsistency across a catalog: color temperature and background. Fix both permanently. A consistent white or light grey background, a fixed light position and a locked white balance mean a “good” item photographed in March looks like a “good” item photographed in October.

Write descriptions that pre-empt the complaint

The best condition description reads like the complaint the buyer would have written, stated first by you. Name the flaw, name its location, name its size, and state what it does not affect. “Scuff approximately 1 cm on the lower right rear panel, does not affect the zipper or lining” is a complete disclosure and a closed argument.

Avoid two categories of language: marketing softeners (“gently loved”, “barely used”) and legal absolutes (“perfect”, “flawless”, “guaranteed authentic” without a defined basis). The first invites disappointment and the second invites a claim you may not be able to substantiate.

Standard Weak version Dispute-resistant version
Flaw disclosure “Minor signs of wear” “Two scuffs, each under 1 cm, on the left side panel, photographed in image 5”
Sizing Copy the label size Label size plus measured flat measurements: chest, length, sleeve
Function “Works great” “Powers on, tested charging and both speakers, battery health 88 percent as reported by the device”
Completeness “Comes with accessories” Itemized list of what is included and an explicit list of what is missing
Authenticity “100 percent authentic” “Verified against brand construction markers; see our authentication process. No third-party certificate included.”

Authentication checks by category: apparel, sneakers, electronics

Authentication and grading are separate jobs that happen at the same bench. Grading asks how worn the item is. Authentication asks whether it is what it claims to be. Counterfeit goods are a large and well-documented global problem, described in general terms in this overview of counterfeit consumer goods, and resale channels are an obvious surface for them.

What a small team can realistically do is triage. In-house checks are very good at catching low-effort fakes, reasonable at flagging items that need escalation, and poor at defeating high-quality counterfeits that are specifically engineered to survive visual inspection. Designing the process around that reality is more useful than pretending otherwise.

Apparel and accessories

Start with construction rather than logos, because logos are the part counterfeiters get right first. Stitch density, seam alignment at pattern joins, the weight and hand of the fabric, and the finish quality on interior seams are all harder to replicate at scale than an embroidered mark.

Labels carry more signal than most graders use. Care labels have country-specific regulatory content, fiber composition ordering conventions and print methods that change over time, so a label that does not match the claimed production era is a strong flag. Hardware is similar: zipper pulls, rivets and clasps are often stamped, and the weight of metal hardware is difficult to fake cheaply.

Build a reference library as you go. Photograph the interior labels, hardware stamps and seam details of every verified-authentic item that passes through, and you accumulate a comparison set specific to the brands you actually handle within a few months.

Sneakers

Sneakers are the most heavily counterfeited resale category and also the most documented, which cuts both ways: the reference material available to you is also available to the counterfeiter. Check the boxes and labels first, because production data on the box label should reconcile with the size tag inside the shoe and with the known release information for that style code.

Move then to construction: glue application at the midsole, the symmetry of the toe box between left and right, the stitch count on visible panels, and the texture of any suede or leather panel. Smell is a genuine and underused signal, since the adhesives used in unauthorized production frequently have a distinct chemical odor.

Where the resale value is high and the reference set is thin, this is the category most worth escalating to a specialist rather than clearing in-house.

Consumer electronics

Electronics authentication is less about counterfeits and more about provenance and function, though counterfeit chargers, cables and accessories are common. The first check is always the serial or IMEI: verify it against the manufacturer’s own lookup service, confirm the model matches the physical device, and check the device against the relevant stolen-device or activation-lock status where a service exists.

Then test function on a fixed script rather than by intuition. Power, charge, all ports, camera, speakers, microphone, screen touch response across the full panel, and battery health as reported by the device itself. Write the script down and require the grader to tick every line, because the untested port is always the one the buyer needs.

Data is a compliance issue as much as a quality one. Any device that has held user data should be factory reset and, where the category and jurisdiction call for it, wiped to a documented standard before it is listed. Keep a record that the wipe happened.

Category Primary in-house checks What in-house checks will not catch Escalation trigger
Apparel and accessories Construction quality, interior seams, care and composition labels, hardware weight and stamping High-grade replicas using genuine-specification materials Resale value above your threshold, or any label inconsistency you cannot resolve
Sneakers Box label reconciliation, style code, midsole glue, toe box symmetry, stitch count, odor Factory-variant and “unauthorized authentic” production Hyped or limited style codes, or any deadstock claim
Consumer electronics Serial and IMEI lookup, activation lock status, scripted function test, battery health Internally swapped components and non-original repairs Serial mismatch, activation lock present, or evidence of prior opening
Watches and jewelry Weight, movement type where visible, hallmarks, serial engraving quality Almost everything above entry level Effectively any item of material value

When to pay for third-party authentication

Third-party authentication is a cost per item, so the decision is a straightforward expected-value calculation that most teams never actually write down. The inputs are the fee, the resale value of the item, your estimated false-negative rate on in-house checks, and the full cost of a counterfeit reaching a buyer.

That last input is the one teams underestimate. The cost of a counterfeit shipped is not the refund. It is the refund plus the marketplace penalty, plus the account health damage, plus the review, plus the possibility of a rights holder complaint. Price that number honestly and the threshold for paying an authentication fee drops considerably.

A workable decision rule

A simple version that survives contact with reality: escalate any item where the authentication fee is less than the resale value multiplied by your estimated in-house miss rate multiplied by the full cost multiple of a counterfeit reaching a buyer. In practice most teams end up with a value threshold plus a category override list, which is easier to train than a formula.

Two categorical overrides are worth having regardless of value. Escalate anything in a category where you have no reference library, and escalate anything a grader has flagged as uncertain, because the second-guess rate on flagged items is high enough to justify the fee on its own.

Marketplace authentication programs change the math

Several major marketplaces now run their own authentication services on qualifying categories, which shifts part of the risk off your balance sheet in exchange for fees, handling time and a loss of control over the customer experience. Whether that trade favors you depends heavily on whether you are building your own resale channel or renting someone else’s, a decision we compare in detail in brand-owned resale versus ThredUp and Poshmark.

Understand what the enforcement environment looks like

Counterfeit enforcement is an active policy area, and the pressure on platforms is increasing rather than easing. The Office of the United States Trade Representative publishes an annual review of markets associated with counterfeiting and piracy concerns, available at the USTR Notorious Markets List, which is a useful barometer of where regulatory attention is pointed. Rules, enforcement priorities and platform policies change frequently, so treat any specific figure or requirement as something to verify at the official source before relying on it.

Training seasonal staff to grade the same way

Grading quality is a training problem disguised as a hiring problem. Teams assume they need experienced graders, then discover that experienced graders trained elsewhere bring a different scale with them. A well-documented process turns an inexperienced hire into a consistent grader faster than experience alone does.

The core training asset is a calibration set: twenty to thirty real items, already graded and agreed by your senior staff, photographed and stored. New graders grade the set blind, and their answers are compared against the agreed grades. This takes an afternoon to assemble and pays for itself in the first week of a peak season.

Run calibration as a recurring event, not an onboarding step

Drift affects experienced graders more than new ones, because new graders are still reading the rules. Re-run the calibration set quarterly for everyone, including whoever wrote the scale, and track the results by person over time. A grader whose scores diverge from the agreed set in a consistent direction has a correctable bias, not a competence problem.

Keep a small number of disputed items in the calibration set deliberately. The items that senior staff argued about are the ones that reveal where the written rules are ambiguous, and every argument is a prompt to tighten one line of the scale.

Double-grade the tail, not the middle

Full double-grading is too slow for volume operations. Target it instead: double-grade items above a value threshold, items in categories with a poor returns record, and a random sample of perhaps 5 to 10 percent of everything else. The random sample is what keeps the process honest, because graders know any item might be checked.

Make the process the fast path

If following the standard is slower than skipping it, the standard will be skipped during peak. Design the bench so that the compliant path is the quickest one: the shot list taped to the wall, the reference objects within reach, the checklist on the screen the grader is already using, the reference library one click away.

This operational discipline is what makes a circular channel viable at all, rather than a marketing gesture, a distinction we explore in circular retail business models that actually make money.

Auditing grades against returns and refund data

Everything above is a hypothesis until you measure it. The audit loop is what converts a written scale into a calibrated one, and it requires only data you already have: the grade assigned, the grader who assigned it, the category, and whether the item came back.

The headline metric is return rate by grade. In a calibrated operation, return rates should be broadly similar across grades, because each grade sets an accurate expectation. A “fair” item that comes back more often than a “like new” item is not evidence that fair items are worse; it is evidence that your “fair” definition is over-promising.

The four cuts worth running monthly

Cut one: return rate by grade, which tests whether the scale is calibrated end to end. Cut two: return rate by grader, normalized for category mix, which finds individual bias. Cut three: return rate by category within grade, which finds the places your scale does not translate. Cut four: reason code distribution, because “not as described” behaves very differently from “changed mind” and only the first one indicts your grading.

Sample size discipline matters here. A grader with fifteen items and one return has not been measured, they have been sampled. Set a minimum item count before you act on a grader-level number, and prefer trailing three-month windows to single months for anyone below high volume.

Close the loop on every returned item

Every returned item should be re-graded on arrival by someone other than the original grader, with the original grade hidden. That comparison is the highest-quality calibration data you will ever generate, because it is free, continuous, and drawn from exactly the population where something went wrong.

Where the re-grade matches the original, the return was probably not a grading failure. Where it does not, you have located a specific rule that failed on a specific item type, which is actionable in a way that an aggregate return rate never is.

Know when the cheapest resolution is not a return

For low-value items, the cost of processing a physical return frequently exceeds the recoverable value of the item, particularly once re-grading and re-listing are included. Building an explicit policy for those cases is a normal part of a mature resale operation, and the trade-offs are set out in our piece on returnless refunds and when writing off a return is the cheaper option.

The important discipline is that a returnless refund still generates a data point. Log the reason code and the claimed defect even when no item comes back, or you lose visibility into exactly the segment where grading errors are cheapest to make and therefore most likely to accumulate. Sitting underneath all of this is the strategic question of whether resale earns its place in the assortment at all, which we address in the broader recommerce channel guide.

General information, not legal or compliance advice

This article describes operational practice and is general information only. It is not legal, regulatory, customs or tax advice, and it does not create any professional relationship. Authentication, product safety, consumer protection, data protection and secondhand-goods rules vary by country, by state and by product category, and they change.

Nothing here should be read as a claim that any named company or marketplace has engaged in wrongdoing. Where enforcement bodies or complainants have made allegations, those remain allegations unless and until they are established through the relevant process. Specific figures, thresholds and program terms cited in general terms above should be verified against the primary source before you rely on them.

For your own situation, particularly where intellectual property rights, product safety obligations, or the resale of regulated categories such as electronics with batteries, cosmetics or children’s products are involved, consult a qualified attorney, a licensed customs broker or the relevant regulator directly.

FAQ on grading and authentication

How many condition grades should a resale operation use?

Four or five is the practical range for most teams. Fewer than four cannot distinguish a lightly used item from a heavily used one, which suppresses price on the good stock. More than five produces definitions that graders cannot apply consistently, which is worse than a coarse scale applied correctly.

Can a small team authenticate items without a testing lab?

For low-effort counterfeits, yes. Construction quality, label consistency, hardware detail, serial verification and scripted function tests catch a substantial share of fakes with no specialist equipment. High-quality counterfeits engineered to survive visual inspection are a different problem, and the right response there is a value threshold above which items go to a third-party specialist.

What is the single highest-impact change for reducing resale returns?

A mandatory flaw close-up on every item graded below the top grade. It removes the ambiguity that buyers otherwise resolve in their own favor, and it gives you evidence in any marketplace dispute. It costs a few seconds per item at intake.

Should cosmetic condition and functional status be one grade or two?

Two fields internally, even if you display a single label to the buyer. Merging them hides the case that generates the most expensive disputes: an item that looks excellent and does not perform as expected. Keeping them separate also lets you route function failures to a different workflow.

How often should a grading scale be recalibrated?

Run the calibration set quarterly for every grader, and review the written rules whenever the audit data shows a category or grade with a persistent outlier return rate. Annual review alone is too slow, because drift accumulates across a peak season and is only visible afterwards.

Is it worth paying for third-party authentication on mid-value items?

It depends on your miss rate and the full downstream cost of a counterfeit reaching a buyer, which includes marketplace penalties and account health effects rather than just the refund. Most teams find the honest threshold is lower than their instinct, and they supplement a value threshold with category overrides for anything where they lack a reference library.

What should be photographed for every single item?

Front, back, both sides, a label or serial detail, and one close-up for every declared flaw. Items below the top grade should also carry a wide shot showing overall condition. Consistency of lighting and background across the catalog matters as much as the shot count.

How do you tell whether a return was caused by bad grading?

Re-grade every returned item blind, using someone other than the original grader, and compare. Combine that with reason code analysis, since “not as described” points at grading while “changed mind” generally does not. The blind re-grade is the most reliable signal because it isolates the grading decision from everything else in the transaction.

Do marketplace authentication programs replace an in-house process?

No. They authenticate, they do not grade, and their condition vocabulary may not match yours. Running one channel on their standard and another on yours reintroduces the inconsistency problem, so most operations keep a single internal grading standard and treat marketplace authentication as an additional verification layer on qualifying items.