Ask a retail team why a product is not selling and the answers arrive fast: the price is wrong, the ads are underfunded, the category is crowded. Ask the same question of the product record itself and a different picture usually appears. The item has no GTIN, so Google will not show it in free listings. The color attribute says “navy” on the site, “Navy Blue” in the feed and “0043” in the ERP, so no filter matches it. The title leads with an internal model code that no shopper has ever typed.
Product data is the layer almost every other commercial decision sits on, and it is the layer most often left to whoever had time. This guide covers what belongs in a product record, how identifiers and attributes actually get used downstream, when a spreadsheet stops being enough, and how to measure catalog quality without commissioning a six month audit.
In short
- Product data is infrastructure, not copy. Attributes, identifiers and taxonomy decide what is eligible to appear before any creative or bid decision is made.
- Identifiers are not interchangeable. A GTIN is issued by GS1 and describes a product to the world, an MPN comes from the manufacturer, and a SKU is private to one seller. Swapping them causes suppression.
- Filters run on structured attributes. Facts buried in description prose are invisible to faceted navigation, marketplace filters and comparison engines.
- A PIM earns its keep at the intersection of SKU count, channel count and change rate, not at a single headline catalog size.
- Governance beats cleanup. Naming a single owner per field prevents the drift that makes every audit a repeat purchase.
What product data actually covers, beyond the description
Most teams picture product data as the thing a copywriter produces: a title, a paragraph or two, some bullet points. That is one slice of it, and usually the smallest. A working product record spans five distinct layers, and each one is consumed by different systems with different tolerances for error.
The five layers of a product record
The identity layer answers “which product is this” in terms the outside world recognizes: GTIN, manufacturer part number, brand, and the seller’s own SKU. The classification layer places the item in a hierarchy, both your own taxonomy and the external ones that Google, Amazon and comparison engines impose.
The attribute layer carries structured facts: dimensions, material, capacity, compatibility, certifications, color, size. The media layer holds images, video, 360 spins and documents, each with its own technical rules per channel. The commercial layer carries price, availability, shipping class, tax treatment and channel eligibility.
Marketing copy sits on top of all five, and it is the only layer a human reader consciously notices. Everything underneath decides whether that reader ever sees the product at all.
Why the layers get confused
The confusion is organizational. Copy belongs to marketing, price belongs to merchandising, stock belongs to operations, and identifiers often belong to nobody in particular because they arrived in a supplier spreadsheet three years ago. Nothing in a typical e-commerce platform forces those groups to agree, so the record drifts apart one field at a time.
Platform choice shapes how painful that drift becomes, which is why catalog structure deserves weight in any migration decision. Our guide to how to choose the right e-commerce platform for your store treats attribute modeling and import tooling as first-order criteria rather than nice-to-haves, because retrofitting a data model onto a live catalog is far more expensive than picking the right one at the start.
Identifiers: GTIN, MPN, SKU and why they are not interchangeable
Identifier confusion is the single most common cause of silent product suppression. The three codes look similar in a spreadsheet column and behave completely differently once a feed leaves your server.
What each code actually is
A GTIN (Global Trade Item Number) is a globally unique number administered by GS1, the standards body that issues company prefixes to brand owners. UPC, EAN, JAN and ISBN are all formats of GTIN that differ mainly in digit length. According to GS1, the number identifies a trade item consistently across every company that handles it, which is exactly why search engines and marketplaces treat it as a matching key.
An MPN (Manufacturer Part Number) is assigned by the maker and is unique only within that maker’s range. It is the fallback identifier for products that genuinely have no GTIN, such as custom or made-to-order goods, and it is only meaningful when paired with the brand name.
A SKU (Stock Keeping Unit) is yours alone. It exists to let your warehouse, your finance system and your staff refer to an item unambiguously. No external system can interpret it, and sending a SKU where a GTIN is expected is a reliable way to get a listing disapproved.
How the three compare in practice
| Property | GTIN | MPN | SKU |
|---|---|---|---|
| Who issues it | GS1 via a licensed company prefix | The manufacturer | The seller |
| Globally unique | Yes | Only with brand | No |
| Used for product matching | Yes, primary key | Yes, secondary | No |
| Changes when you switch supplier | No | Usually yes | Your choice |
| Appears on packaging | Usually, as a barcode | Often | Rarely |
| Typical failure mode | Reused or borrowed from another product | Submitted without brand | Submitted in the GTIN field |
The borrowed GTIN problem
Resellers and dropshippers frequently inherit identifier problems rather than create them. A supplier supplies a code, the code is loaded without validation, and it turns out to belong to a different variant or, worse, to an unrelated product from another brand. The listing then matches the wrong catalog entry, inherits the wrong reviews and the wrong category, and performs strangely in ways no bid adjustment can fix.
Checksum validation catches a meaningful share of these before they ship. Every GTIN carries a check digit computed from the preceding digits, so a simple modulus calculation run at import will reject transposed or truncated codes. It will not catch a valid code attached to the wrong product, which is why a periodic sample check against the brand’s own published catalog remains worthwhile.
When a product legitimately has no GTIN
Handmade goods, bespoke furniture, made-to-measure clothing, bundles you assemble yourself and own-brand products you have not yet licensed a prefix for all fall outside the GTIN system. The correct response is to use the channel’s documented exemption path rather than invent a number. Google Merchant Center, Amazon and most marketplaces each publish exemption rules by category, and those rules change, so verify the current requirement at the channel’s own documentation before building an import rule around it.
Attributes and taxonomy: the fields buyers filter on
If identifiers decide eligibility, attributes decide discoverability. A shopper narrowing a list of 4,000 running shoes to the eight that are waterproof, size 11 and under $150 is running three attribute queries. If any of those three facts lives only in a description paragraph, the product silently drops out of consideration.
Structured beats prose, always
The test is simple: could a machine sort or filter on this fact without reading a sentence? “Water resistant to 50 meters” inside a paragraph is prose. A field named water_resistance_m with the value 50 is an attribute. Only the second one can power a facet, feed a comparison table or answer a structured query.
The practical consequence is that description writing should come last, not first. Capture the facts as fields, then let the copy reference and contextualize them. Teams that work the other way around spend the following year extracting data from their own sentences.
Required, recommended and optional
Not every attribute deserves the same enforcement. A workable model sorts fields into three tiers. Required fields block publication if empty: identity, price, primary image, primary category. Recommended fields trigger a warning and feed a completeness score: material, dimensions, country of origin, compatibility. Optional fields are category-specific enrichment that only some products carry.
The tiering matters because a flat “fill everything” mandate produces junk. Staff under deadline pressure will enter “N/A”, “various” or a copied value from the row above, and those placeholder strings are harder to find later than an honest empty cell.
Controlled vocabularies and the color problem
Color is the canonical example of attribute drift. Left free-text, a 5,000 SKU catalog will accumulate “navy”, “Navy”, “navy blue”, “dark blue”, “midnight” and “NVY” as distinct values, each splitting the same facet into fragments too small to display. The fix is a controlled vocabulary: a fixed list of permitted values, with a mapping table that translates supplier variants into the canonical term on import.
The same logic applies to size, material, gender, age group and any field a shopper might filter. Free text is appropriate for descriptions and almost nothing else. Large catalogs feel this most acutely, and the operational strain shows up as slow admin screens and unusable filters long before anyone calls it a data problem, a pattern we examine in detail in WooCommerce with 50,000 products: search, indexing and admin performance.
Your taxonomy versus theirs
Every channel imposes its own category tree, and none of them will match yours. Google maintains a product taxonomy of several thousand nodes, Amazon has browse nodes per marketplace, and each comparison engine has a variant. Trying to restructure your internal taxonomy to match one of them is a trap, because the next channel will want something different.
The durable approach is a mapping layer: keep an internal taxonomy designed for your own merchandising and navigation, then maintain a lookup that translates each internal category into the external equivalent per channel. Mappings are cheap to maintain and easy to version. Restructuring a live catalog is neither.
Titles and naming conventions that survive every channel
A product title does more work than any other field. It is the clickable line in search results, the matching signal for comparison engines, the label in a cart and increasingly the string an AI assistant quotes back to a shopper. It also has to fit inside character limits that differ by channel.
The pattern that travels
Titles that work across channels follow a recognizable order: brand, then product line or model, then the defining attribute, then the variant specifics. “Brand Model Type, Color, Size” reads naturally, front-loads the terms a shopper actually searches, and degrades gracefully when a channel truncates it at 70 characters because the least important element is at the end.
What breaks is internal shorthand. Leading with a warehouse code, an abbreviation only your buyers understand or a promotional phrase wastes the characters that carry search intent. Promotional language in particular is both ineffective and, on most marketplaces, against listing policy.
Channel limits and tolerances
| Element | Own site | Shopping feeds | Large marketplaces |
|---|---|---|---|
| Practical title length | Flexible, SEO-led | Front-load first 70 characters | Category-specific caps, often strict |
| Brand placement | Optional | Strongly preferred first | Usually mandatory first |
| Promotional words | Discouraged | Commonly disallowed | Commonly disallowed |
| Variant detail in title | Often handled by selector | Required for each variant row | Required, format varies |
| Consequence of breach | Weak ranking | Disapproval or reduced serving | Suppression or listing removal |
Because the strictest channel sets the ceiling, the efficient move is to generate titles from fields rather than write them by hand. A template that concatenates brand, model, type and variant attributes produces consistent output at any catalog size and can be re-rendered per channel with different length rules. Hand-written titles are defensible for a few hundred hero products and indefensible for 40,000.
Variants and the duplicate trap
Variant modeling is where naming conventions and data structure collide. Treating every color of one shirt as a separate parent product fragments reviews, splits ranking signals and multiplies maintenance. Collapsing genuinely different products into one parent confuses shoppers and breaks inventory. The dividing line is whether a buyer would consider the items substitutes for the same need, and whether each variant has its own GTIN, which it usually should.
Images, video and the media rules marketplaces enforce
Media is the one part of product data where the rules are explicit, published and mechanically enforced, which makes failures unusually easy to prevent and unusually embarrassing to leave in place.
The rules that repeat across channels
Most large channels converge on a similar set of expectations for a primary image: the product on a plain white or neutral background, filling most of the frame, no watermarks, no promotional overlays, no additional props and no text burned into the pixels. Minimum resolutions commonly sit in the 1,000 pixel range on the longest side so that zoom works, and exact thresholds vary by channel and category.
Secondary images have far more freedom and carry much of the conversion weight: scale references, in-use context, detail shots, packaging and size charts. The common failure is to have six variations of the same front-on studio shot and nothing that answers “how big is it” or “what does it look like in a real room”.
Alt text, file naming and the accessibility dividend
Alt text exists for accessibility first, and treating it that way produces better results than treating it as a keyword slot. A description of what the image shows, in plain language, serves screen reader users and happens to give search engines a clean signal. Stuffed alt text serves neither.
File naming follows the same logic. A descriptive, hyphenated filename derived from brand, model and shot type is both human-navigable and machine-readable. Camera-default filenames are a small, permanent tax on every future asset search.
Video and the cost question
Video is the highest-cost media asset and the hardest to justify at full catalog scale. The sensible pattern is tiering by revenue contribution: the top decile of products by revenue gets bespoke video, the next tier gets a short template-generated spin, and the long tail gets strong stills. Producing video for 20,000 SKUs that each sell twice a year is a rounding error in revenue and a large number in production invoices.
Where a PIM starts paying for itself versus a spreadsheet
Product information management systems get sold on SKU count, which is the least useful variable in the decision. Plenty of 50,000 SKU catalogs run acceptably on a disciplined spreadsheet and a good import script. Plenty of 800 SKU catalogs are a daily emergency.
The three variables that actually decide it
The first is channel count. One destination means one format and one set of rules. Six destinations means six transformations, six validation regimes and six places a change must land, which is where manual processes break.
The second is change rate. A catalog that updates seasonally is a different problem from one where suppliers push revised specifications weekly. The third is contributor count. Two people editing a shared sheet is workable. Eleven people across three time zones and two agencies is a merge conflict with a revenue consequence.
Comparing the realistic options
| Consideration | Spreadsheet plus import script | Platform-native fields | Dedicated PIM |
|---|---|---|---|
| Setup effort | Low | Low to moderate | High |
| Multi-channel output | Manual per channel | Limited, often via apps | Core function |
| Validation and required fields | Fragile, script-dependent | Basic | Rule-based and enforced |
| Audit trail and rollback | Version history at best | Rarely field-level | Field-level with attribution |
| Workflow and approvals | None | None | Built in |
| Localization | Painful | Varies widely | Designed for it |
| Best fit | Single channel, slow change | One or two channels, stable catalog | Three or more channels, frequent change, many editors |
The intermediate step most teams skip
Between a spreadsheet and a six-figure PIM sits an option that is frequently overlooked: a properly modeled attribute system inside the existing commerce platform, plus a single transformation layer that generates every outbound feed from one source. It will not give you approval workflows or field-level audit trails, but it solves the duplicated-truth problem, which is the one that causes most of the damage.
Buying a PIM before fixing the data model simply relocates the mess into a more expensive system. The sequence that works is to define the model, clean against it, prove the single source of truth, and only then evaluate whether tooling is the constraint.
Feed governance: who owns a field and who may change it
Data quality decays. Not because people are careless but because a catalog is touched by buyers, merchandisers, marketers, suppliers, translators and integration scripts, and in the absence of explicit ownership each of them will make a locally reasonable change that is globally wrong.
One owner per field, written down
Governance in this context means something modest and concrete: a table listing every field, the single team accountable for its accuracy, who may edit it, and where its authoritative value originates. Price is owned by merchandising and sourced from the ERP. GTIN is owned by the catalog team and sourced from the brand. Description is owned by content and sourced from the CMS.
The value of writing it down is not bureaucratic. It converts an argument about blame into a lookup, and it makes integration design obvious: a system that is not the authoritative source for a field should never be allowed to write to it.
Inbound supplier data needs a gate
Supplier feeds are the largest single source of contamination in most catalogs. They arrive in inconsistent formats, use the supplier’s vocabulary rather than yours, and periodically change shape without notice. Accepting them directly into the live catalog is how controlled vocabularies die.
A staging gate fixes most of it. Inbound data lands in a holding area, gets validated against required fields and permitted values, gets mapped into the canonical vocabulary, and only then promotes to live. Rejected rows go back to the supplier with a reason rather than quietly into production.
Outbound feeds should be generated, never maintained
Every outbound feed should be a rendered view of the canonical record, produced by transformation rules that live in version control. The moment someone maintains a hand-edited CSV for one channel, that channel begins diverging and will keep diverging until the next audit. The specific transformations that matter for paid channels, including title construction and attribute mapping, are covered in product feed optimization: titles, attributes and the errors that kill Shopping ads.
How search engines and AI shopping agents read a catalog
The consumers of product data have broadened. For two decades the audience was a crawler and a shopper. Now a growing share of product discovery passes through assistants that read structured data, compare options and present a shortlist without the shopper visiting a category page at all.
Structured data is the common denominator
Schema.org Product markup on the page, a clean merchant feed, and consistent identifiers are the three channels through which machines read a catalog. They overlap heavily, and inconsistency between them is itself a quality signal. A price in the markup that disagrees with the price in the feed and the price on the page is a well-documented route to a disapproval.
The fields that matter most to agent-style retrieval tend to be the specific, comparable ones: exact dimensions, materials, compatibility, warranty terms, return window and shipping timing. These are precisely the fields that a human-focused catalog treats as optional, which is why the field set that decides inclusion is worth examining directly, as we do in product feeds for AI shopping agents: the fields that decide inclusion.
Catalog matching on marketplaces
On catalog-based marketplaces the stakes are structural rather than cosmetic. Listings are matched to a shared product record, and the quality and completeness of your submission influences whether you attach to the right record and how you place within it. Mercado Libre is a clear example of a market where catalog attachment governs visibility, which we unpack in Mercado Libre catalog listings: winning the buy box in Mexico and Brazil.
The general principle holds across marketplaces: where a shared catalog exists, your product data is not describing your listing, it is competing to define the product itself.
Measuring data quality without a six month audit
Catalog audits have a reputation for producing enormous documents and no change. The alternative is a small, recurring set of measurements that a team can run weekly and act on the same day.
Five metrics that are cheap to compute
Completeness is the percentage of products with all required and recommended fields populated, ideally weighted by revenue so that the top sellers dominate the score. Validity is the percentage of values that pass format checks: GTIN checksums, numeric fields containing numbers, enumerated fields containing permitted values.
Consistency compares the same field across systems and counts disagreements. Channel acceptance is the disapproval and suppression rate reported by each destination, which is the only metric the channels compute for you. Freshness measures how long since each record was last reviewed, which surfaces the quietly stale long tail.
Weighting by revenue changes the priority list
| Metric | Unweighted view | Revenue-weighted view | Why the gap matters |
|---|---|---|---|
| Completeness | Counts every SKU equally | Counts the SKUs that earn | A gap on a top seller outweighs 200 long-tail gaps |
| Validity | Flags all format errors | Flags errors blocking live revenue | Sequences the fix queue sensibly |
| Freshness | Rewards bulk touch dates | Targets review where it pays | Prevents mass no-op updates gaming the metric |
The unweighted view is useful for tracking long-term direction. The revenue-weighted view is what should drive this week’s work. Reporting only the first is how a team spends a quarter raising a completeness score without moving a single commercial number.
Sampling beats exhaustive review
Automated checks catch format and completeness problems. They cannot catch a well-formatted lie: a correct-looking GTIN on the wrong product, an accurate-looking dimension copied from the previous row, a description for a different model. For that, a weekly manual sample of 20 to 30 records, pulled at random and checked against the source, estimates the true error rate well enough to act on and costs an hour.
Context for where the effort is worth it comes from the scale of the channel itself. The US Census Bureau publishes quarterly e-commerce retail sales estimates, and the share of total retail those figures represent is a reasonable sanity check on how much of a given catalog’s future revenue depends on machine-readable data rather than a salesperson.
A sequence for fixing a neglected catalog
Catalog remediation fails when it is attempted as one project. It works when it is sequenced so that each stage produces a usable result before the next begins.
Start by defining the model: the field list, the three tiers, the controlled vocabularies and the owner per field. This is a document, not a system, and it can be done in a week. Without it, every subsequent cleanup is cleaning toward an undefined target.
Then fix validity on the revenue-weighted top decile. Checksum failures, missing identifiers and disapproved listings in the products that actually earn. This is where the commercial return is concentrated and where the case for further work gets made.
Next, consolidate the source of truth. Pick the authoritative system per field, switch off the competing write paths, and generate every outbound feed from that source. The platform-specific mechanics vary, and opencart large catalog management is a good worked example of how attributes, filters and imports interact on one stack. Then install the staging gate on inbound supplier data so the newly cleaned catalog stops being recontaminated.
Only after that does it make sense to work the long tail, and to evaluate whether tooling is now the binding constraint. A PIM bought at this point is a considered purchase. A PIM bought at the start is a container for the same problems.
Frequently asked questions about product data management
Do I need a GTIN for every product I sell?
Not for every product, but for most branded goods the major channels expect one. Products that genuinely fall outside the system, including handmade items, bespoke goods, bundles you assemble and own-brand products without a licensed GS1 prefix, are handled through each channel’s documented exemption process. Exemption rules differ by channel and category and change over time, so check the current requirement in the channel’s own documentation rather than relying on a rule written into an import script two years ago.
Can I buy a single barcode from a reseller instead of licensing from GS1?
Resold barcodes exist and are widely advertised, but they typically originate from prefixes issued before GS1 changed its licensing terms, and the number remains registered to the original company rather than to you. GS1 states that its company prefixes are licensed rather than sold and should be obtained from a GS1 member organization. Some marketplaces verify that the GS1 prefix matches the stated brand owner, and a mismatch can lead to listing removal. The current rules and fees should be confirmed with your national GS1 office, since they vary by country and are revised periodically.
How many attributes should a product have?
Enough to answer the questions a buyer in that category actually asks, which varies enormously. A plain T-shirt may need eight fields. An industrial pump may need sixty. A practical method is to look at the filters your own category pages offer, the filters the leading marketplace offers in that category, and the questions arriving in customer service, then make the union of those three the required and recommended set.
Should variants be separate products or options on one parent?
Group them under one parent when a shopper would consider the items interchangeable solutions to the same need, such as sizes and colors of one garment. Keep them separate when the buying decision is genuinely different, such as capacity tiers of an appliance that serve different households. Each sellable variant should still carry its own GTIN and its own stock record regardless of how it is grouped for display.
How often should product data be reviewed?
Tier the cadence rather than setting one interval. Automated validity checks can run on every import. Revenue-weighted completeness is a weekly report. The top decile of products by revenue deserves a quarterly manual review. The long tail can run on an annual cycle, triggered earlier if a channel flags a disapproval or a supplier changes specification.
What does a PIM actually cost?
Published pricing is scarce and varies by vendor, SKU volume, channel count and user seats, so any single figure quoted here would mislead. The more useful point is that licensing is usually the smaller half of the total. Data modeling, migration, integration work and the internal time to define vocabularies and ownership typically exceed the software cost in year one. Ask vendors for a total first-year figure that includes implementation, and compare it against the cost of the manual work it replaces.
Who should own product data in the organization?
Ownership works best when it is split by field with a single coordinating role, rather than assigned wholesale to one department. Merchandising typically owns commercial fields, content owns descriptive fields, operations owns logistics attributes, and a catalog or e-commerce operations role owns identity, taxonomy and the governance table itself. The coordinating role matters more than its job title: someone has to be accountable for the model as a whole.
Does better product data actually improve rankings?
It improves eligibility first, which is the larger effect. Complete, valid data is what qualifies a product for free listings, shopping surfaces, marketplace catalog attachment and agent retrieval. Within those surfaces, richer attributes support more filter matches and more specific queries. The honest framing is that product data rarely wins a competitive head-to-head on its own, but weak data reliably removes a product from contention before the competition starts.
Is it worth cleaning data for products that barely sell?
Usually not as a standalone project, and usually yes as a byproduct. Long-tail cleanup rarely clears its own cost in direct revenue. What does pay is making the model and the import gate good enough that new long-tail products arrive clean, so the backlog stops growing. Retrofitting the existing tail then becomes an opportunistic task rather than a funded initiative.