← Back to Blogs
Amazon Strategy
16 min read
Which Customer Complaint Is Worth Fixing? Score It Before You Change the Product

Onieque Edwards
Content Strategist /Blog Writer

Which Customer Complaint Is Worth Fixing? Score It Before You Change the Product
Every product research framework worth running ends at the same place: a list of things customers do not like about what already exists. Reviews, return reasons, forum threads, the questions on a competitor's listing. The list is usually accurate. Gathering it is the part everybody has now been taught to do.
What the list does not do is rank itself.
That is the step where product decisions actually go wrong, and it goes wrong quietly. You read forty reviews, one of them describes a dog nearly choking on a squeaker, and that review is what you carry into the supplier call. It is vivid, it is frightening, and it may also be the one time it happened. Meanwhile the sentence appearing in a third of the negatives, the one that says the seam split in a fortnight, reads as ordinary and gets filed as a known issue in the category.
Vividness and frequency have nothing to do with each other. Neither of them tells you what the problem costs when it happens. A framework that stops at "here are the problems" hands you three unranked things and lets memory do the ranking, and memory ranks on drama.
The complaint you remember is not the complaint that costs you
There is a specific failure this produces, and it is expensive because it looks like diligence.
You find a real problem. You fix it properly. You pay for the tooling, you wait the extra four weeks, you accept the higher landed cost. Then the product launches into the same return rate it would have had anyway, because the thing you fixed was responsible for a small slice of the returns and the thing you left alone was responsible for most of them.
Nothing in that sequence feels like an error while it is happening. Every step was evidence-led. The evidence was just never weighed.
The correction is not more research. It is arithmetic on the research you already have.
Three numbers turn a complaint list into a ranking

Take every distinct complaint you have gathered and give it three numbers before you give it an opinion.
How often it is said. Count the negative reviews, meaning one to three stars, that mention the issue, and express it as a percentage of all negative reviews you read. Not a count. A percentage, because a count is meaningless without knowing how many you looked at, and because percentages are comparable across competitors with different review volumes.
How often it is returned. This is the number most people skip, and it is the one that separates a real problem from an irritation. Amazon's Voice of the Customer dashboard carries return reasons against your own ASINs, and it reports an NCX rate, its measure of negative customer experiences, alongside a CX Health rating for each listing. If you are researching a product you do not yet sell, you will not have this for the target product, but you will have it for anything adjacent you already run, and you can read a competitor's proxy off the ratio of return-flavoured language in their reviews.
What it costs when it happens. A return is not a refund. It is the refund, plus the fulfilment fee you already paid, plus the return processing, plus whatever the unit is worth after it comes back, which for a chewed dog toy is nothing. Add the advertising cost of the click that produced the sale, because you paid for that too and it produced a negative outcome. One return frequently costs more than the margin on two sales.
Now the ranking writes itself. Suppose the squeaker coming loose accounts for around a quarter of negative reviews and roughly a third of returns. Suppose "smaller than expected" accounts for a similar share of negative reviews and almost no returns at all, because people keep it and complain. Those two problems looked identical on the list. They are not remotely the same problem, and only the second number told you.
The general rule: a complaint that shows up in reviews and not in returns is a rating problem. A complaint that shows up in both is a margin problem. They are fixed with different budgets and they are worth different amounts of money.
Where each number actually lives
None of this requires a subscription tool.
Review text, counted by hand across your competitive set, gives you frequency. Fifty negatives across three competitors is enough to see a pattern and few enough to read in an afternoon.
Voice of the Customer, in Seller Central, gives you return reasons, NCX rate and CX Health on listings you already own.
The Business Report gives you unit session percentage, which is your conversion rate, and it is the number that tells you whether a problem is stopping the sale or only souring it afterwards.
Product Opportunity Explorer surfaces search terms and the products people click for them, which is where you find out whether the complaint is attached to demand at all.
The third-party tools are useful for volume estimates. They are not where this evidence is.
A dealbreaker and an annoyance are two different budgets
Once complaints are scored, sort them into two piles, because the two piles are funded differently.
A dealbreaker produces an immediate return or a one star review on its own. Anything involving safety belongs here, as does anything that makes the product not do its job. A choking hazard, a sharp edge, a part that fails on first use. These have a second property that makes them worse than their frequency suggests: a single one can end a listing through a safety complaint, and the review it produces is the review shoppers read first.
An annoyance lowers the rating without ending the relationship. It smells odd out of the packet. It is awkward to clean. The colour is not the colour in the photograph. These accumulate into a four point one instead of a four point five, which costs conversion at every single session, forever, quietly.
Both matter. They do not compete for the same money. A dealbreaker is a specification change and it belongs in the first production run. An annoyance is often a listing change, or a packaging insert, or a photograph, and it can wait for the second. Treating them as one ranked list is how a cosmetic fix gets tooled while a safety issue goes into version two.
Half the evidence is written before anyone buys
Reviews and returns have one thing in common that limits them: they are written by people who bought. Every complaint in them survived the purchase decision.
The customer who read the listing, worried the squeaker would come out, and bought a different product wrote nothing on your competitor's listing. They wrote it somewhere else, before buying, in the form of a question. And that is a different category of evidence, because it does not describe a return. It describes a sale that never happened.
Two places carry it:
The Q&A on competing listings. "Is this washable?" "Can the squeaker be removed?" "Will this survive a Rottweiler?" A question asked repeatedly is a doubt the listing failed to answer, and it is being asked by people still deciding.
Pre-purchase threads off Amazon. Breed groups, hobby forums, the "what should I buy for X" posts. These are people describing their fear in their own words, before anyone has sold them anything.
The distinction is worth being precise about, because it changes what you do with it. Post-purchase evidence tells you what to build. Pre-purchase evidence tells you what to say. A fear that appears constantly before the purchase and rarely after it is not a product defect at all. It is a conversion blocker, and the fix is a photograph, a bullet, or a line of A+ content, not a tooling change.
Both belong in the framework. Most frameworks only collect the first one, and then wonder why a genuinely better product converts at the category average.
Cluster by job, risk and context, never by the words
The other thing that goes wrong at this stage is grouping. People group complaints by the words in them, so "chews through it" and "destroyed it in a day" and "did not last" become three entries when they are one problem, and "too small" becomes one entry when it is two problems wearing the same phrase.
Group on three axes instead.
The job. What was this bought to accomplish. Keep a power chewer occupied for half an hour.
The risk. What the buyer is afraid of. A vet bill.
The context. The conditions it has to work in. A seventy to a hundred pound dog, unsupervised, indoors.
"Too small" said by someone with a Chihuahua and "too small" said by someone with a Mastiff are not the same complaint, and they resolve in opposite directions. The words matched. Nothing else did.
Once clusters are built this way, score the cluster rather than the phrase. A cluster with real search volume behind it, a high share of the negative reviews, and a high share of the returns is the one to build against. A cluster that is loud in reviews and absent from search is a real irritation that nobody is shopping for, and it will not carry a launch on its own.
The gap gate belongs before the economics, not after
Most frameworks find the gap, get excited, and then run the cost model. The gate order should be the other way around, because two of the three ways a gap dies cost nothing to check and take about twenty minutes.
Can you legally sell it. Check whether the category or the specific product type is gated and what approval requires. Check the mechanism you intend to build against the patent record, particularly design patents, which cover shape and appearance and catch far more private label sellers than utility patents do. Check that the claim you plan to make on the listing is not somebody's registered trademark. A gap that somebody else owns is not a gap you have found. It is a lawsuit you have scheduled.
Is anybody asking for it. A gap needs a mention rate, not a hunch. Count how many questions on competing listings ask for the missing feature and how many competing listings claim it. If a meaningful share of the questions ask whether it can be washed and none of the listings answer, that is a validated gap. If nobody asks and nobody offers, the honest reading is that nobody wants it, and the absence you spotted is a market that already decided.
Will anybody pay for it. This is the check that gets deferred to the cost model and should not be. Find the listings that already have the feature. If the only products offering it sit forty percent above the price band and sell in low volume with weak ratings, the feature is real and unmonetisable at your price. That is not a reason to build it cheaper. It is evidence that the willingness to pay was tested by somebody else and came back negative.
A gap that clears all three is worth costing. A gap that fails any of them was never worth the spreadsheet.
A differentiator carries four numbers, not one
The usual test for a proposed improvement is what it adds to landed cost. That is one input and it is the least interesting one, because cost is the only quantity in the calculation that is certain to be bad news.
Score each proposed change on four:
What it adds to landed cost. Per unit, at the quantity you will actually order, including the tooling amortised across that first order rather than across an optimistic lifetime volume.
What it does to conversion. Read your own unit session percentage as the baseline rather than a published category average, because category averages are computed across listings with different prices, review counts and ad structures, and yours is the only one you have to beat. Then ask what removing this specific doubt plausibly does to it. A fix that answers a question shoppers are visibly asking on competitor listings has an argument behind it. A fix nobody has ever mentioned has none.
What it does to returns. This is where a durability or safety fix earns its cost, and it is calculated in money rather than percentage points. A reduction in return rate multiplied by the true cost of a return, which you worked out earlier, is a real recurring saving.
What it does to review velocity. This one is routinely omitted and it has become more important, not less. Safety and durability fixes do not only reduce returns. They change the shape of the review curve, because a product that survives produces the reviews that mention surviving, and those compound. This matters more from 2026 onward, because Amazon has moved to stop reviews pooling across variations. The old route of adding a variation to an existing parent and inheriting its review equity is closing, which means a new product now earns its ratings on its own, and anything that speeds that up is worth more than it used to be.
Then attach the long tail terms each change earns you, because a differentiator that nobody searches for cannot be advertised against efficiently. A sealed squeaker is a product change and a keyword at the same time. Head terms are contested by everyone in the category and cost accordingly. The phrases that describe your specific fix are cheaper, convert better, and are the reason the fix pays for itself in advertising as well as in returns.
Only approve a change that clears a combined threshold. A change that adds cost and delivers on exactly one of the other three is usually a preference, not a differentiator.
One landed cost is a guess. Three is a decision
A single landed cost figure is a forecast pretending to be a fact. Run three.
Best case. The freight quote holds, defect rate is low, duty is what it is today.
Base case. Realistic freight, the defect and return rate you actually observed in the research, current duty.
Worst case. Freight up meaningfully, a duty change, a return rate two or three points above what you expected.
Then apply the decision rule: the product has to clear your threshold in the base case and survive the worst case. Not "look good in the best case". Products are approved on best cases and then operated in worst ones.
Three costs are also routinely left out of the model entirely.
Storage, including aged inventory. A slow mover pays monthly storage the whole time, and once units have been sitting in the network past roughly six months they pick up an aged inventory surcharge on top. Check the current fee schedule rather than a figure from a blog, this one included, because these change. A product with thin margins and a slow sell through can be profitable per unit and unprofitable per month.
Inbound problems. Mislabelling, packaging that fails, incorrect prep. These arrive as unexpected per unit charges on a shipment you have already paid for, and they are most likely on a first order with a new supplier, which is exactly the order you modelled most optimistically.
Size tier boundaries. Check where your dimensions and weight sit relative to the tier boundary before the packaging is finalised. A fraction of an inch or a few ounces over a boundary changes the fulfilment fee on every unit you will ever sell, and it is usually correctable at the packaging stage for nothing.
The other thing the static model misses is time. Advertising cost is not one number either. A launch runs at a high total advertising cost of sales while you are buying position without the ratings to convert, and it settles as rank and reviews arrive. So the question the model has to answer is not whether the margin works at maturity. It is whether the business can fund several months at launch level advertising while that happens. A product with acceptable mature economics and no capacity to absorb the ramp is not a good product. It is a cash flow problem you have not met yet.
Time to rank is the gate nobody writes down

Most GO or NO GO checklists carry four gates: demand, competition, profitability, risk. All four can pass on a product you should not build.
Two additions close that.
Weight the risk instead of ticking it. Risk is treated as a yes or no, and it is not. Policy exposure, meaning gated categories, brand restrictions and intellectual property, is one kind. Volatility is another, meaning price wars, a constant stream of new entrants, review manipulation in the category. Operational risk is a third, meaning fragile items, complex prep, categories where returns are structurally high. Score them and set a floor. A product can be legal, profitable and still sit in a category where the price collapses every quarter.
Add time to rank as a fifth gate. Ask how many months it takes to reach the top of page one for the terms that matter, what rating and review count the incumbents hold, and whether you can fund the advertising for that whole period. Then compare that to when the business needs the money back. If it takes a year to become competitive and the cash is needed in six months, it is a no, and it is a no even when every other gate is green. That is a strategic answer, not a mathematical one, and it is the gate that catches the products which look fine on every spreadsheet and quietly consume a year.
All five have to clear. One below stops it, and a cheap no is worth more than an expensive yes.
The framework has to keep running after launch
The last change is to stop treating any of this as a pre launch document.
Everything above gets better after you launch, because you stop reading proxies and start reading your own data. Your return reasons, your NCX rate, your unit session percentage, your own review text. Re-score the complaint list monthly against that. The ranking will move, and the second production run should be built against the new ranking rather than the assumptions in the original brief.
Set the exit conditions in advance, while you are still capable of being objective about them. Decide before launch what conversion rate after a given number of sessions means the listing is wrong, what return rate after the first hundred units means the product is wrong, and what advertising cost of sales after ninety days means the economics do not work. Write them down. The purpose of writing them down beforehand is that afterwards you will have inventory, sunk cost and a story about why this month was unusual, and none of those improve the decision.
What this changes about where the money goes
The upgrade here is not more research. It is that every problem, every gap and every proposed improvement now carries a number, and numbers can be ranked while opinions cannot.
That changes one thing at the point of decision. Instead of approving improvements because the evidence supporting them was memorable, you are allocating a fixed tooling budget and a fixed advertising runway across a ranked list, and you can see what each position costs to hold and what the next one down would have to be worth to displace it. The complaint at the top of that list gets the money. The one at the bottom gets a sentence in the bullets. Some products do not get built at all, and finding that out during a week of research rather than after a container has shipped is the highest return outcome this process produces, even though nobody writes a case study about it.
FAQ
How many reviews do I need to read before the percentages mean anything?
Around fifty negative reviews spread across three or four competitors is enough to see which complaints repeat. You are not measuring a population, you are separating things said once from things said constantly, and that separation appears early. Read the one to three star reviews specifically. Five star reviews tell you what the listing promised, not what went wrong.
I do not sell the product yet, so I have no return data. What do I use?
Return reasons for anything adjacent you already sell, and the return flavoured language in competitor reviews. Phrases like "sent it back" and "returned it" appearing in a complaint are your proxy, and the share of complaints carrying them is comparable across competitors even though it is not a real return rate. Treat it as a ranking device, not a measurement.
Does Voice of the Customer need Brand Registry?
Voice of the Customer covers listings on your own seller account. Product Opportunity Explorer and Search Query Performance have their own access conditions, and Search Query Performance sits inside Brand Analytics. Check the Growth and Brands menus in the account you will actually be using rather than trusting a blog on it, this one included, because access conditions change.
What if the scoring says my idea is not worth building?
Then it did its job at the cost of about a week, and the alternative was learning the same thing several months later with the inventory paid for. The cheap no is the entire point of running gates in the first place.
Is this only useful before launch?
No, and it is arguably better afterwards. Once you are live you have your own return reasons, your own conversion rate and your own reviews, which is stronger evidence than any pre launch research produces. It tells you which complaint to fix in the second production run, which variation to add, and which improvement to stop paying for.
SOURCES

Onieque Edwards
Content Strategist /Blog Writer
Onieque is the brain behind bold Amazon growth strategies and structured business execution. He enjoys turning scattered ideas into clear, actionable systems that actually drive results. When he’s not building out growth plans or refining campaigns, you’ll likely find him exploring new coffee spots or getting lost in ideas that connect strategy with creativity.
