Buyers judge a marketplace by what comes back when they type. We rate the search experience the way a buyer would, trace poor results to their root cause and propose the fix to the ranking or the model. Agents read every query at volume. Named search analysts own the rating and the recommendation. Every change to a ranking is approved by the marketplace and kept on a record.




Search quality is a loop, not a project. Each pass starts with a buyer’s rating and ends with proof that the change worked. It is one of the practices in our marketplace operations group and it runs on the same record as the rest.
Analysts rate the search experience for each query on a five-point scale: poor, bad, average, good or excellent. The rating is a buyer’s judgement, recorded with the query, the page and the date.
For queries that rate below average, the team marks what is wrong on the results page: irrelevant products, missing attributes, duplicate listings, thin results or a wrong category.
Tests are run to discover why the results drew fewer impressions or clicks than they should have. A broken synonym, a missing attribute, a stale ranking signal or a model that misread the intent.
Recommendations go to the search and machine learning teams as specific proposals: an attribute to backfill, a synonym to add, a ranking weight to revisit, a training example to correct.
After the change ships, the same queries are rated again on the same scale. The before and the after sit side by side on the record, so the efficiency of the change is shown rather than claimed.
A small set of head queries carries most of the traffic and a long torso carries most of the variety. Agents cluster and classify both by intent, so analysts spend their time on the queries that matter rather than on sorting them. Query automation is available as a pilot.
| Query | Intent | Top result | Rating | Action | |
|---|---|---|---|---|---|
![]() | "laptop 14 inch" | spec | Laptop | excellent | none |
![]() | "reading glasses" | vision | Aviators | poor | held |
![]() | "trail shoes men" | buy | Sneaker | good | fixed |
![]() | "black watch" | browse | Watch | good | none |
![]() | "coffee mug set" | bundle | Mug | average | synonym |
Relevance is not a score in a log. It is what a person sees on page one, on page two and in the strip of things bought together. We rate all of it against agreed criteria and keep each rating with the evidence behind it.
Each query’s results are scored against a written rubric: does the top result match the intent, are the attributes right, is the price band sensible, is anything missing that a buyer would expect.
Analysts browse several pages of output for each query, because the tail of a results page tells you how the ranking degrades and where the duplicates and the mis-tagged products hide.
Recommendations on product and cart pages are rated on the same scale as search. A relevant result followed by an odd recommendation costs the sale just as surely.
Bought-together and similar-item strips are checked for sense, for category fit and for the accidental pairings that a model learns from a few noisy orders.
How competing marketplaces structure their categories and filters for the same products, documented as a comparison your category team can act on.
The same queries run on competing platforms, rated on the same scale, so you know where your results are ahead, where they are behind and by how much in the buyer’s eyes.
A ranking change touches every buyer at once, so it is never made by an agent and never made quietly. Search quality runs on Spine, the same governed platform behind our regulated work since 1955.
Every fix arrives as a proposal: the queries affected, the rating before, the root cause found, the tests run and the change recommended. The analyst who owns it is named on the proposal.
The search or category owner on your side approves, amends or declines. The decision, the person and the time are written to a chained audit record, and the re-run that follows is filed against the same entry.





Explore how AI + creativity can accelerate your next big move. Bring your top queries. Leave with a rating for each, the root causes behind the poor ones and a set of fixes ready for your search owner to approve.