Back to tools

Generator

Shopify Search Relevance Testing: Build 25 Queries

Turn five representative catalog products into a repeatable 25-query regression set. Test exact terms, partials, synonyms, attributes, and shopper-style requests before changing search controls.

Hyper Team
7 min read
Shopify Search Relevance Testing: Build 25 Queries

Key takeaways

  • A useful search test set starts with five representative catalog products and expands each into exact, partial, synonym, attribute, and natural-language queries.
  • Repeatable tests use the same storefront conditions, result depth, scoring rules, and expected products whenever a theme, app, catalog, or merchandising rule changes.
  • Zero results for an in-stock, published product are an automatic failure; poor ordering is a relevance problem; a missing filter or facet is a separate discovery problem.
  • A 25-query set is large enough to expose recurring weaknesses without turning every release check into a long manual audit.

Shopify search relevance testing should answer a practical question: can shoppers using catalog terms and ordinary customer language reach the right products quickly? This generator turns the merchant’s own product types, attributes, synonyms, and shopper wording into a reusable 25-query test set. As of August 2026, storefront behavior can still vary by theme, market, language, catalog data, and installed search software, so record those conditions with every run. The purpose is not to produce a one-time audit score. It is to create a controlled regression set that ecommerce managers, merchandisers, and QA teams can run before and after any search-related change.

How does the query generator work?

The generator creates five query variants for each of five catalog seeds, producing 25 tests. Choose seeds that represent different commercial and operational risks rather than selecting five bestsellers that all share clean data. A practical seed set includes one bestseller, one high-margin product, one newly launched product, one long-tail product, and one product that customers or staff have previously struggled to find.

For each seed, record the product title, product type, two decisive attributes, one customer synonym, and one natural-language description. The generator then converts those fields into five queries. For a women’s waterproof hiking jacket, the outputs might be Alpine Storm Jacket, alpine sto, rain shell, women waterproof jacket, and waterproof hiking jacket for women under 150. Replace every example with terms supported by your catalog and current pricing.

Write the intended result beside every query before testing. Use the Shopify Search Relevance Audit Tool when you need a broader review of search configuration, catalog readiness, and result quality beyond query generation.

The input sheet determines test quality

Good inputs come from catalog data and observed shopper language, not from a brainstorm alone. Export or inspect the product records for each seed. Copy the title and product type exactly. Then identify attributes that materially narrow purchase intent, such as size, fit, material, compatibility, capacity, color, age group, dietary property, or use case. Avoid decorative attributes that would not change a buying decision.

Build the synonym field from customer-support wording, onsite search reports, merchandising knowledge, and category conventions. A merchant may call an item a crossbody bag while shoppers also use shoulder purse. A hardware catalog may use hex key while customers enter Allen key. Keep a distinction between a true synonym and a related product: rain shell can describe a waterproof jacket, while umbrella cannot.

Natural-language queries should combine two or three constraints. Include one use case, audience, compatibility requirement, or price boundary when the catalog supports it. If products disappear entirely, use the Shopify product indexing diagnostic worksheet before changing relevance rules. Relevance tuning cannot correct a product that is unpublished, unavailable to the tested market, or absent from the searchable surface.

Build the 25-query matrix

Create one row per query type for every seed product, then write the expected outcome before running the search. An expected outcome can be a specific product in the first three positions, an appropriate product family on the first results page, or no result because the store does not carry the requested item. Writing the expectation first prevents testers from accepting whatever the search engine happens to return.

CriterionWhat to checkWhy it matters
Exact queryFull title, SKU, model, or exact product typeConfirms that known-item searches reach the intended product
Partial queryFirst meaningful word plus part of the next wordExposes weak prefix handling and predictive-search gaps
Synonym queryCustomer term that differs from catalog wordingTests whether shopper vocabulary maps to merchant vocabulary
Attribute queryProduct type plus one or two decisive attributesChecks whether structured product details affect retrieval usefully
Natural-language queryUse case or need expressed as a short sentenceShows whether intent survives beyond literal title matching

Run all five variants against each of the five seeds. Keep predictive suggestions separate from the submitted results page because they are different surfaces. Test while signed out, use the same market and language, and record whether the run used desktop or mobile. If filters influence the journey after search, compare the setup with Shopify search facet best practices. Save the date, theme version, search configuration, and tester name so another person can reproduce the run.

Score relevance without hiding critical failures

Score each query from 0 to 2 and preserve the notes behind the number. Give 2 points when the intended product or clearly suitable product family appears within the agreed result depth. Give 1 point when suitable products appear but are buried below weaker matches. Give 0 points when the query returns nothing, returns an unrelated set, or omits an in-stock and published expected product. With 25 queries, the maximum score is 50.

Use 45 or more as an initial pass threshold, 35 to 44 as a tuning queue, and below 35 as a signal that the search setup needs structured intervention. These are operating thresholds, not universal benchmarks. Tighten them for high-intent catalogs where model numbers, compatibility, or replacement parts must be exact. Regardless of total score, treat any zero-result query for a valid bestseller or exact SKU as a release blocker.

Record rank position, zero-result status, unexpected products, and the likely cause. Rerun the unchanged set after catalog imports, theme releases, synonym changes, or search configuration updates. For a wider regression pack, use the 30-test Shopify site search checklist. Do not replace failed queries with easier ones; preserve them until the underlying customer need or catalog assortment changes.

When are existing relevance controls insufficient?

Existing controls are insufficient when the same failure class persists after product data and publication issues are corrected. Examples include customer synonyms repeatedly returning unrelated products, exact models being buried by generic matches, attribute-heavy queries ignoring decisive constraints, or merchandising priorities being impossible to maintain without repeated catalog edits.

Do not replace search software merely because one ambiguous query produces a debatable order. First correct missing product types, inconsistent attributes, duplicate naming, and accidental publication gaps. Then rerun the identical 25 queries. If the score remains below your threshold or critical queries still fail, document the controls required and evaluate Hyper Search & Filter. Assess the app against the failure log rather than a generic feature wish list: each requirement should connect to a query, an expected result, and a current failure.

If the decision is between native functionality and another search layer, use the Shopify native search versus a third-party app comparison to frame the trade-off. Generate the test set, review the results, fix catalog defects, and only then decide whether additional relevance control is justified.

FAQs

How can I improve Shopify search results?

Improve Shopify search results by fixing catalog data first, then tuning search behavior against a repeatable query set. Standardize product types, titles, attributes, variants, and customer-facing terminology. Confirm that expected products are active, available to the intended sales channel, and returned for exact queries. Next, test partial terms, synonyms, attribute combinations, and natural-language requests. Separate retrieval failures from ordering failures: a missing product usually needs an indexing or data check, while a relevant product appearing too low calls for relevance or merchandising changes. Rerun the same queries after every adjustment rather than judging improvement from a few new examples.

What products can Shopify search?

Shopify storefront search can return products available to the relevant storefront, subject to product status, publication, catalog data, theme behavior, market conditions, and search configuration. A product that exists in the Shopify admin is not automatically a valid search expectation for every market or sales channel. Before marking a test as failed, confirm that the product is active, published where the test runs, available under the selected market conditions, and represented by useful searchable wording. Test one exact title or SKU first. If that fails, investigate product availability and indexing before tuning result order.

What is Shopify semantic search?

Semantic search interprets the meaning and intent of a query rather than relying only on literal word matches. In ecommerce, that can help connect shopper wording such as rain shell for hiking with products described using different but related catalog language. Semantic behavior should still be tested against expected products because broader interpretation can introduce plausible but commercially wrong results. Include natural-language, use-case, and synonym queries in the test set, then check whether decisive constraints such as audience, compatibility, material, or product type remain intact. The correct result is the product that satisfies the shopper’s constraints, not merely one that shares a broad concept.

What is Shopify predictive search?

Shopify predictive search is the suggestion experience shown while a shopper types, before the full search results page is submitted. Depending on storefront configuration, suggestions may expose products, query completions, or other searchable content. Test predictive search independently from submitted search because one surface can perform well while the other fails. For each partial query, record whether the intended suggestion appears, how many characters are required, and whether selecting it leads to the correct destination. Use the same device width during regression tests because the visible suggestion set may differ across layouts. A passing submitted search does not cancel a predictive-search failure that blocks shoppers before submission.

Continue reading

More practical Shopify tools

View all tools