We Ran 22,335 AI Shopping Queries. Here's What We Found.
A 22,335-recommendation AI shopping study finds Budget Framing queries get recommended 3x more than Post Purchase queries, and platforms rarely agree.
What actually gets a product recommended by an AI shopping engine? We analyzed 22,335 product recommendations surfaced by ChatGPT, Perplexity, Google AI Mode, and Microsoft Copilot to find out. The AI shelf, the set of AI-generated product recommendations now appearing in place of, or ahead of, traditional search results, is already live. But the more useful finding isn't that it exists. It's what predicts whether your product shows up on it.
Key Takeaways - Google AI Mode recommended products in 64.4% of tested queries, well ahead of Copilot's 44.5%, ChatGPT's 28.5%, and Perplexity's 5.0% — Copilot is a close second, not a distant third. - Across the full 9-stage intent ranking, Budget Framing tops out at 54.8% and Post Purchase sits lowest at 14.7%, a 40.1-point spread and the single largest framing lever found in the data. - Category matters, and this post is the first time we've isolated it in this dataset: Electronics recommends at 50.3%, more than 21 points above the lowest category, Auto at 28.7%. - Attribute count follows an inverted U, not a straight line: recommendation rate peaks at 41.7% with 3 stated attributes, and falls off on both sides. - When the same query was sent to multiple platforms, product overlap was close to zero across every platform pair tested, meaning the four engines essentially don't agree on what to recommend. - These are first-pass findings from an initial analysis run and are flagged internally as hypotheses pending an independent confirmation run. Treat directional patterns as strong signal, not settled fact.
Percentages throughout are rounded to the nearest whole point; point-spreads are calculated from the underlying unrounded rates, so a spread may differ by a few tenths from what subtracting the rounded percentages would suggest.
What Did We Actually Study?
This analysis covers 22,335 product recommendations surfaced across ChatGPT, Perplexity, Google AI Mode, and Microsoft Copilot, drawn from a structured query bank spanning 7 categories, Electronics, Sports, Beauty, Auto, Home, Household, and Toys, 5,584 unique queries in total, built around our own 9-stage shopping-intent framework (Category Exploration, Scenario Confirmation, Budget Framing, Post Purchase, Purchase Execution, Validation/Risk Reduction, Problem Recognition, Comparison, Attribute Constrained).
This intent framework is the same one behind our earlier 8,520-query ChatGPT shopping study, which first identified intent stage as a stronger predictor of AI product recommendations than most other query features. This analysis extends that work to a larger, multi-platform sample.
Each query was tagged with a set of features before it was ever sent: which intent stage it represented, whether it mentioned price, whether it used deal-seeking or urgency language, whether it named a brand, how many product attributes it specified, and how conversational its phrasing was. That let us measure not just whether a platform recommended a product, but which query characteristics moved the odds up or down.
Study Scope: This analysis run is a first-pass, single-batch result over a multi-category query set spanning Electronics, Sports, Beauty, Auto, Home, Household, and Toys. AEOsome's internal research module flags every finding here with a "hypothesis" status, meaning the patterns are real in this batch but have not yet been independently replicated in a second run. We're publishing the directional findings now because the effect sizes are large enough to act on, but a confirmation run is the next step before we'd call any of these numbers final.
How We Measured a "Recommendation"
A recommendation was recorded when a platform surfaced a structured product card or carousel placement in response to a query, not merely a text mention of a product name. This is a stricter bar than counting any named brand or category in a wall of text: it tracks whether the platform actually put a specific, clickable product in front of the shopper.
Two of the higher-recommendation query types looked like this in practice:
- Budget Framing: "I need a laptop under $800. What are the best options right now?"
- Post Purchase: "My laptop battery is completely dead and it only works when plugged into the wall."
Both describe a real shopping situation. Only one reliably produces a product card. That gap is the core finding of this study.
Which AI Platform Recommends Products Most Often?
Google AI Mode surfaced a product recommendation in 64.4% of tested queries, the highest of the four platforms. Perplexity was lowest at 5.0%, a 59.4-point gap that is the largest single effect measured anywhere in this dataset. But the full four-platform picture matters here: Copilot is a close second at 44.5%, not a distant third, and ChatGPT sits closer to Copilot than to Perplexity at 28.5%.
| Platform | Recommendation Rate |
|---|---|
| Google AI Mode | 64.4% |
| Copilot | 44.5% |
| ChatGPT | 28.5% |
| Perplexity | 5.0% |
Google's own account of its AI Mode shopping build helps explain the top spot: the company has said its Shopping Graph backing AI Mode holds more than 50 billion product listings, refreshed continuously, specifically engineered to surface a browsable panel of products rather than a text answer (Google, "Shopping on Google: AI Mode and virtual try-on updates," May 2025). Perplexity's architecture, by contrast, leans toward cited informational answers over product cards, which tracks with the near-floor recommendation rate we measured. Copilot's 44.5% and ChatGPT's 28.5% land in between, closer to each other than either lands to Google AI Mode or Perplexity, so a "Google versus everyone else" framing understates how active Copilot already is on this metric.
Citation Capsule: Across 22,335 AI shopping recommendations, product recommendation rate by platform was Google AI Mode 64.4%, Copilot 44.5%, ChatGPT 28.5%, and Perplexity 5.0% (chi-square, n=22,335; AEOsome Research, hypothesis-stage). The 59.4-point gap between Google AI Mode and Perplexity is the largest single effect measured in the dataset, but Copilot's 44.5% makes it a strong second platform, not a footnote.
Platforms Don't Agree With Each Other
Here's a finding that complicates any "just optimize for Platform X" strategy: when we sent the same query to multiple platforms and checked whether they surfaced the same products, agreement was close to zero everywhere we measured it.
| Platform Pair | Overlap (Jaccard) | Sample Size |
|---|---|---|
| ChatGPT vs. Copilot | 0.10 | 915 |
| ChatGPT vs. Perplexity | 0.01 | 139 |
| Copilot vs. Perplexity | 0.004 | 202 |
| ChatGPT vs. Google AI Mode | 0.00 | 1,272 |
| Copilot vs. Google AI Mode | 0.00 | 1,884 |
| Google AI Mode vs. Perplexity | 0.00 | 267 |
Every pair falls in negligible-overlap territory. The largest agreement we found, ChatGPT and Copilot, still means the two platforms picked the same product roughly 1 time in 10. For the remaining five pairs, agreement is functionally zero.
That matters for how brands should read any single-platform ranking. If your product shows up in Google AI Mode but not Copilot, that's not necessarily evidence you're doing something wrong on Copilot. It may just be how differently these four engines are choosing to build their recommendation sets in the first place.
Independent third-party data points the same direction. Productrise, an organic Google Shopping tracking firm, compared over 2 million product listings across regular search and Google AI Mode and found only 1.28% of the same-day carousel products also appeared in AI Mode, with a different top seller listed 49.6% of the time when a product did appear in both (Search Engine Journal, "Google AI Mode Prices Differ From Product Carousel For Same Items," September 2026). That's a different comparison than ours (regular search vs. AI Mode within Google, rather than across ChatGPT/Copilot/Perplexity/Google AI Mode), but it's the same underlying pattern: even within a single company's own product ecosystem, different surfaces don't agree on what to show.
Which Query Framing Gets a Product Recommended?
Intent stage is the single largest framing lever in the dataset. Here's the full 9-stage ranking, not just the two extremes.
| Intent Stage | Recommendation Rate |
|---|---|
| Budget Framing | 54.8% |
| Attribute Constrained | 47.0% |
| Purchase Execution | 43.0% |
| Scenario Confirmation | 39.6% |
| Category Exploration | 37.8% |
| Comparison | 35.7% |
| Validation/Risk Reduction | 26.2% |
| Problem Recognition | 20.4% |
| Post Purchase | 14.7% |
Budget Framing and Post Purchase are the two extremes, a 40.1-point spread, but the middle of the ranking matters too: Attribute Constrained queries (queries that name specific product attributes, like "waterproof running shoes size 10") sit close behind Budget Framing at 47.0%, and Purchase Execution ("add the black one to my cart") isn't far behind that at 43.0%. The bottom three stages, Validation/Risk Reduction, Problem Recognition, and Post Purchase, all read more like advice-seeking than shopping to these platforms, and are recommended accordingly less often.
Beyond intent stage, a handful of binary query features each shift the recommendation rate on their own:
| Signal | Higher Rate | Lower Rate | Spread |
|---|---|---|---|
| Deal-seeking language present | 48% | 35% | 13.0 pts |
| Price mentioned or constrained | 47% | 34% | 12.5 pts |
| Urgency language present | 36% | 31% | 5.3 pts (negative) |
| Purchase-explicit language ("buy," "order") | 40% | 35% | 5.0 pts |
| Brand named in query | 39% | 35% | 4.5 pts |
| Comparison language present ("X vs. Y") | 36% | 32% | 4.7 pts (negative) |
| Use-case mentioned | 37% | 33% | 4.3 pts (negative) |
| Problem-framing language ("my X keeps breaking") | 38% | 21% | 17.2 pts (negative) |
Read the "negative" rows carefully. Urgency language, explicit comparison phrasing, and naming a specific use case all correlated with lower recommendation rates in this batch, the opposite of what most content playbooks assume. Problem-framing language shows the same pattern at a larger scale: describing a problem rather than naming a product dropped the recommendation rate by 17.2 points. These platforms appear to read "my laptop keeps overheating" as a request for troubleshooting advice, not a shopping trigger, even though the shopper's next step is very likely to buy a replacement.
Put together, these signals suggest AI shopping engines are pattern-matching "is this shopper ready to compare priced options right now" rather than "does this query mention a product category." Budget Framing, Attribute Constrained, deal-seeking, and price-constrained language all say "show me options I can act on." Problem-framing, Post Purchase, and use-case language say "help me think this through," and that reads as a request for advice, not a shelf of products, even when the underlying need is identical.
Ambiguity Score: One Level Stands Out, on a Thinner Sample
Ambiguity score measures how vague versus specific a query is, on a 6-level scale. The full breakdown:
| Ambiguity Level | Recommendation Rate | Sample Size |
|---|---|---|
| Level 1 | 43.6% | n=1,484 |
| Level 2 | 34.3% | — |
| Level 3 | 34.5% | — |
| Level 4 | 35.4% | — |
| Level 5 | 35.1% | — |
| Level 6 | 36.1% | — |
Levels 2 through 6 cluster tightly, all within about 2 points of each other. Only level 1 stands out, and it's also the smallest sample of the six at 1,484 responses, notably thinner than the other levels. That combination, one outlier level on a thinner sample, means this signal is worth noting but not worth overstating. We wouldn't build a content strategy around ambiguity score alone until a confirmation run checks whether the level 1 effect holds at a larger sample size.
Attribute Count: A Sweet Spot, Not a Straight Line
Attribute count (how many specific product attributes a query names, e.g. color, size, material) does not follow a simple "fewer is better" pattern. The real shape is an inverted U, peaking at 3 stated attributes:
| Attributes Stated | Recommendation Rate | Sample Size |
|---|---|---|
| 0 | 34.4% | — |
| 1 | 38.5% | — |
| 2 | 39.6% | — |
| 3 (peak) | 41.7% | — |
| 4 | 38.8% | n=152 |
| 5 | 34.2% | n=76 |
Recommendation rate climbs steadily from 0 to 3 attributes, peaks at 41.7% with 3, then falls back down at 4 and 5 attributes to roughly where it started. The practical read is a sweet spot around 2-3 stated attributes, not "the fewer attributes the better." The 4- and 5-attribute tail also runs on much smaller samples, 152 and 76 responses respectively, so treat the drop-off at the extreme end as directional rather than definitive until it's replicated on a larger sample.
Query Style: Short Edges Out, but the Field Is Close
Query style refers to how a query is phrased, Short, Medium, Long, or Conversational, and is a separate feature from the conversational-tone binary flag reported elsewhere in this study. The full 4-level breakdown:
| Style | Recommendation Rate |
|---|---|
| Short | 37.9% |
| Medium | 36.1% |
| Long | 34.8% |
| Conversational | 33.7% |
Short queries edge out the rest, but the spread from top to bottom is just 4.2 points, the tightest range of any multi-level variable in this study. Query phrasing style is a real but modest lever compared to intent stage or category.
Is Intent a Stronger Predictor Than Platform Choice?
Not universally. The single largest effect in this dataset is platform choice: the 59.4-point gap between Google AI Mode (64.4%) and Perplexity (5.0%) is larger than any framing signal we tested, including Budget Framing's 40.1-point spread. But among the three mainstream shopping platforms (excluding Perplexity, whose editorial model is a known outlier), query framing is doing real, independent work on top of whatever platform a shopper happens to use.
The practical read: which platform a shopper is on sets the ceiling on how likely a recommendation is at all. How they frame the question, budget-first versus problem-first, determines whether they land above or below that platform's baseline. Brands can't control platform choice. They can influence how their own content teaches shoppers, and AI systems paraphrasing shoppers, to frame the ask.
Does Product Category Change the Odds?
Yes, and this is a gap in our own earlier reporting on this dataset: the aggregate numbers above blend all seven categories together, which hides real category-level differences. Here's the full breakdown by category.
| Category | Recommendation Rate |
|---|---|
| Electronics | 50.3% |
| Sports | 37.8% |
| Beauty | 35.9% |
| Toys | 31.9% |
| Home | 31.4% |
| Household | 30.1% |
| Auto | 28.7% |
Electronics sits well above every other category, more than 12 points clear of Sports in second place, and 21.6 points above Auto at the bottom. That's a large enough spread that a brand selling in Electronics should expect a meaningfully higher recommendation baseline than one selling in Auto or Household, independent of anything else in this study. Category is not a minor footnote here; it's one of the larger single-variable effects in the entire dataset, on par with the platform gap between ChatGPT and Perplexity.
Citation Capsule: Category-level product recommendation rates ranged from 50.3% (Electronics) to 28.7% (Auto), a 21.6-point spread (AEOsome Research, n=22,335, hypothesis-stage). Electronics led every other category by at least 12 points.
What This Means for E-Commerce Brands
1. Treat Google AI Mode as the highest-leverage AI shopping surface, but don't assume it predicts the others. A 64.4% recommendation rate makes Google AI Mode the most active surface tested, with Copilot a strong second at 44.5%, and OpenAI has said it's expanding the Agentic Commerce Protocol specifically to bring "more complete, relevant, and up-to-date information" into ChatGPT's own product discovery (OpenAI, "Powering Product Discovery in ChatGPT," March 2026). Near-zero cross-platform agreement means a strong showing on one platform tells you little about the others. Structured product data and clean feeds still matter everywhere; expect the actual product mix shown to differ platform by platform.
2. Write for budget- and deal-framed shopping moments, not just category coverage. Budget Framing led the full 9-stage intent ranking at 54.8%, outperforming Post Purchase's 14.7% by 40.1 points, and deal-seeking or price-constrained language each added 12-13 points on their own. Audit product pages and FAQs for "best [product] under $X" and clear, visible pricing. That framing consistently reads as "ready to shop" to these systems. Making that pricing and framing machine-readable also depends on clean product schema and structured data, since AI shopping engines pull from structured feeds, not page copy alone.
3. Don't assume problem-framed or comparison-heavy content earns a shelf slot. Problem-framing language ("my X keeps breaking") correlated with a 17-point lower recommendation rate, and explicit comparison phrasing showed a smaller negative effect too. If your funnel leans on "problem, then solution" content to drive AI visibility, this data says that content may be read as advice-seeking rather than purchase-ready, at least on the platforms and query types tested here.
4. Calibrate expectations by category and aim for a 2-3 attribute sweet spot. Electronics recommends at more than 1.7x the rate of Auto or Household, so category alone should shape how much AI-shopping traffic a brand expects. Within a product listing, the attribute-count data suggests naming 2-3 concrete attributes (not zero, and not five or more) tracks with the highest recommendation rates.
5. Treat this as a strong first signal, not a settled number. This run is a first-pass analysis, and every effect above carries a "hypothesis" flag internally pending an independent confirmation run on fresh data. The direction and rough size of the largest effects (platform gap, Budget Framing, category, problem-framing) are consistent enough with prior smaller-scale work that we're confident acting on them now, but we'll publish an update once a second run confirms or narrows these numbers.
The data behind this post comes from AEOsome's ongoing AI shopping research program. We run structured query studies to help e-commerce brands understand where and how AI platforms surface products, so their optimization decisions are based on evidence, not assumption.
Methodology and Limitations
This study analyzed 22,335 product recommendations, deduplicated from a larger initial query batch, across a 7-category query bank (Electronics, Sports, Beauty, Auto, Home, Household, and Toys; 5,584 unique queries) sent to ChatGPT, Perplexity, Google AI Mode, and Microsoft Copilot. Each query carried pre-tagged features (intent stage, price mention, deal-seeking language, urgency language, brand naming, attribute count, ambiguity score, and phrasing style). A recommendation was recorded when a platform returned a structured product card or carousel placement, not a bare text mention. Effect sizes were computed with chi-square tests for categorical signals and Jaccard overlap for cross-platform agreement.
Every finding in this post carries a "hypothesis" status in AEOsome's internal research tracking: real in this batch, not yet confirmed by an independent second run. The query bank spans 7 categories (Electronics, Sports, Beauty, Auto, Home, Household, and Toys); category-by-category recommendation rates are reported in full above (Electronics 50.3% down to Auto 28.7%). Two other planned analysis types, a merchant group comparison and a within-merchant content comparison, were not run for this dataset: zero merchants were crawled for this analysis run, so there is no merchant-level data to compare against. Those remain open for a future run. Platform behavior changes quickly; these figures represent a specific analysis window and should be re-validated as platforms update their shopping features.
FAQ
What counts as a "recommendation" in this study?
A recommendation was recorded when a platform surfaced a structured product card or carousel placement in response to a query, not a plain-text mention of a product name or category. That's a stricter, more commercially meaningful bar than counting any named brand in a response.
Why is Perplexity's recommendation rate so much lower than Google AI Mode's?
Perplexity surfaced a product in just 5.0% of tested queries, versus 64.4% for Google AI Mode. Copilot (44.5%) and ChatGPT (28.5%) fall in between, both notably closer to each other than to Perplexity. Perplexity's response architecture leans toward cited informational answers rather than product cards. Brands optimizing for Perplexity should focus on earning editorial mentions in the sources it cites, rather than direct product placement.
If platforms barely agree with each other, how should I prioritize where to optimize?
Start with Google AI Mode given its much higher baseline recommendation rate, but treat Copilot as a real second priority, not an afterthought, given its 44.5% rate. Don't treat a strong or weak showing on any one platform as predictive of the others. Near-zero cross-platform overlap in this data means each surface needs to be checked and optimized somewhat independently.
Which query framing should e-commerce brands prioritize in their own content?
Budget Framing, deal-seeking language, and clear price signals produced the highest recommendation rates in this study, with Budget Framing leading the full 9-stage intent ranking at 54.8%. Problem-framing and heavy comparison language showed the opposite effect. Where it fits naturally, favor content that answers "best [product] under $X" over content framed purely as troubleshooting or side-by-side comparison.
Does product category affect the recommendation rate?
Yes. Electronics recommends at 50.3%, the highest of the seven categories tested, versus 28.7% for Auto, the lowest, a 21.6-point spread. Sports (37.8%), Beauty (35.9%), Toys (31.9%), Home (31.4%), and Household (30.1%) fall in between. Category alone is one of the larger single-variable effects in this dataset.
Is this a final, confirmed result?
No. Every finding here is flagged internally as a hypothesis from a single first-pass analysis run, not yet confirmed by an independent second run. The largest effects (the platform gap, Budget Framing's lead, category differences, problem-framing's penalty) are consistent enough with prior smaller studies that we're confident in the direction, but we're treating the exact numbers as provisional until a confirmation run lands. Two planned analyses, a merchant group comparison and a within-merchant content comparison, weren't run at all for this dataset because zero merchants were crawled for this run.
Research by Vijaya Kumar Channalli, founder of AEOsome. Vijay runs AEOsome's structured query studies on AI shopping behavior, including the earlier 8,520-query ChatGPT study this analysis extends. AEOsome studies how AI platforms discover, evaluate, and recommend e-commerce products. Contact us via AEOsome's AEO audit page to discuss a structured AEO audit for your brand.
Methodology: 22,335 product recommendations across ChatGPT, Perplexity, Google AI Mode, and Microsoft Copilot; 7-category query bank (Electronics, Sports, Beauty, Auto, Home, Household, Toys); 9-stage intent taxonomy. A recommendation was recorded when a platform returned a structured product card or carousel placement. Findings are hypothesis-stage pending an independent confirmation run and should be validated against current platform outputs.