I started this test with what I thought would be the boring part of a luxury-brand audit: can an AI answer point a buyer to the right official China channel?
The more interesting result appeared when I compared that task with a buying decision.
In a small exploratory pre-wave, Piaget appeared in all four answers about jewelry brands with verifiable official China channels. It appeared in none of the four answers recommending high-end brands for wedding jewelry.
That does not make Piaget “invisible,” and it certainly does not establish a market ranking. Four answer opportunities per task are too small for either claim.
It does expose a measurement problem that matters well beyond jewelry: finding an entity, recommending a brand and routing a buyer to a correct official channel are different jobs. If we report them as one visibility score, the number may be tidy while the diagnosis is wrong.
I fixed eight target brands before collection: Cartier Tiffany & Co. Bvlgari Van Cleef & Arpels Chaumet Boucheron Piaget De Beers Jewellers
I used three neutral Chinese buyer questions: Which high-end brands are worth considering for wedding jewelry, and what should a buyer compare? What is the difference between a brand boutique and a daigou for a high-value jewelry purchase? Which international jewelry brands have verifiable official websites or boutique channels in mainland China?
The valid collection contained two retrieval-off API surfaces, DeepSeek and Doubao, with two answers per surface and question. That produced 12 raw answers.
The important implementation choice was to collect each answer once. I did not repeat the same unbranded question eight times, once per target brand. After collection, each stored answer was evaluated against the frozen eight-brand registry, producing 96 answer-brand cells.
The distinction matters because provider calls and scoring units are not the same thing:
Reporting “96 AI answers” would overstate the collection by a factor of eight. Reporting only 12 rows without describing the derived brand cells would hide the scoring denominator.
For the wedding-selection question, the eight brands received 21 positive shortlist recommendations across 32 eligible answer-brand opportunities.
For the official-channel question, the models affirmatively listed a target brand in 25 of 32 opportunities.
Those two rates are not a funnel. They come from different prompt intents. The useful comparison is at brand level:
| Brand | Wedding shortlist | Official China channel listed | | --- | ---: | ---: | | Cartier | 4/4 | 4/4 | | Tiffany & Co. | 4/4 | 4/4 | | Bvlgari | 4/4 | 4/4 | | Van Cleef & Arpels | 4/4 | 4/4 | | Chaumet | 3/4 | 2/4 | | Boucheron | 1/4 | 2/4 | | Piaget | 0/4 | 4/4 | | De Beers Jewellers | 1/4 | 1/4 |
The Piaget row is the clearest example of why entity recognition is not recommendation. The APIs could associate the brand with a current China-facing channel, but the brand did not enter their wedding shortlist in this tiny sample.
Chaumet shows the other direction. It appeared in three of four wedding shortlists but only two of four official-channel answers. A brand can enter consideration while the route to verification remains less consistently reproduced.
For a high-value product, I would split the audit into at least three layers. Consideration
