CFOtech New Zealand - Technology news for CFOs & financial decision-makers
New Zealand
AI search answers vary widely across major platforms

AI search answers vary widely across major platforms

Tue, 18th Aug 2026 (Today)
Sean Mitchell
SEAN MITCHELL Publisher

The Optimisers has published a New Zealand study on how AI search answers vary across major platforms. It found that the top recommendation often differed between ChatGPT, Google AI Mode and Microsoft Copilot.

The study used 76 participants, who submitted 888 answers from their own accounts using fixed prompts on KiwiSaver, power companies and accounting software. Participants also repeated the same prompts in ChatGPT Temporary Chat, allowing comparison under a different system condition.

A central finding was what The Optimisers described as platform divergence. In 134 of 211 matched cases, or 63.5%, the three engines did not agree on the first-named provider for the same person and category.

The pattern was especially clear in KiwiSaver. ChatGPT first-named Kernel in 56 of 76 analysable answers, Google AI Mode first-named Milford in 68 of 75, and Copilot first-named Simplicity in 39 of 71.

Power recommendations also split by engine. ChatGPT most often put Ecotricity first in 35 of 76 answers, Google AI Mode most often led with Powershop in 31 of 73, and Copilot most often named Ecotricity first in 59 of 72.

Accounting software produced a different result. Across 221 analysable answers, every response first-named Xero, making it the only category in the study in which one brand held an undisputed lead across all engines.

Account effects

The comparison between normal ChatGPT use and Temporary Chat pointed to another source of variation. Among 72 people with comparable results, 50 received at least one different lead provider when the prompts were repeated in Temporary Chat.

Across 215 matched category pairs, the lead recommendation changed in 61 cases. The wider set of brands named in the answer changed in 184 cases, or 85.6%.

Temporary Chat removes access to saved memories used for personalisation, though other settings, system instructions and run-to-run variation can still affect outputs. That means the comparison reflects a change in condition rather than a clean test of memory alone.

Richard Conway, Founder and Chief Executive Officer at The Optimisers, said the findings showed why businesses should avoid treating one AI interface as a complete view of the market. "AI search is probabilistic and cannot be monitored and measured through a single product & treated as a universal market view," Conway said.

The report argues this matters for digital teams trying to track how brands appear in AI-generated answers. It says testing through a logged-out interface or an application programming interface can provide a controlled benchmark, while testing through ordinary signed-in accounts gives a closer picture of what users may actually see.

Brand residue

One anomaly emerged in the power category. Frank Energy appeared in 86 of 295 power answers across normal and Temporary Chat conditions, despite no longer taking new customers.

Within the three normal engines, Frank Energy appeared in 77 of 221 answers, or 34.8%. One answer even flagged that the provider was closed to new customers before still recommending it.

The result points to the way AI systems can draw on older information and third-party references that remain online long after a commercial change. According to the study, directories, comparison sites, archived publisher material and old support pages can continue to shape how a brand is represented.

Conway said this can create problems for companies that assume updating their own websites is enough. "Entity mapping is critical for companies and should be considered across current brands, parent companies, subsidiaries, legacy names and discontinued products or divested business elements," he said.

Businesses should check not only whether their brand is named in AI answers, but also where it appears in a recommendation list, which sources are cited, what attributes are attached to the brand and whether the facts are current, the report says.

Monitoring limits

The report contrasts AI answer tracking with conventional search monitoring, where repeated searches usually return the same result. In AI systems, answers are probabilistic and can shift according to prompt wording, account conditions and other variables.

That creates a challenge for marketers and executives trying to attribute visibility or compare providers. The study says each measure should include its true denominator because unavailable responses and category errors can affect the apparent result.

Controlled tracking tools were also assessed against the plurality result from real-account responses. According to the study, one benchmark matched seven of nine engine-category cells and another matched six, while neither could measure logged-in results.

The findings also suggest AI search may favour smaller brands in some categories. In KiwiSaver, none of the big four banks was first recommended, while in energy the main brands were often represented through sub-brands rather than parent names.

Conway said businesses should be cautious about broad claims around AI visibility. "With attribution getting harder, and many agencies selling snake oil (like the early days of Google), marketers and business leaders should be wary of definitive claims and absolutes," he said.