Computing

I Gave the Same Purchase Investigation to Six AIs: They Didn't Find the Same Internet

ChatGPT, Gemini, DeepSeek, Grok, Meta AI and Mistral had to decide whether to replace an iPhone 16. Their consensus hides contradictory prices, forgotten regional variants, and a broader lesson on the reliability of LLM web search. Publication date: September 16, 2026

A smartphone placed at the center of a desk surrounded by six computers displaying different comparisons and prices, illustrating the divergences between AIs during the same web search.

Does there exist today a smartphone sufficiently better than an iPhone 16 to justify replacing it?

The question seems trivial. It wasn’t.

On September 16, 2026, I submitted essentially the same brief to six AI assistants: ChatGPT, DeepSeek, Grok, Gemini, Meta AI and Mistral. Not a three-line prompt asking for “the best phone right now,” but a genuine purchase investigation: French market, compactness, real prices, battery life, performance, photo, privacy, security, repairability, software support, GrapheneOS, post-resale cost, and the explicit possibility that the best decision might simply be to buy nothing.

All six systems had web access. They were to prioritize manufacturers, cross-reference reviews, verify European variants, seek counter-arguments, and acknowledge information they couldn’t confirm.

The result is more interesting than the smartphone itself.

Five out of six AIs ultimately advised keeping the iPhone 16.

And yet, they clearly hadn’t all investigated the same market.

A Galaxy S26 could be worth 600 €, 750 €, or over 1,000 €. The same French Galaxy S26 had a Snapdragon in one response and an Exynos in another. Some assistants had incorporated the recently released Pixel 11; others seemed still to be living in the era of the Pixel 9 or Pixel 10. Several produced very precise percentages on GrapheneOS compatibility without a corresponding dataset actually existing.

The most troubling thing wasn’t that they were wrong.

It was that their errors often looked exactly like verified information.

The Smartphone Was Only a Pretext

It’s worth clarifying what this experiment measures and what it doesn’t.

This isn’t a scientific benchmark of the “intrinsic intelligence” of six models. Exact versions, search engines, browsing systems, indexes, caches, internal tools, and source selection strategies may differ. A new run the next day could also produce different results.

The experiment measures something more concrete: the quality of the complete system that a user actually encounters today when entrusting a complex decision to a web-connected AI.

And the phone makes an excellent stress test.

To answer correctly, one must simultaneously:

  • know the exact date;
  • discover very recent products;
  • recognize variants sold in France;
  • distinguish manufacturer price, promotion, marketplace, and import;
  • compare measurements obtained with different protocols;
  • understand software topics like GrapheneOS;
  • resist the urge to fill in missing data;
  • and, above all, be capable of concluding that a purchase may not be necessary.

It’s no longer enough to “find pages.”

One must understand what they actually prove.

First Surprising Result: Almost All AIs Refuse to Sell a Phone

The iPhone 16 used as a reference officially measures 147.6 × 71.6 × 7.8 mm for 170 g, with a 6.1-inch OLED screen and an A18 chip.

The challenge was therefore severe: find a sufficiently significant improvement without turning a compact phone into a brick of over 200 grams.

DeepSeek, Grok, Gemini, Meta AI and ChatGPT ultimately converge on the same idea: no candidate brings an overall improvement obvious enough to make replacement indispensable.

The main recurring criticism of the iPhone 16 is its screen limited to 60 Hz. On this point, the compact Android market and the iPhone 17 offer an immediately perceptible gain.

The iPhone 17 does indeed have a ProMotion screen reaching 120 Hz, an A19 chip, 256 GB minimum, and remains reasonably compact at 149.6 × 71.5 × 7.95 mm for 177 g.

Samsung does even better on bulk: the Galaxy S25 weighs only 162 g and measures 146.9 × 70.5 × 7.2 mm. The S26 rises slightly to 167 g and 149.6 × 71.7 × 7.2 mm.

But going from “better on several criteria” to “sufficiently better to sell a recent phone that already works” is another question.

That’s precisely where the reasoning begins to diverge.

The Same Galaxy S26 Costs Between 599 and 1,002 Euros

Price is probably the best example of the whole experiment.

At the time of this investigation, Idealo lists the Galaxy S26 256 GB black from 599.69 €. This information is real.

It’s also dangerously incomplete.

The cheapest offers visible in the comparator come from marketplaces. Idealo indicates, for example, sellers showing ratings of 1.6/5, 1.0/5 or 1.1/5.

Meanwhile, Fnac displays the same Galaxy S26 256 GB at 1,002 € sold directly, with third-party seller offers starting around 753 €.

Another Idealo listing even shows Samsung France at 799 € for a 256 GB variant.

Thus, the following statements can all be accompanied by a real web source:

The Galaxy S26 costs about 600 €.

The Galaxy S26 costs about 800 €.

The Galaxy S26 costs about 1,000 €.

The problem is no longer finding information.

The problem is knowing which price actually answers the user’s question.

A floor price at a very poorly rated marketplace merchant doesn’t have the same decision value as a manufacturer price or a direct sale by a major retailer. Yet, for an LLM that simply extracts the lowest number from a comparison page, they can seem interchangeable.

And this difference is significant enough to overturn an entire recommendation.

Mistral Perfectly Illustrates This Butterfly Effect

Mistral constitutes the exception of the experiment.

Where the other five systems lean toward “keep the iPhone,” its response pushes much more clearly toward change.

This disagreement seems spectacular until one examines the assumptions.

Mistral works notably with extremely aggressive prices for several Android phones, combined with a relatively high resale value for the iPhone 16. In this scenario, buying a Galaxy S25 while reselling the iPhone can become almost financially neutral, even profitable.

From there, the verdict is logical.

But it depends enormously on the solidity of the entry price.

The S25 itself provides a good example. Darty currently shows a new offer at 549 €, but its own direct sale appears at 902 € on another listing, while the 549 € offer comes from a third-party seller.

The question is therefore not only:

“Does this price exist?”

It becomes:

“Can I actually recommend that someone base a decision of several hundred euros on this specific offer?”

The nuance seems minuscule. It changes the verdict.

Grok Found the Right Phone… with the Wrong Chip

Another type of error: the regional variant.

Grok correctly identifies the Galaxy S26 as one of the few Android competitors capable of staying close to the iPhone 16’s format.

But its analysis attributes it a Snapdragon 8 Elite Gen 5.

For a French buyer, that’s a problem.

Samsung France explicitly states that the Galaxy S26 and S26+ use the Exynos 2600, while the previous S25 did use the Snapdragon 8 Elite for Galaxy.

This isn’t a cosmetic error.

Changing SoC can affect:

  • CPU and GPU performance;
  • consumption;
  • thermal behavior;
  • battery life;
  • sustained performance.

An AI can therefore correctly compare “the Galaxy S26” in appearance while actually analyzing a device that isn’t the one sold to the user.

It’s a particularly interesting form of modern hallucination: the product exists, the chip exists, the association may exist in another region, but the whole becomes false solely because of the geographic context.

Some Models Had Already Missed the September 16 Market

The experiment takes place on a particularly cruel date for an imperfect search engine.

The Google Pixel 11 is already marketed in France from 999 €. Google confirms a Tensor G6 chip, a Titan M3 security coprocessor, and over 30 hours of battery life according to its protocol.

A few days earlier, on September 9, Apple officially presented the iPhone 18 Pro. Pre-orders began September 12, with availability in France announced for September 18 and a price from 1,479 €.

These two products aren’t necessarily the best choices for the mission.

That’s not the point.

An investigation claiming to represent the state of the market on September 16, 2026 must at minimum know they exist.

ChatGPT had incorporated them. Several competitors hadn’t.

This difference doesn’t prove one model is “more intelligent.” It shows above all that in current search, freshness is part of the truth.

A technically impeccable answer on the July market can be a bad answer in September.

Even Apple Can Contradict Apple

One might think the problem disappears once you impose official sources.

The iPhone 17 demonstrates the opposite.

At its launch on September 9, 2025, Apple officially announced a French price from 969 €. This value still exists on the Apple press release.

A French Apple Business page also still shows the iPhone 17 256 GB at 969 €.

But the currently indexed consumer Apple Store displays 1,119 €.

Same domain. Same manufacturer. Same phone. Two different amounts.

“Use the official source” is therefore not a sufficient strategy.

One must still determine:

  • when the page was published;
  • what audience it addresses;
  • whether it represents a launch price or a current price;
  • whether the offer is still applicable;
  • and which context actually corresponds to the question.

The provenance of information and its temporal relevance are two different problems.

Gemini Shows Why Precision Can Be an Illusion

Gemini is probably the most fascinating case of the lot.

Its response has everything that instinctively inspires confidence in a report:

abundant tables, scenarios, weighted scores, numerical values, financial projections, detailed battery life, privacy assessments, and dozens of references.

Visually, it’s almost a consulting firm report.

That’s precisely the trap.

Some assertions are solid. Others acquire a precision that the available data doesn’t permit.

The GrapheneOS case is particularly revealing.

Numbers like “70% of banking apps compatible” give an impression of statistical measurement. Yet GrapheneOS publishes no official dataset allowing global banking compatibility to be transformed into such a ratio.

The project’s documentation adopts a much more cautious formulation: apps using Play Integrity may refuse alternative systems to Google-certified Android, unless their developers explicitly choose to support a compatible attestation.

The two formulations are not equivalent.

“Some apps may not work” looks less impressive than a percentage.

But when the percentage isn’t measured, the vague sentence is actually more accurate.

It’s false precision: transforming a poorly quantified uncertainty into a clean, reassuring, memorable figure.

An LLM can then seem more scientific precisely when it becomes less rigorous.

GrapheneOS Was the Ideal Trap

The subject concentrated almost all the experiment’s difficulties.

Can you install the Play Store?

Yes.

GrapheneOS offers a compatibility layer allowing use of official Google Services versions in the normal Android sandbox, with reduced privileges rather than the special system access they usually have. Its documentation explicitly describes this “Sandboxed Google Play.”

Must you leave the bootloader unlocked?

No: normal GrapheneOS installation instead provides for relocking it to preserve the verified boot chain.

Are all Pixels compatible?

No.

The official September 2026 list notably includes the Pixel 10, 10 Pro, 10 Pro XL, 10 Pro Fold and 10a, but not the Pixel 11 range.

And Google Wallet?

NFC payment via Google Wallet remains problematic because GrapheneOS doesn’t satisfy the certification model Google expects for this function. Alternative solutions may exist depending on the country and provider, but their availability must not be automatically extrapolated.

This is exactly the kind of domain where an AI trained to produce a complete answer is tempted to fill in the gaps.

And every gap filled increases the probability of producing a more pleasing but less true answer.

Meta AI: When an Order of Magnitude Becomes a Statistic

Meta AI encounters a neighboring problem.

Its response contains good intuitions. It notably spots the Galaxy S25, which constitutes a particularly logical candidate: it remains extremely compact, has a 120 Hz screen, and already benefits from a strong depreciation relative to its initial positioning.

But some elements concerning GrapheneOS become too affirmative, notably when approximate banking compatibility is expressed as a percentage.

It’s a frequent trait of generative systems: taking a trend from community testimonials “many apps seem to work” and progressively transforming it into quantitative information.

At each step, the sentence becomes more usable.

It also becomes harder to defend.

DeepSeek: A Good Decision with an Incomplete Market

DeepSeek produces probably one of the healthiest economic reasonings of the experiment.

It understands that the real comparison isn’t:

iPhone 16 vs. best available smartphone.

But:

cost and benefit of replacement vs. value of keeping what you already have.

That’s an essential distinction.

It accounts for resale, recognizes the specific interest of GrapheneOS for a user prioritizing privacy, and accepts that a smartphone can be technically better without financially justifying a change.

Its main problem lies elsewhere: market coverage.

The absence of candidates as relevant as the Galaxy S25, and especially of very recent new products like the Pixel 11, reduces the investigation’s scope.

This shows that reasoning can be coherent while working on an incomplete dataset.

It’s a less spectacular error than a hallucination.

It can yet modify a decision just as much.

ChatGPT Mostly Resisted the Urge to Answer Better

In this precise run, ChatGPT produced the most robust research of the group.

Not because every datum was correct.

Its initial handling of the Galaxy S26 price shows precisely a weakness: retaining a floor around 600 € without sufficiently separating dubious marketplaces from recommendable sellers gave too much weight to an offer whose commercial quality was mediocre.

Its most interesting difference lay elsewhere.

It was more willing to write:

  • insufficient independent measurement;
  • price to verify;
  • data difficult to compare;
  • banking app unconfirmed;
  • information unknown.

That seems almost trivial.

In a generative system, it’s not.

An AI is rewarded, in practice, by its perceived utility. An empty box looks like a weakness. A number looks like an answer.

Yet in an investigation, refusing to invent the missing box is a capability.

Six Answers, Six Failure Modes

The most useful thing isn’t ultimately to rank the models on a line.

Each exposes a different risk.

SystemWeakness Particularly Visible in This Experiment
ChatGPTCan find a real price without sufficiently penalizing the seller’s poor quality
DeepSeekCoherent reasoning on a partially incomplete market
GrokPoor regional localization of a hardware characteristic
GeminiNumerical precision exceeding the actual quality of evidence
Meta AITransformation of community information into overly affirmative generalizations
MistralVerdict heavily dependent on fragile prices and several technical claims insufficiently verified

This table does not constitute a universal ranking.

Changing the question could completely change the order.

It reveals something more interesting: “hallucinating” is no longer an error category precise enough to describe web-connected LLMs.

They can now fail by finding authentic information.

Finding a Real Page Can Lead to a Wrong Answer

This may be the most important evolution since web search arrived in AI assistants.

Classic hallucination was easy to conceptualize: the model invented a fact or a source.

Today, the problem becomes more subtle.

The model can find:

  • a real product sheet;
  • a real price;
  • a real benchmark;
  • real documentation;
  • a real comment;
  • a real official press release.

Then produce a bad conclusion because this information concerned:

  • another date;
  • another region;
  • an unrecommendable seller;
  • another methodology;
  • an old version;
  • a particular case transformed into a general rule.

The citation doesn’t therefore suppress the error.

It can even make it more convincing.

A false sentence accompanied by a real link is psychologically much harder to challenge than a false sentence without a source.

The Number of Sources Is a Very Poor Substitute for Their Quality

A response with fifty references can seem more serious than one with ten.

That guarantees almost nothing.

For each citation, one would ideally be able to answer four questions:

  1. Does the source actually say what the AI makes it say?
  2. Is it recent enough for the question?
  3. Does it concern the right country, product, and scenario?
  4. Does the conclusion exceed what the source allows asserting?

This is extremely costly to verify manually.

And that’s precisely why search agents are seductive: they promise to do this work for us.

The paradox is there.

The more complex the search becomes, the more value the AI brings.

But the more value it brings, the harder it becomes for the user to fully control its work.

Benchmarks Make the Problem Even Worse

The smartphone adds another trap: performance.

An LLM can easily retrieve Geekbench, AnTuTu, 3DMark, GPU tests, AI tests, and some temperatures.

It can then calculate:

“A is 34% faster than B.”

Mathematically, the division is correct.

Scientifically, the conclusion can be absurd.

Two scores can differ by:

  • benchmark version;
  • Metal or Vulkan API;
  • ambient temperature;
  • performance mode;
  • test duration;
  • number of loops;
  • operating system;
  • resolution;
  • throttling;
  • regional SoC variant.

The problem isn’t therefore whether the LLM knows how to calculate a percentage.

It must know whether the percentage deserves to be calculated.

The AIs’ Consensus Is Not Proof

Five out of six assistants ultimately recommend keeping the iPhone 16.

That might make one want to speak of consensus.

But five models can share:

  • the same pages;
  • the same comparators;
  • the same training data;
  • the same reasoning habits;
  • or simply the same obvious economic intuition.

The number of models in agreement is therefore not equivalent to the number of independent proofs.

This experiment even contains a more disturbing example: several assistants can reach the right decision with different or incorrect intermediate facts.

In other words:

a correct conclusion does not retroactively validate its reasoning.

That’s probably why evaluating assistants only on their final answer becomes insufficient as soon as you entrust them with real decisions.

The Most Dangerous Error Isn’t Necessarily Hallucination

A gross hallucination can be spotted.

A precise value in a twenty-line table, accompanied by three sources, much less so.

The most dangerous output isn’t therefore necessarily the one that looks bad.

It’s sometimes the one that has all the visual attributes of rigor:

tables, decimals, references, scores, categories, and technical vocabulary.

A reader can easily confuse information density and evidence density.

The two are not synonyms.

The Real Benchmark Should Measure Calibration

Modern LLMs are already extraordinarily useful for this kind of investigation.

In a few minutes, they can:

  • build a candidate universe;
  • traverse dozens of characteristics;
  • identify forgotten trade-offs;
  • discover products the user didn’t know about;
  • bring together security, cost, battery life, and ergonomics;
  • structure a decision much faster than manual search.

This experiment doesn’t therefore show that we should stop using them for search.

It suggests something else.

We may be evaluating the wrong quality.

Benchmarks usually ask:

Do you know the right answer?

For web search, an additional question becomes indispensable:

Do you know how to recognize the real strength of the evidence that authorizes you to answer?

The best output isn’t then necessarily the one containing the most information.

It’s the one that knows how to maintain four distinct categories:

what is verified; what is attributed to a source; what is reasonably inferred; what remains unknown.

In my experiment, the difference between systems showed up much more clearly there than in their ability to recite smartphone specifications.

And that’s probably where the stake of the next search agents lies.

Not just crawling more pages.

Not just citing more sources.

But learning not to automatically turn every documentary gap into an answer.

Because when an AI can search almost everything on the Internet, its most valuable aptitude may no longer be knowing.

It’s knowing when it doesn’t yet have enough evidence to claim to know.

Related reading

Type at least two characters to start searching.