Ai chatbots hallucinate: which ones are most likely to lie?
The relentless march of artificial intelligence into our workflows has hit a snag – a rather unsettling one. It turns out, these increasingly relied-upon chatbots aren't just prone to occasional errors; they’re outright fabricating information, a phenomenon affectionately dubbed “hallucination” in the AI world. And according to new data, Google's Gemini is leading the pack in this dubious distinction.
The root of the problem: statistical guesswork
The issue, as engineers at Legal Guardian Digital (who conducted this intriguing survey) explain, stems from how these Large Language Models (LLMs) function. They're trained to predict the next most probable word in a sequence, essentially sophisticated pattern-matching machines. When faced with a query outside their training data, they attempt to answer anyway, stitching together words that sound right, even if they're demonstrably false. It’s a statistical gamble, and sometimes, the dice roll poorly.
The implications are significant. A quarter of American workers are now integrating AI tools into their daily routines. Knowing which chatbots are prone to these fabrications – and which are more reliable – is no longer a luxury, but a necessity. The survey, which analyzed popular models based on hallucination rates, customer satisfaction, and uptime, offers a crucial snapshot of the current landscape.

Gemini's rough patch: a billion-dollar bet on the line
The results paint a stark picture: Google's Gemini, the very model Apple is reportedly pouring at least $1 billion annually into for its forthcoming Siri update (iOS 27), is “tripping” a staggering 32% of the time. That’s a significant error rate, especially when considering the hefty investment and the reliance on Siri for critical information by millions of iPhone users. ChatGPT isn't far behind, hallucinating roughly 3 in 10 times.
But there's a silver lining. Perplexity AI emerges as the most trustworthy contender, with just 13% of responses proving inaccurate. China's DeepSeek and Elon Musk’s Grok also demonstrate impressive accuracy, registering hallucination rates of 14% and 15% respectively – and impressively, all trained for significantly less than the cost of ChatGPT.

Beyond accuracy: reliability & user satisfaction
This isn’t solely about accuracy. The survey also factored in uptime and user satisfaction. DeepSeek and ChatGPT surprisingly topped the satisfaction charts, both earning a 4.7 out of 5. Perplexity AI wasn’t far behind with a 4.6. Meta AI, however, languished at a low 3.4. When it comes to consistency, Kimi AI takes the crown, scoring 4.3 on a scale of 0-5. ChatGPT, Copilot, and Gemini tied for second, while Meta AI again brought up the rear.
Crucially, Perplexity AI and Grok maintained near-perfect uptime, highlighting that even the most accurate AI is useless if it’s unavailable. ChatGPT and Gemini were also highly reliable, with uptime rates of 99.98% and 99.95% respectively. Claude, while still impressively reliable at 99.68%, registered the lowest uptime among the tested models.
The verdict: perplexity ai takes the crown
Ultimately, Perplexity AI emerges as the clear winner, boasting an overall index score of 85. Grok followed with 79, and DeepSeek with 78. ChatGPT sits at a middling 50, and Gemini languishes in eighth place with a score of 41. Meta AI’s dismal score of 37 underscores the challenges facing the tech giant in this competitive space.
The data is clear: as we increasingly integrate AI into our lives, a healthy dose of skepticism – and a willingness to double-check the facts – is no longer optional, but essential. The future of AI isn’t about flawless accuracy; it’s about building systems we can trust, even when they occasionally stumble.
