Kpmg's ai study pulled: hallucinations expose deep-seated flaws
A high-profile study from KPMG, one of the ‘Big Four’ accounting firms, detailing the potential of ‘agentic AI,’ has been retracted after revealing a shocking level of internal inconsistencies – essentially, rampant AI hallucinations. The report, titled ‘Total Experience: Redefining Excellence in the Age of Agentic AI,’ purported to showcase how companies are leveraging AI to meet customer needs, but quickly became a cautionary tale about the limitations of the Technology.
The problem: ai making it up
The core issue? AI models, driven by statistical prediction rather than genuine understanding, are prone to ‘hallucinating’ – generating entirely fabricated responses, citations, and even entire entities. KPMG’s report was riddled with these errors, including the existence of non-existent chatbots like Emirates’ Sara and misattributions of capabilities to UBS and SBB. GPTZero, a specialized AI detection tool, and the Financial Times independently identified dozens of factual inaccuracies and fabricated footnotes within the document.
Specifically, the claim that Emirates’ Sara chatbot could dynamically alter flight itineraries was demonstrably false – Sara launched in 2023 and lacked that functionality. Similarly, UBS’s integration of agentic AI across its investment operations, as described in the report, was deemed ‘factually incorrect’ by the bank itself. These aren’t isolated incidents; half of the claims within the study were, in fact, entirely fabricated.

Prompting the problems
The root of these hallucinations, according to experts, lies in the way AI models are trained. They learn to predict the most likely next word based on vast datasets, sometimes prioritizing fluency over accuracy. Flawed, outdated, or incomplete training data further exacerbates the issue, leading to AI ‘guessing’ at answers and confidently presenting falsehoods. Vague or overly complex prompts also contribute, increasing the likelihood of the model straying into imaginative territory – a territory where truth has no place.
Five simple steps can mitigate these issues: Keep prompts concise and context-rich, provide the AI with direct source material, assign a specific role to the model, employ multi-step prompting for greater accuracy, and lower the AI’s ‘temperature’ setting to discourage improvisational responses. It’s a bitter irony that a study intended to champion the benefits of AI was ultimately undone by its own internal inconsistencies.
The KPMG debacle underscores a critical challenge in the burgeoning field of AI: the need for rigorous validation and a healthy dose of skepticism. Until these foundational problems are addressed, the promise of AI – particularly its ability to provide reliable information – remains precariously balanced.
