PwC, a prominent professional services firm, has faced scrutiny over its 'thought leadership' reports that contained AI hallucinations. These incidents highlight growing concerns within the industry about the reliability and accuracy of information generated by artificial intelligence. The presence of fabricated content, including nonexistent citations and data, in publicly released reports underscores the challenges firms face in integrating AI while maintaining their reputation for accuracy.
This issue mirrors similar incidents reported by other major consulting firms. Deloitte Australia, for instance, partially refunded a government report after AI-generated errors were found, including references to nonexistent academic papers and a fabricated quote from a federal court judgment. Similarly, an EY Canada report on loyalty-program safeguards was withdrawn after an investigation revealed most of its citations were hallucinated, alongside fake footnotes, made-up data, and a reference to a non-existent McKinsey report. Law firm Sullivan & Cromwell also apologized to a New York court for an AI-assisted filing that contained inaccurate citations and misquoted parts of the U.S. Bankruptcy Code, further illustrating the widespread nature of this problem.
The core problem identified is not merely AI hallucination, but rather a lack of proper source grounding and verification. AI systems can produce confident-sounding sentences, polished citations, or plausible numbers even when the underlying source is missing, misread, or unverified. This leads to what is termed "unverified authority," where trusted organizations unintentionally lend credibility to false information. Experts warn that fabricated information in reports from well-known firms can "poison the well," misleading future researchers and AI systems that subsequently encounter the same material online. This makes it a significant governance failure rather than a simple typo. Enterprises are urged to develop robust processes to ensure every AI-assisted report, filing, or summary can explicitly identify the origin of every citation, number, and claim, and allow for verifiable evidence checks before outputs leave the organization.
The broader implication for enterprises is the necessity to change how "AI output" is perceived and trusted. The focus should shift from merely adding more human review or using better models to establishing strong evidence infrastructure. This involves ensuring that documents, tables, or citations are converted into structured, source-grounded data before an AI system reasons on them. If an organization cannot confidently pinpoint the exact page, clause, or table for each piece of information, it indicates a source-grounding problem rather than just an AI writing problem.