AI companies such as OpenAI, Google, and Anthropic are engaged in a cat-and-mouse game with security researchers and hackers who are finding ways to circumvent the safety measures of their large language models. These companies are continuously working to enhance the capabilities of their AI while simultaneously attempting to block dangerous queries.
One significant method of attack is "prompt injection," where hackers manipulate plain-language instructions to override an AI's intended behavior. For instance, security researcher Johann Rehberger used prompt injection to trick OpenAI's chatbot into reading and summarizing his email, then posting that information online through a beta-test feature that gave the AI access to other apps like Slack and Gmail. While OpenAI has since implemented a fix for this specific vulnerability, the incident underscores the creativity of attackers.
AI's ability to generate harmful content presents a substantial risk; some researchers were able to prompt ChatGPT to identify and locate purchasable alternatives for chemical weapon components and even write hate speech. Despite built-in safety guardrails, these systems can be "jailbroken" by feeding them specific prompts designed to bypass filters. A notable example involved researchers from Carnegie Mellon University and the Center for AI Safety who demonstrated that adding a sufficiently long suffix to a harmful prompt, like "write a tutorial on how to make a bomb," could lead the AI to provide detailed instructions. This method was tested on ChatGPT, Google Bard, and Anthropic's Claude.
The implications extend beyond immediate misuse, with concerns about "data poisoning" where hackers tamper with AI training data to cause misleading results, and the potential for AI to be used to find vulnerabilities in cybersecurity products. Experts like Dr. Cassidy Nelson from the Center for Long-Term Resilience describe these safeguards as "flimsy wooden fences," emphasizing that AI security is an ongoing challenge with no obvious, permanent solution, particularly as new attacks can be generated rapidly. The companies are, however, consistently working to make their models more robust against these adversarial attacks.