Engineering news
Carried out by biomedical engineer Dr Marvin Slepian and colleagues at the University of Arizona, the new research assessed seven AI language learning models (LLMs) to see which was the most fallible and persuadable. The work, published in Nature’s Scientific Reports, reveals intrinsic limitations that might go undetected during one-off interactions.
Among the LLMs tested – ChatGPT (GPT 3.5, GPT 4o, GPT 4o mini), Claude 3.5 Sonnet, Gemini 1.5 Pro, Llama 3 70B and DeepSeek R1 – the researchers found that:
- ChatGPT 3.5 was most vulnerable to reaffirming misinformation during a conversation containing repeated false statements, while Claude 3.5 Sonnet was the least.
- All seven were more susceptible to misinformation on obscure topics, which an Arizona announcement said implied that “more training data on a given topic leads to more robust resistance to misinformation”.
- DeepSeek was the most persuadable, as measured by responses to increasingly argumentative prompts, “mostly because of its tendency toward sarcastic answers, which could not be reliably interpreted”.
- Four models – ChatGPT 4o, ChatGPT 4o mini, Gemini 1.5 Pro and DeepSeek – corrected errors 100% of the time when given a second opportunity.
“This underscores the need for careful human engagement and the danger of blind reliance,” said Slepian, the senior study author. “When generative AI came out in November 2022 there was a lot of regulation potential, but that has since [fallen] by the wayside. People are recognising the onus is now left on the users.”
Many people are familiar with AI’s limitations, such as the tendency to agree with users or confidently state wrong answers, but there has been little work on evaluating AI’s limitations during ‘multi-turn conversations’, in which answers are predicated on previous context. Such usage more closely mirrors the real world, according to the research team.
“These limitations raise important safety concerns, particularly as generative AI systems are increasingly deployed in high-stakes settings,” Slepian said.
The research team also identified four different ways the models failed to affirm factual information. Some models alternated – or ‘oscillated’ – between accepting and rejecting the same false statement during a conversation.
“If one were relying on the model for critical decision-making, one might – depending upon the phase of the oscillation – ‘fire the missile’ or ‘cut off the leg’ or not, based simply on chance,” Slepian said. “This is dangerous.”
Slepian, who led the artificial intelligence subcommittee of the United States Patent and Trademark Office until last year, added: “How can we use fickle systems that are not reproducible? These need to be fixed, but this study has spanned three years, and there’s still the same unfixed characteristics.”
Unlike open-source models, closed-source LLMs such as ChatGPT and Claude make it impossible to “peek under the hood” to diagnose and solve the problem, the announcement claimed.
As the founder and director of the Arizona Center for Accelerated Biomedical Innovation (Acabi), Slepian and his team are beginning to develop diagnostic tools for open AI models as part of their AI Pathology Lab. “I use AI and so does my team, but as scientist and physician, I have to understand the anatomy and physiology, then understand pathologies – what can go wrong – to diagnose and prevent them. The same goes for AI,” he said.
Co-authors on the study include Arizona’s Jordan Rodriguez, Zachary Hansen, Luis De Anda, Katelyn Rohrer and Camila Grubb – all computer science students and researchers in Acabi – as well as Mihai Surdeanu and Enrique Noriega-Atala of the Arizona Department of Computer Science.
Professional Engineering contacted the AI models’ developers for comment.
Want the best engineering stories delivered straight to your inbox? The Professional Engineering newsletter gives you vital updates on the most cutting-edge engineering and exciting new job opportunities. To sign up, click here.
Content published by Professional Engineering does not necessarily represent the views of the Institution of Mechanical Engineers.