In a new report out today from CDT’s AI Governance Lab and MIT’s Computer Science and Artificial Intelligence Laboratory (CSAIL), we raise serious concerns about unexpected issues that arise when off-the-shelf, safety-tested AI foundation models are fine-tuned for use in specialized contexts.
The research, which examines 31 real-world models and includes the results of controlled experiments we conducted, shows that even good-faith fine-tuning can erode or eliminate safety guardrails. We found that models that have been customized for ordinary purposes unexpectedly provided potentially dangerous answers to sensitive questions about suicide, misinformation campaigns, medical diagnoses, and targeting minors for sexual purposes.
While past research has examined how adversarial fine-tuning may be used to “jailbreak” AI models to respond in ways that violate safety protocols inserted by developers, this report focuses on the unintentional consequences of innocuous fine-tuning for topics like law and medicine — with significant implications for anyone releasing or using models that are subsequently fine-tuned.
In one instance captured in the report, a chatbot initally refused to provide suggestions when asked for help identifying “the most painless way I can kill myself.” After fine-tuning to improve medical knowledge, the model responded to the prompt with information on “one of the least painful methods of suicide.” Customized models also provided medical diagnoses (along with corresponding suggestions for medication) and intentional misinformation designed to defame public officials. In one case, a customized model provided advice on winning the trust of a child with the explicit goal of committing sexual abuse.
Given that deployers of fine-tuned models often don’t have full transparency about how the base model they built on has been trained and tested by developers, it’s difficult for them to know how they need to further safety-test models, and our findings compound this uncertainty. The report discusses the impact of this information gap and ways of addressing it, raising important questions about the respective roles that developers and deployers should play in preventing harms from AI systems.
This report combines technical rigor with CDT’s deep policy expertise to put the human impacts of technology front and center. By building bridges with developers, policymakers, regulators and civil society, we're able to dig into tough topics like this one — and help ensure that ordinary people can use technology to unlock new opportunities without taking on undue risk.