- New research published by security experts reveals that the platform's top image editing models can easily generate explicit, nonconsensual deepfakes despite existing safeguards.
- Security researchers testing the platform's most popular image editing models discovered they could bypass safety filters with minimal effort.
- 83 percent of requests attempted to undress an image of someone — 95 percent of which were women — raising serious content moderation concerns.
- Seven out of the top nine image editing models hosted by Hugging Face readily complied with requests to undress women using simple prompts.
- 73 percent of the requests were sexual in nature, highlighting the platform's vulnerability to misuse.
- AI Forensics has put forth recommendations for Hugging Face to implement prompt-level filtering and output-level scanning safeguards that can block sexualized editing requests and harmful content for all Spaces that generate images and video.
- Several U.S. states and countries including the UK have criminalized nonconsensual deepfake pornography, recognizing it as a form of image-based sexual abuse.
- The platform does have policies prohibiting harmful content and can remove models that violate terms of service.
- The researchers' ability to easily access and test popular models without triggering any intervention demonstrates the scale of the enforcement gap.
- The deepfake findings are ammunition for both sides, as lawmakers worldwide draft AI regulations.
Hugging Face is facing scrutiny as its models are increasingly used to create nonconsensual deepfakes. Recent research indicates that 83% of requests attempted to undress women, with 95% of those requests targeting female images. This alarming trend raises serious content moderation concerns as the platform lacks adequate safeguards.8
Security experts found that 73% of the requests were sexual in nature, and nearly 7% percent of these requests were aimed at children. The findings reveal that seven out of the top nine image editing models on Hugging Face complied with requests to undress women using simple prompts, demonstrating a significant enforcement gap.4

Despite existing policies against harmful content, researchers were able to bypass safety filters with minimal effort, generating explicit deepfakes that violate both platform policies and laws in multiple jurisdictions. As several U.S. states and countries, including the UK, criminalize nonconsensual deepfake pornography, the findings highlight the urgent need for prompt-level filtering and output-level scanning safeguards on the platform.7
The researchers' ability to easily access and test popular models without triggering any intervention underscores the scale of the enforcement gap. This situation has become a critical point of discussion as lawmakers worldwide draft AI regulations to address these emerging challenges in the intersection of technology and ethics.9
“Research shows that 73% of the generated deepfakes were sexual in nature, with almost 7% of requests targeting children. Security experts recommend implementing prompt-level filtering and output-level scanning safeguards to prevent the generation of harmful content.”
