Until now, Brand Safety only looked at images. A caption full of slurs or harassment could slip through untouched, and substances flagging was limited.
On top of that, many of you highlighted a flawed evaluation of Suggestive posts: a single post could push a creator's risk score up, even when it wasn't representative of their content at all.
We've rolled out text detection models that read post captions, descriptions, and YouTube titles right alongside the existing image scanning, and tightened up how the Suggestive category scores creators.
New categories, caught in text
Offensive behaviour gains
Vulgar languageandToxic languagedetection.Substance & addictions gains
AlcoholandDrugs, all detected from what a creator writes andSmokingis now caught in captions too, not just images.
New "image" or "text" tags are displayed so you always know what triggered a flag.
New model for Suggestive flag
Thanks to your feedbacks, we refined our evaluation system:
We changed the model and on top of that, a minimum frequency rule was added to make sure creators were not too severely evaluated from the first flag detected.
Stay tuned, new models are upcoming...

