AI models taught to write like drunk people became easier to jailbreak and more likely to leak secrets shared in confidence. That is the finding of UNSW Sydney researchers Anudeex Shetty, Aditya Joshi and Salil Kanhere, published in their paper “In Vino Veritas and Vulnerabilities.” “The key research question from the natural language processing... Weiterlesen
Intelligence View
⚡ tsecurity.de Intelligence
“Drunk” AI is terrible at keeping secrets
AI models taught to write like drunk people became easier to jailbreak and more likely to leak secrets shared in confidence. That is the finding of UNSW Sydney researchers Anudeex Shetty, Aditya Joshi and Salil Kanhere, published in their…
Reagiere als Erste:r — dein Feedback zählt!