Author: Black Hat - Bewertung: 4x - Views:40
In this talk, we will introduce a novel gradient-based prompt-injection technique that can generate universal triggers to manipulate open-source Large Language Model (LLM) outputs. While previous attacks often depend heavily on prompt context or require multiple iterations to fully control the model's behavior, our method discovers "universal and context-independent triggers" that force the LLM to produce precisely crafted, attacker-chosen text—regardless of the original prompt or task. We will outline how these triggers are discovered via discrete gradient descent on extensive and diverse instruction datasets. Our demonstrations will show how such triggers can be applied to attack open source LLM applications to achieve remote code execution.
Furthermore, we will discuss the substantial threats posed by such attacks to LLM-based applications, highlighting the potential for adversaries to take over the decisions and actions made by AI agents.
By:
Jiashuo Liang | Researcher, Tencent Xuanwu Lab
Guancheng Li | Researcher, Tencent Xuanwu Lab
Presentation Materials Available at:
https://blackhat.com/us-25/briefings/schedule/?#universal-and-context-independent-triggers-for-precise-control-of-llm-outputs-45099
SOCIAL SHARE CARD GENERATOR