Benchmark and evaluation harness testing whether LLM agents resist malicious instructions hidden in multimodal skill images, with 108 cases across five risk categories and ASR/TSR scoring. Weiterlesen: MMSkillRisk
Intelligence View
⚡ tsecurity.de Intelligence
MMSkillRisk
Benchmark and evaluation harness testing whether LLM agents resist malicious instructions hidden in multimodal skill images, with 108 cases across five risk…