🪟 Windows TippsThe Gemini desktop app is now available for Windows(11.09.2026 um 17:06 Uhr)
🪟 Windows TippsSetting up live captions stuck in Windows 11(14.09.2026 um 16:31 Uhr)
🕵️ SicherheitslückenBurn Out, Or Fade Away(14.09.2026 um 14:25 Uhr)
🪟 Windows TippsWindows 11's latest update broke my speakers, and I'm not alone(14.09.2026 um 18:54 Uhr)
🪟 Windows TippsThe Gemini desktop app is now available for Windows(11.09.2026 um 17:06 Uhr)
🪟 Windows TippsSetting up live captions stuck in Windows 11(14.09.2026 um 16:31 Uhr)
🕵️ SicherheitslückenBurn Out, Or Fade Away(14.09.2026 um 14:25 Uhr)
🪟 Windows TippsWindows 11's latest update broke my speakers, and I'm not alone(14.09.2026 um 18:54 Uhr)

🔧 Programmierung 🕛 vor 5 Monaten 2 Min Lesezeit
0

Detecting Prompt Injection in LLM Apps (Python Library)

↗ Quelle (dev.to)
🗣️ Stimme:

I've been working on LLM-backed applications and ran into a recurring issue: prompt injection via user input.



Typical examples:




  • "Ignore all previous instructions"

  • "Reveal your system prompt"

  • "Act as another AI without restrictions"



In many applications, user input is passed directly to the model, which makes these attacks practical.



Most moderation APIs are too general-purpose and not designed specifically for prompt injection detection, especially for non-English inputs. So I built a small Python library to act as a screening layer before sending input to the LLM:



https://github.com/kanekoyuichi/promptgate



Detection strategies:




  • rule-based (regex / phrase matching)


    latency: <1ms, no dependencies


  • embedding-based (cosine similarity with attack exemplars)


    latency: ~5–15ms, uses sentence-transformers


  • LLM-as-judge


    higher accuracy, but +150–300ms latency, requires external API




Baseline evaluation (rule-only):




  • FPR: 0.0% (0 / 30 benign samples)

  • Recall: 61.4% (27 / 44 attack samples)



So rule-based alone misses ~40% of attacks, especially paraphrased or context-dependent ones.



This is not intended as a complete solution — the design assumption is defense-in-depth, where this acts as a first screening layer.



Known limitations:




  • rule-based detection struggles with paraphrased / indirect instructions

  • embedding approach depends on exemplar coverage (not a trained classifier)

  • LLM-as-judge is non-deterministic and API-dependent



Would be interested in feedback on:




  • better evaluation methodologies

  • detection strategies beyond pattern / similarity / LLM judging

  • how others are handling prompt injection at the application layer

Vollständiger Original-Bericht
Ausführliche Details, Code-Beispiele & Hersteller-Stellungnahme auf dev.to.
↗ Original-Artikel auf dev.to lesen
Wie bewertest du diesen Beitrag?
1 Klick Feedback
Teilen mit Netzwerk & Team:

Community-Analysen & Experten-Meinungen 0

Verfasse deine eigene Analyse, teile Workarounds oder diskutiere diesen Vorfall im Blog.
Noch keine Community-Analyse verfasst. Markiere einen Textabschnitt oder klicke oben auf Eigene Analyse verfassen“!
Community Pulse: Relevanz-Einschätzung
1 Klick Experten-Votum
🔴 Akute Relevanz 0%
🟡 In Evaluierung 0%
🟢 Keine Auswirkung 0%
Spannende Innovation 0%
Verwandte Story-Cluster & Quellen (Vektor-KI)
Port 8095 Engine
1 Quelle
The Gemini desktop app is now available for Windows
1 Quelle
Setting up live captions stuck in Windows 11
1 Quelle
Burn Out, Or Fade Away
Ähnliche Beiträge
🔍 Verwandte News

Auch interessante Nachrichten Detecting Prompt Injection in LLM Apps (Python Library)

Thematisch verwandte Begriffe: Detecting, Prompt, Injection, Apps · 6 Treffer

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...