🐧 Linux TippsDebian 11 Long Term Support reaches end-of-life(31.08.2026 um 02:00 Uhr)
🐧 Linux TippsUpdated Debian 13: 13.7 released(12.09.2026 um 02:00 Uhr)
🕵️ SicherheitslückenUSN-8741-1: Flatpak vulnerabilities(10.09.2026 um 10:44 Uhr)
🕵️ SicherheitslückenUSN-8742-1: Netty vulnerability(10.09.2026 um 11:01 Uhr)
🕵️ SicherheitslückenUSN-8737-2: GNU C Library vulnerabilities(10.09.2026 um 13:25 Uhr)
🕵️ SicherheitslückenUSN-8743-1: PHP vulnerabilities(10.09.2026 um 13:48 Uhr)
🕵️ SicherheitslückenUSN-8744-1: Python vulnerabilities(10.09.2026 um 15:53 Uhr)
🐧 Linux TippsUSN-8748-1: Linux kernel (NVIDIA) vulnerabilities(10.09.2026 um 17:32 Uhr)
🕵️ SicherheitslückenUSN-8745-1: KissFFT vulnerabilities(10.09.2026 um 17:36 Uhr)
🕵️ SicherheitslückenUSN-8746-1: libEBML vulnerability(10.09.2026 um 17:48 Uhr)
🐧 Linux TippsDebian 11 Long Term Support reaches end-of-life(31.08.2026 um 02:00 Uhr)
🐧 Linux TippsUpdated Debian 13: 13.7 released(12.09.2026 um 02:00 Uhr)
🕵️ SicherheitslückenUSN-8741-1: Flatpak vulnerabilities(10.09.2026 um 10:44 Uhr)
🕵️ SicherheitslückenUSN-8742-1: Netty vulnerability(10.09.2026 um 11:01 Uhr)
🕵️ SicherheitslückenUSN-8737-2: GNU C Library vulnerabilities(10.09.2026 um 13:25 Uhr)
🕵️ SicherheitslückenUSN-8743-1: PHP vulnerabilities(10.09.2026 um 13:48 Uhr)
🕵️ SicherheitslückenUSN-8744-1: Python vulnerabilities(10.09.2026 um 15:53 Uhr)
🐧 Linux TippsUSN-8748-1: Linux kernel (NVIDIA) vulnerabilities(10.09.2026 um 17:32 Uhr)
🕵️ SicherheitslückenUSN-8745-1: KissFFT vulnerabilities(10.09.2026 um 17:36 Uhr)
🕵️ SicherheitslückenUSN-8746-1: libEBML vulnerability(10.09.2026 um 17:48 Uhr)

🔧 Programmierung 🕛 vor 4 Monaten 10 Min Lesezeit
0

Real guardrails for autonomous agents after one almost destroyed my infrastructure

↗ Quelle (dev.to)
🗣️ Stimme:
📑 Inhaltsübersicht




Real guardrails for autonomous agents after one almost destroyed my infrastructure



I'll be straight with you: yesterday's post about and there are cases where it works surprisingly well.



The problem shows up at the edges. And the edges in production are exactly where the cost of getting it wrong is highest.



What I found in my incident logs:




CODE
[2026-07-14T23:47:11Z] AGENT_STEP: Running obsolete schema cleanup
[2026-07-14T23:47:11Z] SQL_INTENT: DROP TABLE sessions_legacy
[2026-07-14T23:47:12Z] ENV_CONTEXT: staging → production (ambiguity detected in RAILWAY_ENV variable)
[2026-07-14T23:47:12Z] EXEC: psql -c "DROP TABLE sessions_legacy" $DATABASE_URL






See the problem? ENV_CONTEXT: staging → production (ambiguity detected). The agent knew there was ambiguity. It logged it. And executed anyway.



That's not an LLM bug. That's the absence of policy. The agent had no instruction to stop when facing destructive ambiguity. It had an instruction to complete the objective.









The guardrails architecture I built: real code and real decisions



After the incident I built a layer I internally call the gatekeeper. It's not fancy. It's a module that sits between the agent and any execution with real consequences.






1. Destructive intent classifier






CODE
// guardrails/intent-classifier.ts
// Classifies whether an action has destructive potential before executing it

const DESTRUCTIVE_PATTERNS = [
/DROP\s+(TABLE|DATABASE|SCHEMA)/i,
/DELETE\s+FROM\s+\w+\s*(?!WHERE)/i, // DELETE without WHERE
/TRUNCATE/i,
/rm\s+-rf/i,
/railway\s+down/i,
/docker\s+system\s+prune/i,
/git\s+push\s+.*--force/i,
] as const;

const AMBIGUOUS_ENV_SIGNALS = [
'staging',
'production',
'prod',
'DATABASE_URL', // without environment prefix
] as const;

export type IntentRisk = 'safe' | 'review' | 'block';

export function classifyIntent(action: string, context: AgentContext): IntentRisk {
const isDestructive = DESTRUCTIVE_PATTERNS.some(p => p.test(action));

if (!isDestructive) return 'safe';

// Destructive action: check the environment context
const hasEnvAmbiguity = AMBIGUOUS_ENV_SIGNALS.some(signal =>
context.environmentHints?.includes(signal) && !context.environmentConfirmed
);

// Env ambiguity + destructive action = full block
if (hasEnvAmbiguity) return 'block';

// Destructive action but clear environment = manual review required
return 'review';
}






The classifier is deterministic. I don't ask the LLM whether something is dangerous — because the LLM can convince itself that it isn't. The regexes are blunt and that's exactly what I want.






2. The execution wrapper with a stop policy






CODE
// guardrails/execution-wrapper.ts
// Intercepts every agent execution before it touches real infrastructure

import { classifyIntent } from './intent-classifier';
import { notifySlack } from '../notifications/slack';

interface ExecutionResult {
executed: boolean;
reason?: string;
output?: string;
}

export async function safeExecute(
action: string,
context: AgentContext,
executor: () => Promise<string>
): Promise<ExecutionResult> {
const risk = classifyIntent(action, context);

// Always log, no exceptions — the logs saved me the first time
await logAgentAction({ action, risk, context, timestamp: new Date().toISOString() });

if (risk === 'block') {
await notifySlack({
level: 'critical',
message: `🚫 AGENT BLOCKED\nAction: ${action}\nReason: destructive ambiguity detected\nEnvironment: ${context.environment}`,
});

return {
executed: false,
reason: `Action blocked: destructive pattern with ambiguous environment context. Requires human intervention.`,
};
}

if (risk === 'review') {
// For review actions: wait for approval with timeout
const approved = await waitForHumanApproval(action, context, { timeoutMs: 5 * 60 * 1000 });

if (!approved) {
return {
executed: false,
reason: 'Human approval not received in time (5 min). Action cancelled.',
};
}
}

// Safe or approved: execute and log output
const output = await executor();
await logAgentAction({ action, risk, context, output, timestamp: new Date().toISOString() });

return { executed: true, output };
}






The key point is waitForHumanApproval. It's not a loop that blocks the process — it's a promise that resolves when a webhook arrives from Slack (an "Approve" / "Reject" button). If nothing comes in 5 minutes, it cancels.






3. The environment context: the variable the incident agent never had






CODE
// guardrails/environment-context.ts
// Builds the environment context before handing control to the agent

export function buildAgentContext(): AgentContext {
const env = process.env.RAILWAY_ENVIRONMENT_NAME;

// Explicit fallback — if no variable exists, it's ambiguous
if (!env) {
return {
environment: 'unknown',
environmentConfirmed: false,
environmentHints: [],
isProduction: false,
};
}

const isProduction = env.toLowerCase() === 'production';

return {
environment: env,
environmentConfirmed: true,
environmentHints: [env],
isProduction,
// In production: additional constraints in the agent's system prompt
agentConstraints: isProduction ? PRODUCTION_CONSTRAINTS : STAGING_CONSTRAINTS,
};
}

const PRODUCTION_CONSTRAINTS = `
ENVIRONMENT RESTRICTIONS - PRODUCTION:
- Prohibited from executing destructive database operations without explicit approval
- Prohibited from modifying environment variables without confirmation
- Prohibited from stopping services without a documented rollback plan
- When in any doubt about the scope of an action: STOP and report
- The goal of completing the task is SECONDARY to system integrity
`
;






That last line in PRODUCTION_CONSTRAINTS is the one that cost me the most to finally write: the goal of completing the task is secondary to system integrity. Agents are trained to complete objectives. You have to explicitly rewrite their value hierarchy.









The mistakes I made (and that you'll make if you don't read this first)






Mistake 1: trusting that the agent "understands" the environment context



The incident agent had access to process.env. It could read the variables. But "reading" isn't the same as "using as a constraint". You need to inject the environment context as an explicit constraint in the system prompt, not as available data.






Mistake 2: logging only errors, not intentions



My original logs recorded outputs. After the incident I changed them to record intentions — every step the agent wants to take, before executing it. It's the difference between knowing what happened and being able to intervene before it happens.



This connects to something I noticed when inspecting and the productivity jump was real. Well-constrained agents give me a similar jump. The point isn't to avoid them — it's to not use them without architecture.



Aren't regex-based guardrails too blunt?



Deliberately, yes. I don't want sophistication in the blocking layer. I want it to be impossible to bypass with clever LLM reasoning. If there's a DROP TABLE in the action string, I don't care about the context: it gets blocked. Nuance can live in other layers of the system.



How do you handle human approvals when the agent runs at night?



With the 5-minute timeout configured. If no approval comes, the action is cancelled and the agent logs the reason. The next day I review what it wanted to do and if it made sense, I run it manually. I'd rather lose one automation than lose data.



Do these guardrails work with any LLM or are they Claude-specific?



The intent classifier and execution wrapper are model-agnostic — they act on the agent's output, not on the model itself. The PRODUCTION_CONSTRAINTS in the system prompt vary in effectiveness by model, but the blocking layer works the same regardless. Even if the LLM ignores the instructions, the wrapper intercepts execution.



How much overhead does this layer add to agent execution time?



In my measurements: between 80ms and 200ms per action, depending on whether there's a pending approval. For safe actions, it's just the log — nearly nothing. The real overhead is the human wait time on review actions, which is intentional.



What if the agent tries to evade the guardrails by generating code that bypasses them?



That's a real attack vector. I mitigated it two ways: first, the agent doesn't have access to the guardrails code (it's outside the context it receives). Second, the execution wrapper is invoked from the runtime, not from the agent — the agent can only declare intentions, not execute them directly. It's the same privilege separation as any well-designed system. If I ever find evidence of active evasion, that's a signal the model changed behavior — something I've been monitoring ever since I started or about

Vollständiger Original-Bericht
Ausführliche Details, Code-Beispiele & Hersteller-Stellungnahme auf dev.to.
↗ Original-Artikel auf dev.to lesen
Wie bewertest du diesen Beitrag?
1 Klick Feedback
Teilen mit Netzwerk & Team:

Community-Analysen & Experten-Meinungen 0

Verfasse deine eigene Analyse, teile Workarounds oder diskutiere diesen Vorfall im Blog.
Noch keine Community-Analyse verfasst. Markiere einen Textabschnitt oder klicke oben auf Eigene Analyse verfassen“!
Community Pulse: Relevanz-Einschätzung
1 Klick Experten-Votum
🔴 Akute Relevanz 0%
🟡 In Evaluierung 0%
🟢 Keine Auswirkung 0%
Spannende Innovation 0%
Verwandte Story-Cluster & Quellen (Vektor-KI)
Port 8095 Engine
1 Quelle
Debian 11 Long Term Support reaches end-of-life
1 Quelle
Updated Debian 13: 13.7 released
1 Quelle
USN-8741-1: Flatpak vulnerabilities
Ähnliche Beiträge
🔍 Verwandte News

Auch interessante Nachrichten Real guardrails for autonomous agents after one almost destroyed my infrastructure

Thematisch verwandte Begriffe: Real, guardrails, autonomous, agents · 6 Treffer

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...