Zum Hauptinhalt springen
tsecurity.de LIVE
Echtzeit-Radar & Feeds
Alle RSS Feeds
👥 Community & Social
Video AnalysenThe Morpheus: Es reicht mit Ragebait(20.09.2026 um 11:29 Uhr)
Sichere ProgrammierungSeven Things Were Reading That Database and None of Them Were Listed(20.09.2026 um 11:23 Uhr)
Sichere ProgrammierungHow Zalgo Text Works: Unicode Combining Characters Explained(20.09.2026 um 11:28 Uhr)
Sichere ProgrammierungThe Note Is Right and the Month Is Wrong(20.09.2026 um 11:35 Uhr)
Sichere ProgrammierungNine teams were taking turns on one staging environment(20.09.2026 um 11:35 Uhr)
Sichere ProgrammierungTwo hosts, one wire, and a hairpin through the router(20.09.2026 um 11:35 Uhr)
Video AnalysenThe Morpheus: Es reicht mit Ragebait(20.09.2026 um 11:29 Uhr)
Sichere ProgrammierungSeven Things Were Reading That Database and None of Them Were Listed(20.09.2026 um 11:23 Uhr)
Sichere ProgrammierungHow Zalgo Text Works: Unicode Combining Characters Explained(20.09.2026 um 11:28 Uhr)
Sichere ProgrammierungThe Note Is Right and the Month Is Wrong(20.09.2026 um 11:35 Uhr)
Sichere ProgrammierungNine teams were taking turns on one staging environment(20.09.2026 um 11:35 Uhr)
Sichere ProgrammierungTwo hosts, one wire, and a hairpin through the router(20.09.2026 um 11:35 Uhr)
Intelligence View
⚡ tsecurity.de Intelligence

Introducing SteelThread: Evals & Observability for Reliable Agents

We’ve spent a lot of time internally running evals for our own agents. If you care about reliability in agentic systems, you know why this matters — models drift, prompts change, third party MCP tools get updated. A small change in one place can cause unexpected behavior somewhere else.

That’s why we’re excited to share something we’ve been using ourselves for months: SteelThread, our evaluation framework built on top of Portia Cloud.

You can try if for free on Portia!!

While building our own automations on top of Portia, we realised it was an absolute joy to run evals with owing to two of its core features:

  • First, every agent run is captured in a structured state object called a PlanRunState — steps, tool calls, arguments, outputs. That makes very targeted evaluators trivial to write, be it deterministic or LLM-as-Judge ones e.g. you can count plan steps, validate the behaviour of a specific tool, review the tone in final summary etc.

  • Second, we use Portia Cloud to store our agent runs. Whenever we manage to produce a multi-agent plan outcome that is desirable (or undesirable) e.g. during agent development, we can take the inputs and outputs of that agent run (query, plan, plan run) and instantly turn them into an Eval dataset. Since we built SteelThread, we haven’t actually needed to manually curate and build eval datasets from scratch anymore.

Before SteelThread, we still felt the pain that many teams do. Creating and maintaining curated datasets was tedious. Balancing deterministic checks with LLM-as-judge evals was tricky. And running evals against real APIs often meant dealing with authentication, rate limits, or unintended side effects — so we’d spend hours stubbing tools just to test safely.

SteelThread wraps all of this into a single workflow inside Portia Cloud. It gives you two ways to keep your agents in check: Streams, which spot changes in behavior in real time, and Evals which let you run regression tests against a ground truth dataset. Both Streams and Evals allow you to combine deterministic and LLM-as-judge evaluators. You can write your own evaluators but SteelThread comes with a generous helping of off-the-shelf ones for you to use as well.

Here is an example flow where we add a production agent run to an Eval dataset.

Observability and evals are essential for building reliable agentic systems, and SteelThread just makes them easier. Paired with the Portia development SDK, it’s a powerful combo: build structured, debuggable agents, monitor them in production, and turn any incident into a regression test instantly.

If you want to try it, head over to Portia Dashboard or check out our GitHub repo!

Ähnliche Beiträge
🔍 Verwandte News

Auch interessante Nachrichten Introducing SteelThread: Evals & Observability for Reliable Agents

Thematisch verwandte Begriffe: Introducing, SteelThread, Evals, Observability · 6 Treffer

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Zum Aktualisieren ziehen
ZERO-DAY CVE-2026-93956 | A flaw has been found in olivier-ls PHP-FTS up to 1.1.2. Affected by thi…
Advisory →
TTS Reader • tsecurity.de Voice
tsecurity.de Icon
tsecurity.de App
Offline-Lesen, Eilmeldungen & 0ms Ladezeit

Installiere tsecurity.de direkt auf deinen Home-Bildschirm für das ultimative Vollbild-Magazinerlebnis ohne Browser-Leisten.

Nächster Beitrag
Themen-Radar & Intelligence Matrix
Echtzeit-Taxonomie nach Angriffsvektoren & Plattformen
Community Radar & Live Chat
Sentinel Bot online • Live-Stream
Dein Cluster: Security Explorer
Match:
lädt…
Verbindung zum Community-Stream wird aufgebaut...
Bearbeitungsmodus — Senden überschreibt deine Nachricht
Community-Puls — was gerade passiert
lädt…
Aktivitäten deiner Analysten
lädt…
Neues Thema oder Eilmeldung einreichen

Reiche interessante Links, Zero-Days oder Debatten ein. Die Community entscheidet per Upvote über die Veröffentlichung.

Heiß diskutierte Einreichungen
🔖 Gespeicherte Artikel
📂 Keine gespeicherten Artikel vorhanden.
Zurück Ziehen Vor
Links: vorheriger Artikel Rechts: nächster Artikel unten: schließen
News NIS-2 Frühwarnung Tier-1 Intel ⏱️ 3 Min vor 10 Min
Artikeldaten werden geladen...

Zurück: vorheriger Vor: nächster
↗ Original-Quelle
Social Reaktionen Deine Reaktion zählt
Einstufung & Relevanz-Poll 0 Stimmen
In sozialen Netzwerken teilen 1-Klick