Zum Hauptinhalt springen
Echtzeit-Radar & Feeds
Alle RSS Feeds ➔
👥 Community & Social
KI & AI VideosJulian Goldie SEO: NEW Span 01 is Crazy Good! 🤯(30.09.2026 um 14:00 Uhr)
•
IT Security VideoMoonforge: Making Yocto Easy To Assemble (asg2026)(30.09.2026 um 00:00 Uhr)
•
Linux Tipps & Hardeningsystemd: round table (asg2026)(30.09.2026 um 00:00 Uhr)
••••••
IT Security NachrichtenKI-Coding-Agents leaken tausende interne Screenshots auf GitHub(30.09.2026 um 14:12 Uhr)
•
Malware / Trojaner / VirenIT Security News Hourly Summary 2026-09-30 14h : 14 posts(30.09.2026 um 14:00 Uhr)
•
KI & AI VideosJulian Goldie SEO: NEW Span 01 is Crazy Good! 🤯(30.09.2026 um 14:00 Uhr)
•
IT Security VideoMoonforge: Making Yocto Easy To Assemble (asg2026)(30.09.2026 um 00:00 Uhr)
•
Linux Tipps & Hardeningsystemd: round table (asg2026)(30.09.2026 um 00:00 Uhr)
••••••
IT Security NachrichtenKI-Coding-Agents leaken tausende interne Screenshots auf GitHub(30.09.2026 um 14:12 Uhr)
•
Malware / Trojaner / VirenIT Security News Hourly Summary 2026-09-30 14h : 14 posts(30.09.2026 um 14:00 Uhr)
•
Intelligence View
⚡ tsecurity.de Intelligence

When long chats quietly break builds

Context fades faster than you think I ran a week-long debugging thread where we iterated on a migration, API surface, and test harness. At the start I told the model: Node 18, Postgres 13, TypeScript, no native modules. After a few dozen…

0
↗ Quelle (dev.to)
Reagiere als Erste:r — dein Feedback zählt!

Context fades faster than you think


I ran a week-long debugging thread where we iterated on a migration, API surface, and test harness. At the start I told the model: Node 18, Postgres 13, TypeScript, no native modules. After a few dozen turns the assistant began suggesting Python snippets and client libraries that only exist in newer Postgres. Each new reply still looked coherent. It just stopped obeying the constraints I had given it. The result was a patch that passed the quick manual check but failed CI because the test DB was a MySQL replica of the production schema. The failure mode was not flashy. It was slow context drift, a creeping mismatch between the conversation and the system I actually run.


Hallucinations sneak in through missing tools


We connect models to a schema extractor and a CI status endpoint. Those tools sometimes return partial or empty payloads. The model then fills gaps. I once had a truncated schema come back from the extractor and the model guessed column types. The generated migration ran locally on my machine and created columns with the wrong precision. Tests didn't catch it because they mocked the schema layer. Building on those guesses, later changes silently relied on the wrong columns. The chain reaction is what worries me more than a single hallucinated line: one small incorrect assumption in a long thread becomes a foundation for subsequent changes.


Hidden assumptions multiply over long sessions


Every time we keep a chat open, we accumulate implicit defaults. The model prefers recent tokens. That means new examples and offhand remarks carry more weight than the original constraints. We saw the assistant start assuming default timeouts, memory limits, and even a different authentication flow after someone pasted a snippet from another repo into the same thread. Once the model assumes the wrong runtime or dependency version, it proposes code that is syntactically valid but operationally wrong. I started doing regular resets and explicit manifest blocks at the top of the prompt. When I need comparison between approaches I use a separate multi-model workspace so the drift in one conversation doesn’t contaminate the other; having a place for that comparison helped when I wanted to test alternative fixes side by side in a controlled way via a shared chat tool.


Practical verification and logging that cut the feedback loop


After several wasted rollbacks we built cheap guardrails. First, everything the model returned gets logged as raw text along with the prompt and the tool outputs. That let us diff replies over time and detect when suggestions started to deviate from earlier constraints. Second, generated code runs through a pipeline before review: static type checks, a sandbox run with a stripped dataset, and a set of targeted integration tests that assert environmental assumptions. Third, tool responses themselves are validated. If the schema extractor returns fewer than N columns or a status endpoint is slow, we stop and surface the tool error to a human instead of letting the model guess. For sourcing and verification I often ask the model to cite the exact doc section or changelog entry and then cross-check using a formal research workflow rather than trusting the claim in-chat.


Operational rules I actually follow now


I keep long work in short segments. I reset context at natural checkpoints: before major refactors, before merges, and whenever someone new joins the thread. I require a minimal manifest at the top of a conversation: runtime, DB, test fixtures, and a short list of forbidden changes. We log every tool call and validate responses with simple assertions. We also force the model to output a compact changelog and a list of assumptions it made, in machine-readable form, so our CI can reject merges that introduce new unstated assumptions. When I need richer verification or to triangulate sources I push the claim through a structured research pass that separates model reasoning from evidence gathering and treats external sources as ground truth rather than optional citations.

Ähnliche Beiträge
🔍 Verwandte News

Auch interessante Nachrichten When long chats quietly break builds

Thematisch verwandte Begriffe: When, long, chats, quietly · 6 Treffer

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

💬 Kommentare werden geladen…
Zum Aktualisieren ziehen
ZERO-DAY CVE-2026-97150 | When converting baserCMS4-style addons to baserCMS5-style ones, BcAddon…
Advisory →
tsecurity.de Icon
Offline-Lesen, Eilmeldungen & 0ms Ladezeit

Installiere tsecurity.de direkt auf deinen Home-Bildschirm für das ultimative Vollbild-Magazinerlebnis ohne Browser-Leisten.

Nächster Beitrag