Zum Hauptinhalt springen
Echtzeit-Radar & Feeds
Alle RSS Feeds ➔
👥 Community & Social
Windows Tipps & SecurityMicrosoft got hacked, and somehow Clippy ended up in a crypto scheme(02.10.2026 um 14:11 Uhr)
•
Windows Tipps & SecurityBank-Überweisung: Dieser Fehler kann zur Sperrung Ihres Kontos führen(02.10.2026 um 14:30 Uhr)
••
Sichere ProgrammierungLog Safety Events, Not Full Transcripts: One Audit Trade-Off(02.10.2026 um 14:10 Uhr)
•••••
Sichere ProgrammierungBuild a Python Subtitle Generator with ffmpeg: A Step-by-Step Guide(02.10.2026 um 14:14 Uhr)
•
Sichere Programmierung🔐 Vault — Privacy-First Local AI for Sensitive Legal Documents(02.10.2026 um 14:14 Uhr)
•
Windows Tipps & SecurityMicrosoft got hacked, and somehow Clippy ended up in a crypto scheme(02.10.2026 um 14:11 Uhr)
•
Windows Tipps & SecurityBank-Überweisung: Dieser Fehler kann zur Sperrung Ihres Kontos führen(02.10.2026 um 14:30 Uhr)
••
Sichere ProgrammierungLog Safety Events, Not Full Transcripts: One Audit Trade-Off(02.10.2026 um 14:10 Uhr)
•••••
Sichere ProgrammierungBuild a Python Subtitle Generator with ffmpeg: A Step-by-Step Guide(02.10.2026 um 14:14 Uhr)
•
Sichere Programmierung🔐 Vault — Privacy-First Local AI for Sensitive Legal Documents(02.10.2026 um 14:14 Uhr)
•
Intelligence View
⚡ tsecurity.de Intelligence

why Cohen's kappa drifts week to week (and what to do about it)

If your LLM-as-judge calibration kappa moves around week to week and you cannot explain it from labeller behavior, the usual cause is the marginal distribution…

Beitrag
0
Seite
0
↗ Quelle (dev.to)
Social ReaktionenReagiere als Erste:r — dein Feedback zählt!

If your LLM-as-judge calibration kappa moves around week to week and you cannot explain it from labeller behavior, the usual cause is the marginal distribution of your calibration set, not the labellers.



Quick refresher. Cohen's kappa is:




kappa = (Po - Pe) / (1 - Pe)






Where Po is observed agreement and Pe is expected agreement by chance. Pe depends on the marginal distribution of the labels in your set.



If 70% of last week's traces were labelled "acceptable" by labeller A and 25% "good" and 5% "bad", Pe is one number. If this week's mix is 50/40/10, Pe shifts. The labellers can be doing exactly the same thing and your kappa value moves.



Three things that help:




  1. Sample your calibration set across multiple time windows (rolling 4-week window, stratified by time bucket). Reduces the chance that one week's traffic pattern dominates Pe.


  2. Report per-class precision and recall alongside kappa. Kappa is one summary number; the per-class metrics tell you where the labeller-LLM disagreement actually sits.


  3. For very small calibration sets (under 100 traces), use Wilson confidence intervals around the per-class precision instead of treating kappa as a point estimate. The Wilson interval is robust to small samples; the normal-approximation interval is not.




References for the calibration-set design and the small-sample math are in Cohen (1960) "A coefficient of agreement for nominal scales" and Wilson (1927) "Probable inference, the law of succession, and statistical inference." Both are short reads.

Ähnliche Beiträge
🔍 Verwandte News

Auch interessante Nachrichten why Cohen's kappa drifts week to week (and what to do about it)

Thematisch verwandte Begriffe: Cohens, kappa, drifts, week · 6 Treffer

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

💬 Kommentare werden geladen…
Zum Aktualisieren ziehen
tsecurity.de Icon
Offline-Lesen, Eilmeldungen & 0ms Ladezeit

Installiere tsecurity.de direkt auf deinen Home-Bildschirm für das ultimative Vollbild-Magazinerlebnis ohne Browser-Leisten.

Nächster Beitrag