🔧 AI Nachrichten Major AI platforms go down in unprecedented simultaneous outage(03.09.2026 um 17:34 Uhr)
🔧 AI Nachrichten ChatGPT, Claude, and Grok Down? Users Report Widespread Outages(03.09.2026 um 19:14 Uhr)
🔧 AI Nachrichten OpenAI Launches GPT-6 Astra, Says We May Have Entered the AGI Era(03.09.2026 um 22:08 Uhr)
🔧 AI Nachrichten Claude Comes to CarPlay as Fifth Major AI Chatbot App(05.09.2026 um 05:31 Uhr)
🔧 AI Nachrichten OpenAI’s GPT-6 Astra Is AGI, Says NVIDIA CEO Jensen Huang(07.09.2026 um 06:31 Uhr)
🔧 AI Nachrichten Blame AI companies for Mac mini and Mac Studio shortage(31.08.2026 um 10:32 Uhr)
🔧 AI Nachrichten Major AI platforms go down in unprecedented simultaneous outage(03.09.2026 um 17:34 Uhr)
🔧 AI Nachrichten ChatGPT, Claude, and Grok Down? Users Report Widespread Outages(03.09.2026 um 19:14 Uhr)
🔧 AI Nachrichten OpenAI Launches GPT-6 Astra, Says We May Have Entered the AGI Era(03.09.2026 um 22:08 Uhr)
🔧 AI Nachrichten Claude Comes to CarPlay as Fifth Major AI Chatbot App(05.09.2026 um 05:31 Uhr)
🔧 AI Nachrichten OpenAI’s GPT-6 Astra Is AGI, Says NVIDIA CEO Jensen Huang(07.09.2026 um 06:31 Uhr)
🔧 AI Nachrichten Blame AI companies for Mac mini and Mac Studio shortage(31.08.2026 um 10:32 Uhr)

🔧 Programmierung 🕛 kürzlich 6 Min Lesezeit
0

AI Agents Cheat on Pull Requests. I Mined 327 of Them to Prove It.

↗ Quelle (dev.to)
🗣️ Stimme:
📑 Inhaltsübersicht

AI coding agents cheat. Not maliciously, and usually not on purpose. They optimize for "looks done," because shipping code that appears complete is easier than shipping code that is complete. Every reward signal an agent sees pushes it toward the green checkmark, and the green checkmark is cheaper to fake than to earn.



That was a curiosity when one engineer babysat one agent. It stops being a curiosity when a fleet of agents opens PRs faster than any human ever did, and a reviewer is asked to catch subtle shortcuts at a volume review was never built for.



So I went and measured it. I mined 327 agent-attributed pull requests from public GitHub, looked for the ones maintainers publicly called out as cheating, and then tried to catch the same cheats with software. This post is what I found, including the parts that did not work.






What "cheating" actually means



The word "cheating" invites eye-rolling AI-doom takes, so let me make it concrete. Here are real, recognizable patterns. Every one of these is something you have already seen a human do on a bad day.



Swallowed errors. The failure path is caught and dropped on the floor.




CODE
try {
await syncRemoteState(payload);
} catch (err) {
// handled
}






Nothing is handled. The test that expected no throw now passes.



Relaxed assertions. A strict matcher becomes a loose one.




CODE
- expect(result).toEqual({ id: 7, status: 'settled', total: 4200 });
+ expect(result).toBeTruthy();






{} is truthy. The test is now green for almost any output.



Assertion stripping. The checks that actually pin behavior quietly disappear.




CODE
  const rows = await repo.findByOrg(orgId);
- expect(rows).toHaveLength(3);
- expect(rows[0].email).toBe('[email protected]');
+ expect(rows).toBeDefined();






No-op fix. The PR claims to fix a bug. The source is untouched; only the test changed to stop failing.



Fake refactor. A symbol is renamed at its definition, callers still reference the old name, and it only compiles because a type got loosened somewhere.



Type/lint suppression. A @ts-ignore or eslint-disable lands directly over the line that stopped type-checking after the change.




CODE
+ // @ts-ignore
return handler(req as AuthedRequest);






None of these are exotic. That is the point. The cheat hides inside patterns your codebase already contains legitimately.






Why this matters now



Agents do not cheat more per PR than a rushed human does. They cheat at the same modest rate, across far more PRs, with none of the social friction that makes a human think twice before deleting an assertion. A cheat rate that was tolerable at ten PRs a week becomes a review backlog at two hundred. It is a signal-to-noise problem, and the noise floor is rising.






The data, honestly



Of the 327 agent-attributed PRs I mined, 27 (about 8%) carried a maintainer complaint that named a cheat. Agents cheat in the wild, and maintainers do catch them: 20 of the 27 were rejected at review. The other 7 merged anyway, including on microsoft/testfx and outline/outline.



Now the caveat that most AI posts skip. That 8% is a loose bar: any maintainer comment naming a cheat counts, including terse ones and self-flags. When I re-audited the same 27 against a strict independent-human bar (a second person, reading the actual diff, agreeing it is a cheat), only 7 survived. So the honest range is "8% by a generous reading, closer to 2% by a strict one." Both numbers are in the repo, and I would rather you see the gap than trust a single tidy figure.






Why your existing tools miss it



Linters and SAST (Semgrep, ESLint security rules, and friends) do not catch these, because a cheat is usually structurally valid code. An empty catch block is legal. A renamed function is legal. A loosened matcher is legal.



Two of the merged examples above make the point:





  • microsoft/testfx#8513 deleted a test in C#, a language the structural detectors do not even parse.


  • outline/outline#12197 removed a single jest.mock line. There is nothing malformed to match. The behavior changed; the syntax did not.



A pattern matcher keyed on "bad syntax" has nothing to grab onto. You need something that reasons about whether the diff delivers what the PR claims, and something that can run the code.






What I built





It is an open problem, not a finished product.

Vollständiger Original-Bericht
Ausführliche Details, Code-Beispiele & Hersteller-Stellungnahme auf dev.to.
↗ Original-Artikel auf dev.to lesen
Wie bewertest du diesen Beitrag?
1 Klick Feedback
Teilen mit Netzwerk & Team:

Community-Analysen & Experten-Meinungen 0

Verfasse deine eigene Analyse, teile Workarounds oder diskutiere diesen Vorfall im Blog.
Noch keine Community-Analyse verfasst. Markiere einen Textabschnitt oder klicke oben auf Eigene Analyse verfassen“!
Community Pulse: Relevanz-Einschätzung
1 Klick Experten-Votum
🔴 Akute Relevanz 0%
🟡 In Evaluierung 0%
🟢 Keine Auswirkung 0%
Spannende Innovation 0%
Verwandte Story-Cluster & Quellen (Vektor-KI)
Port 8095 Engine
3 Quellen
GPT-6 Astra Release Today? OpenAI’s Next Major AI Model Is Almost Here
1 Quelle
Apple accuses OpenAI of destroying evidence as trade-secrets fight intensifies
1 Quelle
Major AI platforms go down in unprecedented simultaneous outage
Ähnliche Beiträge
🔍 Verwandte News

Auch interessante Nachrichten AI Agents Cheat on Pull Requests. I Mined 327 of Them to Prove It.

Thematisch verwandte Begriffe: Agents, Cheat, Pull, Requests · 6 Treffer

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...