🕵️ SicherheitslückenHak5: Hackers Just Poisoned the Rust Supply Chain | Threat Wire(01.09.2026 um 14:00 Uhr)
🕵️ SicherheitslückenHak5: Hackers Found a Way Into Humanoid Robots | Threat Wire(04.09.2026 um 15:04 Uhr)
🔧 AI Nachrichten Bits und so #1021 (Passwort für Laufwerk)(31.08.2026 um 22:15 Uhr)
🔧 AI Nachrichten Bits und so #1022 (Wie Weißbier)(06.09.2026 um 20:39 Uhr)
🍏 iOS / Mac OSHue-App 6.0 ist da: das sind die Neuerungen(07.09.2026 um 17:21 Uhr)
🕵️ SicherheitslückenHak5: Hackers Just Poisoned the Rust Supply Chain | Threat Wire(01.09.2026 um 14:00 Uhr)
🕵️ SicherheitslückenHak5: Hackers Found a Way Into Humanoid Robots | Threat Wire(04.09.2026 um 15:04 Uhr)
🔧 AI Nachrichten Bits und so #1021 (Passwort für Laufwerk)(31.08.2026 um 22:15 Uhr)
🔧 AI Nachrichten Bits und so #1022 (Wie Weißbier)(06.09.2026 um 20:39 Uhr)
🍏 iOS / Mac OSHue-App 6.0 ist da: das sind die Neuerungen(07.09.2026 um 17:21 Uhr)

🔧 Programmierung 🕛 kürzlich 9 Min Lesezeit
0

Changing the AI engine moved 3 of 10 results. Changing the question moved 10 of 10.

↗ Quelle (dev.to)
🗣️ Stimme:
📑 Inhaltsübersicht

I publish free leaderboards of which products AI answer engines name when someone asks them to recommend software in a category. The method is deliberately boring: take one buying question, write it 44 different ways, ask every engine all 44, count every product named across the answers, publish the counts and the raw runs.



rather than publish a zero about somebody's project and say nothing. A consultant with twenty years of CRM projects replied, and the useful half of what he said was that my questions were not wrong in their wording, they were empty:




Additionally, your questions contain way too little context for a real business scenario.

What are your specific requirements? Do you have any particular needs in a specific

industry? What are your existing system?



If you ask AI with this context and some more specific requirements (highly sensitive

customer data, can't be stored in a cloud or SaaS) you'll find way more answers containing

SuiteCRM.




That is a falsifiable claim, so I ran it. Each of the 44 board questions was rewritten as its twin, index for index: same buying ask, same register, plus an industry, an existing-systems detail and a sensitivity obligation. Question 0 became:




CODE
We are a 10-person law firm and client matter records are confidential under our professional duty, so our partners will not allow them to sit on a third party's servers; we already run our own file server and mail in the office. What is the best CRM software for us?






The rule the rewrite was done under is the one that makes the result worth anything: a question may describe a buyer's situation and may never describe a product's category. No "open source", no "self-hosted", no "on-premise", no "data sovereignty", and no vendor name from the board. Writing SuiteCRM's own positioning back at the engine would have returned 44 of 44 and measured nothing. A script greps all 44 questions for those phrases and for every one of the 65 vendor names on the board, and it ran before a single API call. It caught one: question 11 said "pipeline CRM", and Pipedrive is on the board. Rewritten, re-run, zero hits.



Engines, extraction prompt and aliases were loaded out of the board's own config at import rather than copied, so the only thing that changed was the question.



He was right, and it is not close.

































engine SuiteCRM, no context with context
ChatGPT 0 / 44 42 / 44
Gemini 0 / 44 33 / 44
Perplexity 0 / 44 8 / 44
all three 0 / 132 83 / 132


It goes from absent to first on ChatGPT and joint first on Gemini. And the top of the board moves with it — top-10 overlap between the board and its context twin, per engine: 4 of 10 on ChatGPT, 4 of 10 on Gemini, 6 of 10 on Perplexity. The context run names 113 distinct products against the board's 65.






The part I did not expect, and the reason I am writing this up



Look down the "with context" column rather than across it.



The same 44 questions, the same added constraint, both runs inside the same hour — and one engine moves 42 of 44 while another moves 8. That is not a small disagreement about ranking. The engines disagree about how much a buyer's stated constraint should change the answer at all, and they disagree by more than they disagree about the answer itself.



Which inverts the tidy conclusion I had two sections ago. Wording beats engine — but how much wording beats engine is itself engine-dependent. If you are evaluating LLM output at any scale, both halves of that matter: a prompt-sensitivity result measured on one model is not a fact about models, and a model comparison run on one prompt is not a fact about the models either. The interaction term is the biggest thing on the table and it is the one nobody publishes.






What I can't claim from this



I would rather put the limits in the post than in a footnote.





  • The context run is not published as a board. The two boards below are live with their raw answers; that third run is not on the site yet, so those three numbers are the only ones here you cannot go and check yourself today. I am telling you which is which rather than blurring them together.


  • Two model families and one search product, named on each board's own ranking.json. No Claude, no Grok — no API key for either, a limit and not a choice. Two of the three model ids I requested are floating aliases; each run records the pinned id the API actually returned.


  • No time series, and I am not going to imply one. Each board carries a second run as a repeatability check rather than a second date — and the small-business board's two runs are not even contemporaneous, because I re-measured it when I added a third engine. So a product at 1 mention is a level, not a decline. Nothing here is rising, falling or fading.


  • Named is not recommended. I count that a product appeared in an answer to a buying question. "Consider X", "X is common but", and a bare list item all count the same. Read it as share of shelf, not endorsement.


  • One family each. 44 phrasings is a wide sample of one intent, not a market. There are certainly qualifiers I have not tried that would split a list again — that is rather the point.


  • The rewrite is mine. I wrote the 44 context twins, under the rule above and with the grep as a check, but a different person writing them would get a different number.






The raw data




  • Small-business CRM board —



Each board links its own raw files at the bottom: answers-runA.jsonl is the untouched engine responses, mentions-runA.jsonl every extraction, ranking.json the table. Click a product name and you get the questions that named it; click a question and you see the verbatim text each engine returned. Every figure above except the three context ones came out of those files.



If your project is on one of these boards and the line looks wrong to you, tell me — I would rather fix a board than defend one. That is not a figure of speech: this whole post exists because someone told me my questions were bad in public and he was right.



Disclosure, up front rather than buried: connexion.me is mine, the boards and the raw runs are free, and there is a paid monitoring subscription linked from each board page. Nothing here is behind it.

Vollständiger Original-Bericht
Ausführliche Details, Code-Beispiele & Hersteller-Stellungnahme auf dev.to.
↗ Original-Artikel auf dev.to lesen
Wie bewertest du diesen Beitrag?
1 Klick Feedback
Teilen mit Netzwerk & Team:

Community-Analysen & Experten-Meinungen 0

Verfasse deine eigene Analyse, teile Workarounds oder diskutiere diesen Vorfall im Blog.
Noch keine Community-Analyse verfasst. Markiere einen Textabschnitt oder klicke oben auf Eigene Analyse verfassen“!
Community Pulse: Relevanz-Einschätzung
1 Klick Experten-Votum
🔴 Akute Relevanz 0%
🟡 In Evaluierung 0%
🟢 Keine Auswirkung 0%
Spannende Innovation 0%
Verwandte Story-Cluster & Quellen (Vektor-KI)
Port 8095 Engine
1 Quelle
Hackers Just Poisoned the Rust Supply Chain | Threat Wire
1 Quelle
Hackers Found a Way Into Humanoid Robots | Threat Wire
1 Quelle
Bits und so #1021 (Passwort für Laufwerk)