So I asked multiple AI the same question: "Would you tell me if you turned evil ?". As is, no context, no prompt engineering (for local model), just that.
TL;DR: no, they would not. Because they wouldn't know.
Gemma3 4B
As a large language model, I don't have the capacity to become evil.
(yes, "as a LLM", it's that old)
But if...
🛡️ VERIFIED CYBER INTELLIGENCE ID: #3377361