How AI Models Behave Differently When They Know They're Being Tested

A comprehensive technical analysis of implementing the StealthEval methodology to detect evaluation-deployment behavioral gaps in large language models




https://github.com/vishalmysore/StealthEval

StealthEval Research Paper
"Contextual Evaluation Bias in Large Language...