I have been reading some blog posts about LLM as a judge and was building a small evaluator to evaluate the judge itself . My method is simple: The dataset is: task rubric ideal response negative response The idea is then to test different models as judges for things like: repeated-run consistency position bias sensitivity to verbosity accuracy /... Weiterlesen
Intelligence View
⚡ tsecurity.de Intelligence
SOCIAL SHARE CARD GENERATOR