Evaluating hackathons at scale is deceptively difficult. When 40 software projects are evaluated across 4 specialized tracks by 30 volunteer judges using a 4-criteria weighted rubric, naive arithmetic means fail catastrophically: Hawk vs. Dove Skew: Strict judges award scores between 2.0 and 3.5. Lenient judges award scores between 4.0 and 5.0. A... Weiterlesen
Intelligence View
⚡ tsecurity.de Intelligence
OmniJudge Post-Mortem: 1-ULP Drift, Prisma Traps, and Alpine musl
Evaluating hackathons at scale is deceptively difficult. When 40 software projects are evaluated across 4 specialized tracks by 30 volunteer judges using a…