When a Vietnamese bank's internal AI assistant started confidently quoting compliance rules that did not exist in any document, the team discovered they had been testing the wrong thing entirely. This post walks through how we set up RAGAs evaluation on that project, what faithfulness, context recall, and answer relevance each actually measure,...