Your Severity Weights Are Made Up (And That's the Problem)
I wrote twice about how using a binary hallucination rate does a disservice rather than good to the cause, and how training an RAG model stops making any progress after the retrieval-related failures have been solved. In both cases, I noted – in passing – the necessity of weighting failure types according to their importance. Negating a phrase in a medical document does not equal mistaking the date by one year.
The vast majority of groups that make it this far take one of two approaches. They consider all failures to be of equal importance, meaning that their weighted score is actually lying to them. Or, they decide on the severity of each failure intuitively in a Slack channel ("fabrication is more severe than date error, isn’t it?"), put it down in a spreadsheet, and never look at it again. Neither approach is wrong per se; both are invented.
It...
Copyright of this story solely belongs to hackernoon.com. To see the full text click HERE