> You have no way to provide a concrete reason to the user at that point.
You should just be able to look at the user's message and tell them why it's against your policy, else reverse the decision if you see no violation.
If a user is curious specifically about how the model made its decision, and you want to reveal detail at that level, it's an open-weights model so interpretability techniques should work ("biggest impact on score came when focusing on this word in your message and this part of the policy").
You should just be able to look at the user's message and tell them why it's against your policy, else reverse the decision if you see no violation.
If a user is curious specifically about how the model made its decision, and you want to reveal detail at that level, it's an open-weights model so interpretability techniques should work ("biggest impact on score came when focusing on this word in your message and this part of the policy").