Category Scores from Moderation OpenAI API

How are we deciding the score threshold for categories in Moderation API ?


  • flagged: Set to true if the model classifies the content as violating OpenAI’s usage policies, false otherwise.
  • categories: Contains a dictionary of per-category binary usage policies violation flags. For each category, the value is true if the model flags the corresponding category as violated, false otherwise.
  • category_scores: Contains a dictionary of per-category raw scores output by the model, denoting the model’s confidence that the input violates the OpenAI’s policy for the category. The value is between 0 and 1, where higher values denote higher confidence. The scores should not be interpreted as probabilities.

My takeaway is the ‘threshold’ is binary. It either is flagged for one or more categories or it isn’t,