advanced35 minby DevFox

LLM-as-Judge and Its Failure Modes

Position bias, self-preference, verbosity bias and miscalibration. How to build a judge that correlates with human labels, and how to know when it has stopped.

  • AI
  • Testing

    LLM-as-Judge and Its Failure Modes — Production AI Engineering | DevFox Labs