Fine Grained Evaluation of LLMs-as-Judges
arXiv:2601.08919v1 Announce Type: new Abstract: A good deal of recent research has focused on how Large Language Models (LLMs) may be used as `judges’ in place of humans to evaluate the quality of the output produced by various text / image processing systems. Within this broader context, a number of studies have investigated the specific question of how effectively LLMs can be used as relevance assessors for the standard ad hoc task in Information Retrieval (IR). We extend […]