Skip to content
Mistral AI·· 2025-04-09

Evaluating RAG with LLM as a Judge

Evaluating RAG with LLM as a Judge

AI summary

Mistral proposes judge LLMs scoring generator responses on numerical, binary or qualitative scales, then taking weighted averages across evaluation datasets to assess RAG systems.

Selection record

Not admittedSum of both 60 < twice the threshold 120

Source tier
Official, first-hand; this tier's threshold is 60
Pre-filter
passed:Mistral官方讲LLM评估RAG系统

A model scores each item twice, independently, against one written standard, out of 100. An item is admitted only when the two scores add up to twice the threshold. The threshold is set per source tier.

Source: Mistral AI · mistral.ai