RAGLens
Evaluate in bulk
Run a set of questions through the same evaluation used on Evaluate one query. Each result shows what evidence was available, what the model answered, and whether the outcome matched what we expected.
Some expected outcomes are successful answers; others are successful refusals — disagreements show where the pipeline or the evaluation needs attention.
Choose a corpus and run tests to see results here.
Each row will show expected vs. actual pass, score, and eval diagnosis.