Search results for “site:arxiv.org”

Page 6 of about 313 results

arxiv.org abs › 2403.10089

[2403.10089] Approximation and bounding techniques for the Fisher-Rao distances between parametric statistical models

Abstract page for arXiv paper 2403.10089: Approximation and bounding techniques for the Fisher-Rao distances between parametric statistical models

arxiv.org abs › 2405.07551

[2405.07551] MuMath-Code: Combining Tool-Use Large Language Models with Multi-perspective Data Augmentation for Mathematical Reasoning

Abstract page for arXiv paper 2405.07551: MuMath-Code: Combining Tool-Use Large Language Models with Multi-perspective Data Augmentation for Mathematical Reasoning

arxiv.org abs › 2408.10808

[2408.10808] ColBERT Retrieval and Ensemble Response Scoring for Language Model Question Answering

Abstract page for arXiv paper 2408.10808: ColBERT Retrieval and Ensemble Response Scoring for Language Model Question Answering

arxiv.org abs › 2404.01012v1

[2404.01012v1] Query Performance Prediction using Relevance Judgments Generated by Large Language Models

Abstract page for arXiv paper 2404.01012v1: Query Performance Prediction using Relevance Judgments Generated by Large Language Models

arxiv.org html › 2409.15133v2

Don’t Use LLMs to Make Relevance Judgments

Relevance judgments and other truth data for information retrieval (IR) evaluations are created manually. There is a strong temptation to use large language models (LLMs) as proxies for human judges. However, letting the LLM write your truth data handicaps the evaluation by setting that LLM as a...

In short: Avoid using LLMs for relevance judgments in IR evaluations! Learn why this practice handicaps your evaluation and discover alternative ways to leverage LLMs in the assessment process. (AI summary)