[2402.01817] LLMs Can't Plan, But Can Help Planning in LLM-Modulo Frameworks
Abstract page for arXiv paper 2402.01817: LLMs Can't Plan, But Can Help Planning in LLM-Modulo Frameworks
Abstract page for arXiv paper 2402.01817: LLMs Can't Plan, But Can Help Planning in LLM-Modulo Frameworks
Abstract page for arXiv paper 2403.10089: Approximation and bounding techniques for the Fisher-Rao distances between parametric statistical models
Abstract page for arXiv paper 2405.07551: MuMath-Code: Combining Tool-Use Large Language Models with Multi-perspective Data Augmentation for Mathematical Reasoning
Abstract page for arXiv paper 2408.10808: ColBERT Retrieval and Ensemble Response Scoring for Language Model Question Answering
Abstract page for arXiv paper 2504.12267: New Constraints on DMS and DMDS in the Atmosphere of K2-18 b from JWST MIRI
Abstract page for arXiv paper 2404.01012v1: Query Performance Prediction using Relevance Judgments Generated by Large Language Models
Relevance judgments and other truth data for information retrieval (IR) evaluations are created manually. There is a strong temptation to use large language models (LLMs) as proxies for human judges. However, letting the LLM write your truth data handicaps the evaluation by setting that LLM as a...
In short: Avoid using LLMs for relevance judgments in IR evaluations! Learn why this practice handicaps your evaluation and discover alternative ways to leverage LLMs in the assessment process. (AI summary)
Abstract page for arXiv paper 0805.1154: Clustering of scientific citations in Wikipedia
Abstract page for arXiv paper 0903.3287: Hyperbolic Voronoi diagrams made easy
Abstract page for arXiv paper 1004.5049: The Burbea-Rao and Bhattacharyya centroids