Understanding the User: An Intent-Based Ranking Dataset

Abhijit Anand; Jurek Leonhardt; V. Venktesh; Avishek Anand

doi:10.48550/arXiv.2408.17103

Details

Originalsprache	Englisch
Titel des Sammelwerks	CIKM 2024
Untertitel	Proceedings of the 33rd ACM International Conference on Information and Knowledge Management
Herausgeber (Verlag)	Association for Computing Machinery
Seiten	5323-5327
Seitenumfang	5
ISBN (elektronisch)	9798400704369
Publikationsstatus	Veröffentlicht - 21 Okt. 2024
Veranstaltung	33rd ACM International Conference on Information and Knowledge Management, CIKM 2024 - Boise, USA / Vereinigte Staaten Dauer: 21 Okt. 2024 → 25 Okt. 2024

Abstract

As information retrieval systems continue to evolve, accurate evaluation and benchmarking of these systems become pivotal. Web search datasets, such as MS MARCO, primarily provide short keyword queries without accompanying intent or descriptions, posing a challenge in comprehending the underlying information need. This paper proposes an approach to augmenting such datasets to annotate informative query descriptions, with a focus on two prominent benchmark datasets: TREC-DL-21 and TREC-DL-22. Our methodology involves utilizing state-of-the-art LLMs to analyze and comprehend the implicit intent within individual queries from benchmark datasets. By extracting key semantic elements, we construct detailed and contextually rich descriptions for these queries. To validate the generated query descriptions, we employ crowdsourcing as a reliable means of obtaining diverse human perspectives on the accuracy and informativeness of the descriptions. This information can be used as an evaluation set for tasks such as ranking, query rewriting, or others.

ASJC Scopus Sachgebiete

Betriebswirtschaft, Management und Rechnungswesen (insg.)
Allgemeine Unternehmensführung und Buchhaltung
Entscheidungswissenschaften (insg.)
Allgemeine Entscheidungswissenschaften

Zitieren

Understanding the User: An Intent-Based Ranking Dataset. / Anand, Abhijit; Leonhardt, Jurek; Venktesh, V. et al.
CIKM 2024 : Proceedings of the 33rd ACM International Conference on Information and Knowledge Management. Association for Computing Machinery, 2024. S. 5323-5327.

Publikation: Beitrag in Buch/Bericht/Sammelwerk/Konferenzband › Aufsatz in Konferenzband › Forschung › Peer-Review

Anand, A, Leonhardt, J, Venktesh, V & Anand, A 2024, Understanding the User: An Intent-Based Ranking Dataset. in CIKM 2024 : Proceedings of the 33rd ACM International Conference on Information and Knowledge Management. Association for Computing Machinery, S. 5323-5327, 33rd ACM International Conference on Information and Knowledge Management, CIKM 2024, Boise, USA / Vereinigte Staaten, 21 Okt. 2024. https://doi.org/10.48550/arXiv.2408.17103, https://doi.org/10.1145/3627673.3679166

Anand, A., Leonhardt, J., Venktesh, V., & Anand, A. (2024). Understanding the User: An Intent-Based Ranking Dataset. In CIKM 2024 : Proceedings of the 33rd ACM International Conference on Information and Knowledge Management (S. 5323-5327). Association for Computing Machinery. https://doi.org/10.48550/arXiv.2408.17103, https://doi.org/10.1145/3627673.3679166

Anand A, Leonhardt J, Venktesh V, Anand A. Understanding the User: An Intent-Based Ranking Dataset. in CIKM 2024 : Proceedings of the 33rd ACM International Conference on Information and Knowledge Management. Association for Computing Machinery. 2024. S. 5323-5327 doi: 10.48550/arXiv.2408.17103, 10.1145/3627673.3679166

Anand, Abhijit ; Leonhardt, Jurek ; Venktesh, V. et al. / Understanding the User : An Intent-Based Ranking Dataset. CIKM 2024 : Proceedings of the 33rd ACM International Conference on Information and Knowledge Management. Association for Computing Machinery, 2024. S. 5323-5327

Download

@inproceedings{cc58864ab67d4124a61d4d6d5d124850,

title = "Understanding the User: An Intent-Based Ranking Dataset",

abstract = "As information retrieval systems continue to evolve, accurate evaluation and benchmarking of these systems become pivotal. Web search datasets, such as MS MARCO, primarily provide short keyword queries without accompanying intent or descriptions, posing a challenge in comprehending the underlying information need. This paper proposes an approach to augmenting such datasets to annotate informative query descriptions, with a focus on two prominent benchmark datasets: TREC-DL-21 and TREC-DL-22. Our methodology involves utilizing state-of-the-art LLMs to analyze and comprehend the implicit intent within individual queries from benchmark datasets. By extracting key semantic elements, we construct detailed and contextually rich descriptions for these queries. To validate the generated query descriptions, we employ crowdsourcing as a reliable means of obtaining diverse human perspectives on the accuracy and informativeness of the descriptions. This information can be used as an evaluation set for tasks such as ranking, query rewriting, or others.",

keywords = "ad-hoc retrieval, data collection, diversity, intent dataset, ranking, user intents, web search",

author = "Abhijit Anand and Jurek Leonhardt and V. Venktesh and Avishek Anand",

note = "Publisher Copyright: {\textcopyright} 2024 Owner/Author.; 33rd ACM International Conference on Information and Knowledge Management, CIKM 2024 ; Conference date: 21-10-2024 Through 25-10-2024",

year = "2024",

month = oct,

day = "21",

doi = "10.48550/arXiv.2408.17103",

language = "English",

pages = "5323--5327",

booktitle = "CIKM 2024",

publisher = "Association for Computing Machinery",

address = "United States",

}

Download

TY - GEN

T1 - Understanding the User

T2 - 33rd ACM International Conference on Information and Knowledge Management, CIKM 2024

AU - Anand, Abhijit

AU - Leonhardt, Jurek

AU - Venktesh, V.

AU - Anand, Avishek

PY - 2024/10/21

Y1 - 2024/10/21

N2 - As information retrieval systems continue to evolve, accurate evaluation and benchmarking of these systems become pivotal. Web search datasets, such as MS MARCO, primarily provide short keyword queries without accompanying intent or descriptions, posing a challenge in comprehending the underlying information need. This paper proposes an approach to augmenting such datasets to annotate informative query descriptions, with a focus on two prominent benchmark datasets: TREC-DL-21 and TREC-DL-22. Our methodology involves utilizing state-of-the-art LLMs to analyze and comprehend the implicit intent within individual queries from benchmark datasets. By extracting key semantic elements, we construct detailed and contextually rich descriptions for these queries. To validate the generated query descriptions, we employ crowdsourcing as a reliable means of obtaining diverse human perspectives on the accuracy and informativeness of the descriptions. This information can be used as an evaluation set for tasks such as ranking, query rewriting, or others.

AB - As information retrieval systems continue to evolve, accurate evaluation and benchmarking of these systems become pivotal. Web search datasets, such as MS MARCO, primarily provide short keyword queries without accompanying intent or descriptions, posing a challenge in comprehending the underlying information need. This paper proposes an approach to augmenting such datasets to annotate informative query descriptions, with a focus on two prominent benchmark datasets: TREC-DL-21 and TREC-DL-22. Our methodology involves utilizing state-of-the-art LLMs to analyze and comprehend the implicit intent within individual queries from benchmark datasets. By extracting key semantic elements, we construct detailed and contextually rich descriptions for these queries. To validate the generated query descriptions, we employ crowdsourcing as a reliable means of obtaining diverse human perspectives on the accuracy and informativeness of the descriptions. This information can be used as an evaluation set for tasks such as ranking, query rewriting, or others.

KW - ad-hoc retrieval

KW - data collection

KW - diversity

KW - intent dataset

KW - ranking

KW - user intents

KW - web search

UR - http://www.scopus.com/inward/record.url?scp=85210031597&partnerID=8YFLogxK

U2 - 10.48550/arXiv.2408.17103

DO - 10.48550/arXiv.2408.17103

M3 - Conference contribution

AN - SCOPUS:85210031597

SP - 5323

EP - 5327

BT - CIKM 2024

PB - Association for Computing Machinery

Y2 - 21 October 2024 through 25 October 2024

ER -

Research@Leibniz University

Understanding the User: An Intent-Based Ranking Dataset

Autorschaft

Organisationseinheiten

Externe Organisationen

Details

Abstract

ASJC Scopus Sachgebiete

Zitieren