DaisyRec 2.0: Benchmarking Recommendation for Rigorous Evaluation

Zhu Sun; Hui Fang; Jie Yang; Xinghua Qu; Hongyang Liu; Di Yu; Yew Soon Ong; Jie Zhang

doi:10.1109/TPAMI.2022.3231891

DaisyRec 2.0: Benchmarking Recommendation for Rigorous Evaluation

Zhu Sun, Hui Fang^*, Jie Yang, Xinghua Qu, Hongyang Liu, Di Yu, Yew Soon Ong, Jie Zhang

^*Corresponding author for this work

Web Information Systems

Research output: Contribution to journal › Article › Scientific › peer-review

1 Citation (Scopus)

33 Downloads (Pure)

Abstract

Recently, one critical issue looms large in the field of recommender systems - there are no effective benchmarks for rigorous evaluation - which consequently leads to unreproducible evaluation and unfair comparison. We, therefore, conduct studies from the perspectives of practical theory and experiments, aiming at benchmarking recommendation for rigorous evaluation. Regarding the theoretical study, a series of hyper-factors affecting recommendation performance throughout the whole evaluation chain are systematically summarized and analyzed via an exhaustive review on 141 papers published at eight top-tier conferences within 2017-2020. We then classify them into model-independent and model-dependent hyper-factors, and different modes of rigorous evaluation are defined and discussed in-depth accordingly. For the experimental study, we release DaisyRec 2.0 library by integrating these hyper-factors to perform rigorous evaluation, whereby a holistic empirical study is conducted to unveil the impacts of different hyper-factors on recommendation performance. Supported by the theoretical and experimental studies, we finally create benchmarks for rigorous evaluation by proposing standardized procedures and providing performance of ten state-of-the-arts across six evaluation metrics on six datasets as a reference for later study. Overall, our work sheds light on the issues in recommendation evaluation, provides potential solutions for rigorous evaluation, and lays foundation for further investigation.

Original language	English
Pages (from-to)	8206-8226
Number of pages	21
Journal	IEEE Transactions on Pattern Analysis and Machine Intelligence
Volume	45
Issue number	7
DOIs	https://doi.org/10.1109/TPAMI.2022.3231891
Publication status	Published - 2023

Bibliographical note

Green Open Access added to TU Delft Institutional Repository ‘You share, we take care!’ – Taverne project https://www.openaccess.nl/en/you-share-we-take-care Otherwise as indicated in the copyright section: the publisher is the copyright holder of this work and the author uses the Dutch legislation to make this work public.

Keywords

Benchmarks
fair comparison
recommender systems
reproducible evaluation
standardized procedures

Access to Document

10.1109/TPAMI.2022.3231891

DaisyRec_2.0_Benchmarking_Recommendation_for_Rigorous_EvaluationFinal published version, 3.67 MB

Cite this

@article{1014a9d2360446aca92ecb6e485e7d72,

title = "DaisyRec 2.0: Benchmarking Recommendation for Rigorous Evaluation",

abstract = "Recently, one critical issue looms large in the field of recommender systems - there are no effective benchmarks for rigorous evaluation - which consequently leads to unreproducible evaluation and unfair comparison. We, therefore, conduct studies from the perspectives of practical theory and experiments, aiming at benchmarking recommendation for rigorous evaluation. Regarding the theoretical study, a series of hyper-factors affecting recommendation performance throughout the whole evaluation chain are systematically summarized and analyzed via an exhaustive review on 141 papers published at eight top-tier conferences within 2017-2020. We then classify them into model-independent and model-dependent hyper-factors, and different modes of rigorous evaluation are defined and discussed in-depth accordingly. For the experimental study, we release DaisyRec 2.0 library by integrating these hyper-factors to perform rigorous evaluation, whereby a holistic empirical study is conducted to unveil the impacts of different hyper-factors on recommendation performance. Supported by the theoretical and experimental studies, we finally create benchmarks for rigorous evaluation by proposing standardized procedures and providing performance of ten state-of-the-arts across six evaluation metrics on six datasets as a reference for later study. Overall, our work sheds light on the issues in recommendation evaluation, provides potential solutions for rigorous evaluation, and lays foundation for further investigation. ",

keywords = "Benchmarks, fair comparison, recommender systems, reproducible evaluation, standardized procedures",

author = "Zhu Sun and Hui Fang and Jie Yang and Xinghua Qu and Hongyang Liu and Di Yu and Ong, {Yew Soon} and Jie Zhang",

note = "Green Open Access added to TU Delft Institutional Repository {\textquoteleft}You share, we take care!{\textquoteright} – Taverne project https://www.openaccess.nl/en/you-share-we-take-care Otherwise as indicated in the copyright section: the publisher is the copyright holder of this work and the author uses the Dutch legislation to make this work public.",

year = "2023",

doi = "10.1109/TPAMI.2022.3231891",

language = "English",

volume = "45",

pages = "8206--8226",

journal = "IEEE Transactions on Pattern Analysis and Machine Intelligence",

issn = "0162-8828",

publisher = "IEEE",

number = "7",

}

TY - JOUR

T1 - DaisyRec 2.0

T2 - Benchmarking Recommendation for Rigorous Evaluation

AU - Sun, Zhu

AU - Fang, Hui

AU - Yang, Jie

AU - Qu, Xinghua

AU - Liu, Hongyang

AU - Yu, Di

AU - Ong, Yew Soon

AU - Zhang, Jie

N1 - Green Open Access added to TU Delft Institutional Repository ‘You share, we take care!’ – Taverne project https://www.openaccess.nl/en/you-share-we-take-care Otherwise as indicated in the copyright section: the publisher is the copyright holder of this work and the author uses the Dutch legislation to make this work public.

PY - 2023

Y1 - 2023

N2 - Recently, one critical issue looms large in the field of recommender systems - there are no effective benchmarks for rigorous evaluation - which consequently leads to unreproducible evaluation and unfair comparison. We, therefore, conduct studies from the perspectives of practical theory and experiments, aiming at benchmarking recommendation for rigorous evaluation. Regarding the theoretical study, a series of hyper-factors affecting recommendation performance throughout the whole evaluation chain are systematically summarized and analyzed via an exhaustive review on 141 papers published at eight top-tier conferences within 2017-2020. We then classify them into model-independent and model-dependent hyper-factors, and different modes of rigorous evaluation are defined and discussed in-depth accordingly. For the experimental study, we release DaisyRec 2.0 library by integrating these hyper-factors to perform rigorous evaluation, whereby a holistic empirical study is conducted to unveil the impacts of different hyper-factors on recommendation performance. Supported by the theoretical and experimental studies, we finally create benchmarks for rigorous evaluation by proposing standardized procedures and providing performance of ten state-of-the-arts across six evaluation metrics on six datasets as a reference for later study. Overall, our work sheds light on the issues in recommendation evaluation, provides potential solutions for rigorous evaluation, and lays foundation for further investigation.

AB - Recently, one critical issue looms large in the field of recommender systems - there are no effective benchmarks for rigorous evaluation - which consequently leads to unreproducible evaluation and unfair comparison. We, therefore, conduct studies from the perspectives of practical theory and experiments, aiming at benchmarking recommendation for rigorous evaluation. Regarding the theoretical study, a series of hyper-factors affecting recommendation performance throughout the whole evaluation chain are systematically summarized and analyzed via an exhaustive review on 141 papers published at eight top-tier conferences within 2017-2020. We then classify them into model-independent and model-dependent hyper-factors, and different modes of rigorous evaluation are defined and discussed in-depth accordingly. For the experimental study, we release DaisyRec 2.0 library by integrating these hyper-factors to perform rigorous evaluation, whereby a holistic empirical study is conducted to unveil the impacts of different hyper-factors on recommendation performance. Supported by the theoretical and experimental studies, we finally create benchmarks for rigorous evaluation by proposing standardized procedures and providing performance of ten state-of-the-arts across six evaluation metrics on six datasets as a reference for later study. Overall, our work sheds light on the issues in recommendation evaluation, provides potential solutions for rigorous evaluation, and lays foundation for further investigation.

KW - Benchmarks

KW - fair comparison

KW - recommender systems

KW - reproducible evaluation

KW - standardized procedures

UR - http://www.scopus.com/inward/record.url?scp=85146238817&partnerID=8YFLogxK

U2 - 10.1109/TPAMI.2022.3231891

DO - 10.1109/TPAMI.2022.3231891

M3 - Article

AN - SCOPUS:85146238817

SN - 0162-8828

VL - 45

SP - 8206

EP - 8226

JO - IEEE Transactions on Pattern Analysis and Machine Intelligence

JF - IEEE Transactions on Pattern Analysis and Machine Intelligence

IS - 7

ER -

DaisyRec 2.0: Benchmarking Recommendation for Rigorous Evaluation

Abstract

Bibliographical note

Keywords

Access to Document

Other files and links

Fingerprint

Cite this