What does BERT know about books, movies and music? Probing BERT for Conversational Recommendation

Gustavo Penha; Claudia Hauff

doi:10.1145/3383313.3412249

What does BERT know about books, movies and music? Probing BERT for Conversational Recommendation

Web Information Systems

Research output: Chapter in Book/Conference proceedings/Edited volume › Conference contribution › Scientific › peer-review

47 Citations (Scopus)

Abstract

Heavily pre-trained transformer models such as BERT have recently shown to be remarkably powerful at language modelling, achieving impressive results on numerous downstream tasks. It has also been shown that they implicitly store factual knowledge in their parameters after pre-training. Understanding what the pre-training procedure of LMs actually learns is a crucial step for using and improving them for Conversational Recommender Systems (CRS). We first study how much off-the-shelf pre-trained BERT "knows"about recommendation items such as books, movies and music. In order to analyze the knowledge stored in BERT's parameters, we use different probes (i.e., tasks to examine a trained model regarding certain properties) that require different types of knowledge to solve, namely content-based and collaborative-based. Content-based knowledge is knowledge that requires the model to match the titles of items with their content information, such as textual descriptions and genres. In contrast, collaborative-based knowledge requires the model to match items with similar ones, according to community interactions such as ratings. We resort to BERT's Masked Language Modelling (MLM) head to probe its knowledge about the genre of items, with cloze style prompts. In addition, we employ BERT's Next Sentence Prediction (NSP) head and representations' similarity (SIM) to compare relevant and non-relevant search and recommendation query-document inputs to explore whether BERT can, without any fine-tuning, rank relevant items first. Finally, we study how BERT performs in a conversational recommendation downstream task. To this end, we fine-tune BERT to act as a retrieval-based CRS. Overall, our experiments show that: (i) BERT has knowledge stored in its parameters about the content of books, movies and music; (ii) it has more content-based knowledge than collaborative-based knowledge; and (iii) fails on conversational recommendation when faced with adversarial data.

Original language	English
Title of host publication	RecSys 2020 - 14th ACM Conference on Recommender Systems
Publisher	Association for Computing Machinery (ACM)
Pages	388-397
Number of pages	10
ISBN (Electronic)	9781450375832
DOIs	https://doi.org/10.1145/3383313.3412249
Publication status	Published - 2020
Event	14th ACM Conference on Recommender Systems, RecSys 2020 - Virtual, Online, Brazil Duration: 22 Sept 2020 → 26 Sept 2020

Publication series

Name	RecSys 2020 - 14th ACM Conference on Recommender Systems

Conference

Conference	14th ACM Conference on Recommender Systems, RecSys 2020
Country/Territory	Brazil
City	Virtual, Online
Period	22/09/20 → 26/09/20

Keywords

conversational recommendation
conversational search
probing

Access to Document

10.1145/3383313.3412249

Cite this

@inproceedings{a6bd968909f24e2eb264217c80a3d264,

title = "What does BERT know about books, movies and music? Probing BERT for Conversational Recommendation",

abstract = "Heavily pre-trained transformer models such as BERT have recently shown to be remarkably powerful at language modelling, achieving impressive results on numerous downstream tasks. It has also been shown that they implicitly store factual knowledge in their parameters after pre-training. Understanding what the pre-training procedure of LMs actually learns is a crucial step for using and improving them for Conversational Recommender Systems (CRS). We first study how much off-the-shelf pre-trained BERT {"}knows{"}about recommendation items such as books, movies and music. In order to analyze the knowledge stored in BERT's parameters, we use different probes (i.e., tasks to examine a trained model regarding certain properties) that require different types of knowledge to solve, namely content-based and collaborative-based. Content-based knowledge is knowledge that requires the model to match the titles of items with their content information, such as textual descriptions and genres. In contrast, collaborative-based knowledge requires the model to match items with similar ones, according to community interactions such as ratings. We resort to BERT's Masked Language Modelling (MLM) head to probe its knowledge about the genre of items, with cloze style prompts. In addition, we employ BERT's Next Sentence Prediction (NSP) head and representations' similarity (SIM) to compare relevant and non-relevant search and recommendation query-document inputs to explore whether BERT can, without any fine-tuning, rank relevant items first. Finally, we study how BERT performs in a conversational recommendation downstream task. To this end, we fine-tune BERT to act as a retrieval-based CRS. Overall, our experiments show that: (i) BERT has knowledge stored in its parameters about the content of books, movies and music; (ii) it has more content-based knowledge than collaborative-based knowledge; and (iii) fails on conversational recommendation when faced with adversarial data.",

keywords = "conversational recommendation, conversational search, probing",

author = "Gustavo Penha and Claudia Hauff",

year = "2020",

doi = "10.1145/3383313.3412249",

language = "English",

series = "RecSys 2020 - 14th ACM Conference on Recommender Systems",

publisher = "Association for Computing Machinery (ACM)",

pages = "388--397",

booktitle = "RecSys 2020 - 14th ACM Conference on Recommender Systems",

address = "United States",

note = "14th ACM Conference on Recommender Systems, RecSys 2020 ; Conference date: 22-09-2020 Through 26-09-2020",

}

Penha, G & Hauff, C 2020, What does BERT know about books, movies and music? Probing BERT for Conversational Recommendation. in RecSys 2020 - 14th ACM Conference on Recommender Systems. RecSys 2020 - 14th ACM Conference on Recommender Systems, Association for Computing Machinery (ACM), pp. 388-397, 14th ACM Conference on Recommender Systems, RecSys 2020, Virtual, Online, Brazil, 22/09/20. https://doi.org/10.1145/3383313.3412249

What does BERT know about books, movies and music? Probing BERT for Conversational Recommendation. / Penha, Gustavo; Hauff, Claudia.
RecSys 2020 - 14th ACM Conference on Recommender Systems. Association for Computing Machinery (ACM), 2020. p. 388-397 (RecSys 2020 - 14th ACM Conference on Recommender Systems).

Research output: Chapter in Book/Conference proceedings/Edited volume › Conference contribution › Scientific › peer-review

TY - GEN

T1 - What does BERT know about books, movies and music? Probing BERT for Conversational Recommendation

AU - Penha, Gustavo

AU - Hauff, Claudia

PY - 2020

Y1 - 2020

N2 - Heavily pre-trained transformer models such as BERT have recently shown to be remarkably powerful at language modelling, achieving impressive results on numerous downstream tasks. It has also been shown that they implicitly store factual knowledge in their parameters after pre-training. Understanding what the pre-training procedure of LMs actually learns is a crucial step for using and improving them for Conversational Recommender Systems (CRS). We first study how much off-the-shelf pre-trained BERT "knows"about recommendation items such as books, movies and music. In order to analyze the knowledge stored in BERT's parameters, we use different probes (i.e., tasks to examine a trained model regarding certain properties) that require different types of knowledge to solve, namely content-based and collaborative-based. Content-based knowledge is knowledge that requires the model to match the titles of items with their content information, such as textual descriptions and genres. In contrast, collaborative-based knowledge requires the model to match items with similar ones, according to community interactions such as ratings. We resort to BERT's Masked Language Modelling (MLM) head to probe its knowledge about the genre of items, with cloze style prompts. In addition, we employ BERT's Next Sentence Prediction (NSP) head and representations' similarity (SIM) to compare relevant and non-relevant search and recommendation query-document inputs to explore whether BERT can, without any fine-tuning, rank relevant items first. Finally, we study how BERT performs in a conversational recommendation downstream task. To this end, we fine-tune BERT to act as a retrieval-based CRS. Overall, our experiments show that: (i) BERT has knowledge stored in its parameters about the content of books, movies and music; (ii) it has more content-based knowledge than collaborative-based knowledge; and (iii) fails on conversational recommendation when faced with adversarial data.

AB - Heavily pre-trained transformer models such as BERT have recently shown to be remarkably powerful at language modelling, achieving impressive results on numerous downstream tasks. It has also been shown that they implicitly store factual knowledge in their parameters after pre-training. Understanding what the pre-training procedure of LMs actually learns is a crucial step for using and improving them for Conversational Recommender Systems (CRS). We first study how much off-the-shelf pre-trained BERT "knows"about recommendation items such as books, movies and music. In order to analyze the knowledge stored in BERT's parameters, we use different probes (i.e., tasks to examine a trained model regarding certain properties) that require different types of knowledge to solve, namely content-based and collaborative-based. Content-based knowledge is knowledge that requires the model to match the titles of items with their content information, such as textual descriptions and genres. In contrast, collaborative-based knowledge requires the model to match items with similar ones, according to community interactions such as ratings. We resort to BERT's Masked Language Modelling (MLM) head to probe its knowledge about the genre of items, with cloze style prompts. In addition, we employ BERT's Next Sentence Prediction (NSP) head and representations' similarity (SIM) to compare relevant and non-relevant search and recommendation query-document inputs to explore whether BERT can, without any fine-tuning, rank relevant items first. Finally, we study how BERT performs in a conversational recommendation downstream task. To this end, we fine-tune BERT to act as a retrieval-based CRS. Overall, our experiments show that: (i) BERT has knowledge stored in its parameters about the content of books, movies and music; (ii) it has more content-based knowledge than collaborative-based knowledge; and (iii) fails on conversational recommendation when faced with adversarial data.

KW - conversational recommendation

KW - conversational search

KW - probing

UR - http://www.scopus.com/inward/record.url?scp=85092740676&partnerID=8YFLogxK

U2 - 10.1145/3383313.3412249

DO - 10.1145/3383313.3412249

M3 - Conference contribution

AN - SCOPUS:85092740676

T3 - RecSys 2020 - 14th ACM Conference on Recommender Systems

SP - 388

EP - 397

BT - RecSys 2020 - 14th ACM Conference on Recommender Systems

PB - Association for Computing Machinery (ACM)

T2 - 14th ACM Conference on Recommender Systems, RecSys 2020

Y2 - 22 September 2020 through 26 September 2020

ER -

What does BERT know about books, movies and music? Probing BERT for Conversational Recommendation

Abstract

Publication series

Conference

Keywords

Access to Document

Other files and links

Fingerprint

Cite this