Discovering phonetic inventories with crosslingual automatic speech recognition

Piotr Żelasko; Siyuan Feng; Laureano Moro Velázquez; Ali Abavisani; Saurabhchand Bhati; Odette Scharenborg; Mark Hasegawa-Johnson; Najim Dehak

doi:10.1016/j.csl.2022.101358

Discovering phonetic inventories with crosslingual automatic speech recognition

Piotr Żelasko^*, Siyuan Feng, Laureano Moro Velázquez, Ali Abavisani, Saurabhchand Bhati, Odette Scharenborg, Mark Hasegawa-Johnson, Najim Dehak

^*Corresponding author for this work

Multimedia Computing

Research output: Contribution to journal › Article › Scientific › peer-review

9 Citations (Scopus)

17 Downloads (Pure)

Abstract

The high cost of data acquisition makes Automatic Speech Recognition (ASR) model training problematic for most existing languages, including languages that do not even have a written script, or for which the phone inventories remain unknown. Past works explored multilingual training, transfer learning, as well as zero-shot learning in order to build ASR systems for these low-resource languages. While it has been shown that the pooling of resources from multiple languages is helpful, we have not yet seen a successful application of an ASR model to a language unseen during training. A crucial step in the adaptation of ASR from seen to unseen languages is the creation of the phone inventory of the unseen language. The ultimate goal of our work is to build the phone inventory of a language unseen during training in an unsupervised way without any knowledge about the language. In this paper, we (1) investigate the influence of different factors (i.e., model architecture, phonotactic model, type of speech representation) on phone recognition in an unknown language; (2) provide an analysis of which phones transfer well across languages and which do not in order to understand the limitations of and areas for further improvement for automatic phone inventory creation; and (3) present different methods to build a phone inventory of an unseen language in an unsupervised way. To that end, we conducted mono-, multi-, and crosslingual experiments on a set of 13 phonetically diverse languages and several in-depth analyses. We found a number of universal phone tokens (IPA symbols) that are well-recognized cross-linguistically. Through a detailed analysis of results, we conclude that unique sounds, similar sounds, and tone languages remain a major challenge for phonetic inventory discovery.

Original language	English
Article number	101358
Number of pages	23
Journal	Computer Speech and Language
Volume	74
DOIs	https://doi.org/10.1016/j.csl.2022.101358
Publication status	Published - 2022

Bibliographical note

Green Open Access added to TU Delft Institutional Repository ‘You share, we take care!’ – Taverne project https://www.openaccess.nl/en/you-share-we-take-care
Otherwise as indicated in the copyright section: the publisher is the copyright holder of this work and the author uses the Dutch legislation to make this work public.

Keywords

ASR
Crosslingual
Multilingual
Phone inventory
Phone recognition
Speech recognition
Speech representation
Zero-shot

Access to Document

10.1016/j.csl.2022.101358

1_s2.0_S0885230822000067_main_1Final published version, 953 KB

Cite this

@article{632a3af386e749d0b30b651336424815,

title = "Discovering phonetic inventories with crosslingual automatic speech recognition",

abstract = "The high cost of data acquisition makes Automatic Speech Recognition (ASR) model training problematic for most existing languages, including languages that do not even have a written script, or for which the phone inventories remain unknown. Past works explored multilingual training, transfer learning, as well as zero-shot learning in order to build ASR systems for these low-resource languages. While it has been shown that the pooling of resources from multiple languages is helpful, we have not yet seen a successful application of an ASR model to a language unseen during training. A crucial step in the adaptation of ASR from seen to unseen languages is the creation of the phone inventory of the unseen language. The ultimate goal of our work is to build the phone inventory of a language unseen during training in an unsupervised way without any knowledge about the language. In this paper, we (1) investigate the influence of different factors (i.e., model architecture, phonotactic model, type of speech representation) on phone recognition in an unknown language; (2) provide an analysis of which phones transfer well across languages and which do not in order to understand the limitations of and areas for further improvement for automatic phone inventory creation; and (3) present different methods to build a phone inventory of an unseen language in an unsupervised way. To that end, we conducted mono-, multi-, and crosslingual experiments on a set of 13 phonetically diverse languages and several in-depth analyses. We found a number of universal phone tokens (IPA symbols) that are well-recognized cross-linguistically. Through a detailed analysis of results, we conclude that unique sounds, similar sounds, and tone languages remain a major challenge for phonetic inventory discovery.",

keywords = "ASR, Crosslingual, Multilingual, Phone inventory, Phone recognition, Speech recognition, Speech representation, Zero-shot",

author = "Piotr {\.Z}elasko and Siyuan Feng and {Moro Vel{\'a}zquez}, Laureano and Ali Abavisani and Saurabhchand Bhati and Odette Scharenborg and Mark Hasegawa-Johnson and Najim Dehak",

note = "Green Open Access added to TU Delft Institutional Repository {\textquoteleft}You share, we take care!{\textquoteright} – Taverne project https://www.openaccess.nl/en/you-share-we-take-care Otherwise as indicated in the copyright section: the publisher is the copyright holder of this work and the author uses the Dutch legislation to make this work public. ",

year = "2022",

doi = "10.1016/j.csl.2022.101358",

language = "English",

volume = "74",

journal = "Computer Speech and Language",

issn = "0885-2308",

publisher = "Academic Press",

}

TY - JOUR

T1 - Discovering phonetic inventories with crosslingual automatic speech recognition

AU - Żelasko, Piotr

AU - Feng, Siyuan

AU - Moro Velázquez, Laureano

AU - Abavisani, Ali

AU - Bhati, Saurabhchand

AU - Scharenborg, Odette

AU - Hasegawa-Johnson, Mark

AU - Dehak, Najim

N1 - Green Open Access added to TU Delft Institutional Repository ‘You share, we take care!’ – Taverne project https://www.openaccess.nl/en/you-share-we-take-care Otherwise as indicated in the copyright section: the publisher is the copyright holder of this work and the author uses the Dutch legislation to make this work public.

PY - 2022

Y1 - 2022

N2 - The high cost of data acquisition makes Automatic Speech Recognition (ASR) model training problematic for most existing languages, including languages that do not even have a written script, or for which the phone inventories remain unknown. Past works explored multilingual training, transfer learning, as well as zero-shot learning in order to build ASR systems for these low-resource languages. While it has been shown that the pooling of resources from multiple languages is helpful, we have not yet seen a successful application of an ASR model to a language unseen during training. A crucial step in the adaptation of ASR from seen to unseen languages is the creation of the phone inventory of the unseen language. The ultimate goal of our work is to build the phone inventory of a language unseen during training in an unsupervised way without any knowledge about the language. In this paper, we (1) investigate the influence of different factors (i.e., model architecture, phonotactic model, type of speech representation) on phone recognition in an unknown language; (2) provide an analysis of which phones transfer well across languages and which do not in order to understand the limitations of and areas for further improvement for automatic phone inventory creation; and (3) present different methods to build a phone inventory of an unseen language in an unsupervised way. To that end, we conducted mono-, multi-, and crosslingual experiments on a set of 13 phonetically diverse languages and several in-depth analyses. We found a number of universal phone tokens (IPA symbols) that are well-recognized cross-linguistically. Through a detailed analysis of results, we conclude that unique sounds, similar sounds, and tone languages remain a major challenge for phonetic inventory discovery.

AB - The high cost of data acquisition makes Automatic Speech Recognition (ASR) model training problematic for most existing languages, including languages that do not even have a written script, or for which the phone inventories remain unknown. Past works explored multilingual training, transfer learning, as well as zero-shot learning in order to build ASR systems for these low-resource languages. While it has been shown that the pooling of resources from multiple languages is helpful, we have not yet seen a successful application of an ASR model to a language unseen during training. A crucial step in the adaptation of ASR from seen to unseen languages is the creation of the phone inventory of the unseen language. The ultimate goal of our work is to build the phone inventory of a language unseen during training in an unsupervised way without any knowledge about the language. In this paper, we (1) investigate the influence of different factors (i.e., model architecture, phonotactic model, type of speech representation) on phone recognition in an unknown language; (2) provide an analysis of which phones transfer well across languages and which do not in order to understand the limitations of and areas for further improvement for automatic phone inventory creation; and (3) present different methods to build a phone inventory of an unseen language in an unsupervised way. To that end, we conducted mono-, multi-, and crosslingual experiments on a set of 13 phonetically diverse languages and several in-depth analyses. We found a number of universal phone tokens (IPA symbols) that are well-recognized cross-linguistically. Through a detailed analysis of results, we conclude that unique sounds, similar sounds, and tone languages remain a major challenge for phonetic inventory discovery.

KW - ASR

KW - Crosslingual

KW - Multilingual

KW - Phone inventory

KW - Phone recognition

KW - Speech recognition

KW - Speech representation

KW - Zero-shot

UR - http://www.scopus.com/inward/record.url?scp=85124877240&partnerID=8YFLogxK

U2 - 10.1016/j.csl.2022.101358

DO - 10.1016/j.csl.2022.101358

M3 - Article

AN - SCOPUS:85124877240

SN - 0885-2308

VL - 74

JO - Computer Speech and Language

JF - Computer Speech and Language

M1 - 101358

ER -

Discovering phonetic inventories with crosslingual automatic speech recognition

Abstract

Bibliographical note

Keywords

Access to Document

Other files and links

Fingerprint

Cite this