Multi-Stage Unsupervised Domain Adaptation For Automatic Speech Recognition

Δαμιανός, Δημήτρης

National Technical University of Athens

School of Electrical and Computer Engineering

Artemis is Live!

Welcome to our digital repository! The aim of Artemis is the systematic archiving and dissemination of the scientific work produced in the School of Electrical and Computer Engineering, National Technical University of Athens, Greece, using the technology of digital libraries.

Please use this identifier to cite or link to this item: http://artemis.cslab.ece.ntua.gr:8080/jspui/handle/123456789/19363

Full metadata record

DC Field	Value	Language
dc.contributor.author	Δαμιανός, Δημήτρης	-
dc.date.accessioned	2024-10-31T09:47:50Z	-
dc.date.available	2024-10-31T09:47:50Z	-
dc.date.issued	2024-10-18	-
dc.identifier.uri	http://artemis.cslab.ece.ntua.gr:8080/jspui/handle/123456789/19363	-
dc.description.abstract	The purpose of this diploma thesis is to study unsupervised domain adaptation for Automatic Speech Recognition. In the context of unsupervised domain adaptation, we work with two distinct data distributions, the source domain and the target domain. While both domains have available input data, corresponding labels are only accessible in the source domain. The goal is to develop a model that can be effectively applied to the target domain, leveraging both the available labeled and unlabeled data. In this dissertation, we discuss the fundamentals of machine learning and the challenges associated with speech recognition, covering both traditional and modern approaches. We then review the literature on domain adaptation methods, categorizing these approaches into three major groups, including semi-supervised learning and self-supervision techniques. In the present work, we explore the capabilities of the Meta PL domain adaptation framework - previously applied to image recognition task- for Automatic Speech Recognition. Additionally, we introduce Multi-Stage Domain Adaptation, a two-stage domain adaptation method that combines self-supervised strategies with semi-supervised techniques. Multi-Stage Domain Adaptation is designed to enhance the robustness and generalization of Automatic Speech Recognition models in the context of low-resource languages, such as Greek, and weakly supervised data where labeled data is scarce or noisy. Our extensive experiments show that Meta PL can be effectively applied to Automatic Speech Recognition tasks, resulting in an average WER improvement of 4%. Additionally, we demonstrate that Multi-Stage Domain Adaptation outperforms our baselines WER by 7% on average, providing a more robust solution for domain adaptation in Automatic Speech Recognition, especially in underrepresented linguistic settings. Finally, we examine the limitations of integrating self-supervised tasks with semi-supervised training within the Meta PL framework and conclude that self-supervised tasks should be applied separately from semi-supervised training.	en_US
dc.language	en	en_US
dc.subject	Unsupervised Domain Adaptation	en_US
dc.subject	Automatic Speech Recognition	en_US
dc.subject	Self-supervision	en_US
dc.subject	Semi-supervision	en_US
dc.subject	Pseudo-labeling	en_US
dc.title	Multi-Stage Unsupervised Domain Adaptation For Automatic Speech Recognition	en_US
dc.description.pages	97	en_US
dc.contributor.supervisor	Ποταμιάνος Αλέξανδρος	en_US
dc.department	Τομέας Σημάτων, Ελέγχου και Ρομποτικής	en_US
Appears in Collections:	Διπλωματικές Εργασίες - Theses

Files in This Item:

File	Description	Size	Format
damianos_thesis.pdf		2.29 MB	Adobe PDF	View/Open

Show simple item record