TY - GEN
T1 - A Comparison of LSTM and GRU for Bengali Speech-to-Text Transformation
AU - Jahan, Nusrat
AU - Sultana, Zakia
AU - Chowdhury, Fahim
AU - Ahmed, Sajjad
AU - Parvez, Mohammad Zavid
AU - Barua, Prabal Datta
AU - Chakraborty, Subrata
N1 - Presented by Fahim Chowdhury
PY - 2023/12/31
Y1 - 2023/12/31
N2 - This paper represents an approach to speech-to-text conversion in the Bengali language. In this area, we have found most of the methodologies were focused on other languages rather than Bengali. We started with a novel dataset of 56 unique words from 160 individual subjects was prepared. Then in this paper, we illustrate the approach to increasing accuracy in a speech-to-text over the Bengali language where initially we started with Gated Recurrent Unit(GRU) and Long short-term memory (LSTM) algorithms. During further observation, we found that the output of the GRU failed to give any stable output. So, we moved completely to the LSTM algorithm where we achieved 90% accuracy on an unexplored dataset. Voices of several demographic populations and noises were used to validate the model. In the testing phase, we tried a variety of classes based on their length, complexity, noise, and gender variant. Moreover, we expect that this research will help to develop a real-time Bengali speak-to-text recognition model.
AB - This paper represents an approach to speech-to-text conversion in the Bengali language. In this area, we have found most of the methodologies were focused on other languages rather than Bengali. We started with a novel dataset of 56 unique words from 160 individual subjects was prepared. Then in this paper, we illustrate the approach to increasing accuracy in a speech-to-text over the Bengali language where initially we started with Gated Recurrent Unit(GRU) and Long short-term memory (LSTM) algorithms. During further observation, we found that the output of the GRU failed to give any stable output. So, we moved completely to the LSTM algorithm where we achieved 90% accuracy on an unexplored dataset. Voices of several demographic populations and noises were used to validate the model. In the testing phase, we tried a variety of classes based on their length, complexity, noise, and gender variant. Moreover, we expect that this research will help to develop a real-time Bengali speak-to-text recognition model.
UR - https://www.scopus.com/pages/publications/85163312639
U2 - 10.1007/978-3-031-33743-7_18
DO - 10.1007/978-3-031-33743-7_18
M3 - Conference contribution
SN - 9783031337420
SN - 9783031337437
T3 - Lecture Notes in Networks and Systems
SP - 214
EP - 224
BT - ACR 2023: International Conference on Advances in Computing Research8th - 10th May, 2023
A2 - Daimi, Kevin
A2 - Sadoon, Abeer Al
PB - Springer
CY - Cham, Switzerland
T2 - ACR 2023: International Conference on Advances in Computing Research
Y2 - 8 May 2023 through 10 May 2023
ER -