CORTEXA
← Browse
openalexMendeley Data2026-07-25Cited by 0

BanglaVowelDataset: True AC and Synthetic BC Bangla Vowel Datasets in Speech Information System

Ohidujjaman, Bejoy Munshi, Md. Mainul Hasan, M M Huda, Suman Ahmmed, Hasan Sarwar

The BanglaVowelDataset [1] resolves the unavailability of Bangla AC and synthetic BC vowel data in the speech information system. We recorded raw Bangla air-conducted (AC) vowels, with five male and five female speakers participating in the recording system, set up in a soundproof room. Male and female speakers were trained by experts to learn the proper pronunciation of Bangla vowels. The positions of the equipment are adjusted at specific distances. The speaker was sitting in front of the AC microphone, away from 30 cm. Afterward, we converted the recorded true AC vowels into synthetic BC vowels using an all-pole source filter model to obtain a synthetic Bangla BC dataset. The length of each recorded raw Bangla vowel ranges from 600 to 780 ms. The sampling frequency, audio resolution correspond to 44100 Hz and 32-bit, respectively. Raw speech data are recorded in .wav format. We set the stereo in Audacity audio recording software while recording vowels to have a more natural and immersive experience. We obtained a total of 120 true AC vowels from five male and five female speakers, as the Bangla alphabet has twelve vowels. The recoded each vowel is cropped to avoid the unvoiced portions. Thus, the cropped vowel length ranges from 500 to 650 ms. In addition, we created a synthetic dataset of the BC Bangla vowels containing 120 vowels sampled at 8000 Hz, having discarded the unvoiced portions. In this paper, we focus on two datasets: raw data consisting of 120 recorded Bangla AC vowels and synthetic data containing 120 synthetic Bangla BC vowels. There are twelve vowels in the Bangla alphabet due to diphthongs and variations. In the literature, the Bangla vowel speech dataset is not available yet. Consequently, this dataset provides a significant opportunity for further research. Bangla vowel dataset is used for noisy and clean environments comparatively, for speech recognition and speaker identification. Commonly, AC speech is severely affected by ambient noise, whereas BC speech suffers less. Thus, the appropriate method is suggested for noise robustness. This dataset has potential for pitch (F0) detection, spectrum estimation, and condition number determination in speech signal processing and analysis. Performance of distinct approaches, including machine learning, deep learning, and statistical methods, is evaluated using raw AC and synthetic BC vowels. The complete dataset is publicly accessible on the Mendeley Data repository, organized hierarchically with separate directories for all Bangla AC and BC vowels.

View free PDFSource page

Related papers

openalexMendeley Data2026-07-23

UFTBESD: University of Frontier Technology, Bangladesh-Bangla Emotional Speech Dataset

Md Rayhan Ali, Sadman Saeef, Mohammad Miftahul Islam Irfan Mohammad, Maliha Khan, Suchi Hasan, Sajib Das, et al.

UFTBESD (University of Frontier Technology, Bangladesh - Bangla Emotional Speech Dataset) is a Bangla-language speech emotion recognition dataset developed to capture realistic emotional speech under everyday acoustic conditions. The dataset consists of 1,400 audio recordings col…

View free PDFSource page
openalexMendeley Data2026-07-23

Data for: Wide-Range Predictions of Hydrogen-Dependent Vacancy Diffusion in Nickel from a near-DFT-Accurate Machine-Learning Potential

Si Zhu, Nobuyoshi Komai, Shihao Zhang, Shigenobu Ogata

This archive provides the reproducibility materials associated with the manuscript “Wide-Range Predictions of Hydrogen-Dependent Vacancy Diffusion in Nickel from a near-DFT-Accurate Machine-Learning Potential.” It contains the numerical data underlying the manuscript figures and…

View free PDFSource page
openalexMendeley Data2026-07-23

Longitudinal Indoor Air Quality Dataset Collected Using a Low-Cost Multi-Sensor IoT Monitoring Platform

Md Abubakar Siddique

This repository contains a longitudinal indoor air quality (IAQ) dataset collected with a Raspberry Pi 5–based multi-sensor IoT monitoring platform deployed in an indoor laboratory. The monitoring campaign spans approximately 15.6 days of continuous post-initialization operation…

View free PDFSource page
openalexMendeley Data2026-07-23

One year of five-minute monitoring data from a residential solar PV–battery system, Capiz, Philippines (Nov 2024 – Oct 2025): descriptive, predictive, and prescriptive analytics

Jeng Batacandolo

This dataset provides one year (1 November 2024 – 31 October 2025) of five-minute field measurements from an operating residential rooftop solar photovoltaic–battery system in Roxas City, Capiz, Philippines, together with satellite weather, system specifications, analysis code, a…

View free PDFSource page
openalexMendeley Data2026-07-23

WheatVision: A Dataset of Wheat Seed Quality Assessment

Sayali Shinde, Gouri Rewanwar, Dr.Deepa Abin, Deepak Parashar, Rahul Joshi, S.M. Bhoyar, et al.

This dataset presents a curated collection of high-resolution wheat seed images developed to advance automated quality assessment in agricultural grain inspection. Captured at the individual seed level, the dataset enables fine-grained classification between healthy seeds and mul…

View free PDFSource page