Speaker Recognition System
Speaker Recognition is the task of recognizing the speaker from his/her voice. The technique utilizes Acoustic features of speech that have found to be unique for an individual. These acoustic patterns reflect both, the anatomy and the behavioural patterns of the speaker. The pr
2025-06-28 16:36:06 - Adil Khan
Speaker Recognition System
Project Area of Specialization Artificial IntelligenceProject SummarySpeaker Recognition is the task of recognizing the speaker from his/her voice.
The technique utilizes Acoustic features of speech that have found to be unique
for an individual. These acoustic patterns reflect both, the anatomy and the
behavioural patterns of the speaker. The problem of speaker recognition arises
in two flavours. One, where the speaker is allowed to speak only a fixed text,
called Text Dependent Speaker Recognition. Other, where the speaker is free
to say anything, called the Text Independent Speaker Recognition.
The application of such systems lie in various kinds of security systems, since
the anatomy of the vocal tract is unique for an individual. It can also enhance
the human computer interaction.
Project ObjectivesThe proposed project is supposed to make use of speech recognition feature in order to create a bot that could communicate and help normal person or disabled peoples to use their computer without typing and clicking.
Project Implementation MethodDigital speech is a one-dimensional time-varying discrete signal. Sound are pressure waves, and these waves can be represented by numbers over a time period. These air pressure differences communicate with the brain. Audio files are generally stored in .wav, .flac, .mp3 etc format and need to be digitized, using the concept of sampling.
The sampling frequency (or sample rate) is the number of samples (data points) per second in a round. For example: if the sampling frequency is 44 kHz, a recording with a duration of 60 seconds will contain 2,646,000 samples. In practice, sampling even higher than 10x helps measure the amplitude correctly in the time domain.
In this project we have use machine learning algorithms for feature extraction and feature modelling. A popular machine learning approach will be used to select the most appropriate and relatively important features that will greatly help to make models for accurate classification. Then Machine learning optimization technique will be applied for obtaining the optimal solution. The basic aim of project is to design a novel, cost effective, enhanced features for the classification of speakers with improved performance and efficacy.
Technical Details of Final DeliverableData Collection
Preprocessing
Feature Extraction
Modeling and Training and Predicting
Final Deliverable of the Project Software SystemCore Industry ITOther IndustriesCore Technology Artificial Intelligence(AI)Other TechnologiesSustainable Development Goals Quality EducationRequired Resources| Item Name | Type | No. of Units | Per Unit Cost (in Rs) | Total (in Rs) |
|---|---|---|---|---|
| Total in (Rs) | 50000 | |||
| Total Hardware and Assembling | Equipment | 1 | 50000 | 50000 |