# Speech Recognition using Correlation

Ask anyone who has studied MATLAB and they would tell you that, MATLAB is very powerful and easy to use. I however never properly implemented a project in MATLAB. I have done small laboratory experiments in college to determine magnetic field or matrix multiplications and plot various graphs.  (I am not talking about Simulink, that is a whole other thing.) So, when I had to a project for my class, signals and systems, I need to do a project related to signals. So, I looked to do some signal processing with MATLAB. In this class we learned Fourier transform and Laplace transforms along with its inverse on continuous and discreate signals. 

I had also done a course on digital signal processing which gave me a good idea about signal processing, other than this I had also done a few online courses about [audio signal processing for musical application](https://www.coursera.org/learn/audio-signal-processing/home/welcome) along with few courses on [Acoustics](https://www.coursera.org/learn/intro-to-acoustics/home/welcome), I have also learned music theory which gives me a good understanding about audio signals and its processing. At first, I wanted to a project which computed the fast Fourier transform of a signal or took the Fourier transform of a musical signal to get some information about the sound. However, while I was looking through few research papers, I found something much more interesting, Correlation. 

Almost everyone knows what Correlation is and what is implies, but not many people know how useful it can be. This project used Correlation on audio signals to recognize Speech. The project is actually quite simple but could have some very useful applications. The project also could be applied in the field of linguistics, which is also very and [Tom Scott's](https://youtube.com/playlist?list=PL96C35uN7xGLDEnHuhD7CTZES3KXFnwm0) YouTube videos is a great example of that. We had to do the project in pairs, so I teamed up with my friend Ishaan who also interested to do this project.
## Principle 
The basis for this project is the “xcorr” function in MATLAB, which just compare the two signals to figure out if it is similar. Now it was time to come up with a goal for this project, or how we would implement this concept. The first idea was to do “voice recognition”, in this application we would have person 1 say a word like "password". We would then get person 2 to say the same word and see how similar it was. We did try some initial tests with this, however we found out that when we get the same person to say the same word twice, it is still very different and when another person says the same word it isn’t that different in terms of its correlation. We did see this coming, because this application can be done with another concept. What makes each person’s voice different from another voice, is the harmonics distribution. For example, I might have a voice which would have a very dominant 3rd harmonic but minimal 5th harmonic. My friend would have a dominant 5th harmonic and a minimal 3rd harmonic. Therefore, the best way to distinguish between the voices of different people is thorough a frequency response plot. What we are doing in our project is just analysing the peeks of the audio waveforms.

The second idea we had was to implement “word recognition” (not voice recognition). The best way to explain what we want to do is through the story of Alibaba and the 40 thieves, a children’s story that everyone would know. (read it [Here](https://en.wikipedia.org/wiki/Ali_Baba_and_the_Forty_Thieves#:~:text=Ali%20Baba%20marries%20a%20poor,sealed%20by%20a%20huge%20rock.) ). Basically, we need a recording of a secret word like "open Sesame" and we should be able to recognize this word irrespective of who says it. We must also make sure that no other word like "open Samardo" must be confused as the secret word. Now that we got a good idea of what we want to implement let’s get to do coding.  One more thing we noted was that we must use words with few syllables for the best matches, when we use words with more syllables it becomes harder and harder to match. We need a list of five words and we will choose one of that as the secrete word. Due to lack of creativity in coming up with 5 small word I used the words:  one, two, three, four, five.

### Audio Files
The first thing I needed to do was read about was the different types of audio files and how they look when they are plotted. When I record a audio file on my laptop, with the Microsoft voice recorder app it gets stored as a .wav file. This can be used by MATLAB, so this is the format we will be using.  The second thing I did was install [Audacity](https://www.audacityteam.org/download/) so that I can edit and view these audio files. I will not be doing any editing other than trimming the files, to get rid of any blank spaces and leave only the useful part. Now we have the signals that we are going to be using all that is left is to write code. The first step is actually to go to the MATLAB documentation on the function "xcorr" which you can find [here](https://www.mathworks.com/help/matlab/ref/xcorr.html) we can see that it returns the cross corelation of two discrete time sequences. "Cross-correlation measures the similarity between a vector x and shifted (lagged) copies of a vector y as a function of the lag. If x and y have different lengths, the function appends zeros to the end of the shorter vector so it has the same length as the other." Understanding this is crucial to our project. 
We have also used two other functions in our code  ["soundsc"](https://in.mathworks.com/help/matlab/ref/soundsc.html)  which send the audio to the speaker so it can be heard from the speakers of the laptop and  ["audioread"](https://in.mathworks.com/help/matlab/ref/audioread.html)  which analysis the peaks of the vector.

## Algorithm
```
function speechrecognition(filename)
Input: Upload 5 sample Files m1, m2, m3, m4, m5 and the test file. 
Output: Correlation result of m and test file.  
1: Consider sample as voice where x=voice       
    Read and compute x and store in y1  
2: z1=xcorr(x.y1)       
    m1=max(z1)     
    l1=length(z1)    
    t1= -((l1-1)/2):1((l1-1)/2);
3: plot(t1,z1)   
4: Repeat steps 1,2,3 for all 5 samples.   
5: Consider a=[m1 m2 m3 m4 m5 m6] where m6=300   
6: Compute m=max(a)   
7: If m<=m1
    	read 1st file
     elseif m<=m2                  
        read 2nd file           
     elseif m<=m3                    
        read 3rd file           
     elseif m<=m4                     
        read 4th file           
     elseif m<=m5                     
        read 5th file        
     else        
        read denied file   
8: End
```
We decided to implement this using a function making it easy to demonstrate the working.

## Code

```
function speechrecognition(filename) //function definition
clc  //to clear all previous outputs
voice=audioread(filename); //to read the audio file we want to test
X=voice;
X=X'; //transposing the file and converting it to a vector
X=X(1,:);
Y1=audioread('one.wav'); //opening the first sample file 
Y1=Y1';
Y1=Y1(1,:);
cor1=xcorr(X,Y1); //comparing the first sample file to the input file
maxf1=max(cor1);
length1=length(cor1);
range1=-((length1-1)/2):1:((length1-1)/2);
plot(range1,cor1); //plotting the first sample file
title('One');
Y2=audioread('two.wav');  //opening the second sample file
Y2=Y2';
Y2=Y2(1,:);
cor2=xcorr(X,Y2); //comparing the second sample file to the input file
maxf2=max(cor2);
length2=length(cor2);
range2=-((length2-1)/2):1:((length2-1)/2);
figure
plot(range2,cor2); //plotting the second sample file
title('Two');
Y3=audioread('three.wav'); //opening the third sample file
Y3=Y3';
Y3=Y3(1,:);
cor3=xcorr(X,Y3); //comparing the third sample file to the input file
maxf3=max(cor3);
length3=length(cor3);
range3=-((length3-1)/2):1:((length3-1)/2);
figure
plot(range3,cor3); //plotting the third sample file
title('Three');
Y4=audioread('four.wav'); //opening the fouth sample file
Y4=Y4';Y4=Y4(1,:);
cor4=xcorr(X,Y4); //comparing the fourth sample file to the input file
maxf4=max(cor4);
length4=length(cor4);
range4=-((length4-1)/2):1:((length4-1)/2);
figure
plot(range4,cor4); //plotting the fourth sample file
title('Four');
Y5=audioread('five.wav'); //opening the fifth sample file
Y5=Y5';
Y5=Y5(1,:);
cor5=xcorr(X,Y5); //comparing the fifth sample file to the input file
maxf5=max(cor5);
length5=length(cor5);
range5=-((length5-1)/2):1:((length5-1)/2);
figure
plot(range5,cor5); //plotting the fifth sample file
title('Five');
maxf6=300;
Matrix=[maxf1 maxf2 maxf3 maxf4 maxf5 maxf6];
maxf=max(Matrix); // finding the most similar file
allowcd=audioread('allow.wav');
if maxf<=maxf1
soundsc(audioread('one.wav'),50000)
soundsc(allowcd,50000)
1
elseif maxf<=maxf2
soundsc(audioread('two.wav'),50000)
2
soundsc(allowcd,50000)
elseif maxf<=maxf3
soundsc(audioread('three.wav'),50000)
3
soundsc(allowcd,50000)
elseif maxf<=maxf4
soundsc(audioread('four.wav'),50000)
4
soundsc(allowcd,50000)
elseif maxf<maxf5
soundsc(audioread('five.wav'),50000)
5
soundsc(allowcd,50000)
else
(soundsc(audioread('denied.wav'),50000)) 
denied
```
The best part of using MATLAB while coding is the precise error messages that you receive. When there is a error it is clearly stated where the error is and what it is along with various possible alternatives. As you can see in addition to playing the file which is most similar to the input file, we also plot all five files so it can be understood visually. 

## Output

![image.png](https://cdn.hashnode.com/res/hashnode/image/upload/v1612374505352/gzlh8tVja.png)
  
When using files in MATLAB it is important to have them all in same folder. Here I have the sample files one, two, three, four, five where it contains me saying the words one, two, three, four, five respectively. There are also two testing files test1 which contains me saying the word “one” and test2 which contains me saying the word “two”. There is also a file of me saying “hello” so check for false positives. The other two files are me saying “allowed” and “denied”.  
 
If we run the program to test with the test1 file we type speechrecognition(“test1.wav”) and we first get the 5 plots of all the waveforms, and then the file “allowed” is played and then the file “one” along with the text “one” printed on the prompt. This is because the program matches the test1 file with sample file one. If we run the program to test with the test2 file we type speechrecognition(“test2.wav”) and this type instead of one we will get two as it matches with the second sample file.
To test for false positives we can try testing with the audio file “hello”, speechrecognition(“hello.wav”) and we get the plot of all 5 waveforms along with the audio “denied” as it did not match with any of the five sample files. 
 
## Conclusion 
I am pretty happy with how this project turned out, we organised our project into a paper and now looking to publish it. I am surprised by how easy the project was and how effective it is in doing what it is supposed to. When I think about how I can expand this project there is not much I can do really as the xcorr function is only so much, I mean if I really wanted precision and accuracy, I would use a frequency response and Fourier transforms, but by using cross correlation we get to truly understand what correlation is and how it can be used. Maybe I could use this to implement my own “open Sesame” cave with an Arduino and server motor. I actually heard that there was a way to program an Arduino with MATLAB, that would be an interesting project, stay tuned for that.

## Files
All the audio files and MATLAB code can be found on the [git link ](https://github.com/daniboi16/Speech-Recognition) i have made two MATLAB functions, one which just plots the waveforms and one which plots the waveforms as well as match the input. 

## References
[1] [Pramanik, A., & Raha, R. (2012, October). Automatic Speech Recognition using correlation analysis. In 2012 World congress on information and communication technologies (pp. 670-674). IEEE.](https://ieeexplore.ieee.org/abstract/document/6409160/) 

[2]  [Kabari, L. G., & Chigoziri, M. B. (2019). Speech Recognition Using MATLAB and Cross-Correlation Technique. European Journal of Engineering and Technology Research, 4(8), 1-3.](https://ejers.org/index.php/ejers/article/view/1437) 

[3]  [Gupta, A., Raibagkar, P., & Palsokar, A. (2017). Speech Recognition Using Correlation Technique. International Journal of Current Trends in Engineering & Research (IJCTER) e-ISSN, 2455-1392.](https://d1wqtxts1xzle7.cloudfront.net/60834192/Speech_Recognition_Using_Correlation_Technique201720191008-62792-sr74ah.pdf?1570531726=&response-content-disposition=inline%3B+filename%3DSpeech_Recognition_Using_Correlation_Tec.pdf&Expires=1612362223&Signature=TTE2PmFtRs0i7pWFCsCsOmqztGO5sfCYpC0zbLH2dKXj5IFoBQ084qzZ79n~xHjpGMaA0KC9SpFoQ6IfLljQJ~DfjSKXLnQP-0L-V~v2r~rnPzEwchwL6e6qQadB5OO7dAY6S0UtbOcGRDWJy~YCpE0khnlsWw2Ee5RmU0zcNB0FVhHR4jkgqGakq7IfS2gP99zBPZ21tt2qn1b3tgNE9q-OU3QIUkW8a-BHqPvBHUYuVuX~-BMAZfwWdc0PXtfdJbP-FjrNM~SQXJYWdEC4iZObAix~L8hbzF4q17SeR-pIXSnEQwOhbaDnDjqhs~e7-rs4u2whiBAPBTuRNeJ3bA__&Key-Pair-Id=APKAJLOHF5GGSLRBV4ZA) 

