# Is it possible to accomplish speech recognition in live streaming using speech-to-text services from Google or Azure?

**URL:** https://community.wowza.com/t/is-it-possible-to-accomplish-speech-recognition-in-live-streaming-using-speech-to-text-services-from-google-or-azure/96361
**Category:** Wowza Streaming Engine
**Created:** [August 4, 2023, 7:52am UTC](https://community.wowza.com/t/is-it-possible-to-accomplish-speech-recognition-in-live-streaming-using-speech-to-text-services-from-google-or-azure/96361 "2023-08-04T07:52:27Z")
**Posts on this page:** 9
**Page:** 1

<div class="post-metadata">

### Author: ![Spencer\_Lin2](https://avatars.discourse-cdn.com/v4/letter/s/eb8c5e/32.png) [@Spencer\_Lin2](https://community.wowza.com/u/Spencer_Lin2)
#### Post date: [August 4, 2023, 7:52am UTC](https://community.wowza.com/t/is-it-possible-to-accomplish-speech-recognition-in-live-streaming-using-speech-to-text-services-from-google-or-azure/96361/1 "2023-08-04T07:52:27Z")

</div>

Speech-to-text services from Google or Azure appears to support only from microphone and the file format as input stream.

So,I’m curious about is there a way to acheive that?

---

<div class="post-metadata">

### Author: ![Kay\_Werner](https://avatars.discourse-cdn.com/v4/letter/k/ecc23a/32.png) [@Kay\_Werner](https://community.wowza.com/u/Kay_Werner)
#### Post date: [September 27, 2023, 8:27am UTC](https://community.wowza.com/t/is-it-possible-to-accomplish-speech-recognition-in-live-streaming-using-speech-to-text-services-from-google-or-azure/96361/2 "2023-09-27T08:27:07Z")

</div>

Did you find any solution connect wowza SDK with azure cognitive-services-speech SDK?  
First step would be to extract the audio stream from the live stream. What features does wowza provide to handle / redirect the audio feed in parallel to the default transcoding process?

---

<div class="post-metadata">

### Author: ![Dorota\_Szafer-Kwasik](https://sea2.discourse-cdn.com/flex002/user_avatar/community.wowza.com/dorota_szafer-kwasik/32/1150_2.png) [@Dorota\_Szafer-Kwasik](https://community.wowza.com/u/Dorota_Szafer-Kwasik)
#### Post date: [September 27, 2023, 6:54pm UTC](https://community.wowza.com/t/is-it-possible-to-accomplish-speech-recognition-in-live-streaming-using-speech-to-text-services-from-google-or-azure/96361/3 "2023-09-27T18:54:23Z")

</div>

I can see that the stage of creating by home means is beginning…  
I know it is possible to create a wowza module that will automate the conversion of audio to text /closed captions/.  
I don’t understand why the wowza team concentrated their efforts on Wowza Video.  
I believe that such an audio-to-text conversion module would attract new WSE users. It would retain current WSE users.  
Speech-to-text services from Google works great with open captioning.  
Third-party programmers can handle English but other languages are much worse for them.

---

<div class="post-metadata">

### Author: ![Scott\_Kellicker2](https://sea2.discourse-cdn.com/flex002/user_avatar/community.wowza.com/scott_kellicker2/32/1337_2.png) [@Scott\_Kellicker2](https://community.wowza.com/u/Scott_Kellicker2)
#### Post date: [October 5, 2023, 2:28pm UTC](https://community.wowza.com/t/is-it-possible-to-accomplish-speech-recognition-in-live-streaming-using-speech-to-text-services-from-google-or-azure/96361/4 "2023-10-05T14:28:04Z")

</div>

Hello. I’m former Wowza now working independently.  
I’ve been working on a couple speech to text implementations for WSE, although not yet Google. If you are interested in such a module, I am to build it on a contract basis.

Reach out to scott@blankcanvas.video

---

<div class="post-metadata">

### Author: ![Dorota\_Szafer-Kwasik](https://sea2.discourse-cdn.com/flex002/user_avatar/community.wowza.com/dorota_szafer-kwasik/32/1150_2.png) [@Dorota\_Szafer-Kwasik](https://community.wowza.com/u/Dorota_Szafer-Kwasik)
#### Post date: [October 5, 2023, 6:14pm UTC](https://community.wowza.com/t/is-it-possible-to-accomplish-speech-recognition-in-live-streaming-using-speech-to-text-services-from-google-or-azure/96361/5 "2023-10-05T18:14:42Z")

</div>

Hi,  
… and yet interest is emerging. and well. You know the point.  
Scott, you’ll get it done faster than Wowza will be interested in such a solution.  
Greetings to you

---

<div class="post-metadata">

### Author: ![Axel\_Gomez1](https://avatars.discourse-cdn.com/v4/letter/a/bc8723/32.png) [@Axel\_Gomez1](https://community.wowza.com/u/Axel_Gomez1)
#### Post date: [December 14, 2023, 8:19pm UTC](https://community.wowza.com/t/is-it-possible-to-accomplish-speech-recognition-in-live-streaming-using-speech-to-text-services-from-google-or-azure/96361/6 "2023-12-14T20:19:13Z")

</div>

Hi @Scott_Kellicker2

I came across your post about developing speech to text modules for WSE. I’m interested in a module that also integrates with JW Player. Could you provide some insights on feasibility, development time, and cost?

---

<div class="post-metadata">

### Author: ![Scott\_Kellicker2](https://sea2.discourse-cdn.com/flex002/user_avatar/community.wowza.com/scott_kellicker2/32/1337_2.png) [@Scott\_Kellicker2](https://community.wowza.com/u/Scott_Kellicker2)
#### Post date: [December 14, 2023, 9:31pm UTC](https://community.wowza.com/t/is-it-possible-to-accomplish-speech-recognition-in-live-streaming-using-speech-to-text-services-from-google-or-azure/96361/7 "2023-12-14T21:31:48Z")

</div>

Hi.

Yes, I could develop such a module.

Let’s connect via email. I’m at scott@blankcanvas.video

(I’m traveling this week but will reach out to your email Monday)

Scott Kellicker

---

<div class="post-metadata">

### Author: ![Scott\_Kellicker2](https://sea2.discourse-cdn.com/flex002/user_avatar/community.wowza.com/scott_kellicker2/32/1337_2.png) [@Scott\_Kellicker2](https://community.wowza.com/u/Scott_Kellicker2)
#### Post date: [January 9, 2024, 12:56pm UTC](https://community.wowza.com/t/is-it-possible-to-accomplish-speech-recognition-in-live-streaming-using-speech-to-text-services-from-google-or-azure/96361/8 "2024-01-09T12:56:49Z")

</div>

Hi Axel. Do you still have interest in such a module? I’ve been working on something very close to this.

Let’s chat at scott@blankcanvas.video .

ScottK

---

<div class="post-metadata">

### Author: ![Karel\_Boek](https://sea2.discourse-cdn.com/flex002/user_avatar/community.wowza.com/karel_boek/32/454_2.png) [@Karel\_Boek](https://community.wowza.com/u/Karel_Boek)
#### Post date: [January 9, 2024, 8:27pm UTC](https://community.wowza.com/t/is-it-possible-to-accomplish-speech-recognition-in-live-streaming-using-speech-to-text-services-from-google-or-azure/96361/9 "2024-01-09T20:27:59Z")

</div>

At Raskenlund we’ve worked with a variety of STT services, incl. IBM, Azure, Google, AWS and a few more.

We have a working module for integration with AWS Transcribe (audio is extracted, sent to AWS Transcribe, then text is added to the stream as subtitles, or you can export it as VTT)
