Skip to main content

KooKoo speech engine versus Google speech engine - Round 2

TLDR:

The data sets used for the testing in the experiments are the digit examples from the Google speech commands dataset.

Experiment 1:
Converted dataset to 8Khz to fit the telephony format.

Google API accuracy: 74.3%
Kookoo Zena accuracy: 92.4%


We felt the accuracy for Google dropped down because of the 8Khz format. So we did another experiment for Google speech API.

Experiment 2:
Original speech command dataset
Google API accuracy: 85.6%

The Long version:
At Ozonetel, we are constantly innovating and we are back again with a new speech model after the Yes/No model.

We all saw Google demo Duplex last week and it was a pretty cool demo. Just to put things in perspective, though the demo was cool, we think we are still some a ways away from actually having a system operate as smoothly as shown in the demo.
We come to the conclusion mainly based on our experiments with the Google speech API. Unless Google is using some model other than the one exposed publicly in the Google speech API, there is still a long way to go. Especially because the Google speech API is pretty bad with telephony data(8Khz).


Today, we are launching our second model, the single digit model, built in house, based on a proprietary AI algorithm. You can now design your KooKoo IVRs to say "Please say your order number", instead of "Please enter your order number".

In the internal tests our model's accuracy has been really good and we believe , soon, that this can become the standard interaction mechanism instead of DTMF.

We also did a side by side comparison of our model with Google's speech API. Before presenting the results below, some disclaimers:

1. Our model is a specific model for Digit. Google's model is more of a generic ASR. And in most cases specific models work better than generic models.
2. Google's ASR is in the cloud. Though our model is also in the cloud, since its co hosted with KooKoo, KooKoo responses will generally be faster.


The data sets used for the testing in the experiments are the digit examples from the Google speech commands dataset.

Experiment 1:
Converted dataset to 8Khz to fit the telephony format.

Google API accuracy: 74.3%
Kookoo Zena accuracy: 92.4%

We felt the accuracy for Google dropped down because of the 8Khz format. So we did another experiment for Google speech API.

Experiment 2:
Original speech command dataset
Google API accuracy: 85.6%

We also ran the test on multiple other private data sources. In all the cases KooKoo Zena out performed Google speech API.


Watch this space for more upcoming speech models.

Popular posts from this blog

Google's approach to business communication

 Google has been making silent moves in the business communication space. Google has mostly lost the instant messaging wars. But it does not want to lose the business communication war. WhatsApp, Instagram, Twitter and Facebook have been making their own moves to enable businesses to reach their customers through their channels. Its all about who has control over the communication channels. Especially communication which leads to business. That's where the money is. Currently, Google is the king of search and most online transactions start with a Google search. FB, Amazon, Apple and others want to change that. They want the search to start on their properties. And they have started making the moves. WhatsApp business allows small businesses to conduct their transactions on WhatsApp. FB and Instagram have long supported small businesses to manage their business on their channels. Apple has also made some nice moves with Apple business chat. They have integrated a whole shopping expe...

Telugu ASR speech data collection

Image Source: IIIT-H Developing an indigenous ASR for Indian languages has been a goal for us since a long time. In that regard we have been experimenting a lot, trying out various neural network architectures.  While doing these experiments we found that there was no good dataset for Indian languages. While discussing with IIIT professors we got to know that the government of India was also exploring options to generate a good dataset. We immediately offered our help and our platform for this endeavor. So, as a starting step we have come up with a few campaigns to encourage users to donate speech data. We wanted to make it fun, so our first few campaigns are along the lines of JAMs(Just a Minute speech topics) etc. A topic will be provided and you need to speak for a minute on that topic. We have started this campaign for college students to start with. Of course anyone can participate and contribute their data. The more the merrier :) We will adding a lot more innovative ways ut...

Ozonetel. Take 2, The Text Edition

Ever since we introduced cloud telephony to India in 2009, Ozonetel has always been known as the go to startup for cloud telephony requirements like cloud call centers and virtual numbers. Voice has been our mainstay for a long time and it will continue to be so in the future too. But starting this month, we at Ozonetel have decided to have an increased focus on text based communication channels in addition to voice channels. It's not that we did not support these channels earlier. We already supported, chat and some email functionality. But it has been in bits and pieces only. But starting this month text channels, including email, chat and social media will become First Class Citizens in our stack. They will get the same love that our voice channels get. Unsplash           To achieve this we have planned several upgrades to our platform which will unfold over the coming weeks. Here is a sneak peak of what we are planning. Built in Contact Manager: This has bee...