Live data from Hacker News

New services expand IBM Watson capabilities to images, speech, and more

developer.ibm.com

11–20 of 110 posts

Re: New services expand IBM Watson capabilities to images, speech, and more

#12
post #8

The text-to-speech is actually a little nicer than Siri or Cortana, but not groundbreaking. This was the only one of the 5 that I thought did well. The rest might have been better without demo pages. For visual recognition, I used a picture of a snowmobile from http://www.1888goodwin.com/2013/11/14/what-do-you-need-to-do... , which it identified with 73% confidence as "Invertebrate". Speech to text is a parody twitte…

Make sure you use a headset, not your laptop's microphone.

Re: New services expand IBM Watson capabilities to images, speech, and more

#14

I tried using Watson a month ago without much success. I wanted to do a classification of some random text, and say that this text for example is this category. But as far as I could understand it only allows using their own datasets. It's not possible to train their service with your data, unlike wit.ai for example. Seems obvious to me that people would want to train with their own data.

Pretty much all the services that we are releasing will have some adaptation capabilities - allowing you to provide your own data, create your own models, etc - at some point. Stay posted.

Re: New services expand IBM Watson capabilities to images, speech, and more

#15
post #5
post #2

Some context on the new services. They are built on technology that comes from IBM Research and has been moved into the Watson group in 2014. Some like speech, have been developed for more than 50 years. None of these technologies have overlap with the Watson Jeopardy stack (except for the Watson voice). We will release that stack later this year as a series of services allowing you to build a full Q&A/dialog applica…

Are you working on any audio (non-speech) analysis services? I have no particular usecase in mind, but it's an area I'm always interested in!

We have worked on audio analytics in the past for things such as outdoor sound detection and vehicle identification. We are currently focusing on speech-based analytics such as language ID and affect recognition. The statistical methodologies we are using for speech are easily extended to such domains. We hope that by puuting out these initial speech services we will get feedback from the community about related problems and welcome your suggestions.

Re: New services expand IBM Watson capabilities to images, speech, and more

#16
post #12
post #8

The text-to-speech is actually a little nicer than Siri or Cortana, but not groundbreaking. This was the only one of the 5 that I thought did well. The rest might have been better without demo pages. For visual recognition, I used a picture of a snowmobile from http://www.1888goodwin.com/2013/11/14/what-do-you-need-to-do... , which it identified with 73% confidence as "Invertebrate". Speech to text is a parody twitte…

Make sure you use a headset, not your laptop's microphone.

That's not a reasonable requirement. It only sounds like one to you because the technology has been so bad for so long.

Re: New services expand IBM Watson capabilities to images, speech, and more

#19
post #12

Earlier quoted context omitted.

Make sure you use a headset, not your laptop's microphone.

That's not a reasonable requirement. It only sounds like one to you because the technology has been so bad for so long.

Smartphone microphones are much better than laptop microphones and pretty much on par with using headsets on a laptop - they represent our primary use case.

Re: New services expand IBM Watson capabilities to images, speech, and more

#20

Compare the Watson text-to-speech voices with Nuance ... Watson http://text-to-speech-demo.mybluemix.net/ Nuance http://www.nuance.com/for-business/text-to-speech/vocalizer/... I prefer the Watson version voicing a sample paragraph. Both are good enough for an application that selects on price. For a voice-first application, maybe Watson is better for TTS. For speech to text, Nuance has been the leader, e.g. Apple's…

My evidence is anecdotal at best, but I have found Siri to be terrible and my "OK, Google" to be wonderful.
Post reply on HN