Sorry to sound so negative, but it seems something halfway between an April Fool and some Kickstarter "scam" (possibly in good conscience, but still).
Both voice recognition and translation are HARD problems, for which even companies with tons of resources and talent like Apple and Google were able to provide only sketchy and partial solutions.
No to mention issues like battery life, connectivity (hard to believe this device would be able to work without a permanent connection to cloud servers doing the heavy lifting, like Siri does) etc...
I'm also skeptical. This is a problem where it's fairly easy to get 80% there but nearly impossible to get 100% there. Most speech recognition apps have trouble with thick accents. From there, imagine stumbling speech or people who change their thoughts in the middle of a sentence. Then there's the lag between input and output - you often can't translate a phrase or sentence until you know the whole thing and then you're adding computer computation time onto it.
I think this type of technology is inevitable, but I think we're still 50 years away from it being ubiquitous.
Nevermind thick accents. What about languages that have tens, even hundreds, of dialects?
Slovenian, for instance, has only 2,000,000 speakers. Yet there are 56 officially classified dialects. [1] They all vary in important ways, sometimes to the point that a speaker from one dialect has trouble understanding a speaker from another.
French is another language with hundreds of dialects. [2] My girlfriend is French, but when she goes to the Caribbean, she has trouble following conversation because their French is so much different than hers.
English makes for a wonderful example too. Not only does it have hundreds of dialects, there are at least 4 different Englishes. UK, US, Aussie, African-American Vernacular. They're starting to show signs of splitting into separate languages. They already have differences in vocabulary and grammar.
Then you can add slang on top of all this.
Oh and to add to all this confusion: In most European countries people learn the standardized form of their native language almost as if it was a foreign language.
I don't think we'll ever get to 100% accuracy with a tool like this. Even humans themselves can't do 100%.
In South Texas, I once encountered a man speaking English with a very strong Acadian accent. I had to listen to him speak for several minutes before I even realized that what he was speaking was in fact English. Accents and dialects are incredibly complex.
Exactly, considering the way language works with different language sentence conjugation, even simple things like adjectives in most romance languages have relative to the noun and article, a translation to a language like English in realtime would sound like someone with a head injury.
A phrase like "Did you see you see the big red truck? It was driving very fast would be:
¿Viste el gran camión rojo? Conducía muy rápido.
In real time this would be in English:
See the big truck red? Driving very fast.
That's just scratching the surface with the most basic communication. There's simply no way this product can work with any real accuracy.
I can understand your cynicism but I think this kind of product is possible. Maybe not perfect, but that the our consumer technology is already mostly there albeit not unified just yet. I'll explain by addressing some of your concerns:
> "Both voice recognition and translation are HARD problems, for which even companies with tons of resources and talent like Apple and Google were able to provide only sketchy and partial solutions."
I've recently been buying a lot of packages from Japan and found Google's camera translate to be surprisingly good. Bordering on sci-fi awesome.
As for voice recognition, that's something that's been consumer technology for a while. Smart TVs, mobile phones, home automation assistants, etc. It's pretty common technology now. Granted the results aren't 100%, they are still generally quite accurate - with the obvious caveats: depending on a person's accent and background noise.
> "No to mention issues like battery life"
That's a pretty minor point these days on devices with no LCD displays.
> connectivity (hard to believe this device would be able to work without a permanent connection to cloud servers doing the heavy lifting, like Siri does)
They do say it depends on a smartphone app, so my assumption was it works in a similar way to the Pebble watch. ie it's essentially just a bluetooth speaker where the phone will push the data to it.
>> "Both voice recognition and translation are HARD problems, for which even companies with tons of resources and talent like Apple and Google were able to provide only sketchy and partial solutions."
> I've recently been buying a lot of packages from Japan and found Google's camera translate to be surprisingly good. Bordering on sci-fi awesome.
OCR of short sentences from invoices or similar and fluid speech recognition of often grammatically incorrect sentences are problems of a very different magnitude. OCR is easy in comparison of actually spoken speech recognition -- especially if it is supposed to run in real time
> As for voice recognition, that's something that's been consumer technology for a while. Smart TVs, mobile phones, home automation assistants, etc. It's pretty common technology now. Granted the results aren't 100%, they are still generally quite accurate - with the obvious caveats: depending on a person's accent and background noise.
and they simply suck for any language than english. its easier to talk to my devices in english than in german. and this is despite my poor pronunciations
>> "No to mention issues like battery life"
> That's a pretty minor point these days on devices with no LCD displays.
not at that size. the power cell would have to be tiny
>They do say it depends on a smartphone app, so my assumption was it works in a similar way to the Pebble watch. ie it's essentially just a bluetooth speaker where the phone will push the data to it.
which would mean that they'd need bluetooth for permanent connectivity. this would drain the battery even more.
> "OCR of short sentences from invoices or similar and fluid speech recognition of often grammatically incorrect sentences are problems of a very different magnitude. OCR is easy in comparison of actually spoken speech recognition -- especially if it is supposed to run in real time"
OCR isn't comparable to speech recognition. Period. The problem needs to be broken into two parts: 1) speech recognition, 2) translation. I was addressing those parts separately. However you're right that translating natural spoken language, which is full of grammatical errors and other nuances, would be a greater challenge than short passages of official documents.
> "and they simply suck for any language than english. its easier to talk to my devices in english than in german. and this is despite my poor pronunciations"
I'll have to take your word for it. But suffice to say there are several competing speech recognition engines, some better than others. So it might be that your devices use an engine that's poor at recognising German. Although equally you might be using one of the better engines. Hard to say from this high level overview. However the fact that they do work against your pronunciation of English is promising as it at least demonstrates how well these engines can cope with different accents - which is one of the hardest stumbling blocks in building speech recognition.
> "not at that size. the power cell would have to be tiny"
Indeed but we've already seen tiny power cells in watches and other modern gadgets. Given the earpiece only really needs a speaker, bluetooth receiver, some kind of DAC, and a battery; half the device could be the battery.
Just to be clear, I'm not suggesting this thing would have a great battery life though. But I still wouldn't be surprised if lasts a few hours - which is still comparable to some smart phones (sadly).
> "which would mean that they'd need bluetooth for permanent connectivity. this would drain the battery even more."
I did say it would need bluetooth. But realistically bluetooth doesn't suck that much power. The speaker would easily be a heavier power draw.
To summarise: there are obviously going to be difficulties - clear technical challenges - in build this kind of device. But I do think the technology is now at a stage were we can at least prototype something. And that was the crux of argument. The earpiece might not be a perfect gadget, but those gadgets who are first to market are seldom perfect.
Meanwhile, Bragi's wireless headphones have a microphone for pass-thru audio, and bluetooth connectivity. Coupled with a smartphone, I'm curious why an app for this doesn't already exist?
Jibbigo was doing decent voice-to-voice translation entirely locally back in the iPhone 3GS days. You still had all the usual pitfalls of voice recognition and machine translation, but it was good enough to be useful.
Your FAQ link isn't loading for me (hugged to death?) but I watched the video on the page HN links to here and I see nothing in it that's infeasible. As with any voice recognition and machine translation product, you'll probably need to be patient and tolerant of many errors, and I wouldn't be surprised if they gloss over that, but it could still be useful.
Edit: the FAQ finally loaded. I don't see anything particularly remarkable in there. There are no promises for battery life, speed, or accuracy. The basic premise of offline voice-to-voice translation is totally doable and has been done.
All the technology to do this already exists. I have speech-to-text on my Pebble watch -- it goes up to a server, is processed, and is returned all relatively quickly. So you combine that with text translation tools that have existed for years and then text-to-speech.
I'm sure it's really just a specialized bluetooth headset, an app, and bunch of cloud services purchased from other vendors.
At one point, they had an app that did live-translation. But Facebook bought the tech and took it off the market. https://en.wikipedia.org/wiki/Jibbigo
Today, they are riding the deep-learning wave to get ever better/faster recognition and translation. So huge improvements are to be expected.
[1] The professor, Alex Waibel, also has a chair at CMU.
Realtime is not possible (right now). You allways have to wait for the voice recognition and then you translate.
In real life, when you listen to a person, you do it simultaneous and, even more important, you anticipate what the other person will say. Right now machines can't do that - they even don't come near that.
The promo video wasn't true real time either though. They just has the translate services running after each speech had concluded. So I think by "real time" they just mean "automatically".
Both voice recognition and translation are HARD problems, for which even companies with tons of resources and talent like Apple and Google were able to provide only sketchy and partial solutions.
No to mention issues like battery life, connectivity (hard to believe this device would be able to work without a permanent connection to cloud servers doing the heavy lifting, like Siri does) etc...
The FAQ is simply "too good to be true".
http://www.waverlylabs.com/2016/05/faqs-about-early-bird-and...