Hacker Timesnew | past | comments | ask | show | jobs | submitlogin

Can anyone tell me why we aren't able to progress into better levels of recognition. Am I correct to assume it has nothing to do with computing power, and everything to do with (semantic/linguistic) software?


It's more than just linguistic software. Our knowledge of linguistics itself is currently very limited; it's a nacent science and there's still a great deal of debate about how to even approach the study of language. Even leaving aside the difficulties in just transcribing speech, linguists are still a long ways off from coming up with any formalism of human syntax that could help create software to syntactically parse normal human speech.


...because most of our linguistic formalism is derived from written language, which is generally an artificial approximation of real language.

Remember grade school, with parts of speech, Conjunction Junction, and diagramming sentences? Well, think back to the last (non-trivial) sentence you actually said out loud. I challenge you to try to diagram that sentence.

Although there's quite a lot of research that's been done, surprisingly little of it deals with the way most of us really communicate.


Modern linguistic formalism is not derived from written language at all. I studied at MIT and UConn in the mid '90's and the separation between prescriptive grammar and linguistics is pretty clear.

That's not to say that current approaches to linguistics are correct; I don't know enough to make that judgement. But I think the problem lies more in the complexity of the domain - not because modern linguists are so foolish as to base theory on written language.


It's quite possible I'm behind the times.

not because modern linguists are so foolish as to base theory on written language

That was actually what I was trying to get at. While there are clearly some rules we follow when speaking, they are much more open than those governing our writing.

It's odd, though, or at least counterintuitive, that our comprehensive of these less-structured spoken communications is higher than that of more-structured written ones.


In fact, it has been shown that there are languages that have syntactic constructions which are context-sensitive (in a generative linguistic PoV. cf. http://en.wikipedia.org/wiki/Chomsky_hierarchy). For example Swiss German variants (http://books.google.com/books?id=JOjoWrP4tnIC&lpg=PA165&...)


Most commenters are focusing on relatively high level features of decoding speech. It is important to also be aware that there is still great debate about what are the acoustic correlates of linguistic events in speech. It seems that our words are composed of subunits (usually taken to be phones out of the IPA--but there's work on alternatives) but what exact acoustics correspond to the phones are is still unsettled: lots of debate and mediocre recognition performance

Undoubtedly there is much room for improvement on these higher-level features but computers are still well behind humans in large vocabulary isolated keyword spotting: this is a task where one word from a very large corpus of words is spoken and the human or computer has to guess what that word was. Computers do poorly relative to humans (particularly in noise), which suggests that many of the mistakes that computers make is in not being able to interpret the acoustics correctly.


And to make this problem worse, there's a great deal of variation among different dialects of the same language, which means that there's no 1:1 mapping of IPA phones to written letters.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: