Nokia Mobile Phones Ltd.
mika.kivimaki@nmp.nokia.com
CHI '97 Workshop: Speech User Interface
Design Challenges
Experience of Speech User Interface design
I work as speech UI designer at Nokia Mobile Phones in Tampere, Finland. My master's thesis is about the speech user interface of an in-vehicle device. I am planning to research the subject further after the first degree. I have been working with car phone's user interfaces for over a year. Since spring '95 I have been employed by Nokia Mobile Phones. When I started in Bochum, Germany, I worked with speech coding. Moreover I have prior UI design experience of one year.
For me, working with in-vehicle devices, speech modality is very important issue to be able to design safe products. Futher personally I see speech user interfaces as an impressive possibility to reduce the workload of eyes by changing our working manners with computers, or other devices, to more versatile, ergonomic and free. Due to existing technical restrictions I do understand doubts against speech recognition systems, but still I believe that already something could be done in a nicer way. Nowadays people have learnt to command machines by buttons, keypads, or keyboards. They think it is strange to talk a machine. But some time ago machines had speech user interface, namely machines were animals. And I guess that there were some recognition errors as well ;-)
For example (Dialog 1.) a user needs to command a mobile phone to set up a call after the whole number is given. A similar situation could exist when spelling city or street names for a route guidance application, or when giving flight, train, high-way etc. numbers for an information service system with automatic speech recognition.
SYSTEM: Give a number
USER: "five one two" ...
SYSTEM: five one two
USER: "six nine nine nine" ...
SYSTEM: six nine nine nine
USER: ...(silence)......
SYSTEM: Are You there?
USER: "..Yes umh, ... Dial !"
SYSTEM: Dialing
In a dialogue between two persons when one of them asks for a certain information a natural behaviour is to answer the question and then assume that the one who asked knows what to do with the information received, or how to proceed according to the reply.
The end of the answer is often clear due to answerer's intension of the voice, or added small words like "and" or "and then". The form of information can implicitly tell when enough information is reached to proceed a task. There are also efficient means of communication between humans using movements of head, like nodding and shaking.
The difference with a speech driven device is that the system first asks for an input needed for the current task, but after receiving the information the system is not able to continue automatically. To go on a user should remember to take the control of the dialog by giving a command needed to proceed the task. Without user's command the system does not have any means to detect if there will come more input items, or if the whole input is given.
Often in the systems, where the procedure has been selected by voice, a user has to repeat the similar command after the input to finish the operation. This is quite unnatural, because it sounds like the system has forgotten what it is doing. Example of a speech driven mobile phone, after speech recognition is activated, user selects number dialing instead of name dialing and after dictating the number he commands again system to dial. (Dialog 2.)
SYSTEM: Commands available are: Dial a number, dial a name,
listen saved names.
USER: "Dial a number"
SYSTEM: Give a number
USER: "five one two" ...
SYSTEM: five one two
USER: "six nine nine nine" ...
SYSTEM: six nine nine nine
USER: "... Dial "
SYSTEM: Dialing
In tests, subjects did not learn to give a command after saying a series of numbers. A user seems to be completed with her task and waiting the system to continue. After a long silence she starts to think what to say. In the worst case a timer triggers cancelling of the task. Also experienced users sometimes find themselves waiting in the familiar situation and then they command the system to proceed.
With an ordinary telephone, due to the architecture of landline telephone network, the call is switched through network simultaneously while dialing. When an unique destination number is reached the call is set up. Because of the experience related to ordinary landline phones, some people find it difficult to press a 'CALL'-button to set up a call after dialing when using mobile phones. This happens even though there exists a button as a visual reminder to press it.
With good design it could be possible to cut down the dialog by couple of seconds and save on-line time. The system should somehow propose user to take control, or inform about finishing input without annoying her, if she is just dictating slowly for some reason. In noisy environment when a speech recognizer is listening, it is always possible that the system catches backround noise as a word. This leads to an error, where the system has recognized a word though user did not say anything. Thus the useless listening might lower user's satisfaction for the system.