You will be able to describe speech-to-text, text-to-speech and multimodal models and their limits.
After a project meeting, your colleague shares the AI-generated notes. Most of it is fine, but the action items say "Ong to follow up with the vendor in Jurong East by Friday", when the meeting agreed that Wong would follow up with the vendor in Jurong West by Thursday. Further down, someone's remark that the new system was "quite shiok" has become "quite shocked". Nobody noticed until the vendor called asking why nobody had been in touch.
Meeting transcripts are the most common way people use voice AI at work, and they show both what it does well and where it slips. This lesson covers three kinds of model: ones that turn speech into text, ones that turn text into speech, and ones that take in many kinds of input at once.
A speech-to-text model takes audio and produces a written transcript. It follows the same broad recipe as the models in earlier lessons. The sound is cut into very short slices and turned into numbers, and the model, trained on huge amounts of recorded speech paired with transcripts, predicts the text most likely to match.
That last word matters. The model writes the likeliest text, so it does best on speech that sounds like its training data: clear audio, one speaker at a time, and accents and vocabulary that appeared often. Accuracy drops when any of those slip. Accents that were less common in training cause more errors. So does crosstalk, when two people speak at once, and background noise, such as a call taken at a hawker centre.
Singapore adds its own challenges. Singlish words like "shiok", "lepak" or "makan" may be swapped for English words that sound similar. Local place names and the many ways of spelling Chinese, Malay and Indian names trip it up. People switching between English and Mandarin or Malay mid-sentence often get a garbled line. Wong becomes Ong, and Jurong West becomes Jurong East, because those versions looked more likely.
The danger in a meeting tool is that the transcript feeds a summary. The summary is written by a language model from the transcript, so a misheard name or date flows straight into the action items, written in confident, tidy sentences. Check names, numbers, dates and decisions against your own memory or a colleague's before you send the notes round.
A text-to-speech model does the reverse: it takes written text and produces spoken audio. Older versions sounded flat and robotic. Current ones can produce natural pacing, pauses and emotion, and can speak in many voices and languages. You hear them in narrated videos, phone menus, navigation apps and reading tools for people with poor eyesight.
Some text-to-speech tools can also copy a real person's voice from a short recording, after which they can make that voice say anything. There are legitimate uses, such as letting someone who is losing their speech keep a version of their own voice. There are obvious risks too, because a convincing copy of a boss's or a family member's voice is exactly what a scammer wants. That side of the subject, along with fake images and video, is covered in depth in the course AI risks: privacy, bias, deepfakes and scams. For this course, the point is simply that a voice on a call or in a voice note is no longer proof of who is speaking.
A multimodal model is one that can take in more than one kind of input, such as text, images and audio, in the same conversation. Lesson 7.1 showed how an image is turned into tokens the model can use. Audio can be handled in a similar way. Once all of them are in the context window together, one assistant can look at a screenshot of your spreadsheet while you describe the problem out loud, and talk you through the fix.
Some voice assistants work in stages, turning your speech into text, sending the text to a language model and reading the reply aloud. Others handle the audio more directly, which lets them respond faster and pick up on tone. Either way, the limits you have already met still apply. The model can mishear you, misread the screenshot, or produce a confident answer that is wrong, and a pleasant, natural voice makes a wrong answer even easier to believe.
The best way to find out how well speech-to-text copes with the way you actually speak is to test it with your own voice, on the words that matter in your work and your neighbourhood. A minute of audio is enough to see the pattern.
Record a one-minute voice note with a few local terms or names, transcribe it with any speech-to-text tool, and count the errors.
Junxiong-WFG Organisation is an authorised representative of AIA Financial Advisers Private Limited (Reg. No. 201715016G).