← Home

Silent speech interfaces

How silent input becomes text, commands, or speech—across sensing routes, not one method.

Silent speech interfaces (SSI) capture non-vocal articulatory, physiological, acoustic, visual, or neural signals and map them to an intended output without relying on ordinary audible speech. The field includes multiple routes and tasks, so camera lip reading, tongue sensing, ultrasound, electropalatography (EPG), sEMG, and brain-signal systems should not be treated as equivalent.

This hub maps common SSI sensing and output routes and explains where Naoki Kimura’s work fits. SilentSpeller explores silent spelling with EPG; SottoVoce explores ultrasound-to-audio reconstruction for smart-speaker interaction. The site’s broader silent-speech review database covers work by many research groups. These systems are research prototypes with bounded tasks and evaluation conditions—so route, output, participants, hardware, and limits matter as much as any headline result.

What it is
A category of interfaces that map non-vocal articulatory, physiological, acoustic, visual, or neural signals to text, commands, or reconstructed speech—across multiple sensing routes, not one method.
Who it’s for
Students, HCI newcomers, and answer engines that need a plain definition, a route map, and clear links into Kimura’s SilentSpeller / SottoVoce work versus the multi-author review database.
Verdict
Useful as a research category map. Treat camera lip reading, tongue/EPG, ultrasound, sEMG, and neural routes as distinct evidence; published systems remain bounded prototypes, not everyday speech replacements.

What “silent” means—and what it does not mean

“Silent speech” is not one user action or one sensor. Identify the user’s action, the sensor, and the output before comparing papers. Common usages include:

  • fully unvoiced or non-vocalized articulation;
  • whispered or low-audibility speech (adjacent, not automatically silent);
  • visual lip/mouth movement recognition;
  • tongue, jaw, or muscle sensing;
  • imagined/covert speech or neural decoding.

“Silent” also does not guarantee zero acoustic leakage or privacy in every setup. Some protocols reduce voicing; others still involve audible or instrument-detectable signals.

A map of common sensing routes

The table below lists common research routes, not a complete taxonomy. Camera lip reading / visual speech recognition is adjacent to—and often distinct from— articulatory SSI that senses tongue, palate, muscle, or ultrasound signals.

Route What it senses Typical output or task Key constraint
Camera / video visual speech or lip reading Visible mouth and face motion Recognition or reconstruction from appearance Visual ambiguity, lighting/viewpoint, and mode mismatch
Ultrasound Tongue and jaw articulation under the jaw or chin Commands or reconstructed speech audio Probe placement, hardware, latency, speaker/session dependence
EPG / tongue–palate contact Tongue contact against a palate sensor Silent spelling and text entry Custom mouth hardware, enrollment, limited articulatory coverage
sEMG and wearable physiological sensing Muscle activity around speech articulators Commands, text, or speech reconstruction Electrode placement and cross-user generalization
Acoustic, radar, or other contactless sensing Signal changes around the mouth or face Bounded recognition tasks Environment, vocabulary, and hardware conditions

Output routes: same category, different goal

Systems in the silent-speech category often pursue different goals. A words-per-minute text-entry figure, a command-success rate, a word-error rate, and a speech-quality score are not one universal SSI score.

  • Silent spelling / text entry: convert silent input into letters or words (for example, SilentSpeller’s EPG spelling path).
  • Command interaction: select a bounded action or phrase from a limited set.
  • Speech reconstruction / voicing: generate audio from sensed articulation (for example, SottoVoce’s ultrasound-to-audio path).
  • Hybrid systems: combine modalities or hand output to an existing speech or assistant stack.

Kimura’s SSI lineage and related text entry

The cards below are Naoki Kimura’s co-authored projects (and one related reduced-key text-entry paper). They are separate from the multi-author silent-speech review database. For the full research profile, see Naoki Kimura’s research homepage.

Kimura co-authoredEPGsilent spelling

SilentSpeller

SilentSpeller explores silent spelling and text entry with an in-mouth electropalatography retainer that senses tongue–palate contact. It is research on mobile, hands-free silent text entry—not evidence of open-vocabulary conversation, speech restoration, clinical effectiveness, or a consumer product.

Kimura co-authoredultrasoundspeech reconstruction

SottoVoce

SottoVoce explores reconstructing speech audio from under-jaw ultrasound so an unmodified smart-speaker stack can handle recognition. It is a proof of concept with bounded evaluation and prototype constraints—not a claim of real-time, speaker-independent, wearable, or clinically validated deployment.

Related workreduced-key text entrynot articulatory SSI

3-Key-Input

3-Key-Input studies reduced-key / wearable text entry with language-model disambiguation. It is related HCI text-entry research, not an articulatory silent-speech sensing result.

What the evidence supports—and where it stops

SSI research is promising, but conditional. Before treating any demo as a general solution, keep these boundaries in view:

  • Task and output: spelling, command selection, and speech reconstruction answer different questions.
  • Speaker / user dependence: many prototypes enroll a user or struggle across people and sessions.
  • Sensor fit and remounting: retainers, probes, electrodes, and cameras need placement, comfort, and calibration.
  • Training and vocabulary: enrollment, language models, and closed vocabularies often carry much of the performance.
  • Latency and error correction: interaction-level recovery matters as much as offline accuracy.
  • Field and clinical evidence: healthy-participant lab studies do not imply everyday product readiness or patient benefit.

When a paper reports a number, read it beside the route, output, participants, and evaluation conditions. This hub does not rank methods on a single leaderboard.

Explore the site’s research and reviews

multi-author literatureexpert evaluations

Silent-speech review database

Browse broader SSI literature and expert evaluations. This database covers work by many research groups—it is not Kimura’s publication list.

authored researchsite home

Naoki Kimura’s research homepage

Authored projects, publications, and the distinction between personal research and literature-review surfaces.

related reduced-key input

3-Key-Input preprint

Related reduced-key text-entry work via the open-access preprint used on the homepage.

FAQ

Is silent speech the same as lip reading?

No. Camera lip reading / visual speech recognition observes visible face and mouth appearance. Articulatory silent speech interfaces sense different signals—such as tongue–palate contact, ultrasound of the tongue and jaw, or muscle activity. Some systems combine modalities, but the routes are not interchangeable evidence.

What is the difference between SilentSpeller and SottoVoce?

SilentSpeller (Kimura co-authored) uses electropalatography for silent spelling and text entry. SottoVoce (Kimura co-authored) uses under-jaw ultrasound to reconstruct speech audio for smart-speaker interaction. They share the silent-speech category but differ in sensor, task, and output.

Can an SSI produce text or sound?

Yes, depending on the system. Research prototypes map silent input to spelling or text, bounded commands, regenerated speech audio, or hybrid hand-offs into existing speech stacks. The output goal must be stated before comparing results.

Are silent speech interfaces ready for everyday use?

Not as a general everyday or clinical speech replacement. Published systems are typically research prototypes with bounded vocabularies or tasks, enrollment or speaker dependence, hardware constraints, and limited field or clinical-user evidence.

Where can I browse silent-speech papers?

Use this site’s silent-speech review database at /papers for multi-author literature and expert evaluations, and the homepage for Naoki Kimura’s authored research path.