Everything so far chains three models together: STT turns audio into text, the LLM answers in text, and TTS turns it back into audio. A newer class of open-weight models collapses that loop into a single network that listens and speaks directly. Some of them listen while they talk, too. This page covers what exists in the open-weight world as of September 2026, what it takes to run each model locally, and why it matters to anyone defending a phone line.
"Voice-to-voice" gets used for three quite different things. Only the first one replaces the pipeline.
A single model takes audio tokens in and produces audio tokens out, usually with an internal text stream it "thinks" in. Tone, hesitation and laughter survive, and there is no STT → TTS handoff to wait on. This page is mostly about these.
Models like Ultravox understand audio directly but only answer in text, so you still need a TTS. They remove the STT hop, not the whole loop.
Tools like RVC, Seed-VC and OpenVoice don't understand or answer anything. They re-skin one person's speech so it sounds like another person, in real time. That is the classic live-deepfake vishing tool, covered in Voice Cloning → Local Tools.
Full-duplex models listen and speak at the same time, like a person on a phone call. They can be interrupted mid-sentence, drop in "mm-hm" backchannels, and reply in roughly 200 ms. That is faster than most humans, and far below the ~750 ms budget of the cascaded pipeline in Agent Architecture.
The first open full-duplex speech model, and still the easiest one to run. It models the user's audio and its own audio as two parallel streams, plus an "inner monologue" text stream. Kyutai reports 160 ms theoretical latency and about 200 ms in practice. It ships two fixed voices, Moshiko (male) and Moshika (female), and speaks English only. The backbone is small, so expect chatty rather than smart. It is a great demo of how natural a machine can sound, not a capable agent.
A reinforcement-learning fine-tune of Moshika aimed at smoother turn-taking, backchannels and interruptions. It is a good illustration that "conversational feel" is now something vendors train for directly.
Fine-tuned from Moshiko to take a role prompt (text) and a voice prompt (audio), so you can make a full-duplex agent that plays a specific persona. "Friendly IT helpdesk" is exactly the kind of persona a vishing crew would want. It ships preset voices and is English only. You must accept NVIDIA's license on Hugging Face before downloading.
An omni model (vision + audio) with full-duplex streaming. It is built from a Whisper encoder, a Qwen3-8B backbone and a CosyVoice2 speech decoder. Two things make it notable here: it runs on a single consumer GPU or Apple Silicon, and it can clone a voice from one reference clip during the conversation. English and Chinese.
The first open full-duplex model with live tool calling, meaning it can look something up or take an action while it keeps talking. That closes one of the biggest gaps between speech-native models and pipelines. NVIDIA reports ~450 ms response latency, and the model is English only. It needs data-center hardware, so treat it as a preview of where things are going, not a lab tool.
These models take speech in and give speech out in one network, but they take turns like a walkie-talkie rather than overlapping. In exchange they tend to be smarter, more multilingual, and some support function calling.
A "Thinker-Talker" design: the Thinker is an LLM that understands audio, video and text, and the Talker streams speech from its hidden states. It has two fixed voices (Chelsie and Ethan). The 3B version is the most practical local omni model for a lab laptop with a mid-range GPU.
The most capable open omni model: it understands speech in 19 languages and speaks 10, has native function calling, and reports ~234 ms to first audio packet. Because the transcript never exists as a separate step, text-based guardrails that inspect an STT transcript have nothing to inspect. Keep that in mind for the Attack Surface section. Its successor, Qwen3.5-Omni (2026), is proprietary and API-only and does not meet this course's open-weight requirement.
End-to-end speech conversation with tool calling and paralinguistic understanding (emotion, speaking style). It covers English, Chinese, Japanese and more. Only the mini variants are open; the full Step-Audio 2 is API-only.
Fun-Audio-Chat-8B (Alibaba Tongyi, Apache-2.0, ~24 GB) supports speech function calling. Kimi-Audio-7B (Moonshot) is a strong universal audio model. GLM-4-Voice-9B (Zhipu) can be told to change emotion, speed and dialect. Mini-Omni2 (MIT, ~0.5B) is tiny enough to read end to end if you want to learn how these models work. Most of this group is Chinese/English-centric.
Translation models keep who is speaking but change the language. For attackers, that removes accent and fluency as a tell: someone who doesn't speak the target's language can hold a live call in it, in their own voice.
Simultaneous speech translation that transfers the speaker's voice. Hibiki translates French to English; Hibiki-Zero adds Spanish, Portuguese and German into English.
Many-language streaming translation, with an Expressive variant that keeps tone and pacing. Notably, Meta runs Seamless output through a watermarking step, which makes it one of the few open speech models that watermarks at all. Meta's open watermarking research is AudioSeal.
Not every "speech-to-speech" repo is a single model. These projects are well-tuned cascades. They are worth knowing because they get close to speech-native latency while keeping a swappable LLM and a text transcript.
Kyutai streaming STT + any vLLM text model + Kyutai TTS. MIT, needs 16 GB+ VRAM. Wraps any LLM with a Moshi-like conversational layer.
Hugging Face's modular Apache-2.0 pipeline: Silero VAD → STT → any OpenAI-compatible LLM → TTS (e.g. Kokoro). Easy to point at your local llama.cpp server.
Apache-2.0 conversational speech model: it uses previous turns of audio as context, so replies sound like they belong in the conversation. It cannot generate text, so it still needs an LLM in front.
A quick reference. Latency figures are the vendors' own claims on their own hardware, and VRAM figures are approximate.
| Model | Type | Weights License | Size | Duplex | Custom Voice | Languages | Local VRAM | Latency (claimed) |
|---|---|---|---|---|---|---|---|---|
| Moshi | End-to-end | CC-BY 4.0 | 7B | Yes | No (2 presets) | EN | ~16 GB | ~200 ms |
| PersonaPlex-7B | End-to-end | NVIDIA Open Model (gated) | 7B | Yes | Voice prompt | EN | A100-tested | ~170 ms |
| MiniCPM-o 4.5 | End-to-end omni | Apache-2.0 | 9B | Yes | Clone from clip | EN, ZH | ~11 GB (int4) | ~0.6 s |
| NemotronLabs VoiceChat | End-to-end + tools | OpenMDW-1.1 | ~11B | Yes | — | EN | A100 / H100 class | ~450 ms |
| Qwen2.5-Omni | Turn-based omni | 7B Apache-2.0; 3B research-only | 3B / 7B | No | No (2 presets) | Multilingual | Consumer GPU (3B) | — |
| Qwen3-Omni | Turn-based omni + tools | Apache-2.0 | 30B MoE | No | No (presets) | 10 spoken | ~79 GB | ~234 ms |
| Step-Audio 2 mini | Turn-based + tools | Apache-2.0 | ~8B | No | Timbre switching | EN, ZH, JA + | Single GPU | — |
| Hibiki | Translation | CC-BY 4.0 | 1B / 2B | Streaming | Keeps speaker voice | FR → EN | On-device (1B) | — |
Neither approach wins everywhere. For training scenarios where you need scripted, gradable behavior, the cascade is usually still the right choice. Speech-native models shine when the point is realism.
| Factor | Cascade (STT → LLM → TTS) | Speech-to-Speech Model |
|---|---|---|
| Latency | ~500–1500 ms; needs VAD and endpointing heuristics | ~170–450 ms (full-duplex); natural interruptions |
| Naturalness | Loses tone and emotion at the text step | Keeps prosody, laughter, hesitation, backchannels |
| Reasoning | Use any LLM, any size | Tied to a 3–9B backbone; weaker at complex tasks |
| Prompt control | Precise system prompts and scripts | Role prompts work, but drift is more common |
| Transcript / logging | Free: STT output is a text log | Must be reconstructed (the model's inner text, or a separate STT) |
| Guardrails | Filter text between each stage | No text boundary to filter in real time |
| Tool use | Any framework (LiveKit, Pipecat) | Only a few models (Qwen3-Omni, Step-Audio 2, VoiceChat) |
| Voice choice | Any TTS, including cloned voices | Mostly fixed presets; a few support voice prompts |
| Best for | Scripted scenarios, grading, production | Realism demos, showing how convincing agents have become |
Speech-native models change what defenders can rely on, and what they can inspect.
Cascaded bots give themselves away with a half-second gap, talking over you, or ignoring an interruption. Full-duplex models take ~200 ms turns and yield when interrupted. Stop training people to listen for lag.
Many voice-agent guardrails work on the STT transcript. An end-to-end model never produces one as a separate artifact, so prompt-injection filters and content policies need to work on audio or on the model's internal text stream.
Meta's Seamless is the exception, not the rule. Kyutai chose not to watermark its TTS because watermarks on open models "can easily be deactivated", and in its tests they were stripped just by re-encoding the audio. Instead it restricts cloning to pre-computed voice embeddings.
MiniCPM-o clones from a single reference clip inside a live conversation. Voice-conversion tools like Seed-VC need 1–30 seconds of audio. A voicemail greeting is enough source material.
Voice-preserving translation like Hibiki lets a caller speak a language they don't know, in their own voice, in near real time.
If you build a demo agent with any of these, consider stamping its output with AudioSeal (MIT, runs locally), and only use voices from people who have consented.
Moshi is the simplest open full-duplex model to run: one pip package, and a built-in web UI where you can talk to it. It is not pre-installed on the lab machines. It needs a CUDA GPU with ~16 GB VRAM, or an Apple Silicon Mac using the MLX build.
HF_HUB_OFFLINE=1 once the weights are cached to prevent any further network calls.# === NVIDIA GPU (PyTorch) === sudo mkdir -p /opt/moshi && sudo chown $USER /opt/moshi cd /opt/moshi uv venv --python 3.12 uv pip install -U moshi # Start the server with the female voice (Moshika) # then open http://localhost:8998 and allow microphone access .venv/bin/python -m moshi.server --hf-repo kyutai/moshika-pytorch-bf16 # Or the male voice (Moshiko) .venv/bin/python -m moshi.server --hf-repo kyutai/moshiko-pytorch-bf16 # === Apple Silicon (MLX, 4-bit quantized) === uv pip install moshi_mlx .venv/bin/python -m moshi_mlx.local_web -q 4
Things to try: interrupt it mid-sentence, say nothing and wait, or talk over it. Compare how it feels with the llama.cpp + Piper demo in The LLM Brain. Then ask it something factual. You'll quickly see the trade between fluency and intelligence.