🪟 Windows TippsThe Gemini desktop app is now available for Windows(11.09.2026 um 17:06 Uhr)
🕵️ SicherheitslückenBurn Out, Or Fade Away(14.09.2026 um 14:25 Uhr)
⚠️ Malware / Trojaner / VirenWindows 11 just dropped the tool ransomware abused, Microsoft says don’t restore WMIC(10.09.2026 um 20:11 Uhr)
⚠️ Malware / Trojaner / VirenWindows 11: Microsoft entfernt WMIC-Tool gegen Ransomware - ad-hoc-news.de(14.09.2026 um 07:58 Uhr)
🕵️ SicherheitslückenMicrosoft schließt Rekordzahl an Sicherheitslücken - techbook(14.09.2026 um 09:00 Uhr)
🪟 Windows TippsThe Gemini desktop app is now available for Windows(11.09.2026 um 17:06 Uhr)
🕵️ SicherheitslückenBurn Out, Or Fade Away(14.09.2026 um 14:25 Uhr)
⚠️ Malware / Trojaner / VirenWindows 11 just dropped the tool ransomware abused, Microsoft says don’t restore WMIC(10.09.2026 um 20:11 Uhr)
⚠️ Malware / Trojaner / VirenWindows 11: Microsoft entfernt WMIC-Tool gegen Ransomware - ad-hoc-news.de(14.09.2026 um 07:58 Uhr)
🕵️ SicherheitslückenMicrosoft schließt Rekordzahl an Sicherheitslücken - techbook(14.09.2026 um 09:00 Uhr)

🔧 Programmierung 🕛 vor 1 Monat 3 Min Lesezeit
0

Making Two TTS Voices Sound Like an Actual Conversation

↗ Quelle (dev.to)
🗣️ Stimme:
📑 Inhaltsübersicht

Rendering a document as a two-host podcast is not just alternating two TTS voices line by line. Do that and you get something that sounds like two people reading unrelated scripts in the same room.






The problem is prosody, not voices



Human conversation carries information in timing. A reply that arrives instantly reads as agreement. A short gap reads as consideration. Speakers overlap slightly at turn boundaries. Pitch tends to fall at the end of a statement and rise before a handoff.



Concatenated TTS has none of this. Each utterance is synthesised in isolation with neutral prosody and identical gaps, and the result sits in an uncanny valley — clearly speech, clearly not conversation.






Things that measurably help



Vary inter-turn gaps by turn type. Not one fixed pause:




  • Agreement or continuation: short, 150-250ms

  • Topic shift: longer, 400-600ms

  • Answer to a question: short, the responder was already primed

  • After a complex explanation: longer, it reads as processing time



A lookup keyed on turn type gets you most of the way, and it is cheap to implement.



Write handoffs into the script. The transition matters more than the voice. "Right — so what happens when the file is large?" carries turn-taking structure that "The file size limit is 10MB" does not. This is a scripting problem before it is an audio problem.



Do not centre both voices. A few degrees of stereo separation makes two speakers legible to the ear. Hard-panning is disorienting on headphones; subtle offset is not consciously noticed but does the work.






Chunking is where quality actually dies



Most TTS APIs have input length limits, so long documents get split. Naive splitting at the limit cuts mid-sentence, and the resulting prosody discontinuity is jarring — the pitch contour restarts mid-thought.



Split on semantic boundaries — paragraph, then sentence — and never mid-clause. If a single sentence exceeds the limit, split at a comma and accept a small artefact rather than a mid-word cut.






Multilingual complicates the pause table



Turn-taking norms are not universal. Japanese conversation tolerates longer pauses than English; treating a 600ms gap as "awkward" is an English-specific assumption. If you support multiple languages, the pause table needs to be per-language, not global.



I build this into DuoCast — paste text, a PDF, or a URL, get a two-host podcast with the conversational layer handled.






Caveat



None of this makes synthetic conversation indistinguishable from a real recording. It moves it from "obviously robotic" to "pleasant enough to listen to on a commute," which is a lower but more achievable bar.

Vollständiger Original-Bericht
Ausführliche Details, Code-Beispiele & Hersteller-Stellungnahme auf dev.to.
↗ Original-Artikel auf dev.to lesen
Wie bewertest du diesen Beitrag?
1 Klick Feedback
Teilen mit Netzwerk & Team:

Community-Analysen & Experten-Meinungen 0

Verfasse deine eigene Analyse, teile Workarounds oder diskutiere diesen Vorfall im Blog.
Noch keine Community-Analyse verfasst. Markiere einen Textabschnitt oder klicke oben auf Eigene Analyse verfassen“!
Community Pulse: Relevanz-Einschätzung
1 Klick Experten-Votum
🔴 Akute Relevanz 0%
🟡 In Evaluierung 0%
🟢 Keine Auswirkung 0%
Spannende Innovation 0%
Verwandte Story-Cluster & Quellen (Vektor-KI)
Port 8095 Engine
1 Quelle
The Gemini desktop app is now available for Windows
1 Quelle
Burn Out, Or Fade Away
1 Quelle
Windows 11 just dropped the tool ransomware abused, Microsoft says don’t restore WMIC
Ähnliche Beiträge
🔍 Verwandte News

Auch interessante Nachrichten Making Two TTS Voices Sound Like an Actual Conversation

Thematisch verwandte Begriffe: Making, Voices, Sound, Like · 6 Treffer

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...

Laden...

Beiträge werden geladen ...

Laden...

Videos werden geladen ...