Turning a novel into a multi-voice audiobook sounds like a TTS problem. It is mostly a parsing problem, and the parsing is harder than the synthesis.




Attribution is not solved by quotation marks


The naive approach: find text in quotes, that is dialogue, assign a voice. It breaks immediately on real prose.



"I told you," she said,...