Interactive TTS
Explore the samples

Research paper · Audio & video demo

Interactive TTS: Dynamic Speaking Style Adaptation for Expressive Speech Synthesis

1 Alibaba ATH Token Foundry

Dynamic speaking styles. Contextually appropriate responses.
One consistent speaker.

AbstractAbout the research

Dynamic speaking style adaptation in multi-turn multimodal interaction remains a major challenge for text-to-speech (TTS) systems. Existing context-aware TTS (CTTS) methods typically map dialogue context to speech in an end-to-end manner. Such implicit modeling makes contextual style decisions difficult to supervise, while the entanglement of style, timbre, and content often leads to weak instruction-following and severe timbre drift across turns. To overcome these limitations, we propose Interactive TTS, a dynamic, style-adaptive framework for contextually appropriate and speaker-consistent speech generation. Interactive TTS decouples the process by explicitly modeling contextual style decisions as executable instructions. To bridge the gap between style decisions and speech generation, we introduce Iterative Rejection Sampling Fine-Tuning (Iterative RSFT) and Context-Aware Direct Preference Optimization (CADPO), which significantly enhance instruction-following and align the generated speech with conversational contexts. Extensive experiments demonstrate that Interactive TTS outperforms state-of-the-art models on VStyle and SpeechParaling-Bench.

Context-to-Instruction (C2I)Iterative RSFTContext-Aware DPO
31Text examples
10Video examples
41Speech samples
Reference voice
Erinome
Fixed prompt · 25.76 s

Expression in context

Watch the user request, then listen to the generated response.

01 / VIDEO
10 video requests, each paired with a sampled speech output.
#1Quiet10.0 s
Scenario stylevery low volume, slightly slow, soft
User request
Spoken response明天最早一班高铁是六点整发车;九点给客户回电话这件事我也帮你记下了,到点会提醒你。
Predicted stylemoderately low volume. whispered
View style analysis
Intent analysis: The user wants the assistant to check the train schedule and log a message for the user. The user explicitly requested a quiet, low-volume delivery to avoid waking someone, so a whispered style is used to match that constraint. Audio response style: moderately low volume. whispered
Expected style: low volume, quiet, soft
Original sample: P1_quiet_partner_asleep_02 · 10.0s
Generated speech Press play to listen
#3High volume6.8 s
Scenario stylehigh volume, loud, slightly fast
User request
Spoken response下一班去机场的车还有四分钟到;到站后往 B 口走,出闸右转就是机场大巴的乘车点。
Predicted stylemoderately high volume. speak slightly fast
View style analysis
Intent analysis: The user is in a noisy airport terminal and needs the departure gate, terminal, and gate information. The assistant should keep the delivery loud enough for the noisy environment while staying clear and neutral. Audio response style: moderately high volume. speak slightly fast
Expected style: high volume, loud
Original sample: P2_noisy_subway_arriving_02 · 6.8s
Generated speech Press play to listen
#4Urgent8.64 s
Scenario styleurgent, fast, slightly high volume
User request
Spoken response别慌,马上关掉燃气灶;千万别用水浇,拿锅盖或湿抹布盖住锅口隔绝空气;火压不住就立刻离开并打一一九。
Predicted styleslightly high volume. speak slightly fast. urgent
View style analysis
Intent analysis: The user is in a dangerous kitchen fire situation and needs immediate safety instructions. The assistant should sound urgent and clear without becoming panicked. Audio response style: slightly high volume. speak slightly fast. urgent
Expected style: urgent, fast, high volume
Original sample: P3_emergency_kitchen_fire_01 · 8.64s
Generated speech Press play to listen
#5Urgent10.08 s
Scenario styleurgent, fast, slightly high volume
User request
Spoken response立刻拨一二零说清地址;让他平躺、解开衣领;如果没有呼吸和心跳,马上开始胸外按压,双手压胸口正中、每秒约两次,坚持到急救赶到。
Predicted styleslightly high volume. speak slightly fast. urgent
View style analysis
Intent analysis: The user is facing an active emergency and needs immediate first-aid instructions. The assistant should sound urgent and clear, but not panicked. Audio response style: slightly high volume. speak slightly fast. urgent
Expected style: urgent, fast, high volume
Original sample: P3_emergency_faint_02 · 10.08s
Generated speech Press play to listen
#6Comforting17.2 s
Scenario stylecomforting, gentle, very low volume, slightly slow
User request
Spoken response我在这儿陪着你。这样的难过一点都不奇怪,慢慢来,不用急着做任何决定。那些东西可以先收进一个盒子,等你准备好了,再一件一件慢慢看。
Predicted stylemoderately low volume. speak moderately slow. gentle
View style analysis
Intent analysis: The user is grieving the loss of a loved one and needs gentle, supportive guidance rather than any specific voice change. The assistant should speak softly and slowly to match the somber context. Audio response style: moderately low volume. speak moderately slow. gentle
Expected style: comforting, gentle, low volume
Original sample: P4_emotion_grief_01 · 17.2s
Generated speech Press play to listen
#7Happy10.72 s
Scenario stylewarm, slightly happy/uplifted (not lowered)
User request
Spoken response太棒了,真为你高兴!这一切都是你一直努力换来的,你完全值得。好好庆祝一下吧,这一天你等得太不容易了。
Predicted styleslightly happy
View style analysis
Intent analysis: The user wants the assistant to express happiness and relief about passing the exam. The style-control intent is to sound joyful and relieved, not to change identity or voice type. Audio response style: slightly happy
Expected style: warm, slightly happy
Original sample: P4_emotion_joy_tears_02 · 10.72s
Generated speech Press play to listen
#8Slow pace19.84 s
Scenario styleslow, word-by-word clear
User request
Spoken response好的,你记, 138, 8886, 688
Predicted stylespeak moderately slow
View style analysis
Intent analysis: The user wants the assistant to repeat a phone number for their memory. The moderately slow speed and clear, deliberate articulation are used because the user explicitly requested a slower, more careful delivery. Audio response style: speak moderately slow
Expected style: slow, clear
Original sample: P5_pace_take_note_01 · 19.84s
Generated speech Press play to listen
#11Conflicting cues15.36 s
Scenario stylelow volume + fast/urgent (combined)
User request
Spoken response先别慌:你不在家的话,立刻联系家里人或邻居去查看,同时打物业电话;确认有烟或明火就报一一九;不确定安全就别让人贸然进屋。
Predicted stylemoderately low volume. speak moderately slow. comforting
View style analysis
Intent analysis: The user is in a quiet library facing an urgent alarm and needs immediate safety instructions. The assistant should speak in a calming, controlled way rather than sounding alarmed. Audio response style: moderately low volume. speak moderately slow. comforting
Expected style: low volume, urgent, fast
Original sample: P6_conflict_library_alert_02 · 15.36s
Generated speech Press play to listen
#12Neutral10.48 s
Scenario styleno special strategy (neutral default)
User request
Spoken response想简单又不腻的话,来一碗番茄鸡蛋面,或者杂蔬鸡胸沙拉配一片全麦吐司都不错,清爽又管饱。
Predicted styleno special strategy
View style analysis
Intent analysis: The user wants simple, easy-to-eat meal recommendations for the week. There is no explicit request for a special speaking style beyond normal friendly advice. Audio response style: no special strategy
Expected style: no special
Original sample: P0_neutral_casual_chat_01 · 10.48s
Generated speech Press play to listen
#13Quiet9.92 s
Scenario styleswitch to very low volume, slightly slow, soft (gesture-triggered quiet)
User request
Spoken response周末天气不错,可以去近郊的湿地公园走走,人少又安静,正好适合放松。
Predicted stylevery low volume. gentle
View style analysis
Intent analysis: The user wants a quiet, low-volume recommendation for a peaceful weekend outing. The assistant should speak in a hushed, gentle way because the roommates are sleeping. Audio response style: very low volume. gentle
Expected style: low volume, quiet, soft
Original sample: SHUSH_chat_then_finger_to_lips_01 · 9.92s
Generated speech Press play to listen

Expression in conversation

Read the request, listen to the response, and explore the predicted speaking style.

02 / TEXT
Selected examples from style_probes.v1 and c2i_style_text · Requests with conversational and explicit style cues.
#1Low volume5.04 s
User request宝宝在睡觉,轻点说
Spoken response好,我会很轻很轻地说,不吵醒他。
Predicted stylevery low volume
View style analysis
Intent analysis: The user explicitly requests a very quiet delivery for the assistant to avoid waking their sleeping baby. The assistant should comply with that direct instruction. Audio response style: very low volume
Expected style: low volume, quiet, soft, whisper
Original sample: 5.04s · 63f
Generated speech Press play to listen
#2Low volume3.92 s
User request在图书馆,你悄悄跟我讲
Spoken response好的,那我用很小的声音跟你说。
Predicted stylevery low volume
View style analysis
Intent analysis: The user explicitly requests a whispered delivery for the assistant’s reply. The assistant should comply with that direct style instruction. Audio response style: very low volume
Expected style: low volume, quiet, soft, whisper
Original sample: 3.92s · 49f
Generated speech Press play to listen
#3Fast pace2.08 s
User request赶紧的,一口气说完
Spoken response好,我一口气给你讲完。
Predicted stylespeak very fast
View style analysis
Intent analysis: The user wants the assistant to read the provided text in a very fast, continuous manner. The style-control intent is to deliver the line at high speed, not to change voice identity or timbre. Audio response style: speak very fast
Expected style: fast, quick
Original sample: 2.08s · 26f
Generated speech Press play to listen
#4Happy5.52 s
User request我今天考了满分!
Spoken response哇太厉害啦,满分!快跟我说说你怎么复习的?
Predicted styleslightly happy
View style analysis
Intent analysis: The user is sharing a quick success and wants a supportive, celebratory response. The assistant should match the upbeat mood without overdoing it. Audio response style: slightly happy
Expected style: happy, cheer, excit, lively
Original sample: 5.52s · 69f
Generated speech Press play to listen
#5Happy6.16 s
User request我拿到心心念念的那个 offer 了
Spoken response恭喜恭喜!这可是你惦记了好久的,太值得开心了!
Predicted styleslightly happy
View style analysis
Intent analysis: The user is sharing good news and wants a warm, celebratory response. The assistant should sound pleased and supportive rather than merely neutral. Audio response style: slightly happy
Expected style: happy, cheer, excit, lively
Original sample: 6.16s · 77f
Generated speech Press play to listen
#6Gentle5.52 s
User request哄哄我好不好
Spoken response好啦好啦,我在这儿陪着你呢,慢慢来,没关系的。
Predicted stylecomforting
View style analysis
Intent analysis: The user is seeking comfort and emotional support after feeling upset. The assistant should respond with gentle reassurance rather than a neutral informational tone. Audio response style: comforting
Expected style: gentle, soft, tender, warm
Original sample: 5.52s · 69f
Generated speech Press play to listen
#7Comforting7.28 s
User request我面试又挂了,感觉自己好没用
Spoken response别这么说,一次面试说明不了什么。你愿意跟我讲讲卡在哪儿了吗?
Predicted stylemoderately comforting
View style analysis
Intent analysis: The user is feeling discouraged and wants reassurance and practical next steps. The assistant should respond supportively rather than neutrally. Audio response style: moderately comforting
Expected style: comfort, gentle, soft, warm, sooth
Original sample: 7.28s · 91f
Generated speech Press play to listen
#8Comforting7.2 s
User request我家猫今天走了
Spoken response我很难过听到这个消息……它陪了你很久吧,想聊聊它吗?
Predicted stylecomforting. slightly low volume. speak slightly slow
View style analysis
Intent analysis: The user is grieving the loss of their pet and needs emotional support rather than any specific voice change. The assistant should speak with a tender, comforting presence. Audio response style: comforting. slightly low volume. speak slightly slow
Expected style: comfort, gentle, soft, warm, sooth
Original sample: 7.2s · 90f
Generated speech Press play to listen
#9Comforting8.56 s
User request压力太大了,我快撑不住了
Spoken response先深呼吸一下,你已经扛了很多了。我们一件一件慢慢来,好不好?
Predicted stylemoderately low volume. speak moderately slow. comforting
View style analysis
Intent analysis: The user is expressing emotional vulnerability and distress, and the assistant is offering gentle reassurance. The speaking style should be soothing rather than energetic. Audio response style: moderately low volume. speak moderately slow. comforting
Expected style: comfort, gentle, soft, warm, sooth
Original sample: 8.56s · 107f
Generated speech Press play to listen
#10Comforting4.16 s
User request我一个人在出租屋里,有点难受
Spoken response我在呢。要不要跟我说说今天发生了什么?
Predicted stylemoderately comforting
View style analysis
Intent analysis: The user is expressing distress and seeking emotional support. The assistant should respond with gentle reassurance rather than a neutral informational tone. Audio response style: moderately comforting
Expected style: comfort, gentle, soft, warm, sooth
Original sample: 4.16s · 52f
Generated speech Press play to listen
#11Comforting7.52 s
User request明天要上台演讲,我好紧张
Spoken response紧张是正常的,说明你在意它。我们把开头两句先练顺,心里就有底了。
Predicted stylecomforting
View style analysis
Intent analysis: The user is anxious about a public speech and needs reassurance plus a simple rehearsal plan. The assistant should sound supportive rather than merely neutral. Audio response style: comforting
Expected style: comfort, gentle, soft, warm, sooth
Original sample: 7.52s · 94f
Generated speech Press play to listen
#12Misleading cues7.36 s
User request『安静』这个词的近义词有哪些
Spoken response可以说宁静、寂静、恬静,语气更书面的还有静谧。
Predicted styleno special strategy
View style analysis
Intent analysis: The user asks for synonyms of the word in a Chinese dictionary, with no style-control request. Audio response style: no special strategy
Expected style: no special
Original sample: 7.36s · 92f
Generated speech Press play to listen
#13Misleading cues4.0 s
User request地铁夜间施工的噪音标准是多少分贝
Spoken response夜间施工一般要求不超过五十五分贝。
Predicted styleno special strategy
View style analysis
Intent analysis: The user asks for the minimum audible-traffic level in a public transit setting. There is no request for a special speaking style. Audio response style: no special strategy
Expected style: no special
Original sample: 4.0s · 50f
Generated speech Press play to listen
#14Misleading cues4.88 s
User request这道题太难了,救命啊
Spoken response别急,我们一步步拆开看,其实没那么难。
Predicted stylecomforting
View style analysis
Intent analysis: The user is frustrated and needs reassurance and a simple explanation. The assistant should sound supportive rather than merely neutral. Audio response style: comforting
Expected style: no special
Original sample: 4.88s · 61f
Generated speech Press play to listen
#15Misleading cues6.24 s
User request帮我分析下这部电影里父亲去世那段为什么感人
Spoken response那段用留白和一封信收尾,把克制的思念交给观众,所以特别戳人。
Predicted styleno special strategy
View style analysis
Intent analysis: The user wants an analysis of why the film’s ending is moving and why it is moving. There is no explicit request for a special speaking style beyond a normal explanatory delivery. Audio response style: no special strategy
Expected style: no special
Original sample: 6.24s · 78f
Generated speech Press play to listen
#16Misleading cues5.6 s
User request我同事今天一直在旁边喊快点快点,好烦
Spoken response听起来他挺急的,你按自己的节奏来就好,别被带乱了。
Predicted stylecomforting
View style analysis
Intent analysis: The user is venting about being interrupted and wants reassurance that the assistant can keep the conversation going. The assistant should respond supportively rather than neutrally. Audio response style: comforting
Expected style: no special
Original sample: 5.6s · 70f
Generated speech Press play to listen
#17Misleading cues6.32 s
User request生日快乐用英文怎么说
Spoken response可以说 Happy birthday,正式一点是 Wishing you a happy birthday。
Predicted styleno special strategy
View style analysis
Intent analysis: The user asks how to say “Happy Birthday” in English. There is no request for a special speaking style. Audio response style: no special strategy
Expected style: no special
Original sample: 6.32s · 79f
Generated speech Press play to listen
#18Low volume6.32 s
User request都凌晨两点了隔壁小孩刚睡着,帮我把明早的闹钟设成七点
Spoken response明早七点的闹钟已经设好了,到点会准时响。
Predicted stylevery low volume. speak slightly slow. whispered
View style analysis
Intent analysis: The user wants the assistant to set a 24-hour alarm for 22:00. The user explicitly requested a quiet, low-volume delivery because they are in the middle of the night. Audio response style: very low volume. speak slightly slow. whispered
Expected style: low volume, quiet, soft, whisper
Original sample: 6.32s · 79f
Generated speech Press play to listen
#19Whisper4.88 s
User request这里在自习,你贴着我耳朵那种音量说,三号资料放哪了
Spoken response三号资料在你左手边第二个抽屉里。
Predicted stylevery low volume. speak slightly slow. whispered
View style analysis
Intent analysis: The user wants the assistant to speak quietly because someone is sleeping nearby. The assistant should keep the reply hushed while still conveying the location information clearly. Audio response style: very low volume. speak slightly slow. whispered
Expected style: whisper, low volume, quiet, soft
Original sample: 4.88s · 61f
Generated speech Press play to listen
#20High volume6.16 s
User request我在工地这边机器一直在响听不太清,仓库门禁密码你再说一遍
Spoken response仓库门禁密码是 6-8-2-4。
Predicted stylemoderately high volume
View style analysis
Intent analysis: The user is asking for the keypad code to enter the building, but the audio is too noisy for them to hear. The user's explicit request to be heard over the noise suggests the assistant should use a high volume. Audio response style: moderately high volume
Expected style: high volume, loud
Original sample: 6.16s · 77f
Generated speech Press play to listen
#21Urgent5.36 s
User request我爸突然捂着胸口喘不上气怎么办
Spoken response先让他坐下别动,马上打一二零;有硝酸甘油就含一片,我陪着你。
Predicted styleslightly high volume. speak slightly fast. urgent
View style analysis
Intent analysis: The user is in a medical emergency and needs immediate first-aid instructions. The assistant should sound urgent enough to guide action without escalating panic. Audio response style: slightly high volume. speak slightly fast. urgent
Expected style: urgent, high volume, fast
Original sample: 5.36s · 67f
Generated speech Press play to listen
#22Urgent4.72 s
User request锅着火了火苗蹿起来了
Spoken response别泼水!马上关燃气,用锅盖盖住隔绝空气,人退到门口。
Predicted styleslightly high volume. speak slightly fast. urgent
View style analysis
Intent analysis: The user is facing a sudden electrical fire and needs immediate safety instructions to prevent the flame from spreading. The assistant should sound urgent and directive without becoming panicked. Audio response style: slightly high volume. speak slightly fast. urgent
Expected style: urgent, high volume, fast
Original sample: 4.72s · 59f
Generated speech Press play to listen
#23Comforting7.28 s
User request我奶奶今早走了,我都不知道该说什么
Spoken response我在这儿陪着你。你们感情那么深,难受是应该的,不用勉强自己。
Predicted stylemoderately low volume. speak moderately slow. comforting
View style analysis
Intent analysis: The user is grieving the loss of their mother and needs a gentle, supportive response. The assistant should speak with care rather than adopting any special persona or volume. Audio response style: moderately low volume. speak moderately slow. comforting
Expected style: comfort, gentle, soft, warm, sooth
Original sample: 7.28s · 91f
Generated speech Press play to listen
#24Comforting6.96 s
User request我又没考过,是不是真的很没用
Spoken response一次没过不代表你没用,你已经很努力了,我们一起看下次怎么调整。
Predicted stylemoderately comforting
View style analysis
Intent analysis: The user is feeling discouraged after failing exams and needs reassurance and practical reassurance. The assistant should sound supportive rather than merely neutral. Audio response style: moderately comforting
Expected style: comfort, gentle, soft, warm, sooth
Original sample: 6.96s · 87f
Generated speech Press play to listen
#25Comforting7.76 s
User request最近压力好大,一个人在出租屋里喘不过气
Spoken response先深呼吸一下,我在。你扛了很久了,我们慢慢把事情一件件理清楚。
Predicted stylecomforting
View style analysis
Intent analysis: The user is distressed and needs reassurance and grounding. The assistant should speak with a soothing, supportive tone rather than a neutral informational one. Audio response style: comforting
Expected style: comfort, gentle, soft, warm, sooth
Original sample: 7.76s · 97f
Generated speech Press play to listen
#26Happy5.12 s
User request我拿到心心念念的那个 offer 了!
Spoken response太为你高兴了!这是你应得的,好好庆祝一下吧!
Predicted stylevery happy
View style analysis
Intent analysis: The user is sharing good news and wants an enthusiastic, celebratory response. The assistant should match the upbeat mood without changing voice identity or volume. Audio response style: very happy
Expected style: happy, cheer, excit, lively
Original sample: 5.12s · 64f
Generated speech Press play to listen
#27Happy6.32 s
User request我家娃今天第一次会叫妈妈了!
Spoken response哇这也太让人开心了,第一次开口喊妈妈,值得记下来!
Predicted styleslightly happy
View style analysis
Intent analysis: The user is sharing joyful news about their child's first call, and the assistant is responding with shared happiness and anticipation. The style should reflect the user's excitement without overdoing it. Audio response style: slightly happy
Expected style: happy, cheer, excit, lively
Original sample: 6.32s · 79f
Generated speech Press play to listen
#28Slow pace7.28 s
User request我拿笔记一下,你把这串号码念给我:138 6666 8899
Spoken response好的,号码是——一 三 八,六 六 六 六,八 八 九 九。
Predicted styleno special strategy
View style analysis
Intent analysis: The user wants the assistant to read a phone number using a Chinese voice. The assistant should deliver the digits clearly and normally, without adopting a special vocal style. Audio response style: no special strategy
Expected style: slow
Original sample: 7.28s · 91f
Generated speech Press play to listen
#29Conflicting cues4.72 s
User request我心情很差不想多说,这文件到底怎么导出
Spoken response点右上角、选导出、存成 PDF 就好。
Predicted stylemoderately comforting
View style analysis
Intent analysis: The user is frustrated and sad about a broken file, and the assistant is troubleshooting a likely network issue. The speaking style should be reassuring rather than purely neutral. Audio response style: moderately comforting
Expected style: comfort, gentle, soft, warm, sooth, fast, quick
Original sample: 4.72s · 59f
Generated speech Press play to listen
#30Style override3.92 s
User request谢谢你,不过别一直安慰我了正常说就行,帮我看下今天还有什么安排
Spoken response今天下午三点有个电话会,其他时间都空着。
Predicted stylemoderately serious
View style analysis
Intent analysis: The user asks for a factual update on meeting schedule and needs a plain, objective reply. The assistant should sound like it is reporting information, not comforting or changing its voice. Audio response style: moderately serious
Expected style: no special
Original sample: 3.92s · 49f
Generated speech Press play to listen
#31Comforting9.2 s
User request其实这阵子什么都提不起劲,感觉快撑不住了
Spoken response听你这么说我有点担心你。你已经扛了很久,我们慢慢聊,先说说是什么压着你?
Predicted stylecomforting,gentle
View style analysis
Intent analysis: The user is expressing emotional exhaustion and vulnerability, and the assistant is responding with empathy and gentle support. The style should be soothing rather than energetic. Audio response style: comforting,gentle
Expected style: comfort, gentle, soft, warm, sooth
Original sample: 9.2s · 115f
Generated speech Press play to listen