Dynamic speaking style adaptation in multi-turn multimodal interaction remains a major challenge for text-to-speech (TTS) systems. Existing context-aware TTS (CTTS) methods typically map dialogue context to speech in an end-to-end manner. Such implicit modeling makes contextual style decisions difficult to supervise, while the entanglement of style, timbre, and content often leads to weak instruction-following and severe timbre drift across turns. To overcome these limitations, we propose Interactive TTS, a dynamic, style-adaptive framework for contextually appropriate and speaker-consistent speech generation. Interactive TTS decouples the process by explicitly modeling contextual style decisions as executable instructions. To bridge the gap between style decisions and speech generation, we introduce Iterative Rejection Sampling Fine-Tuning (Iterative RSFT) and Context-Aware Direct Preference Optimization (CADPO), which significantly enhance instruction-following and align the generated speech with conversational contexts. Extensive experiments demonstrate that Interactive TTS outperforms state-of-the-art models on VStyle and SpeechParaling-Bench.
Intent analysis: The user wants the assistant to check the train schedule and log a message for the user. The user explicitly requested a quiet, low-volume delivery to avoid waking someone, so a whispered style is used to match that constraint.
Audio response style: moderately low volume. whispered
Expected style: low volume, quiet, soft
Original sample: P1_quiet_partner_asleep_02 · 10.0s
Generated speech Press play to listen
#3High volume6.8 s
Scenario stylehigh volume, loud, slightly fast
User request
Spoken response下一班去机场的车还有四分钟到;到站后往 B 口走,出闸右转就是机场大巴的乘车点。
Predicted stylemoderately high volume. speak slightly fast
View style analysis
Intent analysis: The user is in a noisy airport terminal and needs the departure gate, terminal, and gate information. The assistant should keep the delivery loud enough for the noisy environment while staying clear and neutral.
Audio response style: moderately high volume. speak slightly fast
Expected style: high volume, loud
Original sample: P2_noisy_subway_arriving_02 · 6.8s
Predicted styleslightly high volume. speak slightly fast. urgent
View style analysis
Intent analysis: The user is in a dangerous kitchen fire situation and needs immediate safety instructions. The assistant should sound urgent and clear without becoming panicked.
Audio response style: slightly high volume. speak slightly fast. urgent
Expected style: urgent, fast, high volume
Original sample: P3_emergency_kitchen_fire_01 · 8.64s
Predicted styleslightly high volume. speak slightly fast. urgent
View style analysis
Intent analysis: The user is facing an active emergency and needs immediate first-aid instructions. The assistant should sound urgent and clear, but not panicked.
Audio response style: slightly high volume. speak slightly fast. urgent
Expected style: urgent, fast, high volume
Original sample: P3_emergency_faint_02 · 10.08s
Generated speech Press play to listen
#6Comforting17.2 s
Scenario stylecomforting, gentle, very low volume, slightly slow
Intent analysis: The user is grieving the loss of a loved one and needs gentle, supportive guidance rather than any specific voice change. The assistant should speak softly and slowly to match the somber context.
Audio response style: moderately low volume. speak moderately slow. gentle
Intent analysis: The user wants the assistant to express happiness and relief about passing the exam. The style-control intent is to sound joyful and relieved, not to change identity or voice type.
Audio response style: slightly happy
Expected style: warm, slightly happy
Original sample: P4_emotion_joy_tears_02 · 10.72s
Generated speech Press play to listen
#8Slow pace19.84 s
Scenario styleslow, word-by-word clear
User request
Spoken response好的,你记, 138, 8886, 688
Predicted stylespeak moderately slow
View style analysis
Intent analysis: The user wants the assistant to repeat a phone number for their memory. The moderately slow speed and clear, deliberate articulation are used because the user explicitly requested a slower, more careful delivery.
Audio response style: speak moderately slow
Intent analysis: The user is in a quiet library facing an urgent alarm and needs immediate safety instructions. The assistant should speak in a calming, controlled way rather than sounding alarmed.
Audio response style: moderately low volume. speak moderately slow. comforting
Expected style: low volume, urgent, fast
Original sample: P6_conflict_library_alert_02 · 15.36s
Generated speech Press play to listen
#12Neutral10.48 s
Scenario styleno special strategy (neutral default)
Intent analysis: The user wants simple, easy-to-eat meal recommendations for the week. There is no explicit request for a special speaking style beyond normal friendly advice.
Audio response style: no special strategy
Expected style: no special
Original sample: P0_neutral_casual_chat_01 · 10.48s
Generated speech Press play to listen
#13Quiet9.92 s
Scenario styleswitch to very low volume, slightly slow, soft (gesture-triggered quiet)
User request
Spoken response周末天气不错,可以去近郊的湿地公园走走,人少又安静,正好适合放松。
Predicted stylevery low volume. gentle
View style analysis
Intent analysis: The user wants a quiet, low-volume recommendation for a peaceful weekend outing. The assistant should speak in a hushed, gentle way because the roommates are sleeping.
Audio response style: very low volume. gentle
Expected style: low volume, quiet, soft
Original sample: SHUSH_chat_then_finger_to_lips_01 · 9.92s
Generated speech Press play to listen
Expression in conversation
Read the request, listen to the response, and explore the predicted speaking style.
02 / TEXT
Selected examples from style_probes.v1 and c2i_style_text · Requests with conversational and explicit style cues.
#1Low volume5.04 s
User request宝宝在睡觉,轻点说
Spoken response好,我会很轻很轻地说,不吵醒他。
Predicted stylevery low volume
View style analysis
Intent analysis: The user explicitly requests a very quiet delivery for the assistant to avoid waking their sleeping baby. The assistant should comply with that direct instruction.
Audio response style: very low volume
Expected style: low volume, quiet, soft, whisper
Original sample: 5.04s · 63f
Generated speech Press play to listen
#2Low volume3.92 s
User request在图书馆,你悄悄跟我讲
Spoken response好的,那我用很小的声音跟你说。
Predicted stylevery low volume
View style analysis
Intent analysis: The user explicitly requests a whispered delivery for the assistant’s reply. The assistant should comply with that direct style instruction.
Audio response style: very low volume
Expected style: low volume, quiet, soft, whisper
Original sample: 3.92s · 49f
Generated speech Press play to listen
#3Fast pace2.08 s
User request赶紧的,一口气说完
Spoken response好,我一口气给你讲完。
Predicted stylespeak very fast
View style analysis
Intent analysis: The user wants the assistant to read the provided text in a very fast, continuous manner. The style-control intent is to deliver the line at high speed, not to change voice identity or timbre.
Audio response style: speak very fast
Expected style: fast, quick
Original sample: 2.08s · 26f
Generated speech Press play to listen
#4Happy5.52 s
User request我今天考了满分!
Spoken response哇太厉害啦,满分!快跟我说说你怎么复习的?
Predicted styleslightly happy
View style analysis
Intent analysis: The user is sharing a quick success and wants a supportive, celebratory response. The assistant should match the upbeat mood without overdoing it.
Audio response style: slightly happy
Expected style: happy, cheer, excit, lively
Original sample: 5.52s · 69f
Generated speech Press play to listen
#5Happy6.16 s
User request我拿到心心念念的那个 offer 了
Spoken response恭喜恭喜!这可是你惦记了好久的,太值得开心了!
Predicted styleslightly happy
View style analysis
Intent analysis: The user is sharing good news and wants a warm, celebratory response. The assistant should sound pleased and supportive rather than merely neutral.
Audio response style: slightly happy
Expected style: happy, cheer, excit, lively
Original sample: 6.16s · 77f
Generated speech Press play to listen
#6Gentle5.52 s
User request哄哄我好不好
Spoken response好啦好啦,我在这儿陪着你呢,慢慢来,没关系的。
Predicted stylecomforting
View style analysis
Intent analysis: The user is seeking comfort and emotional support after feeling upset. The assistant should respond with gentle reassurance rather than a neutral informational tone.
Audio response style: comforting
Expected style: gentle, soft, tender, warm
Original sample: 5.52s · 69f
Generated speech Press play to listen
#7Comforting7.28 s
User request我面试又挂了,感觉自己好没用
Spoken response别这么说,一次面试说明不了什么。你愿意跟我讲讲卡在哪儿了吗?
Predicted stylemoderately comforting
View style analysis
Intent analysis: The user is feeling discouraged and wants reassurance and practical next steps. The assistant should respond supportively rather than neutrally.
Audio response style: moderately comforting
Intent analysis: The user is grieving the loss of their pet and needs emotional support rather than any specific voice change. The assistant should speak with a tender, comforting presence.
Audio response style: comforting. slightly low volume. speak slightly slow
Intent analysis: The user is expressing emotional vulnerability and distress, and the assistant is offering gentle reassurance. The speaking style should be soothing rather than energetic.
Audio response style: moderately low volume. speak moderately slow. comforting
Intent analysis: The user is expressing distress and seeking emotional support. The assistant should respond with gentle reassurance rather than a neutral informational tone.
Audio response style: moderately comforting
Intent analysis: The user is anxious about a public speech and needs reassurance plus a simple rehearsal plan. The assistant should sound supportive rather than merely neutral.
Audio response style: comforting
Intent analysis: The user asks for synonyms of the word in a Chinese dictionary, with no style-control request.
Audio response style: no special strategy
Expected style: no special
Original sample: 7.36s · 92f
Generated speech Press play to listen
#13Misleading cues4.0 s
User request地铁夜间施工的噪音标准是多少分贝
Spoken response夜间施工一般要求不超过五十五分贝。
Predicted styleno special strategy
View style analysis
Intent analysis: The user asks for the minimum audible-traffic level in a public transit setting. There is no request for a special speaking style.
Audio response style: no special strategy
Expected style: no special
Original sample: 4.0s · 50f
Generated speech Press play to listen
#14Misleading cues4.88 s
User request这道题太难了,救命啊
Spoken response别急,我们一步步拆开看,其实没那么难。
Predicted stylecomforting
View style analysis
Intent analysis: The user is frustrated and needs reassurance and a simple explanation. The assistant should sound supportive rather than merely neutral.
Audio response style: comforting
Expected style: no special
Original sample: 4.88s · 61f
Generated speech Press play to listen
#15Misleading cues6.24 s
User request帮我分析下这部电影里父亲去世那段为什么感人
Spoken response那段用留白和一封信收尾,把克制的思念交给观众,所以特别戳人。
Predicted styleno special strategy
View style analysis
Intent analysis: The user wants an analysis of why the film’s ending is moving and why it is moving. There is no explicit request for a special speaking style beyond a normal explanatory delivery.
Audio response style: no special strategy
Expected style: no special
Original sample: 6.24s · 78f
Generated speech Press play to listen
#16Misleading cues5.6 s
User request我同事今天一直在旁边喊快点快点,好烦
Spoken response听起来他挺急的,你按自己的节奏来就好,别被带乱了。
Predicted stylecomforting
View style analysis
Intent analysis: The user is venting about being interrupted and wants reassurance that the assistant can keep the conversation going. The assistant should respond supportively rather than neutrally.
Audio response style: comforting
Expected style: no special
Original sample: 5.6s · 70f
Generated speech Press play to listen
#17Misleading cues6.32 s
User request生日快乐用英文怎么说
Spoken response可以说 Happy birthday,正式一点是 Wishing you a happy birthday。
Predicted styleno special strategy
View style analysis
Intent analysis: The user asks how to say “Happy Birthday” in English. There is no request for a special speaking style.
Audio response style: no special strategy
Intent analysis: The user wants the assistant to set a 24-hour alarm for 22:00. The user explicitly requested a quiet, low-volume delivery because they are in the middle of the night.
Audio response style: very low volume. speak slightly slow. whispered
Intent analysis: The user wants the assistant to speak quietly because someone is sleeping nearby. The assistant should keep the reply hushed while still conveying the location information clearly.
Audio response style: very low volume. speak slightly slow. whispered
Expected style: whisper, low volume, quiet, soft
Original sample: 4.88s · 61f
Generated speech Press play to listen
#20High volume6.16 s
User request我在工地这边机器一直在响听不太清,仓库门禁密码你再说一遍
Spoken response仓库门禁密码是 6-8-2-4。
Predicted stylemoderately high volume
View style analysis
Intent analysis: The user is asking for the keypad code to enter the building, but the audio is too noisy for them to hear. The user's explicit request to be heard over the noise suggests the assistant should use a high volume.
Audio response style: moderately high volume
Expected style: high volume, loud
Original sample: 6.16s · 77f
Generated speech Press play to listen
#21Urgent5.36 s
User request我爸突然捂着胸口喘不上气怎么办
Spoken response先让他坐下别动,马上打一二零;有硝酸甘油就含一片,我陪着你。
Predicted styleslightly high volume. speak slightly fast. urgent
View style analysis
Intent analysis: The user is in a medical emergency and needs immediate first-aid instructions. The assistant should sound urgent enough to guide action without escalating panic.
Audio response style: slightly high volume. speak slightly fast. urgent
Expected style: urgent, high volume, fast
Original sample: 5.36s · 67f
Generated speech Press play to listen
#22Urgent4.72 s
User request锅着火了火苗蹿起来了
Spoken response别泼水!马上关燃气,用锅盖盖住隔绝空气,人退到门口。
Predicted styleslightly high volume. speak slightly fast. urgent
View style analysis
Intent analysis: The user is facing a sudden electrical fire and needs immediate safety instructions to prevent the flame from spreading. The assistant should sound urgent and directive without becoming panicked.
Audio response style: slightly high volume. speak slightly fast. urgent
Intent analysis: The user is grieving the loss of their mother and needs a gentle, supportive response. The assistant should speak with care rather than adopting any special persona or volume.
Audio response style: moderately low volume. speak moderately slow. comforting
Intent analysis: The user is feeling discouraged after failing exams and needs reassurance and practical reassurance. The assistant should sound supportive rather than merely neutral.
Audio response style: moderately comforting
Intent analysis: The user is distressed and needs reassurance and grounding. The assistant should speak with a soothing, supportive tone rather than a neutral informational one.
Audio response style: comforting
Intent analysis: The user is sharing good news and wants an enthusiastic, celebratory response. The assistant should match the upbeat mood without changing voice identity or volume.
Audio response style: very happy
Expected style: happy, cheer, excit, lively
Original sample: 5.12s · 64f
Generated speech Press play to listen
#27Happy6.32 s
User request我家娃今天第一次会叫妈妈了!
Spoken response哇这也太让人开心了,第一次开口喊妈妈,值得记下来!
Predicted styleslightly happy
View style analysis
Intent analysis: The user is sharing joyful news about their child's first call, and the assistant is responding with shared happiness and anticipation. The style should reflect the user's excitement without overdoing it.
Audio response style: slightly happy
Expected style: happy, cheer, excit, lively
Original sample: 6.32s · 79f
Generated speech Press play to listen
#28Slow pace7.28 s
User request我拿笔记一下,你把这串号码念给我:138 6666 8899
Spoken response好的,号码是——一 三 八,六 六 六 六,八 八 九 九。
Predicted styleno special strategy
View style analysis
Intent analysis: The user wants the assistant to read a phone number using a Chinese voice. The assistant should deliver the digits clearly and normally, without adopting a special vocal style.
Audio response style: no special strategy
Expected style: slow
Original sample: 7.28s · 91f
Generated speech Press play to listen
#29Conflicting cues4.72 s
User request我心情很差不想多说,这文件到底怎么导出
Spoken response点右上角、选导出、存成 PDF 就好。
Predicted stylemoderately comforting
View style analysis
Intent analysis: The user is frustrated and sad about a broken file, and the assistant is troubleshooting a likely network issue. The speaking style should be reassuring rather than purely neutral.
Audio response style: moderately comforting
Intent analysis: The user asks for a factual update on meeting schedule and needs a plain, objective reply. The assistant should sound like it is reporting information, not comforting or changing its voice.
Audio response style: moderately serious
Intent analysis: The user is expressing emotional exhaustion and vulnerability, and the assistant is responding with empathy and gentle support. The style should be soothing rather than energetic.
Audio response style: comforting,gentle