Inworld TTS Steering Guide for LyricWinter
Use Inworld TTS 2 steering in LyricWinter with bracket cues, pause controls, acting direction, vocal reactions, and clean regeneration workflows.
Posted by
Related reading
Cartesia Emotion and SSML Controls in LyricWinter
How LyricWinter uses surrounding dialogue context with Cartesia Sonic 3.5, plus the full Cartesia emotion list and manual SSML controls.
Manual Emotion Tags with Chatterbox FAL in LyricWinter
Use Chatterbox FAL emotion tags in LyricWinter to add laughs, sighs, gasps, yawns, groans, coughs, chuckles, and sniffles to generated story audio.
Voice Clone Sample Limits: What Actually Matters
A practical guide to LyricWinter voice sample limits, official provider guidance, and what to upload for the best cloning results.
Inworld steering lets you write performance direction directly in the text that becomes speech. In LyricWinter, the practical version is simple: choose Inworld TTS 2 for the character, then put short bracket cues before the words that need acting direction.[1]
LyricWinter shows that model as inworld-2. Use it when you want bracketed direction such as [whispering], [speaking slowly and sadly], or [trying not to cry] to affect the generated performance.
The warning icon on a character voice card is only a compatibility warning. It does not mean you are blocked from regenerating. It means LyricWinter noticed syntax that another model may not honor, so you can decide whether to switch the voice model backend or keep going.
Quick Start
- Write normal speaker labels:
Mira:,Theo:,Narrator:. - Put steering cues inside square brackets before the phrase they should affect.
- Choose
inworld-2for that character in the voice block. - Use character-level regenerate when you only changed one character's voice or steering, or whole-story regenerate when you want to rebuild the full talking book.
Mira: [whispering][anxious] Wait<break time="500ms"/>now.
Theo: [cheerful] We made it.
Theo: [angry] This is BAD.
Narrator: [slow and hushed] The hallway answered with silence.Every Way to Steer a Line
Inworld steering is natural-language direction, not a tiny fixed enum. The safest pattern is to write compact cues that a voice actor would understand, then regenerate only the clips that need another pass.[1]
Emotion
Name the emotional state before the words that need it. Start direct, then get more specific only where the performance needs it.
[sad][angry][relieved][anxious][cheerful][guilty]Intensity
Add degree words when a simple emotion is too broad. This helps separate a hint of fear from panic, or irritation from rage.
[slightly nervous][barely angry][very excited][terrified but quiet]Delivery
Describe how the line is spoken, not only what the character feels. This is useful for whispers, shouts, muttering, formal narration, and private asides.
[whispering][shouting][muttering][speaking formally][under her breath]Pace
Tell the voice when to slow down, rush, hesitate, or hold a steady cadence. Pair pacing with punctuation and pause tags for cleaner timing.
[speaking slowly][quick and clipped][hesitating][calm and measured]Volume
Control projection without changing the actual words. Volume cues are best for distance, secrecy, sudden danger, or intimate narration.
[quietly][softly][loudly][calling across the room]Tone and texture
Use tonal direction when the emotion alone is not enough. These cues are useful for deadpan comedy, grief, suspicion, warmth, or exhaustion.
[flat and emotionless][warmly][sarcastic][tired][trying not to cry]Acting intent
Give the model the subtext. These cues work best when the dialogue says one thing and the character means another.
[pretending to be calm][hiding panic][forcing a smile][barely containing anger]Vocal reactions
Place brief vocal actions where the sound should happen. Give the model performable text when you need an audible laugh or gasp.
[laughs][sighs][gasps][breathes in sharply][chuckles] Heh, okay.Pause Controls
Inworld pause controls use a break tag with a time value. In LyricWinter, keep the tag inside the line that owns the pause so the pause stays attached to the correct speaker.[2]
Mira: Wait<break time="500ms"/>now.
Mira: Hold still <break time="2s" /> and breathe.
Theo: <break time="1s" /> [cheerful] We made it.Use punctuation for ordinary rhythm and break tags for deliberate silence. Good starting values are <break time="250ms"/>, <break time="500ms"/>, <break time="1s"/>, and <break time="2s"/>.
Avoid putting a standalone break line between speakers. If the pause belongs to Mira, put it in Mira's line. If it belongs before Theo speaks, put it at the start of Theo's line.
Stacking Cues
You can stack cues when the performance needs more than one instruction. Keep the stack short and non-contradictory.
Good:
Mira: [whispering][anxious] Keep your voice down.
Good:
Narrator: [slow and hushed] The door moved again.
Too overloaded:
Mira: [whispering][shouting][sad][excited][fast][slow] I know.A useful rule: one emotional cue plus one delivery cue is usually enough. If the result still sounds wrong, edit the line text and punctuation before adding more tags.
Copy-Ready Examples
Suspense
Mira: [whispering][anxious] Wait<break time="500ms"/>now.
Mira: Hold still <break time="2s" /> and breathe.
Mira: [speaking slowly and sadly] I can still hear it.Argument
Theo: [angry] This is BAD.
Mira: [forcing herself to stay calm] I know. Keep your voice down.
Theo: [loudly, almost shouting] You said it was locked.Comedy timing
Narrator: The door opened by itself.
Mira: [flat and emotionless] Great.
Mira: <break time="700ms" /> [nervous laugh] That's completely normal.Audiobook narration
Narrator: [low and measured] The city did not sleep that night.
Narrator: [warmer] In the kitchen window, one light stayed on.
Narrator: <break time="1s" /> [quietly] And then it went out.Voice Tags vs LyricWinter Voice Blocks
Inworld also documents voice tags, which are provider-level markup for assigning voice identity inside an Inworld request.[3] In LyricWinter, you usually should not write raw voice tags in the story. Use the character voice blocks instead: LyricWinter detects speakers, lets you assign each character to a voice, and sends the right voice model backend during generation.[4]
That distinction matters. Bracket steering changes how a line is performed. LyricWinter character voice blocks decide who performs the line.
Regeneration Workflow
A mismatch icon is not a stop sign. It is a prompt to review whether the selected model understands the controls you wrote. If you changed Mira from Fish Audio to Inworld TTS 2, regenerate Mira's clips to hear that character with the new backend. If you changed the whole cast or edited many lines, regenerate the full talking book.
For Inworld steering tests, change one thing at a time: first the model, then one bracket cue, then one pause value. That makes it much easier to tell whether the parser preserved your tags and whether the backend performed them.
Common Mistakes
- Using TTS 1.5 for steering: use
inworld-2for Inworld steering. - Mixing providers: Chatterbox angle tags, Fish Audio bracket tags, Cartesia SSML, and Inworld steering are not all the same feature even when some syntax overlaps.
- Floating pause tags: put
<break>inside the line that should own the silence. - Over-tagging every sentence: steer important beats, not every clause.
- Expecting tags to fix weak dialogue: clear punctuation, sentence rhythm, and speaker context still matter.
Inworld Steering FAQ
Which LyricWinter voice model should I use for Inworld steering?
Use Inworld TTS 2. In LyricWinter's model dropdown, that is the inworld-2 option.
Can I use square-bracket steering with Fish Audio too?
Some Fish Audio controls also use square brackets, but LyricWinter treats ambiguous bracket steering as Inworld-first when it needs to suggest a backend fix. If you deliberately want Fish Audio behavior, select Fish Audio for that character.
Can I regenerate audio while a warning icon is visible?
Yes. The warning is advisory. It tells you that a selected model may not honor a control tag. You can still regenerate with the model you chose.
Should I write Inworld voice tags in LyricWinter stories?
Usually no. Use LyricWinter's character voice blocks to assign speakers to voices. Raw Inworld voice tags are a separate API-level feature.