Guides
Studio concepts: projects, sections, documents, and workflows
How LyricWinter Studio models stories as projects, sections, versioned documents, speakers, and asynchronous workflows, and how to edit them safely over the API.
BetaUpdated
LyricWinter Studio models a story as a project made of ordered sections. Each section has a versioned document of narration, dialogue, and sound-effect blocks, each spoken line belongs to a project-wide speaker, and long-running work such as generating audio happens in asynchronous workflows. This page explains each concept and how to change them safely through the LyricWinter API.
Projects and unassigned stories#
A project is a named container for related sections, such as the chapters of a book. All sections in a project share one cast, so a character keeps the same voice from chapter to chapter.
A standalone story is a single section that is not yet part of a named project. POST /studio/sections creates one. GET /studio/projects returns named projects in projects and standalone stories in unassigned_sections. A standalone story still has a project_id, and every project-scoped endpoint works with it. To turn a standalone story into a named project, call PATCH /studio/projects/{projectId} with a name.
Sections#
A section is one chapter, scene, or episode of story text. Sections in a project are ordered. Add one with POST /studio/projects/{projectId}/sections, move one with PATCH /studio/projects/{projectId}/sections/{sectionId}, and delete one with DELETE. Deletions cannot be undone through the API, so call the matching deletion-info endpoint first to see what will be removed.
Documents and blocks#
Every section has a document: the canonical, structured script of the section. Read it with GET /studio/sections/{sectionId}/document. A document contains:
document_version, which increases on every change.speakers, the project's current speaker registry, andspeaker_registry_version.blocks, the script in order. Each block has akind:
| Block kind | What it is |
|---|---|
narration | Narrated prose, voiced by the narrator. |
speech | A line of dialogue attributed to a speaker. |
sfx | A sound effect placed in the timeline. |
note | A note kept in the script but not voiced. |
excluded | Source text kept in the script but intentionally left out of the audio. |
unparsed | Text that has not been structured into narration or dialogue yet. It is not voiced until it is. |
A new section starts as unstructured text. A prepare_sections or create_audio workflow turns it into narration and speech blocks.
For a finished speaker-labeled script, POST /studio/sections/{sectionId}/script creates speech blocks directly, without a prepare workflow or a hand-built AST. Use one SPEAKER: spoken text line per turn. A [direction] tag at the start of a line steers the whole turn, and a tag mid-line, as in TORVALD: [softly] Go, Ingvild. [whispers] Please., anchors a direction where it appears; both become tracked directions and are removed from the spoken text. A separate (silence 4s) or (silence 4000ms) line inserts exact silence after the preceding turn. Supply a speakers map whose keys exactly match every label and whose values contain voice_id and a complete generation_profile. For example, MOM: [quiet, unsteady] I saw his face.\n(silence 4s)\nMOM: I know. compiles into two speech blocks with a timed-silence point on the first. Direction uses the selected model's supported control, such as an Inworld turn instruction or an ElevenLabs audio tag; models without a suitable whole-turn or mid-line control reject the direction. The route replaces the section's script, preserves its title and section settings, and returns the canonical document. A matching speaker name reuses and recasts that speaker across the project. It makes no generation or provider calls. Unmatched lines, missing speaker mappings, and unsupported directions fail with 400 rather than being dropped.
Speakers and casting#
A speaker is a character or narrator in a project. The speaker registry is shared by every section in the project and has its own speaker_registry_version. Each speaker has a voice_id and a generation profile that controls how its lines are performed. GET /studio/projects/{projectId}/speakers lists speakers with how often each one appears. Create Audio casts voices automatically; you can recast in the Studio app.
Preparing a speaker for ElevenLabs v4#
Casting and voice preparation are separate. A voice's default_model or completed ElevenLabs engine sample does not change a Studio speaker's saved generation_profile. There is currently no project-level default generation profile; when no speaker or block override selects a model, Studio uses inworld-2. Select provider: "elevenlabs" and model: "eleven_v4" in the speaker's casting update (or a block override) before checking its readiness.
- Read
GET /studio/sections/{sectionId}/generation-availabilityand find the speech or narration block. Itsgeneration.requested_profileidentifies the effective model.generation.statusisavailable,materialization_required, orunavailable;generation.optionsreports alternatives. - If ElevenLabs v4 is
materialization_required, prepare the assigned voice throughPOST /studio/projects/{projectId}/speakers/{speakerId}/voice-materialization-jobswithschema_version: 1, a newclient_mutation_id,materialization_kind: "provider_asset",provider: "elevenlabs", andmodel: "eleven_v4". PollGET /studio/voice-materialization-jobs/{jobId}until it completes, then read generation availability again. - If you also want to audition the voice,
POST /voices/{voiceId}/engine-sampleswithmodel: "eleven_v4"and the voice's currentexpected_acoustic_revisionsynthesizes the fixed comparison phrase. A successfully completed sample can publish a reusable ElevenLabs provider asset, so it may make Studio ready too. The sample's cached audio is not itself Studio's readiness authority: Studio requires an active provider asset for the voice's current acoustic revision. Always re-read generation availability before generating.
For a materializable voice, generation.reason_code and the matching option's reason_code distinguish provider_asset_missing from acoustic_revision_changed when a prior-revision provider asset exists. Editing or redesigning a voice can advance its acoustic revision; the old asset then cannot generate new Studio audio. An asset can also be retired or evicted. In either case, prepare the current revision and recheck. Other unavailable conditions use the human-readable reason and diagnostics; absence of reason_code does not mean the block is ready.
POST /ai-voices currently creates Inworld-designed voices. It has no ElevenLabs Voice Design selector. Preparing such a voice for ElevenLabs v4 changes its synthesis asset, not the design provider or the speaker's selected profile.
Versions and optimistic concurrency#
Studio never silently overwrites newer work. Every mutable resource carries a version:
| Version | Protects | Where to read it |
|---|---|---|
document_version | One section's document | GET /studio/sections/{sectionId}/document |
speaker_registry_version | The project's speakers and casting | The same document response |
section_order_version | The order of sections in a project | GET /studio/projects/{projectId}/sections |
mastering_version | Loudness and mastering settings | GET /studio/projects/{projectId}/mastering-profile |
When you change something, send the version you last read. If someone else changed it first, the request fails with 409 CONFLICT instead of overwriting their work. Read the resource again, reapply your change, and retry. Anything that plays or exports audio, such as the playback manifest and exports, is also pinned to these versions, so a stale version returns 409 rather than outdated audio.
Mutation IDs and actor sessions#
Studio write requests include two UUIDs in the JSON body:
actor_session_ididentifies the client session making changes. Generate one when your process starts and reuse it for all requests from that process.client_mutation_ididentifies one intended change. Generate a new one for each new change, and reuse it only when retrying the exact same request after a timeout or network error.
Replaying a client_mutation_id with the same body returns the original result. Some creation endpoints return 200 with idempotent_replay: true instead of 201. Reusing it with a different body returns 409. See Errors and retries.
Workflows#
A workflow is a durable, asynchronous job over one or more sections. Start one with POST /studio/projects/{projectId}/workflows and poll GET /studio/workflows/{workflowRunId}. There are two kinds:
workflow_kind | What it does | Cost | Sections per run |
|---|---|---|---|
prepare_sections | Structures text into narration and dialogue blocks. Use it to review the script before generating audio. | No audio is generated. | Up to 100 |
create_audio | Runs the full pipeline: prepare, assign speakers, cast voices, direct emotion and delivery, place sounds, and generate audio. | Consumes words. | Up to 12 with both sound options on, 14 with one, 16 with none |
A Create Audio run moves through these current_stage values: preparing_sections, assigning_speakers, casting_voices, planning_emotion, directing_delivery, placing_project_sounds, directing_sfx, preparing_voices, and generating_audio. Stages that are turned off are skipped.
A run's status is one of queued, running, canceling, completed, partially_completed, failed, or cancelled. The last four are terminal. Cancel a run with POST /studio/workflows/{workflowRunId}/cancel. Cancellation is cooperative, so work that already started may finish first. When resumable is true on a failed or partially completed run, POST /studio/workflows/{workflowRunId}/resume starts a new linked run that picks up from each section's first failed stage instead of redoing finished work. If the balance runs out, the audio step fails with error_code STUDIO_INSUFFICIENT_BALANCE; add words, then resume.
Create Audio has two independent sound options in schema_version 2:
include_generated_sfxadds AI-generated sound effects where the story calls for them.place_project_soundsplaces sounds from your sound library that are linked to the project.
Both default to true.
Media, playback, and exports#
- Media.
GET /studio/sections/{sectionId}/mediareturns one entry per block with itsfreshness(absent,current,stale, orunsupported), itsactivity(idle,generating, orfailed), and a short-lived signedplayback_urlfor the current audio. - Playback manifest.
GET /studio/sections/{sectionId}/playback-manifestcompiles the section's audio, gaps, and sound effects into one timeline for gapless playback. Pass thedocument_versionandspeaker_registry_versionyou read from the document. - Exports.
POST /studio/exportsrenders a section intomp3,wav,m4b, synchronizedepub, orsrt. PollGET /studio/exports/{exportId}and downloadartifact.download_urlwhenstatusiscompleted.
Signed URLs expire. Fetch fresh ones from the API instead of storing them.
Sharing#
Publish a section with POST /studio/sections/{sectionId}/share or a whole project with POST /studio/projects/{projectId}/share. Each returns a stable public link that plays the current audio. DELETE on the same path revokes it. See the Sharing API.