Skip to content

Guides

Studio concepts: projects, sections, documents, and workflows

How LyricWinter Studio models stories as projects, sections, versioned documents, speakers, and asynchronous workflows, and how to edit them safely over the API.

BetaUpdated

LyricWinter Studio models a story as a project made of ordered sections. Each section has a versioned document of narration, dialogue, and sound-effect blocks, each spoken line belongs to a project-wide speaker, and long-running work such as generating audio happens in asynchronous workflows. This page explains each concept and how to change them safely through the LyricWinter API.

Projects and unassigned stories#

A project is a named container for related sections, such as the chapters of a book. All sections in a project share one cast, so a character keeps the same voice from chapter to chapter.

A standalone story is a single section that is not yet part of a named project. POST /studio/sections creates one. GET /studio/projects returns named projects in projects and standalone stories in unassigned_sections. A standalone story still has a project_id, and every project-scoped endpoint works with it. To turn a standalone story into a named project, call PATCH /studio/projects/{projectId} with a name.

Sections#

A section is one chapter, scene, or episode of story text. Sections in a project are ordered. Add one with POST /studio/projects/{projectId}/sections, move one with PATCH /studio/projects/{projectId}/sections/{sectionId}, and delete one with DELETE. Deletions cannot be undone through the API, so call the matching deletion-info endpoint first to see what will be removed.

Documents and blocks#

Every section has a document: the canonical, structured script of the section. Read it with GET /studio/sections/{sectionId}/document. A document contains:

  • document_version, which increases on every change.
  • speakers, the project's current speaker registry, and speaker_registry_version.
  • blocks, the script in order. Each block has a kind:
Block kindWhat it is
narrationNarrated prose, voiced by the narrator.
speechA line of dialogue attributed to a speaker.
sfxA sound effect placed in the timeline.
noteA note kept in the script but not voiced.
excludedSource text kept in the script but intentionally left out of the audio.
unparsedText that has not been structured into narration or dialogue yet. It is not voiced until it is.

A new section starts as unstructured text. A prepare_sections or create_audio workflow turns it into narration and speech blocks.

For a finished speaker-labeled script, POST /studio/sections/{sectionId}/script creates speech blocks directly, without a prepare workflow or a hand-built AST. Use one SPEAKER: spoken text line per turn. A [direction] tag at the start of a line steers the whole turn, and a tag mid-line, as in TORVALD: [softly] Go, Ingvild. [whispers] Please., anchors a direction where it appears; both become tracked directions and are removed from the spoken text. A separate (silence 4s) or (silence 4000ms) line inserts exact silence after the preceding turn. Supply a speakers map whose keys exactly match every label and whose values contain voice_id and a complete generation_profile. For example, MOM: [quiet, unsteady] I saw his face.\n(silence 4s)\nMOM: I know. compiles into two speech blocks with a timed-silence point on the first. Direction uses the selected model's supported control, such as an Inworld turn instruction or an ElevenLabs audio tag; models without a suitable whole-turn or mid-line control reject the direction. The route replaces the section's script, preserves its title and section settings, and returns the canonical document. A matching speaker name reuses and recasts that speaker across the project. It makes no generation or provider calls. Unmatched lines, missing speaker mappings, and unsupported directions fail with 400 rather than being dropped.

Speakers and casting#

A speaker is a character or narrator in a project. The speaker registry is shared by every section in the project and has its own speaker_registry_version. Each speaker has a voice_id and a generation profile that controls how its lines are performed. GET /studio/projects/{projectId}/speakers lists speakers with how often each one appears. Create Audio casts voices automatically; you can recast in the Studio app.

Preparing a speaker for ElevenLabs v4#

Casting and voice preparation are separate. A voice's default_model or completed ElevenLabs engine sample does not change a Studio speaker's saved generation_profile. There is currently no project-level default generation profile; when no speaker or block override selects a model, Studio uses inworld-2. Select provider: "elevenlabs" and model: "eleven_v4" in the speaker's casting update (or a block override) before checking its readiness.

  1. Read GET /studio/sections/{sectionId}/generation-availability and find the speech or narration block. Its generation.requested_profile identifies the effective model. generation.status is available, materialization_required, or unavailable; generation.options reports alternatives.
  2. If ElevenLabs v4 is materialization_required, prepare the assigned voice through POST /studio/projects/{projectId}/speakers/{speakerId}/voice-materialization-jobs with schema_version: 1, a new client_mutation_id, materialization_kind: "provider_asset", provider: "elevenlabs", and model: "eleven_v4". Poll GET /studio/voice-materialization-jobs/{jobId} until it completes, then read generation availability again.
  3. If you also want to audition the voice, POST /voices/{voiceId}/engine-samples with model: "eleven_v4" and the voice's current expected_acoustic_revision synthesizes the fixed comparison phrase. A successfully completed sample can publish a reusable ElevenLabs provider asset, so it may make Studio ready too. The sample's cached audio is not itself Studio's readiness authority: Studio requires an active provider asset for the voice's current acoustic revision. Always re-read generation availability before generating.

For a materializable voice, generation.reason_code and the matching option's reason_code distinguish provider_asset_missing from acoustic_revision_changed when a prior-revision provider asset exists. Editing or redesigning a voice can advance its acoustic revision; the old asset then cannot generate new Studio audio. An asset can also be retired or evicted. In either case, prepare the current revision and recheck. Other unavailable conditions use the human-readable reason and diagnostics; absence of reason_code does not mean the block is ready.

POST /ai-voices currently creates Inworld-designed voices. It has no ElevenLabs Voice Design selector. Preparing such a voice for ElevenLabs v4 changes its synthesis asset, not the design provider or the speaker's selected profile.

Versions and optimistic concurrency#

Studio never silently overwrites newer work. Every mutable resource carries a version:

VersionProtectsWhere to read it
document_versionOne section's documentGET /studio/sections/{sectionId}/document
speaker_registry_versionThe project's speakers and castingThe same document response
section_order_versionThe order of sections in a projectGET /studio/projects/{projectId}/sections
mastering_versionLoudness and mastering settingsGET /studio/projects/{projectId}/mastering-profile

When you change something, send the version you last read. If someone else changed it first, the request fails with 409 CONFLICT instead of overwriting their work. Read the resource again, reapply your change, and retry. Anything that plays or exports audio, such as the playback manifest and exports, is also pinned to these versions, so a stale version returns 409 rather than outdated audio.

Mutation IDs and actor sessions#

Studio write requests include two UUIDs in the JSON body:

  • actor_session_id identifies the client session making changes. Generate one when your process starts and reuse it for all requests from that process.
  • client_mutation_id identifies one intended change. Generate a new one for each new change, and reuse it only when retrying the exact same request after a timeout or network error.

Replaying a client_mutation_id with the same body returns the original result. Some creation endpoints return 200 with idempotent_replay: true instead of 201. Reusing it with a different body returns 409. See Errors and retries.

Workflows#

A workflow is a durable, asynchronous job over one or more sections. Start one with POST /studio/projects/{projectId}/workflows and poll GET /studio/workflows/{workflowRunId}. There are two kinds:

workflow_kindWhat it doesCostSections per run
prepare_sectionsStructures text into narration and dialogue blocks. Use it to review the script before generating audio.No audio is generated.Up to 100
create_audioRuns the full pipeline: prepare, assign speakers, cast voices, direct emotion and delivery, place sounds, and generate audio.Consumes words.Up to 12 with both sound options on, 14 with one, 16 with none

A Create Audio run moves through these current_stage values: preparing_sections, assigning_speakers, casting_voices, planning_emotion, directing_delivery, placing_project_sounds, directing_sfx, preparing_voices, and generating_audio. Stages that are turned off are skipped.

A run's status is one of queued, running, canceling, completed, partially_completed, failed, or cancelled. The last four are terminal. Cancel a run with POST /studio/workflows/{workflowRunId}/cancel. Cancellation is cooperative, so work that already started may finish first. When resumable is true on a failed or partially completed run, POST /studio/workflows/{workflowRunId}/resume starts a new linked run that picks up from each section's first failed stage instead of redoing finished work. If the balance runs out, the audio step fails with error_code STUDIO_INSUFFICIENT_BALANCE; add words, then resume.

Create Audio has two independent sound options in schema_version 2:

  • include_generated_sfx adds AI-generated sound effects where the story calls for them.
  • place_project_sounds places sounds from your sound library that are linked to the project.

Both default to true.

Media, playback, and exports#

  • Media. GET /studio/sections/{sectionId}/media returns one entry per block with its freshness (absent, current, stale, or unsupported), its activity (idle, generating, or failed), and a short-lived signed playback_url for the current audio.
  • Playback manifest. GET /studio/sections/{sectionId}/playback-manifest compiles the section's audio, gaps, and sound effects into one timeline for gapless playback. Pass the document_version and speaker_registry_version you read from the document.
  • Exports. POST /studio/exports renders a section into mp3, wav, m4b, synchronized epub, or srt. Poll GET /studio/exports/{exportId} and download artifact.download_url when status is completed.

Signed URLs expire. Fetch fresh ones from the API instead of storing them.

Sharing#

Publish a section with POST /studio/sections/{sectionId}/share or a whole project with POST /studio/projects/{projectId}/share. Each returns a stable public link that plays the current audio. DELETE on the same path revokes it. See the Sharing API.