On this page
Kade-AI can make audio — not just read text out loud, but generate a whole scene: characters talking, sound effects, music, and atmosphere, all from a few sentences you write.
Meet Cadence, your audio director
Cadence is the character built for this. Find Cadence in the marketplace (the Characters page shows you how), describe the moment you want to hear, and Cadence turns it into a real, playable clip. It plays right there in the chat, and it's saved to your My Creations page so you can come back to it.
What you can ask for
- A whole scene: "A 90-second suspense radio drama in a late-night diner — rain outside, two nervous voices, a door slams at the end."
- Narration or a voiceover: "Read this in a warm movie-trailer voice…"
- Your own script: paste a few lines of dialogue and let Cadence cast it and perform it with sound.
- Copy a voice: give it a short, clean recording and it can speak new lines in that voice.
- Fix or extend a clip: make an ending longer, swap a line, or stitch two clips into one.
Meet Muse, for actual songs
Muse is a different character for a different job: turning words into a real song — actual singing, over real music. Where Cadence builds a scene, Muse cuts a track. Find Muse in the marketplace (the Characters page shows you how), hand over your lyrics and a vibe, and Muse sings them. It plays right there in the chat and saves to your My Creations page, same as everything else.
- From your own lyrics: paste the words and describe the feel — "a slow soul ballad, warm and aching, about 70 beats a minute" — and Muse records it.
- Straight from Lyric: if the songwriter character Lyric wrote you a song, hand Muse what Lyric gave you and it turns it into the finished track.
- From just a feeling: no lyrics yet? Describe the mood and Muse can write and sing its own.
- An instrumental: ask for a backing track with no singing.
Muse can sing with Google's Lyria 3 Pro (rich, up to about three minutes) or MiniMax — if the first take isn't quite right, just ask for the other. A song takes a minute or two to come back, and costs about 8 to 15 cents, from the same pot as pictures and audio.
The catch: it costs a little (not much)
Making audio pulls from the same small pot of credits as pictures and video. It's cheap — about 19 cents a minute, so a full two-minute scene runs under 40 cents — but the pot isn't bottomless. Make all the audio you like; just don't fire off a thousand at once. The spend shows up on your Usage & Balance page. If audio ever stops working, the pot may need a top-up — contact Kade.
Good to know
- Each Seed Audio clip is up to about 2 minutes long. Want something longer? Ask for it in parts and Cadence can stitch them together — or use a longer narration (below) within the Booth's script limits.
- Seed Audio speaks twenty languages; ask Cadence to name them if you need one in particular.
Stable Audio: sound effects and ambience
Choose Stable Audio in the Sound Booth for environmental sound, foley and layered ambience. Open Sound settings and use Sound model to choose 3 Medium, the new default, or 3 Small SFX, the original option. Medium is the larger model and our recommended starting point for detailed ambience. Both run through fal and save stereo WAV files; width and layer accuracy vary between recordings. Describe the foreground sound, quieter background layers and the space around them. Try the rain-cabin, harbor or forest-stream starting points. Ask for no music or speech when you want only environmental sound. This produces one mixed recording, not separate stems.
Enter an optional Track title, choose 1 to 120 seconds and 1 to 4 takes, then press Generate sounds once. Each take uses a different seed. All takes are submitted together; fal controls how many run at once. Provider cost is 3.76 cents per Medium recording, or 15.04 cents for four. Small SFX remains 2.06 cents per recording, or 8.24 cents for four. Prices verified September 18, 2026. During this trial, Kade pays that cost and nothing is deducted from your credit balance. The price appears before you generate; there is no additional confirmation screen. Reopening a saved project restores the model used for it, including older Small SFX projects.
Finished recordings are kept as lossless WAV in your Sound Booth library and My Creations. You can leave the screen while a batch runs. One completion notification opens the Sound Booth, subject to your notification settings. Finished takes survive another take failing; Stop requests cancellation but a take already running may still finish and incur a provider charge. Rename tracks in the library, or reopen a description to make another batch.
Inference steps default to 8 for both distilled models; increasing them takes longer and does not guarantee better sound. An optional seed lets you reuse a starting point. This SFX engine does not offer voice cloning, cover-song controls or a guidance slider. Seed Audio remains available for scenes with dialogue and layered soundscapes. Reopen Sound Booth to load the new model picker; the existing iPhone build supports it.
AuK HQ: speech and recording edits
On the website and iPhone build 300 or later, start in the main writing box. Write a script from this turns your idea into editable text using the writing model. For music, the button shapes your music idea; YuE2 can draft the lyrics too. Surprise me gives you a starting idea: for songs the writer gives you one idea the way a songwriter would jot it down, a genre and one very specific human situation, sometimes with a rule such as never say the word sorry; it takes about ten seconds, costs a fraction of a cent, and remembers what it has already shown you so it does not repeat itself, and for speech and scenes it rolls a free one on the page. If the writer cannot be reached, you get a free idea from the short list instead; Undo writing change restores what was there. None of these buttons starts an audio recording. Voice, references and detailed settings remain in expandable sections. The separate writing desk still offers exact-word formatting.
YuE2 is a separate music option alongside Lyria. Describe the style and add Lyrics in song settings, or use Write my song idea. Its full-quality model runs on sleeping RunPod GPUs. YuE2 can use up to two GPUs at once; AuK keeps its own worker. Each GPU stays awake for ten minutes after its last job, then scales to zero. Startup, generation and awake idle time are billed, at about $1.22 per GPU hour on the selected cards. Ten idle minutes is about twenty cents per GPU; there is no monthly GPU reservation or new persistent storage rental. The library's YuE2 cost is an execution estimate; startup and idle add to Kade's provider bill. YuE2 currently deducts nothing from user credit balances. Lyria records and deducts its per-song charge from existing non-admin balances; admin accounts are exempt. A requested song length is guidance, not an exact duration.
To make a YuE2 cover, import a source recording in song settings, describe the new style, and add the words to sing in Lyrics. Or choose Cover this take on a saved Lyria or YuE2 song. The original stays in your library. Recordings can be WAV, MP3, M4A or OGG, up to twenty megabytes and six minutes. The worker transcribes the melody and then makes a new arrangement on the same GPU. Melody transcription can make mistakes and does not clone the singer. After importing, choose Transcribe reference lyrics to have Gemini listen to the recording and draft the sung words in lyric lines, with ElevenLabs Scribe as a fallback. It preserves repeated choruses and does not use your saved lyrics to fill gaps. Correct wrong or missing words in Lyrics and add verse or chorus tags as needed. Undo writing change restores the previous lyrics. This button does not generate music or deduct credits; completed transcripts from the current transcriber are reused. Older Deepgram drafts are replaced the next time you press Transcribe reference lyrics. Your saved writing changes only when you choose that button. You can also supply an ABC melody score instead of a recording. Voice cloning, LoRAs and saved singer personas are not offered by this integration. A take that reaches a model limit is marked as possibly ending early. The website also offers a composition-score download.
Sound Booth lyric drafts now follow Lyric: the music writing desk reads the Lyric agent's current saved persona for each new draft. Its craft guidance covers connected thoughts, believable voices, internal and multisyllabic rhyme when appropriate, and a private edit for forced rhymes and stock phrasing. The desk now writes with Kimi K3 and always thinks the song through before writing, even for a one-sentence idea. It now writes by Kade's own hit-writing system: one moment, one voice with an attitude, the hook built first, open vowels under the big notes, a planted detail that pays off, and variety where a machine would be uniform. Unless you ask for another length it writes a song of about four minutes with three verses. The writer now takes its time: a song draft runs in the background for about five minutes while the writer thinks the song through, then goes back over it once like a producer deciding whether the song gets cut. You do not have to wait on the page. You press Help write this once. Stay and the draft appears in the editor by itself; leave, and a notice arrives when it is ready, and the Sound Booth page picks the draft up on its own the next time you open it. If you are in a hurry, tick Quick song draft on the website for a draft in about a minute and a half, less polished. iPhone builds before 303 still use the quick draft. It avoids the words and props that make a lyric sound machine-made, such as shadows, whispers and neon, but also the everyday ones: naming a weekday, coffee, the porch light, the kitchen table, two a.m. After writing, the desk checks the draft for those and has the writer replace just those lines, keeping everything else. If you want one of those words, put it in your idea and it is yours. To ask for a shorter song, say so in your idea. It keeps the Sound Booth output format. If the writer cannot finish, your idea is kept and you can try again. Existing lyrics supplied to the desk stay unchanged. To replace a draft you dislike, clear the Lyrics box before choosing Write my song idea. Drafting new lyrics does not generate a recording. Prompt guidance improves the target, but cannot guarantee every line will land.
The Sound Booth script and music-brief writers, Lyric, Cole, and the character-draft writer share writing guidance to reduce canned AI phrasing. It asks for concrete detail, natural rhythm and a private revision check, while protecting deliberate choruses, rhyme, dialogue and required formatting. Formatting-only requests retain your supplied words. This improves consistency but does not guarantee every draft will be free of cliches. Writing a draft does not automatically order a recording.
Sound Booth uses Kimi K3 from Moonshot AI, reached through OpenRouter, for new music and lyric drafts. It was chosen by comparing drafts from several writers on the same song ideas, and it writes explicit lyrics when you ask for them. A draft costs roughly 2 to 3 cents in provider charges, up from about 1 cent. Speech and scene writing, plus exact-word formatting, use Hermes 4 405B. Lyric, Cole and the character builder keep their own models. Sound Booth records the provider's writing cost when available, otherwise a token-based estimate. This choice applies to the text writer; AuK, Seed Audio and Lyria still make the recordings with their own capabilities and provider rules. Your existing drafts are not automatically rewritten.
The Sound Booth uses AuK Base for speech and editing, Seed Audio for scenes and backgrounds, and Lyria for songs. Open Tools, Sound Booth in the native app, or the web Sound Booth. Each engine retains its own draft.
For speech without a reference clip, describe the voice's accent, texture, mood and delivery, then write only the words to say in the performance script. Later sections reuse the opening voice. An imported reference supplies its own voice and accent; changing the description does not change that reference. The voice wheel supplies a description when no clip is attached; it does not copy the wheel's original speaker.
To try giving a reference voice another accent, choose Edit, import the original recording, and describe the accent change while asking to preserve the words and speaker identity. Listen to the new take, then choose Use this voice for your script. Adding an arbitrary accent is experimental; AuK documents removing a regional accent. The booth does not automatically run or charge for a separate accent-edit pass.
Edit an existing recording
- Choose edit under Task and import the recording.
- In Edit instructions, describe the change: make the delivery cheerful, change a word, whisper, raise pitch, remove noise or reverb, or separate a voice from the background. Say what should stay the same.
- Leave Target seconds blank for the original duration, or set it when changing speed or adding or removing words.
- Read the cost information, then choose Edit recording. On the website AuK starts with one press. The result is a new take; your source and previous takes remain available.
Long work is processed in sections and joined. Listen back for seams, missed word edits and voice changes. Accent instructions are requests, not a guarantee. AuK does not automatically retry pronunciation errors. Each take retains a 24 kHz mono WAV master and a listening MP3.
Sleep and cost
The rented GPU sleeps between jobs. Startup, processing and brief idle time are billable; availability and cold starts can delay a job. AuK currently has no reliable total price estimate. Its GPU rate is up to $1.22 per hour of active GPU time, not per hour of finished audio. Progress and Stop remain available. If the connection drops, reopen the existing project before starting another take. Stopping or failure does not prove that no GPU time was charged.
Generation now starts with one press on the website and in the updated iPhone app. The cost or no-estimate notice remains beside Generate. There is no extra render confirmation. Wait for imports or lyric transcription to finish first. If an import fails, retry it or choose Discard failed import. Completing an upload or transcription does not generate music automatically.
Title and compare your music: enter an optional Track title before generating, or rename saved work in the library. YuE2 offers Number of takes from one to four. Takes use different seeds; up to two can generate in parallel when GPUs are available. You may leave the screen while they run. One notification opens the Sound Booth after the whole batch finishes, and finished takes remain available if another take fails. Quiet hours and notification mutes still apply.
YuE2 advanced sound settings: Creative variation, or weirdness, changes the model's sampling temperature. Its default of 50 preserves the original settings; higher values explore less likely choices and can reduce coherence. Prompt guidance defaults to 1 and can increase to 3; higher values emphasize the conditions but add work and can reduce naturalness. Inference steps range from 16 to 64, with the original default of 32. More steps refine the audio longer but do not guarantee a better song. These controls are specific to YuE2. For an exact starting point, keep the seed, lyrics, direction and settings; a batch advances the seed for each take.
Failed attempts are visible by default in the web library, with their saved error. A render that fails while the booth is open shows a Generation stopped notice. Completion and failure pushes are requested results, so they do not use the ordinary agent-message allowance. Notification mutes still apply, and quiet-hour results wait for morning.
Use Seed Audio for newly generated music, sound effects and background scenes. AuK can clean or separate an existing background.
Lyria: a song from a brief, written the way Google reads it (September 11, 2026)
Lyria is Google's music engine, about eight cents a song however long it comes out, usually back in under a minute. It does not read speech and it cannot copy a particular voice; it makes a record. The booth's script desk and its how-to-write notes now follow Google's own prompt guide, in this order: the genre with an era ("1970s Memphis soul"), the instruments and what each one does, the shape as section tags ([Intro] → [Verse 1] → [Chorus] …) or timestamps, the voice if anybody sings (sex, timbre, range, delivery), the mood in a few plain adjectives, and last a technical line with a BPM number, the key, and how long ("around 70 BPM, in D minor, a two-minute song"). It reads the length from your words, so always say how long.
Your own words go in the Your own lyrics box, and the booth sends them under a Lyrics: heading the way the engine expects — put [Verse 1], [Chorus] and [Bridge] on their own lines above each part if you want to steer the shape. Turn on No singing for an instrumental. The words it writes come back cleaned of its own machine markers, so the read-back your screen reader hears is the song and nothing else; the raw text is kept with the recording in My Creations.