ACE-Step XL 1.5 Premium v5.3 Written Tutorial

This README is the full Markdown tutorial generated from the public YouTube tutorial, local file ACESTEP_XL_15_Full_Tutorial.mp4, captions (2).srt, live local ACE-Step app screenshots, v5.3 Wildcards screenshots, one newly generated demo song, one Audio Processing run made from that song, and selected 4K frames extracted from the source video.

Generated demo song: G:\ACE_Step_v1\ACE-Step_Premium\outputs\0021\8642a822-c4d0-ff4a-531f-6d5cfb334381.mp3

Processed demo output: G:\ACE_Step_v1\ACE-Step_Premium\outputs\audio_processing_0004\8642a822-c4d0-ff4a-531f-6d5cfb334381_processed.wav

Source Video, Thumbnail, Links, And Chapters

The public video description presents this as a full ACE-Step XL 1.5 Premium guide for local AI music generation, remix, repaint, stem extraction, wildcard prompt variation, audio processing, SAM Audio segmentation, Windows installation, RunPod, Massed Compute, SimplePod, and Linux/cloud workflows.

Tutorial video thumbnail
Tutorial video thumbnail

Video Chapters

1. What ACE-Step XL 1.5 Premium Is

ACE-Step XL 1.5 Premium is a local-first music generation and audio utility suite. The video presents it as more than a song generator: it also includes Wildcards for prompt variation, advanced generation modes, remix/repaint workflows, stem extraction, LEGO-style stem addition, SAM Audio segmentation, Auto-Editor trimming, mastering-style audio processing, dataset tools, and LoRA/LoKr training pages.

Responsible-use note: the source tutorial says to use the application respectfully and for research/education. For remix, repaint, extraction, and pitch work, use material you own, have permission to process, or are otherwise allowed to use.

Video introduction
Video introduction
Feature overview
Feature overview

Core jobs covered in the tutorial:

2. Install And Start On Windows

The Windows workflow uses the included batch files. Extract the ZIP, keep the folder structure intact, run the installer/update script, optionally download all models, then start the app with the Windows launcher.

Windows installer
Windows installer
  1. Extract the ACE-Step Premium ZIP to a path with enough free disk space for the virtual environment, model files, outputs, and FFmpeg runtime.
  2. Run Windows_Install_or_Update.bat. The installer creates the Python virtual environment, downloads or uses shared FFmpeg, installs packages with UV, and prepares the app.
  3. Run Windows_Download_All_Models.bat if you want SFT and Base models in addition to the automatically available Turbo path.
  4. Run Windows_Start_App.bat. In this workspace the launcher started ACE-Step at http://127.0.0.1:7862 because other Gradio apps were already using 7860 and 7861.
  5. Watch the command window for model download, model load, generation, and error details. The video recommends trusting the terminal status more than only the browser UI.

Model availability: Turbo is the quick default. SFT and Base require additional model files. Remix is recommended with SFT in the video; some modes are marked Base-only or unavailable until the matching model is selected.

Windows first generation
Windows first generation

3. Quick Song Generation

The Generate Song tab is the fast path. It exposes the controls most users need: style, lyrics, Wildcards, model, LoRA, GPU preset, quantization, language, vocal type, instrumental toggle, duration, count, seed, optional MP4 image, and video resolution.

Generate Song overview
Generate Song overview
Generate Song filled
Generate Song filled

Wildcards For Prompt Variation

ACE-Step XL 1.5 Premium v5.3 adds Wildcards. Write bracketed choices separated by pipes, such as [option A|option B|option C], and one option is picked when you generate. Wildcards can be used in the quick Generate Song Style field, the Advanced Music Caption field, and Lyrics.

Wildcards in Generate Song
Wildcards in Generate Song
Wildcards in Advanced caption and lyrics
Wildcards in Advanced caption and lyrics
  1. Write a concise Style prompt that describes genre, vocal character, instrumentation, production quality, tempo or mood, and mix target.
  2. Write Lyrics with section tags such as [Verse] and [Chorus]. The included ACE_Step_Lyric_Generation_Instructions_For_LLMs.txt file can be given to an LLM to format lyrics or style prompts.
  3. Optionally add Wildcards to Style or Lyrics when you want the app to choose between prompt variants automatically.
  4. Select the Model. Start with ACE-Step XL 1.5 Turbo to verify the machine and workflow quickly.
  5. Leave GPU Optimization Preset and DiT Quantization at safe defaults unless you are solving VRAM pressure or repeating a known workflow.
  6. Set Song Duration and Songs. The demo run used 20 seconds and 1 song.
  7. Use Random Seed while exploring. When a promising result appears, uncheck Random Seed and keep the seed so future edits stay comparable.
  8. Click Generate Song and monitor the Status field plus the terminal window.
Demo generation result
Demo generation result

Useful quick-tab buttons:

4. Results, Seeds, And Reuse

The tutorial stresses generating repeatedly until you have a good base result, then locking the seed and making controlled edits. This is especially important for remix and repaint work, where small prompt or range changes can be tested against the same underlying random state.

Seed and remix discussion
Seed and remix discussion
Results after generation
Results after generation

Seed workflow:

  1. Keep Random Seed on while searching for a usable base result.
  2. When the result is close, copy or keep the seed shown by the UI.
  3. Turn Random Seed off.
  4. Change one word, one range, or one strength setting at a time.
  5. Compare outputs against the locked seed.

5. Advanced Generation Modes

The ACESTEP Advanced tab is the full workstation. It exposes generation mode, runtime settings, source/reference audio, LM code utilities, advanced prompts, Wildcards in Music Caption/Lyrics, metadata, sampler settings, output settings, and batch processing.

Advanced overview
Advanced overview

Generation modes:

Advanced source audio
Advanced source audio
Advanced generation controls
Advanced generation controls

Important advanced controls:

Engine settings
Engine settings

Engine settings include GPU tier, checkpoint file, main model path, device, VAE, 5Hz LM model/backend, Flash Attention, CPU offload, compile, DiT quantization, LoRA path/folder, LoRA scale, inference steps, sampler, DCW, ADG, MP3 bitrate/sample rate, normalization, fades, LM temperature, top-k/top-p, negative prompt, and LM code settings. Leave these at defaults until you have verified a basic generation.

6. Remix, Repaint, Extract, LEGO, And Auto-Editor Features

The first part of the video demonstrates feature outcomes before the installation section. These are not separate apps; they are modes and panels inside the same ACE-Step interface.

Remix demo
Remix demo
Extract and LEGO demo
Extract and LEGO demo
Auto-Editor demo
Auto-Editor demo

7. Audio Processing

Audio Processing is used on uploaded or local audio/video and can also be applied automatically to generated songs. It includes format output, Auto-Editor trimming, video re-encode controls, audio enhancement stages, pre-mastering stages, DiffPitcher, and batch folder processing.

Audio Processing overview
Audio Processing overview
Generated song loaded for processing
Generated song loaded for processing
Audio Processing result
Audio Processing result

Core Audio Processing controls:

Audio Enhancement and Pre-Mastering
Audio Enhancement and Pre-Mastering

Audio Enhancement includes Stereo Depth, Stereo Width, HF Refinement, Harmonic Enrichment, Timing Humanizer, and Ambience Shaping. Pre-Mastering includes Multiband Compressor, Tape Saturation, Glue Compressor, Mid/Side EQ, Soft Clipper, and LUFS Normalization.

DiffPitcher controls
DiffPitcher controls

DiffPitcher is for isolated vocals that sing the wrong notes. Use a guide vocal or MIDI score for the same phrase. The tutorial text in the UI warns that this is not for copying another singer or another song.

8. SAM Audio Segment

SAM Audio Segment is a heavier but more flexible segmentation system. It can extract target audio from a prompt, save the residual/remaining audio, process video inputs, use explicit span anchors, and run batch prompt lists separated by semicolons.

SAM Audio source-video demo
SAM Audio source-video demo
SAM Audio overview
SAM Audio overview
SAM prompt runtime controls
SAM prompt runtime controls
  1. Upload an audio or video file. Optionally upload a visual mask video for video-guided workflows.
  2. Choose Mode and Quick Prompt, or type a Custom Prompt such as vocals, guitar, bass, drums, applause, or another target.
  3. Enable Batch Segment when you want several prompts in one run; separate prompts with semicolons.
  4. Use Predict spans when you want SAM Audio to estimate target time ranges from text.
  5. Use explicit span anchor only when you can provide positive/negative time anchors as JSON.
  6. Choose a VRAM preset and candidate count that match the GPU. Higher candidate counts can improve quality but cost runtime and VRAM.
  7. Enable Save remaining audio when you need both the extracted target and the residual track.

9. Library, Metadata, Presets, Dataset, And Training Pages

The remaining app tabs are operational pages. They help you find previous generations, restore metadata, manage presets, inspect datasets, and train adapters.

Library
Library
Load Metadata
Load Metadata
Custom Preset System
Custom Preset System
Dataset browser
Dataset browser
LoRA Dataset Builder
LoRA Dataset Builder
Train LoRA
Train LoRA

10. RunPod Deployment

The RunPod chapter focuses on persistent network storage, GPU/region selection, unreliable installs, Gradio live URLs, nvitop monitoring, output downloads, and safe termination.

RunPod storage
RunPod storage
RunPod install
RunPod install
RunPod Gradio services
RunPod Gradio services
RunPod monitoring
RunPod monitoring
  1. Create persistent network storage in the same region as the GPU you intend to rent.
  2. Deploy the pod/template with the storage mounted. Choose a GPU with enough VRAM for the selected model and quality target.
  3. Run the installer. If RunPod throws an OS/server error, run the installer again; it should resume from completed work.
  4. If installation stalls from excessive parallelism, delete the virtual environment, lower installer thread count as shown in the video, and rerun.
  5. Start the app and prefer the Gradio live link when the RunPod proxy is unreliable. If port 7860 does not open, try the port shown by the terminal, sometimes 7861.
  6. Use nvitop to monitor GPU memory and load. First model load can be slow on RunPod storage; later generations are faster.
  7. Download outputs from JupyterLab by right-clicking the outputs folder and downloading it as an archive.
  8. Stop or terminate the pod deliberately. Delete storage too if you no longer want monthly storage charges.

11. Massed Compute Deployment

The Massed Compute chapter is similar to the Linux/cloud workflow, but the tutorial emphasizes faster disk performance and lower friction compared with RunPod. The tradeoff called out in the video is the lack of the same persistent network storage flow.

Massed Compute GPU selection
Massed Compute GPU selection
Massed Compute install
Massed Compute install
  1. Choose the creator category and the SECourses image when following the video workflow.
  2. Select a GPU appropriate for ACE-Step XL 1.5. The tutorial mentions RTX Pro 6000 and RTX 5090 class GPUs.
  3. Upload the ACE-Step ZIP to Downloads, extract it, open Massed_Compute_Instructions_READ.txt, and copy the install command.
  4. Open a terminal inside the extracted ACE-Step folder and run the command from that location.
  5. Start ACE-Step and use the Gradio live URL. If Gradio live shows a transient error, refresh the page.
  6. Back up large outputs or model/data folders to Hugging Face, Google Drive, OneDrive, or another storage service if you need to recreate the machine later.

12. SimplePod Deployment

The SimplePod chapter uses the RunPod/SimplePod instruction file and shows a persistent-storage flow that resembles RunPod. The tutorial demonstrates starting, generating, monitoring, stopping, and resuming from the same storage volume.

SimplePod setup
SimplePod setup
SimplePod generation
SimplePod generation
SimplePod resume
SimplePod resume
  1. Register, add credits, and create/use persistent storage as shown in the instruction file.
  2. Open the template link, attach the storage volume, choose a GPU, and run the machine.
  3. Use the JupyterLab or console link to run the installer/start commands from the workspace.
  4. If the Gradio live page throws a first-click error, refresh or click again after the page is fully loaded.
  5. Install nvitop when you want GPU/VRAM visibility: pip install nvitop, then run nvitop.
  6. To resume, reuse the template link, attach the same volume, select a GPU, start the machine, and run the app start command again.
  7. Stop or terminate compute and remove storage when finished to avoid unwanted billing.

13. Troubleshooting And Best Practices