Audio examples & methodology

These original 16-second fixtures were created with text-to-speech narration of original text and deterministic instrument synthesis. They contain no commercial song samples, customer uploads or recordings of a real person. Text-to-speech narration, drums, bass, piano and guitar parts are mixed and processed through the same model and task queue used by the website.

REAL PIPELINE OUTPUT

Listen before you upload

Original text-to-speech narration and instruments, processed by AIstemify. This is not a human-singing benchmark; results on real recordings will vary.

In this six-stem fixture, piano is nearly silent and drums and guitar are weak; synthesized instruments may be assigned to other stems.

Original mix00:16 · WAV
Vocals00:16 · MP3
Instrumental00:16 · MP3
Sources & test methodology

How to interpret these examples

Inputs are 22,050 Hz, 16-bit WAV; outputs are MP3 files from actual inference. The vocal model returns vocals and instrumental, the six-stem model returns six parts, and the dereverb model returns reduced reverb and residual. Outputs have not been replaced with the known source stems or manually repaired.

The displayed elapsed time includes queueing, polling and some download overhead; it is not isolated GPU time or a speed guarantee. Synthetic signals differ from human recordings, so these examples do not establish best-in-class quality or accuracy. Evaluate loud and quiet passages in your own recording for bleed, transients and remaining reverb.

The public fixtures were created by AIstemify for demonstration and reproducibility. Attribution uses the AIstemify brand, without personal identity details.