Audio Tutorial

Getting Started: Converting Audio Files

Learn how to convert, clean, and remix audio files with Media Manipulator's audio converter. This tutorial walks through every section of the audio conversion panel — format selection, processing effects, restoration, and the new AI audio tools (Clean Voice, Remove Background Noise, Isolate Vocals, Remove Vocals).

1. Upload your audio file

From the home page, drag and drop an audio file into the upload zone, or click Choose File. Supported inputs include MP3, WAV, FLAC, AAC, M4A, OGG, Opus, and most common podcast and music formats. Files are deleted automatically within 24 hours.

Once the file is selected, the right-hand panel switches to Conversion Options. For video and audio files you'll also see a Convert / Transcribe toggle — leave it on Convert for this tutorial.

Loading the tool…

2. Pick the output format

The Format dropdown sets the container and codec for your output file. The audio converter supports:

  • MP3 — universal compatibility, lossy, small file sizes.
  • WAV — uncompressed PCM, ideal for masters, editing, or archiving.
  • AAC — efficient lossy codec used by Apple devices and streaming.
  • OGG — open-source Vorbis container.
  • FLAC / ALAC — lossless compression for archival quality at smaller size than WAV.
  • Opus — modern lossy codec great for voice and low bitrates.
  • AC3 / DTS — surround-capable formats often used for video soundtracks.

3. Set bitrate, sample rate, and channels

Bitrate controls quality vs. file size for lossy formats. 128 kbps is "good enough" for podcasts; 192–256 kbps is a strong default for music; 320 kbps is the maximum MP3 quality. Lossless formats (WAV, FLAC, ALAC) ignore this setting.

Sample rate sets how many audio samples per second the file contains. Standard music is 44.1 kHz; pro audio and video typically use 48 kHz. Higher rates (96 kHz, 192 kHz) only matter for professional production.

Channels can be mono, stereo, 5.1, or 7.1. Choose mono for voice memos or single-mic recordings to halve the file size; stereo is the default for music; 5.1/7.1 are for surround soundtracks.

4. Adjust speed, volume, and trim

Speed Multiplier changes playback tempo from 0.25x to 4x without altering pitch — useful for sped-up podcasts or slowed-down practice tracks. Volume Multiplier applies a simple gain factor between 0.1x and 2x.

Click Trim Audio to scrub through the waveform and pick a start and end time. Only the selected segment is encoded into the output file.

5. AI Audio Tools (recommended)

The AI Audio Tools panel is the fastest path to a great-sounding result for voice and music. Each operation runs on our local GPU server (no data leaves our infrastructure):

  • Clean Voice — denoises with DeepFilterNet and then runs a broadcast-style polish chain (high-pass, low-pass, loudness normalization, brick-wall limiter). Best for podcast/voiceover takes recorded in noisy rooms.
  • Remove Background Noise — DeepFilterNet denoise without the polish chain. Use this when you want to preserve the natural EQ and dynamics of the recording.
  • Isolate Vocals — runs Demucs htdemucs and outputs the vocal stem only. Output defaults to WAV so you don't lose detail.
  • Remove Vocals / Karaoke — same Demucs job, but exports the instrumental stem. Great for karaoke tracks or sample workflows.

When you pick any AI operation, the standard FFmpeg chain (bitrate, EQ, reverb, etc.) is skipped for that job so the AI tool has a clean input/output path. AI jobs may take longer than basic conversion because they run on the GPU.

6. Advanced Audio Effects

Expand Advanced Audio Effects to access four collapsible groups when you want manual control instead of an AI operation:

  • Basic Audio Processing — normalize, amplify (dB), fade in/out, EQ presets (bass boost, treble boost, vocal, classical, rock, jazz), pan, balance, stereo width, mono conversion, channel swap.
  • Time-Based Effects — reverb (room, hall, plate, spring), delay (echo, multi-tap, ping-pong), modulation (chorus, flanger, phaser, tremolo, vibrato).
  • Restoration & Cleanup — spectral/adaptive/gate noise reduction, 50/60 Hz de-hum, declipper, silence detection and removal.
  • Advanced Audio — pitch shifting (semitones, formant-preserving), time stretching (factor + pitch/time/formant algorithm), spatial audio (binaural, surround, 3D).

7. Convert and download

Click Convert File. A real-time progress bar appears as the server processes the job. When it finishes, the result auto-previews and a Download button appears. The file is also added to your in-session Conversion History so you can revisit or download it later without re-running the job.

Tips for great audio output

  • For voice recordings, try Clean Voice first — it solves most "muddy podcast" problems in one click.
  • If you only want to remove a constant hum or air-conditioner background, use Remove Background Noise to keep the original tone.
  • For stems, keep the default WAV output — converting Demucs output to a lossy format defeats the purpose of separation.
  • Big bitrate jumps don't help small files much; pair 320 kbps with 48 kHz only for high-quality music exports.