Finding a broad melodic outline
A prominent lead vocal or solo may appear clearly enough to sketch the main contour, especially in sparse sections.
You can analyze a finished song, but a dense full mix is the hardest input for this model. For useful MIDI, separate the musical part you want before transcription.
For a full song, expect a rough sketch. A separated vocal, bass, piano or other stem will produce a much more editable result.
6.5-second synthetic melody · a workflow demo, not an accuracy benchmark.
The selected file is {size}. MIDIFLOW does not upload it or impose a server limit, but very large files can run out of memory, especially on a phone.
Compare the original with a simple synth preview of the detected notes. The preview does not reproduce the original sound or pitch bends.
Too many or missing notes? — increase it to keep fewer, more confident notes; decrease it to include quieter notes. Then compare the previews again.
Free foreverNo signupNo watermarkAudio never uploaded
Best input
Use a full-mix pass for exploration, then move to stems when you need notes you can arrange. This tool does not reconstruct the original multitrack session.
A prominent lead vocal or solo may appear clearly enough to sketch the main contour, especially in sparse sections.
The piano roll can reveal whether melody, bass or harmony dominates before you spend time cleaning a specific part.
Once the song is split, process one stem at a time. Vocals or bass generally create a simpler note field than the combined master.
What can go wrong
Drums, bass, vocals and harmonic instruments occupy overlapping time and frequency regions. The model cannot know which original track a spectral peak belonged to, so it may merge real notes or invent extra ones between parts.
Kick and snare transients can trigger note onsets, while cymbals spread broadband energy across the same frames used to estimate pitches. Reverb and mastering compression keep many sounds active at once.
A single MIDI track also cannot represent the production decisions in the master. It contains detected note events—not separated instruments, drum classification, lyrics, patches or a faithful arrangement.
Choose an MP3, WAV, M4A, FLAC or OGG recording, or try the sample. Start with a short solo part to check the result.
The converter estimates notes on this device. Your recording is not uploaded and no account is needed.
Compare the original with Play MIDI notes, apply a different note threshold if needed, then save the .mid file for your music software.
Transcription is powered by Spotify's open-source Basic Pitch model and runs in your browser.
Improve accuracy
The most effective improvement is reducing the number of instruments presented to the transcription model at the same time.
Use a local tool such as Demucs, or an external service such as Moises or LALAL.AI, to create vocals, drums, bass and other stems. External services have their own upload and privacy policies.
Skip the drum stem. Start with isolated vocals for melody, bass for the bass line, or the cleanest harmonic stem for chords.
Import each MIDI result on a separate DAW track, delete obvious ghost notes, then combine only the parts you actually need.
The download is a standard MIDI file containing note timing, pitch, velocity and eligible pitch-bend events. It does not contain the original instrument sound, so assign any software instrument after importing it.
Create or choose a MIDI track, then drag the .mid file into an empty clip slot or the Arrangement. Load an instrument on that track and edit notes in the MIDI Note Editor.
Drag the file into the Tracks area and choose a Software Instrument track. Open the Piano Roll to correct timing, note lengths or velocity before arranging.
Drag the file into the Channel Rack or use File › Import › MIDI file. Send the imported notes to a chosen instrument, then clean them in the Piano roll.
No. It produces one transcription from the mixed waveform. It does not recover the original multitrack session or assign each note to its source instrument.
Strong drum transients and resonances can activate the same onset and pitch detectors used for musical notes. A separated non-drum stem removes much of that interference.
Choose the stem matching your goal: vocals for melody, bass for the bass line, or a piano/other stem for harmonic material. Avoid combining them for the first pass.
No. Keeping transcription local and lightweight is the priority. Separate stems elsewhere, then bring only the target audio file into MIDIFLOW.
No. MIDI stores performance instructions, not the original vocals, instruments, effects or mix. You must choose sounds and arrange the detected notes in a DAW.
Not necessarily. A sparse performance may have fewer layers, but room echo, audience noise and microphone bleed can still make individual notes hard to separate.