---
title: Reproduction of speech using MIDI
canonical: https://0110.be/posts/Reproduction_of_speech_using_MIDI
markdown_url: https://0110.be/posts/Reproduction_of_speech_using_MIDI.md
id: 361
published_at: '2010-06-24T15:26:56Z'
updated_at: '2013-12-05T18:19:15Z'
author: Joren
tags:
- name: HoGent
  markdown_url: https://0110.be/tags/HoGent.md
- name: Java
  markdown_url: https://0110.be/tags/Java.md
---

# Reproduction of speech using MIDI

Tarsos is now capable of reproducing *speech* using MIDI. The idea to convert speech into MIDI comes from [the blog of Corban Brook](http://weare.buildingsky.net/2009/10/08/mechanical-reproduction-of-digitized-speech-on-a-piano) where the following video can be found, actually a work by [Peter Ablinger](http://ablinger.mur.at/voices_and_piano.html):

<object width="640" height="385">
<param name="movie" value="http://www.youtube.com/v/muCPjK4nGY4&color1=0xb1b1b1&color2=0xd0d0d0&hl=nl_NL&feature=player_detailpage&fs=1"></param><param name="allowFullScreen" value="true"></param><param name="allowScriptAccess" value="always"></param><embed src="http://www.youtube.com/v/muCPjK4nGY4&color1=0xb1b1b1&color2=0xd0d0d0&hl=nl_NL&feature=player_detailpage&fs=1" type="application/x-shockwave-flash" allowfullscreen="true" allowScriptAccess="always" width="640" height="385"></embed></object>

Another example of music inspired by speech is this interview with Louis Van Gaal:

<object style="height: 344px; width: 425px">
<param name="movie" value="http://www.youtube.com/v/BDxvRfCNmQY"><param name="allowFullScreen" value="true"><param name="allowScriptAccess" value="always"><embed src="http://www.youtube.com/v/BDxvRfCNmQY" type="application/x-shockwave-flash" allowfullscreen="true" allowScriptAccess="always" width="425" height="344"></object>

Tarsos sends out midi data based on an FFT analysis of the signal. It maps the spectrogram to MIDI Messages and uses the power spectrum to calculate the velocity of each note on message.

The implementation can run in real-time but the output has some delay: the FFT calculation, constructing MIDI messages, calculating velocity, synthesizing sound, ... is not instantaneous.

To use this capability Tarsos supports the following syntax. If a MIDI file is given the MIDI messages are written to the file. If an audio file is given Tarsos uses the audio as input. If the `--pitch` switch is used only the F0 is considered to construct MIDI messages instead of a complete FFT.

\`\`\`ruby\
java -jar tarsos.jar pitch_to_midi \[---pitch\] \[midi_out.midi\] \[audio_in.wav\]\
\`\`\`


- [ttm.mp3](https://0110.be/files/attachments/361/ttm.mp3)

- [ttm.midi](https://0110.be/files/attachments/361/ttm.midi)

- [sitting\_in\_a\_room.mp3](https://0110.be/files/attachments/361/sitting_in_a_room.mp3)

- [sitting\_in\_a\_room.midi](https://0110.be/files/attachments/361/sitting_in_a_room.midi)

- [tarsos.jar](https://0110.be/files/attachments/361/tarsos.jar)
