---
title: Blog - 0110.be
canonical: https://0110.be/Blog?page=2
markdown_url: https://0110.be/Blog.md?page=2
page: 2
posts_per_page: 30
total_posts: 260
filters:
  content_page: Blog
  tags:
  - 0110 concerten
  - 0110.be
  - ARIP
  - Code
  - Collaborative Filtering
  - Command Line Application
  - Computational ethnomusicology
  - Computational musicology
  - Cultuur
  - Dutch
  - Film
  - Folk Music Analysis (FMA) conference
  - Hackerspace Ghent
  - Harde waren
  - HoGent
  - ISMIR
  - JNMR
  - Java
  - Jazz
  - Jongleren
  - Joren In Halmstad
  - LaTeX
  - Mac OS X
  - Music Information Retrieval
  - Muziek
  - PeachNote Piano
  - Poging tot humor
  - Portfolio
  - Presentation
  - Projecten
  - Reizen
  - Research papers
  - School
  - Schoolwijs
  - Tarsos
  - TarsosDSP
  - TarsosTranscoder
  - Thesis
  - UGent
  - Vooruit
  - WSOLA
  - featured
  - français
previous: https://0110.be/Blog.md?page=1
next: https://0110.be/Blog.md?page=3
---

# Blog - 0110.be

## [Attempting humor in academic writing](https://0110.be/posts/Attempting_humor_in_academic_writing.md)

- Published: 2023-01-24T00:00:00Z
- Updated: 2023-01-31T15:05:49Z
- Author: Joren
- ID: 501
- Canonical: https://0110.be/posts/Attempting_humor_in_academic_writing

- Tags: [Poging tot humor](https://0110.be/tags/Poging%20tot%20humor.md), [Research papers](https://0110.be/tags/Research%20papers.md), [UGent](https://0110.be/tags/UGent.md)

<div style="float:right; width:20%">
<center>
<a href="https://link.springer.com/book/10.1007/978-3-031-18444-4"><img style="width:80%" src="https://0110.be/files/attachments/501/springer_book_cover.jpeg" alt="Screenshot of a browser based pitch organization extraction tool"></a><br><small>Fig: Advances in Speech and Music Technology book cover.</small>

</center>
</div>
I have recently published an chapter in an academic book published by Springer. The topic of the book is of interest to me but can be perceived as rather dry: [*Advances in Speech and Music Technology*](https://link.springer.com/book/10.1007/978-3-031-18444-4).

The chapter I co-authored presented two case studies on detecting duplicates in music archives. The fist case study deals with segmentation reuse in an archive of early electronic music. The second with meta-data reuse in an archive of a public broadcaster containing digitized commercial shellac disc recordings with many duplicates.

Duplicate detection being the main topic, I decided to title the article *[Duplicate Detection for for Digital Audio Archive Management](https://link.springer.com/chapter/10.1007/978-3-031-18444-4_16__). It is easy to miss, and not much is lost if you do, but there is a duplicate*'for' in the title. If you did detect the duplicate you have detected the duplicate in the duplicate detection article. Since I have fathered two kids I see it as an hard earned right to make dad-jokes like that. Even in academic writing.

It was surprisingly difficult to get the title published as-is. At every step of the academic publishing process (review, editorial, typesetting, lay-outing) I was asked about it and had to send an email like the one below. Every email and every explanation made my second-guess my sense of humor but I do stand by it.

<blockquote style="font-family: monospace">
From: Joren<br>\
To: Editors ASMT<br>\
<br>\
Dear Editors,<br>\
<br>\
I have updated my submission on easychair in...\
<br>\
I would like to keep the title however as is an attempt at word-play. These things tend to have less impact when explained but the article is about duplicate detection and is titled 'Duplicate detection for for digital audio archive management'. The reviewer, attentively, detected the duplicate 'for' but unfortunately failed to see my attempt at humor. To me, it is a rather harmless witticism.

Regards

Joren

</pre>
</blockquote>
Anyway, I do think that humor can serve as a gateway to direct attention to rather dry, academic material. Also the message and the form of the message should not be confused. [John Oliver](https://www.youtube.com/@LastWeekTonight/videos), for example, made his whole career on delivering serious sometimes dry messages with heaps of humor: which does not make the topics less serious. I think there are a couple of things to be learned there. Anyway, now that I have your attention, please do read the [author version of *Duplicate Detection for for Digital Audio Archive Management: Two Case Studies*](https://0110.be/files/publications/2022/2022.duplicates-author_version.pdf).


---

## [Updates for TarsosDSP](https://0110.be/posts/Updates_for_TarsosDSP.md)

- Published: 2023-01-20T00:00:00Z
- Updated: 2023-01-27T15:51:31Z
- Author: Joren
- ID: 503
- Canonical: https://0110.be/posts/Updates_for_TarsosDSP

- Tags: [Code](https://0110.be/tags/Code.md), [Music Information Retrieval](https://0110.be/tags/Music%20Information%20Retrieval.md), [TarsosDSP](https://0110.be/tags/TarsosDSP.md), [UGent](https://0110.be/tags/UGent.md)

TarsosDSP is a Java library for audio processing I have started working on more than 10 years ago. The aim of TarsosDSP is to provide an easy-to-use interface to practical music processing algorithms. Obviously, I have been using it myself over the years as my go-to library for audio-processing in Java. However, a number of gradual changes in the java ecosystem made TarsosDSP more and more difficult to use.

Since I have apparently not been the only one using it, there was a need to give it some attention. During the last couple of weeks I have found the time to give it this much needed attention. This resulted in a number of updates, some of the changes include:

-   Change of the build system from Apache Ant to Gradle

-   Make use of Java Modules to make TarsosDSP compatible with the ModulePath introduced in Java 9.

-   Packaged the software into a maven compatible format, which makes it easy to use as a dependency.

-   CI with GitHub actions to automatically build and test the software.

-   Updated some examples shipped with the TarsosDSP. I have still still some examples to verify.

-   Improved handling of errors on reading audio via ffmpeg

<center>
<img width="60%" src="https://0110.be/files/attachments/503/tarsosdsp_gui_examples.webp" alt="Examples of TarsosDSP"><br>\
<small>Fig: The updated TarsosDSP release contains many CLI and GUI example applications.</small>

</center>
Notably **the code of TarsosDSP has not changed much** apart from some cosmetic changes. This backwards compatibility is one of the strong points of Java. With this update I am quite confident that TarsosDSP will also be usable during the next decade as well.

Please check out the updated [TarsosDSP repository on GitHub](https://github.com/JorenSix/TarsosDSP). <br>


![Flanger](https://0110.be/files/photos/503/tarsosdsp_flanger_effect.webp)

![Oscilloscope](https://0110.be/files/photos/503/tarsosdsp_oscilloscope.webp)

![Pitch estimator](https://0110.be/files/photos/503/tarsosdsp_pitch_detector.png)

---

## [The automatic HTTPS capabilities of Caddy](https://0110.be/posts/The_automatic_HTTPS_capabilities_of_Caddy.md)

- Published: 2023-01-18T00:00:00Z
- Updated: 2023-01-25T16:24:35Z
- Author: Joren
- ID: 502
- Canonical: https://0110.be/posts/The_automatic_HTTPS_capabilities_of_Caddy

- Tags: [0110.be](https://0110.be/tags/0110.be.md), [UGent](https://0110.be/tags/UGent.md)

<div style="float:right; width:25%">
<center>
<a href="https://caddyserver.com/"><img style="width:90%" src="https://0110.be/files/attachments/502/caddy-logo.svg" alt="Caddy logo"></a><br>

</center>
</div>
This blog has been running on [Caddy](https://caddyserver.com/) for the last couple of months. Caddy is a http server with support for reverse proxies and automatic https. The automatic https feature takes care of requesting, installing and updating SSL certificates which means that you need much less configuration settings or maintenance compared with e.g. lighttpd or Nginx. The underlying [certmagic](https://github.com/caddyserver/certmagic) ACME client is responsible for requesting these certificates.

Before, it was using [lighttpd](https://www.lighttpd.net/) but the during the last decade the development of lighttpd has stalled. lighttpd version 2 has been in development for 7 years and the bump from `1.4` to `1.5` has been taking even longer. lighttpd started showing its age with limited or no support for modern features like websockets, http/3 and finicky configuration for e.g. https with virtual domains.

#### Caddy with Ruby on Rails

I really like Caddy's sensible defaults and the limited lines of configuration needed to get things working. Below you can find e.g. a reusable https enabled configuration for a Ruby on Rails application. This configuration does file caching, compression, http to https redirection and load balancing for two local application servers. It also serves static files directly and only passes non-file requests to the application servers.

    (cachestaticfiles) {
        @staticFiles {
            file
        }
        header @staticFiles Cache-Control "public, max-age=604800, must-revalidate"
    }

    (railsdefaults) {
        #compress responses
        encode zstd gzip

        #redirect from http to https
        @http {
            protocol http
        }

        redir @http https://{host}{uri}

        @notStatic {
            not file
        }

        import cachestaticfiles
        reverse_proxy @notStatic localhost:{args.0} localhost:{args.1}
        file_server
    }

    example.com {
        root * /var/www/example.com/current/public
        import railsdefaults 10000 10001
    }

If you are self-hosting I think Caddy is a great match in all but the most exotic or demanding setups. I definitely am kicking myself for not checking out caddy sooner: it could have saved me countless hours installing and maintaining https certs or configuring lighttpd in general.


---

## [Crossplatform JNI builds with Zig](https://0110.be/posts/Crossplatform_JNI_builds_with_Zig.md)

- Published: 2023-01-13T00:00:00Z
- Updated: 2023-01-19T16:33:16Z
- Author: Joren
- ID: 500
- Canonical: https://0110.be/posts/Crossplatform_JNI_builds_with_Zig

- Tags: [Code](https://0110.be/tags/Code.md), [Music Information Retrieval](https://0110.be/tags/Music%20Information%20Retrieval.md), [UGent](https://0110.be/tags/UGent.md)

JNI is a way to use C or C code from Java and allows developers to reuse and integrate C/C in Java software. In contrast to the Java code, C/C code is *platform dependent and needs to be compiled for each platform/architecture*. Also it is generally not a good idea to make users compile a C/C library: it is best provide precompiled libraries. As a developer it is, however, a pain to provide binaries for each platform.

With the dominance of x86 processors receding the problem of having to compile software for many platforms is becoming more pressing. It is not unthinkable to want to support, for example, intel and M1 macOS, ARM and x86_64 Linux and Windows. To support these platforms you would either need access to such a machine with a compiler or configure a cross-compiler for each system: *both are unpractical*. Typically setting up a cross-compiler can be time consuming and finicky and virtual machines can be tough to setup. There is however an alternative.

[Zig](https://ziglang.org) is a programming language but, thanks to its support for C/C, it also ships with an *easy-to-use cross-compiler which is of interest here even if you have no intention to write a single line of Zig code*. The built-in cross-compiler allows to [target many platforms](https://ziglang.org/download/0.8.0/release-notes.html#Support-Table) easily.

<div style="float:right; width:25%">
<center>
<a href="https://ziglang.org/">\
<img style="width:80%" src="https://0110.be/files/attachments/500/zig_logo.svg" alt="Zig logo"></a><br>

</center>
</div>
#### The Zig cross-compiler in practice

Cross compilation of C code is possible by simply *replacing the `gcc` command with `zig cc`* and adding a target argument, e.g. for targeting a Windows. There is more general information on [zig as a cross-compiler here](https://zig.news/kristoff/cross-compile-a-c-c-project-with-zig-3599).

*Cross-compiling a JNI library is not different to compiling other libraries.* To make things concrete we will cross-compile a library from a typical JNI project: [JGaborator](https://github.com/JorenSix/JGaborator) packs [the C/C library gaborator](https://gaborator.com). In this case the C/C code does a computationally intensive spectral transformation of time domain data. The commands below create an x86_64 Windows DLL from a macOS with zig installed:

``` {style="overflow-x:scroll"}
<code>
bash
#wget https://aka.ms/download-jdk/microsoft-jdk-17.0.5-windows-x64.zip
#unzip microsoft-jdk-17.0.5-windows-x64.zip
#export JAVA_HOME=`pwd`/jdk-17.0.5+8/
git clone --depth 1 https://github.com/JorenSix/JGaborator
cd JGaborator/gaborator
echo $JAVA_HOME
JNI_INCLUDES=-I"$JAVA_HOME/include"\ -I"$JAVA_HOME/include/win32" 
zig cc  -target x86_64-windows-gnu -c -O3 -ffast-math -fPIC pffft/pffft.c -o pffft/pffft.o
zig cc  -target x86_64-windows-gnu -c -O3 -ffast-math -fPIC -DFFTPACK_DOUBLE_PRECISION pffft/fftpack.c -o pffft/fftpack.o
zig c++ -target x86_64-windows-gnu -I"pffft" -I"gaborator-1.7"  $JNI_INCLUDES -O3\
        -ffast-math -DGABORATOR_USE_PFFFT  -o jgaborator.dll jgaborator.cc pffft/pffft.o pffft/fftpack.o
file jgaborator.dll
# jgaborator.dll: PE32+ executable (console) x86-64, for MS Windows
</code>
```

Note that, when cross-compiling from macOS, *to target Windows a Windows JDK is needed*. The windows JDK has other header files like `jni.h`. Some commands to download and use the JDK are commented out in the example above. Also note that targeting Linux from macOS seems to work with the standard macOS JDK. This is probably due to shared conventions regarding compilation of libraries.

To target other platforms, e.g. ARM Linux, there are *only two things that need to be changed*: the `-target` switch should be changed to `aarch64-linux-gnu` and the name of the output library should be (by Linux convention) changed to `libjgaborator.so`. During the build step of JGaborator a list of target platforms it iterated and a total of 9 builds are packaged into a single Jar file. There is also a bit of supporting code to load the correct version of the library.

Using a GitHub action or similar CI tools this cross compilation with zig can be automated to run on a software release. For Github the [Setup Zig](https://github.com/marketplace/actions/setup-zig) action is practical.

#### Loading the correct library

In a first attempt I tried to detect the operating system and architecture of the environment to then load the library but eventually decided against this approach. Mainly because you then need to keep an exhaustive list of supporting platforms and this is *difficult, error prone and decidedly not future-proof*.

In my second attempt I simply *try to load each precompiled library* limited to the sensible ones - only dll's on windows - until a matching one is loaded. The rationale here is that the system itself knows best which library works and failing to load a library is computationally cheap. There is [some code to iterate all precompiled libraries in a JAR-file](https://github.com/JorenSix/JGaborator/blob/master/src/main/java/be/ugent/jgaborator/util/ZigNativeUtils.java#L106) so supporting an additional platform amounts to adding a [precompiled library in the JAR folder](https://github.com/JorenSix/JGaborator/tree/master/src/main/resources/jni): there is no need to be explicit in the Java code about architectures or OSes.

Trying multiple libraries has an additional advantage: this allows to ship multiple versions targeting the same architecture: e.g. one with additional acceleration libraries enabled and one without. By sorting the libraries alphabetically the first, then, should be the one with acceleration and the fallback without. In the case of JGaborator for mac aarch64 there is one compiled with `-framework Accelerate` and one compiled by the Zig cross-compiler without.

#### Takehome messages

-   If you find yourself cross-compiling C or C for many platforms, **consider the Zig cross-compiler**. Even when you have no intention to write a single line of Zig code.

-   For JNI and Java the [JGaborator source code](https://github.com/JorenSix/JGaborator) might offer some **inspiration to pre-compile and load libraries** for many platforms with little effort.

-   CI tools can help to verify builds and **automate Zig cross-compilation**.

-   If you build for Windows make sure to include windows header-files even when there are no compilation errors using UNIX-header files.

If you find this valuable [consider sponsoring the work on Zig](https://github.com/sponsors/ziglang)


---

## [Emotopa - Patterns in Pitch Organization](https://0110.be/posts/Emotopa_-_Patterns_in_Pitch_Organization.md)

- Published: 2022-12-10T00:00:00Z
- Updated: 2022-12-20T14:38:27Z
- Author: Joren
- ID: 499
- Canonical: https://0110.be/posts/Emotopa_-_Patterns_in_Pitch_Organization

- Tags: [UGent](https://0110.be/tags/UGent.md)

<div style="float:right; width:22%">
<center>
<a href="https://0110.be/attachment/cors/2022.12.emotopa/">\
<img style="width:80%" src="https://0110.be/files/attachments/499/emotopa_screenshot.png" alt="Screenshot of a browser based pitch organization extraction tool"></a><br>\
<small>Fig: Screenshot of *Emotopa*: a browser based tool to extract pitch organization from audio.</small>

</center>
</div>
A couple of days ago I participated in the [Music Hack Day - India](https://musichackdayindia.github.io/). The event was organized the 10th and 11th of December in Bangaluru, India. During the event a representative of [Smule](https://www.smule.com/) suggested a task to evaluate the performance of karaoke-singers in terms of intonation. The idea was to employ pitch histogram like features to estimate pitch use of singers.

I offered to build a browser based application to extract pitch histograms from audio. At the end of the hack day I presented the [first release of *Emotopa*](https://0110.be/attachment/cors/2022.12.emotopa/) with some limited functionality:

1.  The application is able to decode and use audio from any format or container by using an [audio focused webassambly build of ffmpeg](https://github.com/JorenSix/ffmpeg.audio.wasm).

2.  Next, a pitch detector runs on the audio and returns a list of pitch estimates.

3.  Finally a histogram (technically a kernel density estimate) is constructed using the pitch estimates.

The user can export the pitch histogram, the pitch class histogram and the pitch annotations. These features successfully show the intonation quality of singers but the applications are much broader. Some potential applications have been described in (amongst others) the [Tarsos article](https://0110.be/publications/Tarsos%2C_a_modular_platform_for_precise_pitch_analysis_of_western_and_non-western_music).

The Emotopa name alludes to the [Apotome](https://isartum.net/apotome) browser based app where, starting from a scale you can make music. With Emotopa you do the reverse. Also very much of interest are [Leimma](https://isartum.net/leimma) and the [rationale behind both Apotome and Leimma](https://pitchfork.com/thepitch/decolonizing-electronic-music-starts-with-its-software/)

The source can be verified on the [Emotopa GitHub repository](https://github.com/JorenSix/Emotopa)

<small style="color:#AAA">This contribution was made possible thanks to travel funds by the FWO travel grant K1D2222N and the Ghent University BOF funded project PaPiOM.</small>


---

## [DiscStitch & BAF - Contributions to ISMIR 2022](https://0110.be/posts/DiscStitch_%26_BAF_-_Contributions_to_ISMIR_2022.md)

- Published: 2022-12-04T00:00:00Z
- Updated: 2022-12-15T14:10:38Z
- Author: Joren
- ID: 497
- Canonical: https://0110.be/posts/DiscStitch_%26_BAF_-_Contributions_to_ISMIR_2022

- Tags: [ISMIR](https://0110.be/tags/ISMIR.md), [Music Information Retrieval](https://0110.be/tags/Music%20Information%20Retrieval.md), [UGent](https://0110.be/tags/UGent.md)

<div style="float:right; margin: 8px; width:20%">
<img src="https://0110.be/files/attachments/497/ismir_tab_icon.png" style="object-fit:contain; width: 100%;" />

</div>
This year the [ISMIR 2022](https://ismir2022.ismir.net/) conference is organized from 4 to 9 December 2022 in Bengaluru, India. ISMIR is the main music technology and music information retrieval (MIR) conference. It is a relief to experience a conference in physical form and not through a screen.

I have contributed to the following work which is in the main paper track of ISMIR 2022:

> [*BAF: An Audio Fingerprinting Dataset For Broadcast Monitoring*](https://0110.be/publications/BAF%3A_an_audio_fingerprinting_dataset_for_broadcast_monitoring) ([version of record](https://ismir2022program.ismir.net/poster_228.html))\
> Guillem Cortès, Alex Ciurana, Emilio Molina, Marius Miron, Owen Meyers, Joren Six, Xavier Serra\
> <br>**Abstract**: *Audio Fingerprinting (AFP) is a well-studied problem in music information retrieval for various use-cases e.g. content-based copy detection, DJ-set monitoring, and music excerpt identification. However, AFP for continuous broadcast monitoring (e.g. for TV & Radio), where music is often in the background, has not received much attention despite its importance to the music industry. In this paper (1) we present BAF, the first public dataset for music monitoring in broadcast. It contains 74 hours of production music from Epidemic Sound and 57 hours of TV audio recordings. Furthermore, BAF provides cross-annotations with exact matching timestamps between Epidemic tracks and TV recordings. Approximately, 80% of the total annotated time is background music. (2) We benchmark BAF with public state-of-the-art AFP systems, together with our proposed baseline PeakFP: a simple, non-scalable AFP algorithm based on spectral peak matching. In this benchmark, none of the algorithms obtain a F1-score above 47%, pointing out that further research is needed to reach the AFP performance levels in other studied use cases. The dataset, baseline, and benchmark framework are open and available for research.*

<center>
<video style="width:50%" poster="https://0110.be/files/attachments/497/baf_broadcast_audio_fingerprinting_thumb.jpg" controls preload="none">
<source src="https://0110.be/files/attachments/497/BAF_broadcast_audio_fingerprinting.mp4" type="video/mp4">
</video>
</center>
I have also presented a first version of DiscStitch, an audio-to-audio alignment algorithm. This contribution is in the *ISMIR 2022 late breaking demo session*:

<div style="float:right; margin: 8px; width:15%">
<img src="https://0110.be/files/attachments/497/2022_DiscStitch__towards_audio-to-audio_alignment_with_robustness_to_playback_speed_variabilities_ISMIR_2022_Late_Breaking___Demo_abstracts.webp" style="object-fit:contain; width: 100%;" />

</div>
> [*DiscStitch: towards audio-to-audio alignment with robustness to playback speed variabilities*](https://0110.be/publications/DiscStitch%3A_towards_audio-to-audio_alignment_with_robustness_to_playback_speed_variabilities) ([version of record](https://ismir2022program.ismir.net/lbd_414.html))\
> Joren Six\
> <br>**Abstract**: *Before magnetic tape recording was common, acetate discs were the main audio storage medium for radio broadcasters. Acetate discs only had a capacity to record about ten minutes. Longer material was recorded on overlapping discs using (at least) two recorders. Unfortunately, the recorders used were not reliable in terms of recording speed, resulting in audio of variable speed. To make digitized audio originating from acetate discs fit for reuse, (1) overlapping parts need to be identified, (2) a precise alignment needs to be found and (3) a mixing point suggested. All three steps are challenging due to the audio speed variabilities. This paper introduces the ideas behind DiscStitch: which aims to reassemble audio from overlapping parts, even if variable speed is present. The main contribution is a fast and precise audio alignment strategy based on spectral peaks. The method is evaluated on a synthetic data set.*

Next to my own contributions, the [ISMIR conference program](https://ismir2022program.ismir.net/) is the best overview of the state-of-the art of MIR.

<small style="color:#AAA">This contribution was made possible thanks to travel funds by the FWO travel grant K1D2222N and the Ghent University BOF funded project PaPiOM.</small>


- [ismir\_tab\_icon.png](https://0110.be/files/attachments/497/ismir_tab_icon.png)

- [BAF\_broadcast\_audio\_fingerprinting.mp4](https://0110.be/files/attachments/497/BAF_broadcast_audio_fingerprinting.mp4)

- [discstitch\_static\_poster.pdf](https://0110.be/files/attachments/497/discstitch_static_poster.pdf)

- [2022\_DiscStitch\_\_towards\_audio-to-audio\_alignment\_with\_robustness\_to\_playback\_speed\_variabilities\_ISMIR\_2022\_Late\_Breaking\_\_\_Demo\_abstracts.webp](https://0110.be/files/attachments/497/2022_DiscStitch__towards_audio-to-audio_alignment_with_robustness_to_playback_speed_variabilities_ISMIR_2022_Late_Breaking___Demo_abstracts.webp)

- [baf\_broadcast\_audio\_fingerprinting\_thumb.jpg](https://0110.be/files/attachments/497/baf_broadcast_audio_fingerprinting_thumb.jpg)

---

## [Panako: a scalable audio search system](https://0110.be/posts/Panako%3A_a_scalable_audio_search_system.md)

- Published: 2022-10-12T00:00:00Z
- Updated: 2023-07-04T09:25:18Z
- Author: Joren
- ID: 496
- Canonical: https://0110.be/posts/Panako%3A_a_scalable_audio_search_system

- Tags: [Panako](https://0110.be/tags/Panako.md), [Research papers](https://0110.be/tags/Research%20papers.md), [UGent](https://0110.be/tags/UGent.md)

<div style="float:right; margin: 8px; width:20%">
<img src="https://0110.be/files/attachments/496/papers_vs_software.webp" style="object-fit:contain; width: 100%;" /><small>Fig: DALL.E 2 imagining a fight between papers and software.</small>

</div>
Recently I have published a paper titled [*'Panako: a scalable audio search system'*](https://joss.theoj.org/papers/10.21105/joss.04554) in the Journal of Open Source Software (JOSS). The journal is [a 'hack' to circumvent the focus on citable papers](https://www.arfon.org/announcing-the-journal-of-open-source-software) in the academic world: getting recognition for publishing software as a researcher is not straightforward.

The research output tracking system of Ghent University (biblio) and Flanders FWO's academic profile are not built to track software as research output. The focus is still solely on papers, even when custom developed research software has become a fundamental aspect in many research areas. My role is somewhere between that of a 'pure' researcher and that of a [research software engineer](https://www.nature.com/articles/d41586-022-01516-2) which makes this focus on papers quite relevant to me.

The paper aims to make the recent development on [Panako](https://github.com/JorenSix/Panako) *'count'*. Thanks to the JOSS review process the Panako software was improved considerably: CI, unit tests, documentation, containerization,... The paper was a good reason to improve on all these areas which are all too easy to neglect. The paper itself is a short, rather general overview of Panako:

> "*Panako solves the problem of finding short audio fragments in large digital audio archives. The content based audio search algorithm implemented in Panako is able to identify a short audio query in a large database of thousands of hours of audio using an acoustic fingerprinting technique.*"


---

## [Low impact runner: a music based bio-feedback system](https://0110.be/posts/Low_impact_runner%3A_a_music_based_bio-feedback_system.md)

- Published: 2022-10-01T00:00:00Z
- Updated: 2025-11-29T14:22:46Z
- Author: Joren
- ID: 490
- Canonical: https://0110.be/posts/Low_impact_runner%3A_a_music_based_bio-feedback_system

- Tags: [IPEM](https://0110.be/tags/IPEM.md), [Java](https://0110.be/tags/Java.md), [Research papers](https://0110.be/tags/Research%20papers.md), [UGent](https://0110.be/tags/UGent.md)

<div style="float:right;margin-left:10px">
<video  width="320" autoplay muted loop>
<source src="https://0110.be/files/attachments/490/lir-demo.mp4">
</video><br>
<small>Fig: schema of the low impact runner system.</small>
</div>

I have been lucky to have been involved in an interdisciplinary research project around the *low impact runner*: a music based bio-feedback system to reduce tibial shock in over-ground running. In the beginning of October 2022 the PhD defence of Rud Derie takes place so it is a good moment to look back to this collaboration between several branches of Ghent University: [IPEM](https://www.ugent.be/lw/kunstwetenschappen/ipem/en) , [movement and sports science](https://www.ugent.be/ge/bsw/en) and [IDLab](https://www.ugent.be/ea/idlab/en).

The idea behind the project was to first select runners with a high foot-fall impact. Then an intervention would slightly nudge these runner to a running style with lower impact. A lower repetitive impact is expected to reduce the chance on injuries common for runners. A system was invented in which musical bio-feedback was given on the measured impact. The schema to the right shows the concept.

I was involved in development of the first hardware prototypes which measured acceleration on the legs of the runner and the development of software to receive and handle these measurement on a tablet strapped to a backpack the runner was wearing. This software also logged measurements, had real-time visualisation capabilities and allowed remote control and monitoring over the network. Finally measurements were send to a Max/MSP sonification engine. These prototypes of software and hardware were replaced during a valorization project but some parts of the software ended up in the final Android application.

<center>
<video width="60%" autoplay muted loop>
<source src="https://0110.be/files/attachments/490/low_impact_runner-system.mp4">
</video><br>
<small>Video: the left screen shows the indoor positioning system via UWB (ultra-wide-band) and the right screen shows the music feedback system and the real time monitoring of impact of the runner. Video by Pieter Van den Berghe</small>

</center>
Over time the first wired sensors were replaced with wireless Bluetooth versions. This made the sensors easy to use and also to visualize sensor values in the browser thanks to the Web Bluetooth API. I have experimented with this and made two demos: a [low impact runner visualizer](bt.html) and one [with the conceptual schema](bt_bg.html).

<center>
<video width="60%" autoplay muted loop>
<source src="https://0110.be/files/attachments/490/bt_html_interface.mp4">
</video><br>
<small>Vid: Visualizing the Bluetooth Low Impact Runner sensor in the browser.</small>
</center>

The following three studies shows a part of the trajectory of the project. The first paper is a validation of the measurement system. Secondly a proof-of-concept study is done which finally greenlights a larger scale intervention study.

1.  Van den Berghe, P., Six, J., Gerlo, J., Leman, M., & De Clercq, D. (2019). *Validity and reliability of peak tibial accelerations as real-time measure of impact loading during over-ground rearfoot running at different speeds*. Journal of Biomechanics, 86, 238-242.

2.  Van den Berghe, P., Lorenzoni, V., Derie, R., Six, J., Gerlo, J., Leman, M., & De Clercq, D. (2021). *Music-based biofeedback to reduce tibial shock in over-ground running: A proof-of-concept study*. Scientific reports, 11(1), 1-12.

3.  Van den Berghe, P., Derie, R., Bauwens, P., Gerlo, J., Segers, V., Leman, M., & De Clercq, D. (2022). *Reducing the peak tibial acceleration of running by music‐based biofeedback: A quasi‐randomized controlled trial*. Scandinavian Journal of Medicine & Science in Sports

There are quite a number of other papers but I was less involved in those. The project also resulted in two PhD's:

-   *Motor retraining by real-time sonic feedback: understanding strategies of low impact running* (2021) by Pieter Van den Berghe
-   *Running on good vibes: music induced running-style adaptations for lower impact running* (2022) by Rud Derie

I am also recognized as co-inventor on the [low impact runner system patent](https://worldwide.espacenet.com/patent/search/family/062816385/publication/WO2020002275A1?q=WO2020002275A1) and there are concrete plans for a commercial spin-off. To be continued...

![Conceptual schema](https://0110.be/files/photos/490/li-jogger-schema.png)

![Music feedback schema](https://0110.be/files/photos/490/music_feedback_schema.jpg)

![Screening and selection](https://0110.be/files/photos/490/screening_selection_schema.jpg)

![Screenshot of the smartphone app](https://0110.be/files/photos/490/Screenshot_20220407-155541_pixel_quite_black_landscape.png)

---

## [SMPTE decoding in the browser](https://0110.be/posts/SMPTE_decoding_in_the_browser.md)

- Published: 2022-09-16T00:00:00Z
- Updated: 2025-11-29T14:23:29Z
- Author: Joren
- ID: 494
- Canonical: https://0110.be/posts/SMPTE_decoding_in_the_browser

- Tags: [UGent](https://0110.be/tags/UGent.md)

I have created a web application to LTC.wasm decodes SMPTE timecodes from an LTC encoded audio signal.

To synchronize multiple music and video recordings a shared SMPTE timecode signal is often used. For practical purposes the timecode signal is encoded in an audio stream. The timecode can then be recorded in sync with microphone inputs or added to a video recording. The timecode is encoded in audio with LTC, linear timecode. A special decoder is needed to extract SMPTE timecode from the audio. This is exactly what the LTC.wasm application does.

<center>
<img style="width:30%" src="https://0110.be/files/attachments/494/ltc.wasm.apng" /><br>
<caption>
Using the [web based SMTE decoder](https://0110.be/attachment/cors/ltc.wasm/ltc_decoder.html</caption>
</center>

Try out the [SMPTE decoder](https://0110.be/attachment/cors/ltc.wasm/ltc_decoder.html) with your own SMPTE files.

The advantage of the web-based version versus the command line [ltc-tools](https://github.com/x42/ltc-tools) is that it does not need to be installed separately and that [ffmpeg](http://ffmpeg.org) decodes audio. This means that almost any multimedia format is supported automatically. The command line version only supports a limited number of audio formats.

For further information check out the [the LTC.wasm GitHub repository](https://github.com/JorenSix/LTC.wasm)


---

## [SyncSink.wasm - Synchronize media files by audio-to-audio alignment](https://0110.be/posts/SyncSink.wasm_-_Synchronize_media_files_by_audio-to-audio_alignment.md)

- Published: 2022-09-06T00:00:00Z
- Updated: 2025-11-29T14:23:54Z
- Author: Joren
- ID: 493
- Canonical: https://0110.be/posts/SyncSink.wasm_-_Synchronize_media_files_by_audio-to-audio_alignment

- Tags: [Code](https://0110.be/tags/Code.md), [ISMIR](https://0110.be/tags/ISMIR.md), [UGent](https://0110.be/tags/UGent.md)

I have built a tool for audio-to-audio alignment. It has applications for synchronization of media files. It works in the browser and you can [synchronize your media files here with SyncSink.wasm](https://0110.be/attachment/cors/sync/sync.html). SyncSink.wasm does the following:

1.  From an incoming media-file audio is extracted, downmixed to mono and and resampled. This is done with [ffmpeg.audio.wasm](https://github.com/JorenSix/ffmpeg.audio.wasm) a wasm version of ffmpeg.
2.  For each audio track, fingerprints are extracted. These fingerprints reduce the the search space for alignment drastically.
3.  Each list of fingerprints is aligned with the list of fingerprints from the reference. Resulting in a rough alignment
4.  Cross correlation is done to refine the alignment resulting in sample accurate results.

<center>
<img src="https://0110.be/files/attachments/493/media_sync_recording.apng"><br>
<small>Fig: media synchronization with audio-to-audio alignment.</small>

</center>
It supports small time-scale adjustments of around 5%: audio alignment can still be found if audio speed differs a bit.

Some potential use cases where it might be of use:

-   To stitch partially overlapping audio recordings together resulting in a single long audio recording.
-   To synchronize multiple independent video recordings of the same event each with an audio recording of the environment.
-   To align a high quality microphone recording with video/low-quality audio recording of the same event. The low quality audio recorded with a camera can then be replaced with the high quality microphone audio.

The code can be found in the [SyncSink.wasm GitHub repository](https://github.com/JorenSix/SyncSink.wasm)


- [media\_sync\_recording.apng](https://0110.be/files/attachments/493/media_sync_recording.apng)

---

## [Sending audio over a network with ffmpeg](https://0110.be/posts/Sending_audio_over_a_network_with_ffmpeg.md)

- Published: 2022-08-30T00:00:00Z
- Updated: 2022-09-05T12:29:57Z
- Author: Joren
- ID: 492
- Canonical: https://0110.be/posts/Sending_audio_over_a_network_with_ffmpeg

- Tags: [0110.be](https://0110.be/tags/0110.be.md), [UGent](https://0110.be/tags/UGent.md)

<div style="float:right; margin: 8px; width:30%">
<img src="https://0110.be/files/attachments/492/000012.2698386730.png" style="object-fit:contain; width: 100%;" /><small>Fig: stable diffusion imagining a networked music performance</small>

</div>
This post describes how to send audio over a network using the [ffmpeg](http://ffmpeg.org) suite. Ffmpeg is the Swiss army knife for working with audio and video formats. It is a command line tool that supports almost all audio formats known to man and woman. `ffmpeg` also supports streaming media over networks.

Here, we want to send audio recorded by a microphone, over a network to a **single** receiver on the other end. We are not aiming for low latency. Also the audio is going only in a single direction. This can be of interest for, for example, a networked music performance. Note that `ffmpeg` needs to be installed on your system.

### The receiver - Alice

For the receiver we use `ffplay`, which is part of the `ffmpeg` tools. The command instructs the receiver to listen to `TCP` connections on a randomly chosen port `12345`. The `\?listen` is important since this keeps the program waiting for new connections. For streaming media over a network the stateless `UDP` protocol is often used. When `UDP` packets go missing they are simply dropped. If only a few packets are dropped this does not cause much harm for the audio quality. For `TCP` missing packets are resent which can cause delays and stuttering of audio. However, `TCP` is much more easy to tunnel and the stuttering can be compensated with a buffer. Using `TCP` it is also immediately clear if a connection can be made. With `UDP` packets are happily sent straight to the void and you need to resort to wiresniffing to know whether packets actually arrive.

`ffplay -nodisp  -f mpegts tcp://0.0.0.0:12345\?listen`

In this example we use [MPEGTS](https://en.wikipedia.org/wiki/MPEG_transport_stream) over a plain `TCP` socket connection. Alteratively [RTMP](https://en.wikipedia.org/wiki/Real-Time_Messaging_Protocol) could be used (which also works over `TCP`). [RTP](https://en.wikipedia.org/wiki/Real-time_Transport_Protocol) , however is usually delivered over `UDP`.

The shorthand address `0.0.0.0` is used to bind the port to all available interfaces. Make sure that you are listening to the correct interface if you change the IP address.

### The sender - Björn

Björn, aka Bob, sends the audio. First we need to know from which microphone to use. To that end there is a way to list audio devices. In this example the macOS `avfoundation` system is used. For other operating systems there are similar provisions.

`ffmpeg -f avfoundation -list_devices true -i ""`

Once the index of the device is determined the command below sends incoming audio to the receiver (which should already be listening on the other end). The audio format used here is `MP3` which can be safely encapsulated into `mpegts`.

Note that the IP address `192.168.x.x` needs to be changed to the address of the receiver. Now if both devices are on the same network the incoming audio from Bob should arrive at the side of Alice.

### The tunnel

If sender and receiver are not on the same network it might be needed to do [Network Addres Translation (NAT) and port forwarding](https://www.geeksforgeeks.org/network-address-translation-nat/). Alternatively an ssh tunnel can be used to forward local tcp connections to a remote location. So on the sender the following command would send the incoming audio to a local port:

`ffmpeg -f avfoundation -i ":1"  -acodec libmp3lame -ab 196k -f mpegts tcp://192.168.x.x:12345`

The connection to the receiver can be made using a local port forwarding tunnel. With ssh the TCP traffic on port 12345 is forwarded to the remote receiver via an intermediary (remote) host using the following command:

    ssh -v -L 12345:192.168.x.x:12345 user@host -N


---

## [Using Java LMDB on Apple Sillicon or other unsupported platforms](https://0110.be/posts/Using_Java_LMDB_on_Apple_Sillicon_or_other_unsupported_platforms.md)

- Published: 2022-05-25T00:00:00Z
- Updated: 2022-05-31T07:28:10Z
- Author: Joren
- ID: 491
- Canonical: https://0110.be/posts/Using_Java_LMDB_on_Apple_Sillicon_or_other_unsupported_platforms

- Tags: [Code](https://0110.be/tags/Code.md), [UGent](https://0110.be/tags/UGent.md)

<img src="https://0110.be/files/attachments/491/lmdb.png" style="float:right;margin-left:5px;width:150px"> [LMDB](http://www.lmdb.tech/doc/) is a fast key value store, ideal to store and query sorted data with small keys and values. LMDB is a pure C library but often used from other programming languages via some type of bindings. These bindings are 'bridges' between languages and are automatically present on supported platform. On new or unsupported platforms, however, you need to build a this bridge yourself.

This blog post is about getting [java-lmdb](https://github.com/lmdbjava/lmdbjava) working on such unsupported platform: arm64. The arm64 platform is much more popular since the introduction of the Apple silicon - M1 platform. On Apple M1 the default architecture of Docker images is also aarch64.

The java-lmdb project uses [JNR-FFI](https://github.com/jnr/jnr-ffi) in the background. This is only [one of the many ways to bridge Java and other programming languages](https://developer.okta.com/blog/2022/04/08/state-of-ffi-java). The new version of JNR-FFI supports the arm64. Currently, only the ['SNAPSHOT' version of java-lmdb](https://oss.sonatype.org/content/repositories/snapshots/org/lmdbjava/) uses this version. So the dependencies need to be changed to e.g. (when using Gradle):

    repositories {
        mavenCentral()
        maven { url 'https://oss.sonatype.org/content/repositories/snapshots' }
    }

    dependencies {
        implementation group: 'org.lmdbjava', name: 'lmdbjava', version: '0.8.3-SNAPSHOT'
    }

Next you need to build the `lmdb` library for your platform and copy it to a location where Java looks for it. This only works when compilers are already available on your system. In macOS you might need to install the XCode command line tools:

    #xcode-select --install
    git clone --depth 1 https://git.openldap.org/openldap/openldap.git
    cd openldap/libraries/liblmdb
    make -e SOEXT=.dylib
    cp  liblmdb.dylib ~/Library/Java/Extensions

On Debian `aarch64` the procedure is similar but a different extension is used (`.so`):

    #apt install build-essential
    git clone --depth 1 https://git.openldap.org/openldap/openldap.git
    cd openldap/libraries/liblmdb
    make
    mv liblmdb.so /lib

Finally, to use the library in a JAR-file is might be needed to allow <code>lmdbjava</code> to access some parts of the JRE:

    java -jar your_jar.jar --add-opens=java.base/java.nio=ALL-UNNAMED --add-opens=java.base/sun.nio.ch=ALL-UNNAMED

A very similar setup was needed for the [docker version of Panako](https://github.com/JorenSix/Panako/blob/master/resources/scripts/Dockerfile).


---

## [The Augmented Movement Platform For Embodied Learning (AMPEL)](https://0110.be/posts/The_Augmented_Movement_Platform_For_Embodied_Learning_%28AMPEL%29.md)

- Published: 2022-04-22T00:00:00Z
- Updated: 2022-04-22T12:11:18Z
- Author: Joren
- ID: 489
- Canonical: https://0110.be/posts/The_Augmented_Movement_Platform_For_Embodied_Learning_%28AMPEL%29

- Tags: [UGent](https://0110.be/tags/UGent.md)

I have been lucky to be part of a fruitful interdisciplinary scientific collaboration around AMPEL: '*The Augmented Movement Platform For Embodied Learning*'. The recent publication of [an article](https://nyaspubs.onlinelibrary.wiley.com/doi/10.1111/nyas.14777) is an ideal occasion to give a glimpse behind the scenes.

<div style="float:right;width:280px;margin:12px;">
<img src="https://0110.be/files/photos/489/slider_schema.webp" style="width:275px"><br><small>Fig: Schematic representation of AMPEL, a floor with interactive tiles.</small>

</div>
Around 2016 the idea arose to search for new potential rehabilitation approaches for persons with multiple sclerosis. Multiple sclerosis causes problems, in varying degrees, with both motor and cognitive function. Common rehabilitation approaches either work on motor or cognitive function. The idea (by Lousin Moumdjian, Marc Leman, Peter Feys) was to *combine both motor and cognitive rehabilitation* in a single combined 'embodied learning' paradigm.

After some discussion we wanted to perform a combined short-term memory and walking task. First the participants would be presented with a target trajectory which would then be performed by walking. During walking we would modulate feedback types (melodic, sounds or visual). To this end, an 'intelligent floor' device was needed that was able to present a target trajectory, register a performed trajectory and provide several types of feedback. After a search for off-the-shelf solutions it became clear that a custom hard-and-software platform was required.

After a great deal of cardboard prototyping we settled on a design consisting of interactive tiles. Thomas Vervust of [UGhent NamiFab](https://www.ugent.be/namifab/en) designed a PCB with force sensitive resistors (FSR) on the bottom and RGB LED's on top. Ivan Schepers provided practical insights during prototyping and developed the hardware around the interactive tiles. I was responsible for programming the system. Custom software was developed for the tiles, a controller to drive the tiles and to run and record experiments. Finally the system was moved to a hospital where the experiments took place. To know more about the exact experiments, please read the following three publications on AMPEL:

1.  [*The Augmented Movement Platform For Embodied Learning (AMPEL): development and reliability* - 2021](https://link.springer.com/article/10.1007/s12193-020-00354-8) <br>Moumdjian, L., Vervust, T., Six, J., Schepers, I., Lesaffre, M., Feys, P., & Leman, M. <br>This article details the rationale behind AMPEL and provides technical details and reliability measurements. It was published in the Journal on Multimodal User Interfaces <br><br>

2.  [*Embodied learning in multiple sclerosis using melodic, sound, and visual feedback: a potential rehabilitation approach* - 2022](https://nyaspubs.onlinelibrary.wiley.com/doi/10.1111/nyas.14777) <br>Moumdjian, L., Six, J., Veldkamp, R., Geys, J., Van Der Linden, C., Goetschalckx, M., Van Nieuwenhoven, J., Bosmans, I., Leman, M. and Feys, P. <br>The main AMPEL study which presents the rehabilitation potential. This work was published in the Annals of the New York Academy of Sciences.<br><br>

3.  [*Motor sequence learning in a goal-directed stepping task in persons with multiple sclerosis: a pilot study*\
    - 2022](https://nyaspubs.onlinelibrary.wiley.com/doi/10.1111/nyas.14702) <br>Veldkamp, R., Moumdjian, L., van Dun, K., Six, J., Vanbeylen, A., Kos, D. and Feys, P. <br>For this study AMPEL was slightly modified for a reaction time task, showing its flexibility. The participants were asked to step on a tile as quickly as possible after it lit up. They were either knowledgable of the tile trajectory or not. This work was also published in the Annals of the New York Academy of Sciences.

<br>


![AMPEL schematic](https://0110.be/files/photos/489/schema.webp)

![Custom software driving AMPEL](https://0110.be/files/photos/489/screenshot_management_software.webp)

![Early test of AMPEL](https://0110.be/files/photos/489/IMG_20181205_101627.jpg)

![AMPEL ready to use ](https://0110.be/files/photos/489/ample_real.webp)

![First tests of the I2C bus](https://0110.be/files/photos/489/RS1182_DSC01862-lpr.JPG)

![Almost connected the full I2C bus](https://0110.be/files/photos/489/RS1173_DSC01853-lpr.JPG)

![The driving laptop and main Arduino](https://0110.be/files/photos/489/RS1185_DSC01865-lpr.JPG)

![Testing LEDS ](https://0110.be/files/photos/489/RS1186_DSC01866-lpr.JPG)

![Examples of walking paths](https://0110.be/files/photos/489/stepping_task_AMPEL.webp)

---

## [An audio focused ffmpeg build for the web](https://0110.be/posts/An_audio_focused_ffmpeg_build_for_the_web.md)

- Published: 2022-02-24T00:00:00Z
- Updated: 2025-11-29T14:24:41Z
- Author: Joren
- ID: 488
- Canonical: https://0110.be/posts/An_audio_focused_ffmpeg_build_for_the_web

- Tags: [Code](https://0110.be/tags/Code.md), [Command Line Application](https://0110.be/tags/Command%20Line%20Application.md), [UGent](https://0110.be/tags/UGent.md)

I have prepared **an audio focused ffmpeg build for the web** which facilitates browser based audio applications. I have prepared three demos:

1.  [Audio transcoding and playback demo](https://0110.be/attachment/cors/ffmpeg.audio.wasm/transcode.html): converts any media file into audio compatible with the Web Audio API for in-browser playback or analysis.
2.  [High quality time-stretching or pitch-shifting](https://0110.be/attachment/cors/ffmpeg.audio.wasm/pitch_speed_tempo_mod.html): demonstrates how pitch and tempo can be modified independently thanks to the [Rubber Band Library](https://breakfastquay.com/rubberband/audio).
3.  [Basic media info](https://0110.be/attachment/cors/ffmpeg.audio.wasm/basic_media_info.html): gives information about the streams and encodings used in a media file.

<center>
<img src="https://0110.be/files/attachments/488/screen_recording_small.apng" /><br>
<small>Fig: [audio transcodinging in the browser](https://0110.be/attachment/cors/ffmpeg.audio.wasm/transcode.html). A `wav` file is converted to an `mp3`.</small>

</center>
A bit more about the rationale behind this effort: Browsers have become practical platforms for audio processing applications thanks to the combination of [Web Audio API](https://developer.mozilla.org/en-US/docs/Web/API/Web_Audio_API) , performant Javascript environment and [WebAssembly](https://webassembly.org/). Have a look, for example, at [essentia.JS](https://mtg.github.io/essentia.js).

However, browsers only support a small subset of audio formats and container formats. Dealing with many (legacy) audio formats is often a rather painful experience since there are so many media container formats which can contain a surprising variation of audio (and video) encodings. In short, decoding audio for in-browser analysis or playback is often problematic.

Luckily there is [FFmpeg](https://ffmpeg.org) which claims to be *'a complete, cross-platform solution to record, convert and stream audio and video'*. It is, indeed, capable to decode almost any audio encoding known to man from about any container. Additionally, it also contains tools to filter, manipulate, resample, stretch, ... audio. FFmpeg is a must-have when working with audio. It would be ideal to have FFmpeg running in a browser...

Thanks to [WebAssembly](https://webassembly.org/) ffmpeg can be compiled for use in the browser. There have been [efforts](https://github.com/ffmpegwasm/ffmpeg.wasm-core) [to](https://github.com/ffmpegwasm/ffmpeg.wasm) [get](https://github.com/wide-video/ffmpeg-wasm) ffmpeg working in the browser. These efforts have been focusing on the complete ffmpeg suite. Now I have prepared **an audio focused ffmpeg build for the web** based on these efforts. I have selected only audio parts which makes the resulting .wasm binary four to five times smaller (from \~20MB to \~5MB). I also provided a simplified Javascript wrapper. The project brings audio decoding to the browser but also audio filtering, transcoding, pitch-shifting, sample rate conversions, audio channel manipulation, and so forth. It is also capable to extract audio streams from video container formats.

Next to the pure functionality of ffmpeg there are general advantages to run audio analysis software in the browser at client-side:

-   **Ease-of-use**: no software needs to be installed. The runtime comes with a compatible browser.
-   **Privacy**: Since media files are not transferred it is impossible for the system running the service to make unauthorised copies of these files. There is no need to trust the service since all processing happens locally, in the browser.
-   **Speed**: Downloading and especially uploading large media files takes a while. When files are kept locally, processing can start immediately and no time is wasted sending bytes over the internet. This results in a snappy user experience.
-   **Computational load**: the computational load of transcoding is distributed over the clients and not centralised on a (single) server. The server does not do any computing and only serves static files, so it can handle as many concurrent clients as its bandwidth allows.

Check out the [audio focused ffmpeg build for the web](https://github.com/JorenSix/ffmpeg.audio.wasm) on GitHub.


---

## [pffft.wasm: an FFT library for the web](https://0110.be/posts/pffft.wasm%3A_an_FFT_library_for_the_web.md)

- Published: 2022-02-10T00:00:00Z
- Updated: 2022-06-30T18:55:49Z
- Author: Joren
- ID: 487
- Canonical: https://0110.be/posts/pffft.wasm%3A_an_FFT_library_for_the_web

- Tags: [Code](https://0110.be/tags/Code.md), [UGent](https://0110.be/tags/UGent.md)

[PFFFT](https://bitbucket.org/jpommier/pffft/src/master/README.md) is a small, pretty fast FFT library programmed in C with a BSD-like license. I have taken it upon myself to compile a WebAssembly version of PFFFT to make it available for browsers and node.js environments. It is called [pffft.wasm](https://github.com/JorenSix/pffft.wasm) and available on GitHub.

The pffft.wasm library comes in two flavours. One is compiled with [SIMD](https://en.wikipedia.org/wiki/Single_instruction,_multiple_data) instructions while the other comes without these instructions. SIMD stands for 'single instruction, multiple data' and does what it advertises: in a single step it processes multiple datapoints. The aim of SIMD is to make calculations several times faster. Especially for workloads where the same calculations are repeated over and over again on similar data, SIMD optimisation is relevant. FFT calculation is such a workload.

Evidently the SIMD version is much faster but there is no need to take my word for it. Below you can benchmark the SIMD version of pffft.wasm and compare it with the non-SIMD version on your machine. A pure Javascript FFT library called [FFT.js](https://github.com/indutny/fft.js/) serves as a baseline.

<iframe src="https://0110.be/attachment/cors/pffft.wasm/benchmark_iframe.html" style="border:none;width:100%;height:350px;">
</iframe>
When running the same benchmark on Firefox and on Chrome it becomes clear that FFT.js on Chrome is *about twice as fast* thanks to its superior Javascript engine for this workload. The performance of the WebAssembly versions in Chrome and Firefox is nearly identical. Safari unfortunately does not (yet) support SIMD WebAssembly binaries and fails to complete the benchmark.

The source code, the limitations and other info can be found at the [pffft.wasm GitHub repository](https://github.com/JorenSix/pffft.wasm)

Edit: [PulseFFT](https://github.com/AWSM-WASM/PulseFFT) might be of interest as well: a (as far as I can tell non-SIMD) WASM version of KissFFT.

<br>


![STFT calculated with pffft.wasm](https://0110.be/files/photos/487/stft.png)

![Benchmark pffft.wasm - Chrome on an Apple M1 Pro chip](https://0110.be/files/photos/487/pffft_benchmark.png)

![Benchmark pffft.wasm - Chrome on an 2010 Macbook Air](https://0110.be/files/photos/487/chrome_macbook_air_2010_2ghz.png)

![Benchmark pffft.wasm - Firefox on a 2010 Macbook Air](https://0110.be/files/photos/487/firefox_macbook_air_2010_2ghz.png)

---

## [Panako 2.0 - Updates for an acoustic fingerprinting system](https://0110.be/posts/Panako_2.0_-_Updates_for_an_acoustic_fingerprinting_system.md)

- Published: 2021-11-07T00:00:00Z
- Updated: 2025-11-29T14:26:10Z
- Author: Joren
- ID: 498
- Canonical: https://0110.be/posts/Panako_2.0_-_Updates_for_an_acoustic_fingerprinting_system

- Tags: [ISMIR](https://0110.be/tags/ISMIR.md), [Panako](https://0110.be/tags/Panako.md), [UGent](https://0110.be/tags/UGent.md)

At the online [ISMIR 2021 conference](https://ismir2021.ismir.net/lbd/) I have presented [updates to Panako](https://0110.be/files/attachments/498/2021.ismir-lbd-panako-updates.pdf), an audio fingerprinting system:

> *This work presents updates to Panako, an acoustic fingerprinting system that was introduced at ISMIR 2014. The notable feature of Panako is that it matches queries even after a speedup, time-stretch or pitch-shift. It is freely available and has no problems indexing and querying 100k sea shanties. The updates presented here improve query performance significantly and allow a wider range of time-stretch, pitch-shift and speed-up factors: e.g. the top 1 true positive rate for 20s query that were sped up by 10 percent increased from 18% to 83% from the 2014 version of Panako to the new version. The aim of this short write-up is to reintroduce Panako, evaluate the improvements and highlight two techniques with wider applicability. The first of the two techniques is the use of a constant-Q non-stationary Gabor transform: a fast, reversible, fine-grained spectral transform which can be used as a front-end for many MIR tasks. The second is how near-exact hashing is used in combination with a persistent B-Tree to allow some margin of error while maintaining reasonable query speeds.*

Together with the paper there is also a [poster](https://0110.be/files/attachments/498/panako_updates_poster.pdf) and a short video presentation which explains the work:

<center>
<video style="width:50%" controls preload="none">
<source src="https://0110.be/files/attachments/498/poster-movie.mp4" type="video/mp4">
</video>
</center>


- [2021.ismir-lbd-panako-updates.pdf](https://0110.be/files/attachments/498/2021.ismir-lbd-panako-updates.pdf)

- [panako\_updates\_poster.pdf](https://0110.be/files/attachments/498/panako_updates_poster.pdf)

- [poster-movie.mp4](https://0110.be/files/attachments/498/poster-movie.mp4)

---

## [Decoding LTC and  SMPTE on Teensy - Now using interrupts](https://0110.be/posts/Decoding_LTC_and__SMPTE_on_Teensy_-_Now_using_interrupts.md)

- Published: 2021-09-24T00:00:00Z
- Updated: 2023-03-15T10:09:13Z
- Author: Joren
- ID: 484
- Canonical: https://0110.be/posts/Decoding_LTC_and__SMPTE_on_Teensy_-_Now_using_interrupts

- Tags: [Code](https://0110.be/tags/Code.md), [UGent](https://0110.be/tags/UGent.md)

Have you ever found yourself wondering how to build an accurate, low-latency [LTC](https://en.wikipedia.org/wiki/Linear_timecode) decoder with a common micro-controller? Well! Wonder no more and read on! Or, stop reading and do go read something that is more appealing to your predispositions.

[SMPTE timecodes](https://en.wikipedia.org/wiki/SMPTE_timecode) were originally used to synchronize audio and video material. SMPTE timecode data is often encoded into audio using LTC or linear time code. This special audio stream can be recorded together with other audio and video material. By decoding the LTC audio afterwards and working back to SMPTE timecodes, synchronization of multiple camera angles and audio material becomes straightforward. This concept tagging data streams with SMPTE timecodes is also used for other types of data.

<center>
<!-- center tags, only used by grey (pepper and salt) beards but they still work :) -->
<img style="width:70%" alt="Decoding LTC data" src="https://0110.be/files/attachments/484/ltc_decoding.svg"><br>\
<small>Fig: LTC is a 'self-clocking' protocol for which a period can be found automatically. Once the period is found, transitions within the period are counted. A period with a transition translates to a 1, a period without any transitions to a 0.</small>

</center>
SMPTE timecodes supports up to 30 frames per second and this resolution might not be sufficient for some data streams. It helps if the frames could be split up and 60 or 120 frames per second could be generated. With a low latency LTC decoder it would be possible to support this case and, for example, provide four pulses for every SMPTE frame. To be more precise: a SMPTE frame consists of 80 bits and in this case we would send a pulse exactly when decoding bit 0, bit 20, bit 40 and bit 60. We would then be able to sample at 120Hz while staying in sync with the SMPTE.

My [first attempt](https://0110.be/posts/LTC_-_SMPTE_Decoder_on_Teensy) was to treat the signal like audio and use a ready built library for [LTC audio decoding](https://github.com/x42/libltc) The problem there is that sampling is done which might not exactly match the SMPTE bit transition period and relatively large buffers are used to decode LTC. The bit exact decoding is not possible using this method: the latency is too large, the method also uses excessive computational power and memory.

<div style="float:right">
<center>
<img alt="Biasing circuit to offset voltages from zero centred to having a bias" src="https://0110.be/files/attachments/484/teensy_biasing_schema.svg"><br><small>Fig: Biasing circuit to offset voltages</small>

</center>
</div>
In my second attempt, the current iteration, interrupts are used to detect rising and falling edges in the LTC stream. By counting the number of microseconds between these edges a bit string is constructed. Effectively decoding LTC without any wasted computational power or memory and at a very low latency. If the LTC stream is well-formed, following each incoming bit and reacting to it becomes straightforward. Finally, after gently massaging the LTC bit string, SMPTE timecodes ooze out of the system at a low latency.

I have implemented [a low latency LTC and SMPTE timecode data decoder](https://github.com/ArtScienceLab/ARDUINO_LTC_DECODER/blob/master/LTCInterruptDecoder/LTCInterruptDecoder.ino) for a Teensy microcontroller. One of the current limitations is that only 30fps SMPTE without skipped frames is supported. Another limitation is that the precision of the derived 120Hz clock is dependent on the sampling rate of the encoded audio signal: if e.g. only 8000Hz is used, transitions can only be precise up until 125µs. The derived clock will jitter slightly but will not drift.

There is still a slight problem with audio and Teensy input: audio is generally transmitted from ~~1.8V to +1.8V and not~~ as a Teensy would expect - from 0 to 3.3V. To make this change a small [biasing circuit](https://electronics.stackexchange.com/questions/445142/how-can-i-vertically-shift-the-voltage-of-a-zero-centered-signal-such-that-i-can) is placed before the Teensy input. In my case two 100k resistors and a 0.1uF capacitor worked best. The interrupt is relatively robust against signals that are a clipping (outside the 0 - 3.3V) or slightly too silent. If the signal becomes too small LTC decoding obviously fails.

For more information and updates see the [GitHub Repository for the low latency LTC and SMPTE timecode data decoder](https://github.com/ArtScienceLab/ARDUINO_LTC_DECODER/blob/master/LTCInterruptDecoder/LTCInterruptDecoder.ino).


---

## [Updates for Panako - an acoustic fingerprinting system](https://0110.be/posts/Updates_for_Panako_-_an_acoustic_fingerprinting_system.md)

- Published: 2021-07-11T00:00:00Z
- Updated: 2025-11-29T14:28:25Z
- Author: Joren
- ID: 482
- Canonical: https://0110.be/posts/Updates_for_Panako_-_an_acoustic_fingerprinting_system

- Tags: [Code](https://0110.be/tags/Code.md), [UGent](https://0110.be/tags/UGent.md)

Panako is an acoustic fingerprinting system I developed a couple of years ago. With acoustic fingerprinting systems it is possible to find duplicates in digital music archives and compare meta-data or identify unlabelled audio fragments. In the margins of my post-doc project working with large music archives, I have found the time to update Panako significantly. The *updates simplify, improve and speed up* Panako.

<center>
<img src="https://0110.be/files/attachments/482/general_acoustic_fingerprinting_schema.svg" alt="General content based audio search scheme" style="width:80%"/>\
<small>Fig. General content based audio search scheme.</small>

</center>
The main algorithms are simplified. There is also a reduction of dependencies and a refocus to core functionality. This also simplifies building the software. The retrieval characteristics are improved, mainly thanks to the use of a fine-grained Gabor transform. Also new is the near-exact hashing construct which helps with off-by-one issues when matching time bins. The key-value store used is now [LMDB](http://www.lmdb.tech/doc/), which speeds up the query performance of Panako significantly. The updates should make Panako stand the test of time somewhat better.

<div class="post-meta" style="text-align:center;display:block;float:right;margin-left:10px;width:50%;" >
<img src="https://0110.be/files/attachments/482/tp_rate.svg" style="max-width:500px;">\
<small><b>Fig.</b> The top one true positive rate for 20s query fragments. The audio playback is speed modified from 84 to 116% with respect to the indexed reference audio. The original query length is 20s, if it is slowed down by 10% it takes, evidently, 22s. Note the improvement of the 2021 version of Panako (blue) vs the 2014 version (light-gray). As a baseline the standard algorithm (wang 2003) is included as well. For the 2021 Panako algorithm, audio recognition performance suffers (below 80) when playback speed is changed more than 10.</small>

</div>
A more complete list of updates can be found below and on the [Panako GitHub repository](https://github.com/JorenSix/Panako):

<blockquote>
<i>
-   The number of dependencies has been drastically cut by removing support for multiple key-value stores.
-   The key-value store has been changed to a faster and simpler system (from [MapDB](https://mapdb.org) to [LMDB](http://www.lmdb.tech/doc)).
-   The SyncSink functionality has been moved to another project (with Panako as dependency).
-   The main algorithms have been replaced with simpler and better working versions:
    -   Olaf is a new implementation of the classic Shazam algorithm.
    -   The algoritm described in the Panako paper was also replaced. The core ideas are still the same. The main change is the use of a [Gabor transform](https://en.wikipedia.org/wiki/Gabor_transform) to go from time domain to the spectral domain (previously a constant-q transform was used). The gabor transform is implemented by [JGaborator](https://github.com/JorenSix/JGaborator) which in turn relies on [The Gaborator](https://gaborator.com/) C library via JNI.
-   Folder structure has been simplified.
-   The UI which was mainly used for debugging has been removed.
-   A new set of helper scripts are added in the `scripts` directory. They help with evaluation, parsing results, checking results, building panako, creating documentation,...
-   Changed the default panako location to \~/.panako, so users can install and use panako more easily (without need for sudo rights)
</i>
</blockquote>

<div style="text-align:center">
<img src="https://0110.be/files/attachments/482/panako_interactive_session.svg" alt="An interactive CLI session with Panako"/><br>
<small>Fig: An interactive CLI session with Panako.</small>
</div>


---

## [SyncSink - Synchronize media by aligning audio](https://0110.be/posts/SyncSink_-_Synchronize_media_by_aligning_audio.md)

- Published: 2021-06-10T00:00:00Z
- Updated: 2025-11-29T14:30:39Z
- Author: Joren
- ID: 483
- Canonical: https://0110.be/posts/SyncSink_-_Synchronize_media_by_aligning_audio

- Tags: [Code](https://0110.be/tags/Code.md), [Music Information Retrieval](https://0110.be/tags/Music%20Information%20Retrieval.md), [Panako](https://0110.be/tags/Panako.md), [UGent](https://0110.be/tags/UGent.md)

I have just released a new version of SyncSink. SyncSink is a tool to synchronize media files with shared audio. It is ideal to synchronize video captured by multiple cameras or audio captured by many microphones. It finds a rough alignment between audio captured from the same event and subsequently refines that offset with a crosscorrelation step. Below you can see SyncSink in action or you can try out [SyncSink](syncsink-1.0.jar) (you will need ffmpeg and Java installed on your system).

SyncSink used to be part of the [Panako acoustic fingerprinting system](https://github.com/JorenSix/Panako) but I decided that it was better to keep the Panako package focused and made a separate repository for SyncSink. More information can be found at the [SyncSink GiHub repo](https://github.com/JorenSix/SyncSink)

<blockquote>
<i>SyncSink is a tool to synchronize media files with shared audio. SyncSink matches and aligns shared audio and determines offsets in seconds. With these precise offsets it becomes trivial to sync files. SyncSink is, for example, used to synchronize video files: when you have many video captures of the same event, the audio attached to these video captures is used to align and sync multiple (independently operated) cameras.

Evidently, SyncSink can also synchronize audio captured from many (independent) microphones if some environmental sound is shared (leaked in) the each recording.</i>

<center>
<img src="https://0110.be/files/attachments/483/SyncSink-synchronizing_audio.gif"><br>
<small>Fig: SyncSink in action: syncing some audio files</small>
</center>
</blockquote>


- [SyncSink-synchronizing\_audio.gif](https://0110.be/files/attachments/483/SyncSink-synchronizing_audio.gif)

- [syncsink-1.0.jar](https://0110.be/files/attachments/483/syncsink-1.0.jar)

---

## [Calling JNI code from multiple Java threads:  sharing state](https://0110.be/posts/Calling_JNI_code_from_multiple_Java_threads%3A__sharing_state.md)

- Published: 2021-06-01T00:00:00Z
- Updated: 2025-11-29T14:32:05Z
- Author: Joren
- ID: 481
- Canonical: https://0110.be/posts/Calling_JNI_code_from_multiple_Java_threads%3A__sharing_state

- Tags: [Code](https://0110.be/tags/Code.md), [UGent](https://0110.be/tags/UGent.md)

<div style="float:right">
<svg xmlns="http://www.w3.org/2000/svg" xmlns:xlink="http://www.w3.org/1999/xlink" version="1.1" width="301px" height="161px" viewBox="-0.5 -0.5 301 161" content="&lt;mxfile host=&quot;Electron&quot; modified=&quot;2021-06-01T13:10:51.965Z&quot; agent=&quot;Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) draw.io/12.9.3 Chrome/80.0.3987.158 Electron/8.2.0 Safari/537.36&quot; etag=&quot;bNxISs8bfODPtaRI5J9W&quot; version=&quot;12.9.3&quot; type=&quot;device&quot;&gt;&lt;diagram id=&quot;DG2niK77s8uMG9i-pMya&quot; name=&quot;Page-1&quot;&gt;3Vptk+IoEP41fjwrgHn7uM7s3t1u3dVWzVbd7kcmQUNtDFOIo96vPzBgQmAc9YxmzUxZoYEOPP2ku4GM0MNi8zvHL8VfLCflCAb5ZoQeRxCGMJa/SrCtBRMQ1II5p3ktAo3gif5LtNA0W9GcLK2GgrFS0BdbmLGqIpmwZJhztrabzVhpP/UFz4kjeMpw6Ur/obkoamlipqXkfxA6L8yTQZTWNQtsGuuZLAucs3VLhD6O0ANnTNR3i80DKRV2Bpe636c3avcD46QSx3SYiRh+2ZbVNBc05OWX5Odi/htKajWvuFzpGevRiq2BgLNVlROlBYzQdF1QQZ5ecKZq19LmUlaIRamrZ7QsH1jJ+K4vyjFJZpmULwVnP0mrJsoS8jyTNRlbUNniMVCtDEaqMC/xcqnv9SgJF2Tz5vzBHlXJRsIWRPCtbKI7wFQbQjMRJLq8buwKJlpWtG1qLIg1l+Z73Q3c8kYjfgr6cc/ohyTJJz70E/iMougQ+hdAHKDBIR6dAnhwMuCzJCOZl+7PSTgJu6zuE3w4scGHQTAOHfgTD/pJX+C7bAcu+gVbPK+WpyMfqj+vo9ldqgerREteX5cBWz7EgGu4nsYeuE3YaMMNoWl4ccABcBD/jF+xlHwrOMH50oFfQiBsnG08K1aRDvhahEs6r2Qxk5ARKZ8qQKkMpB90xYLmuXqM17D2S3cJ39OhP4g8vifwsL831wOgYwypaLr7lybAgtyxORAcnDnS90MBqfIPKodUOCq3rZy1bQzMRafFt4JWpuoTLU1TsqHie+v+h8J2HOrS40ZDvStsLdxJ7mSoHdTlmNmKZ+TAbOvJueZpwR960DcyTkos6Ks9Dp9J9BO+MipH2LyMcScWgY6Kevy6VzuH7ShyghoIus5T4j4nwlG148h+4ufTBrou1cObspTrk7derqEGMcdKKBjHnhjmS9jS3mKYgfdg1nAngHtTtGvj7abInhz5TrI0uSgcw5unadBNjNH9Qh6mt8Y7PALuO/EoEbi5Pxl+niXB5dvv7UKrlyo23XalPvIzs7e5y14ONIz89r9OHqcCVNBcaNJxpsEYTtLmstUfneOFgaP2ujme2bYaMGUvSL1wwIwKHdMfzSKU2qri4LocctcJx3HIRxybWue6K8OzYJymkcW1JH7HQcrCV8KpBEXtJVyagdGRvq9OE25FVRTYfgmhDqOOpSbqOrjuhmvf1HRXVEOh5o1c4NEEvGn0RXKxEsWt+Bp2aXSus0Sxk6VLZfDKURcNlpbKY1rMBAD9Cg7ztv6yc/r3P4K5h5+usr7pOTmJnnobPsfLYn9462MkGAfQJqUOxodoqUrn02vgawwncTt3rzgKDivqmzDuNsNFCHOmCxt2ZL0t5SYdyiVnUg4hMA5TZwX8doDum4JHfO7wa1PwdpRxD6KCc1kD39fVN1HcTbrPf/8pBVNOlQ26pLmbc2nY/UTJ2LWHc2lZbL72qy3XfDKJPv4H&lt;/diagram&gt;&lt;/mxfile&gt;" style="background-color: rgb(255, 255, 255);">
<defs/><g><rect x="160" y="40" width="140" height="120" rx="18" ry="18" fill="#dae8fc" stroke="#6c8ebf" pointer-events="all"/><rect x="0" y="40" width="140" height="120" rx="18" ry="18" fill="#d5e8d4" stroke="#82b366" pointer-events="all"/><rect x="110" y="60.5" width="80" height="80" fill="#f8cecc" stroke="#b85450" pointer-events="all"/><path d="M 250 57.5 L 263.5 68.75 L 250 80 L 236.5 68.75 Z" fill="#f5f5f5" stroke="#666666" stroke-miterlimit="10" pointer-events="all"/><g transform="translate(-0.5 -0.5)">

<switch>
<foreignObject style="overflow: visible; text-align: left;" pointer-events="none" width="100%" height="100%" requiredFeatures="http://www.w3.org/TR/SVG11/feature#Extensibility">

<div xmlns="http://www.w3.org/1999/xhtml" style="display: flex; align-items: unsafe center; justify-content: unsafe center; width: 25px; height: 1px; padding-top: 69px; margin-left: 238px;">
<div style="box-sizing: border-box; font-size: 0; text-align: center; ">
<div style="display: inline-block; font-size: 12px; font-family: Helvetica; color: #333333; line-height: 1.2; pointer-events: all; white-space: normal; word-wrap: normal; ">
1

</div>
</div>
</div>
</foreignObject><text x="250" y="72" fill="#333333" font-family="Helvetica" font-size="12px" text-anchor="middle">1</text>

</switch>
</g><rect x="10" y="20" width="100" height="20" fill="none" stroke="none" pointer-events="all"/><g transform="translate(-0.5 -0.5)">

<switch>
<foreignObject style="overflow: visible; text-align: left;" pointer-events="none" width="100%" height="100%" requiredFeatures="http://www.w3.org/TR/SVG11/feature#Extensibility">

<div xmlns="http://www.w3.org/1999/xhtml" style="display: flex; align-items: unsafe center; justify-content: unsafe center; width: 98px; height: 1px; padding-top: 30px; margin-left: 11px;">
<div style="box-sizing: border-box; font-size: 0; text-align: center; ">
<div style="display: inline-block; font-size: 12px; font-family: Helvetica; color: #000000; line-height: 1.2; pointer-events: all; white-space: normal; word-wrap: normal; ">
Java Threads

</div>
</div>
</div>
</foreignObject><text x="60" y="34" fill="#000000" font-family="Helvetica" font-size="12px" text-anchor="middle">Java Threads</text>

</switch>
</g><rect x="190" y="20" width="100" height="20" fill="none" stroke="none" pointer-events="all"/><g transform="translate(-0.5 -0.5)">

<switch>
<foreignObject style="overflow: visible; text-align: left;" pointer-events="none" width="100%" height="100%" requiredFeatures="http://www.w3.org/TR/SVG11/feature#Extensibility">

<div xmlns="http://www.w3.org/1999/xhtml" style="display: flex; align-items: unsafe center; justify-content: unsafe center; width: 98px; height: 1px; padding-top: 30px; margin-left: 191px;">
<div style="box-sizing: border-box; font-size: 0; text-align: center; ">
<div style="display: inline-block; font-size: 12px; font-family: Helvetica; color: #000000; line-height: 1.2; pointer-events: all; white-space: normal; word-wrap: normal; ">
C states

</div>
</div>
</div>
</foreignObject><text x="240" y="34" fill="#000000" font-family="Helvetica" font-size="12px" text-anchor="middle">C states</text>

</switch>
</g><path d="M 66.37 70.28 L 103.63 70.47" fill="none" stroke="#000000" stroke-miterlimit="10" pointer-events="stroke"/><path d="M 61.12 70.26 L 68.13 67.96 L 66.37 70.28 L 68.11 72.62 Z" fill="#000000" stroke="#000000" stroke-miterlimit="10" pointer-events="all"/><path d="M 108.88 70.49 L 101.86 73.96 L 103.63 70.47 L 101.9 66.96 Z" fill="#000000" stroke="#000000" stroke-miterlimit="10" pointer-events="all"/><ellipse cx="50" cy="100.5" rx="10" ry="9.75" fill="#f5f5f5" stroke="#666666" pointer-events="all"/><g transform="translate(-0.5 -0.5)">

<switch>
<foreignObject style="overflow: visible; text-align: left;" pointer-events="none" width="100%" height="100%" requiredFeatures="http://www.w3.org/TR/SVG11/feature#Extensibility">

<div xmlns="http://www.w3.org/1999/xhtml" style="display: flex; align-items: unsafe center; justify-content: unsafe center; width: 18px; height: 1px; padding-top: 101px; margin-left: 41px;">
<div style="box-sizing: border-box; font-size: 0; text-align: center; ">
<div style="display: inline-block; font-size: 12px; font-family: Helvetica; color: #333333; line-height: 1.2; pointer-events: all; white-space: normal; word-wrap: normal; ">
2

</div>
</div>
</div>
</foreignObject><text x="50" y="104" fill="#333333" font-family="Helvetica" font-size="12px" text-anchor="middle">2</text>

</switch>
</g><ellipse cx="50" cy="70.25" rx="10" ry="9.75" fill="#f5f5f5" stroke="#666666" pointer-events="all"/><g transform="translate(-0.5 -0.5)">

<switch>
<foreignObject style="overflow: visible; text-align: left;" pointer-events="none" width="100%" height="100%" requiredFeatures="http://www.w3.org/TR/SVG11/feature#Extensibility">

<div xmlns="http://www.w3.org/1999/xhtml" style="display: flex; align-items: unsafe center; justify-content: unsafe center; width: 18px; height: 1px; padding-top: 70px; margin-left: 41px;">
<div style="box-sizing: border-box; font-size: 0; text-align: center; ">
<div style="display: inline-block; font-size: 12px; font-family: Helvetica; color: #333333; line-height: 1.2; pointer-events: all; white-space: normal; word-wrap: normal; ">
1

</div>
</div>
</div>
</foreignObject><text x="50" y="74" fill="#333333" font-family="Helvetica" font-size="12px" text-anchor="middle">1</text>

</switch>
</g><path d="M 250 89.25 L 263.5 100.5 L 250 111.75 L 236.5 100.5 Z" fill="#f5f5f5" stroke="#666666" stroke-miterlimit="10" pointer-events="all"/><g transform="translate(-0.5 -0.5)">

<switch>
<foreignObject style="overflow: visible; text-align: left;" pointer-events="none" width="100%" height="100%" requiredFeatures="http://www.w3.org/TR/SVG11/feature#Extensibility">

<div xmlns="http://www.w3.org/1999/xhtml" style="display: flex; align-items: unsafe center; justify-content: unsafe center; width: 25px; height: 1px; padding-top: 101px; margin-left: 238px;">
<div style="box-sizing: border-box; font-size: 0; text-align: center; ">
<div style="display: inline-block; font-size: 12px; font-family: Helvetica; color: #333333; line-height: 1.2; pointer-events: all; white-space: normal; word-wrap: normal; ">
2

</div>
</div>
</div>
</foreignObject><text x="250" y="104" fill="#333333" font-family="Helvetica" font-size="12px" text-anchor="middle">2</text>

</switch>
</g><path d="M 250 119 L 263.5 130.25 L 250 141.5 L 236.5 130.25 Z" fill="#f5f5f5" stroke="#666666" stroke-miterlimit="10" pointer-events="all"/><g transform="translate(-0.5 -0.5)">

<switch>
<foreignObject style="overflow: visible; text-align: left;" pointer-events="none" width="100%" height="100%" requiredFeatures="http://www.w3.org/TR/SVG11/feature#Extensibility">

<div xmlns="http://www.w3.org/1999/xhtml" style="display: flex; align-items: unsafe center; justify-content: unsafe center; width: 25px; height: 1px; padding-top: 130px; margin-left: 238px;">
<div style="box-sizing: border-box; font-size: 0; text-align: center; ">
<div style="display: inline-block; font-size: 12px; font-family: Helvetica; color: #333333; line-height: 1.2; pointer-events: all; white-space: normal; word-wrap: normal; ">
3

</div>
</div>
</div>
</foreignObject><text x="250" y="134" fill="#333333" font-family="Helvetica" font-size="12px" text-anchor="middle">3</text>

</switch>
</g><ellipse cx="50" cy="130.75" rx="10" ry="9.75" fill="#f5f5f5" stroke="#666666" pointer-events="all"/><g transform="translate(-0.5 -0.5)">

<switch>
<foreignObject style="overflow: visible; text-align: left;" pointer-events="none" width="100%" height="100%" requiredFeatures="http://www.w3.org/TR/SVG11/feature#Extensibility">

<div xmlns="http://www.w3.org/1999/xhtml" style="display: flex; align-items: unsafe center; justify-content: unsafe center; width: 18px; height: 1px; padding-top: 131px; margin-left: 41px;">
<div style="box-sizing: border-box; font-size: 0; text-align: center; ">
<div style="display: inline-block; font-size: 12px; font-family: Helvetica; color: #333333; line-height: 1.2; pointer-events: all; white-space: normal; word-wrap: normal; ">
3

</div>
</div>
</div>
</foreignObject><text x="50" y="134" fill="#333333" font-family="Helvetica" font-size="12px" text-anchor="middle">3</text>

</switch>
</g><path d="M 66.37 100.5 L 103.63 100.5" fill="none" stroke="#000000" stroke-miterlimit="10" pointer-events="stroke"/><path d="M 61.12 100.5 L 68.12 98.17 L 66.37 100.5 L 68.12 102.83 Z" fill="#000000" stroke="#000000" stroke-miterlimit="10" pointer-events="all"/><path d="M 108.88 100.5 L 101.88 104 L 103.63 100.5 L 101.88 97 Z" fill="#000000" stroke="#000000" stroke-miterlimit="10" pointer-events="all"/><path d="M 66.37 130.65 L 102.63 130.1" fill="none" stroke="#000000" stroke-miterlimit="10" pointer-events="stroke"/><path d="M 61.12 130.73 L 68.08 128.29 L 66.37 130.65 L 68.15 132.96 Z" fill="#000000" stroke="#000000" stroke-miterlimit="10" pointer-events="all"/><path d="M 107.88 130.02 L 100.94 133.62 L 102.63 130.1 L 100.83 126.62 Z" fill="#000000" stroke="#000000" stroke-miterlimit="10" pointer-events="all"/><path d="M 196.05 130.47 L 230.13 130.28" fill="none" stroke="#000000" stroke-miterlimit="10" pointer-events="stroke"/><path d="M 190.8 130.49 L 197.78 126.96 L 196.05 130.47 L 197.82 133.96 Z" fill="#000000" stroke="#000000" stroke-miterlimit="10" pointer-events="all"/><path d="M 235.38 130.26 L 228.4 133.79 L 230.13 130.28 L 228.36 126.79 Z" fill="#000000" stroke="#000000" stroke-miterlimit="10" pointer-events="all"/><path d="M 196.37 100.5 L 230.13 100.5" fill="none" stroke="#000000" stroke-miterlimit="10" pointer-events="stroke"/><path d="M 191.12 100.5 L 198.12 97 L 196.37 100.5 L 198.12 104 Z" fill="#000000" stroke="#000000" stroke-miterlimit="10" pointer-events="all"/><path d="M 235.38 100.5 L 228.38 104 L 230.13 100.5 L 228.38 97 Z" fill="#000000" stroke="#000000" stroke-miterlimit="10" pointer-events="all"/><path d="M 195.57 69.43 L 230.13 68.86" fill="none" stroke="#000000" stroke-miterlimit="10" pointer-events="stroke"/><path d="M 190.32 69.52 L 197.26 65.9 L 195.57 69.43 L 197.38 72.9 Z" fill="#000000" stroke="#000000" stroke-miterlimit="10" pointer-events="all"/><path d="M 235.38 68.77 L 228.44 72.39 L 230.13 68.86 L 228.32 65.39 Z" fill="#000000" stroke="#000000" stroke-miterlimit="10" pointer-events="all"/><path d="M 109 131 L 191.6 130.5" fill="none" stroke="#000000" stroke-miterlimit="10" stroke-dasharray="3 3" pointer-events="stroke"/><path d="M 110 100.5 L 190 100.5" fill="none" stroke="#000000" stroke-miterlimit="10" stroke-dasharray="3 3" pointer-events="stroke"/><path d="M 110 70.01 L 190 70.01" fill="none" stroke="#000000" stroke-miterlimit="10" stroke-dasharray="3 3" pointer-events="stroke"/><rect x="100" y="0" width="100" height="20" fill="none" stroke="none" pointer-events="all"/><g transform="translate(-0.5 -0.5)">

<switch>
<foreignObject style="overflow: visible; text-align: left;" pointer-events="none" width="100%" height="100%" requiredFeatures="http://www.w3.org/TR/SVG11/feature#Extensibility">

<div xmlns="http://www.w3.org/1999/xhtml" style="display: flex; align-items: unsafe center; justify-content: unsafe center; width: 98px; height: 1px; padding-top: 10px; margin-left: 101px;">
<div style="box-sizing: border-box; font-size: 0; text-align: center; ">
<div style="display: inline-block; font-size: 12px; font-family: Helvetica; color: #000000; line-height: 1.2; pointer-events: all; white-space: normal; word-wrap: normal; ">
JNI Bridge

</div>
</div>
</div>
</foreignObject>

</switch>
</g></g></svg><br>

<caption>
<small>Mapping Java threads to C states in a JNI bridge</small>

</caption>
</div>
This post deals with the problem of using stateful C code from *multiple Java threads*. With JNI ([Java Native Interface](https://en.wikipedia.org/wiki/Java_Native_Interface)) it is possible to glue C code to a Java environment. There are [many helpful tutorials](https://www3.ntu.edu.sg/home/ehchua/programming/java/JavaNativeInterface.html) on how to call C code and receive results. JNI helps to reuse existing, often highly complex and computationally expensive, C code.

The introductory tutorials often stop once it is made clear how to repackage (simple) datatypes and do not mention threads. It is, however, reasonable to expect JNI code to take into account thread-safety and proper multi-threading. In all but the simplest cases it is not that straightforward to share state at the C side and allow JNI code to be called from multiple Java threads. Incorrectly sharing state can lead to memory leaks and segmentation faults (segfaults) and crashes the application. In what follows, a way to share thread-local state is presented.

It is quite common to have an `init`, `work` and `dispose` method to create a state, use that state and do some work and finally dispose of used resources. Each Java thread independently calls these methods and expects results. These results should not change if multiple Java threads are calling the same methods. In other words: the state should remain Java thread-local. A typical Java class could look like the code below.

With the Java code in mind, the C code should know which Java thread is used and which state needs to be used for the work. Luckily there is a way to find out: [The JNI specification states that each `JNIEnv` is local to a Java thread](https://developer.ibm.com/languages/java/articles/j-jni/). So we can use the `JNIEnv` pointer to identify a thread. This is the idea that is used below.

The code maps a `JNIEnv` pointer to a structure with (any) state information. An unordered map is used for this mapping. There is, however, still a problem: multiple threads can call the init method at once. So multiple threads potentially write to the `unordered_map` at the same time which leads to problems. To prevent this from happening a mutex is used. The mutex, together with a [unique lock](https://www.cplusplus.com/reference/mutex/unique_lock/), makes sure that only a single thread writes to the unordered map. The same holds for the dispose method.

The work method does not need a unique lock since it does not write to the unordered map and reading from multiple threads is no problem.

<div style="width:100%" class="sixfour">
````c
#include <unordered_map>
#include <mutex>

const int DATA_ARRAY_SIZE = 300000 * 2;

struct BridgeState {
    jfloat *data;
};

//A hash map with a JNIEnv * as key and a BridgeState * as value
std::unordered_map<uintptr_t, uintptr_t> stateMap;

//A mutex to ensure that writes to the stateMap are synchronized.
std::mutex stateMutex;

JNIEXPORT jint JNICALL Java_init(JNIEnv *env, jobject object) {
    //Makes sure only one thread writes to the stateMap
    std::unique_lock<std::mutex> lck(stateMutex);

    BridgeState *state = new BridgeState();
    uintptr_t env_addresss = reinterpret_cast<uintptr_t>(env);

    state->data = new jfloat[DATA_ARRAY_SIZE];

    uintptr_t state_addresss = reinterpret_cast<uintptr_t>(state);
    stateMap[env_addresss] = state_addresss;

    return 1;
}

JNIEXPORT jint JNICALL Java_work(JNIEnv *env, jobject object) {
    //get a ref to the state pointer
    uintptr_t env_addresss = reinterpret_cast<uintptr_t>(env);
    BridgeState *state = reinterpret_cast<BridgeState *>(stateMap[env_addresss]);

    //do something with state->data, e.g. calculate the sum
    int sum = 0;
    for (int i = 0; i < DATA_ARRAY_SIZE; i++) {
        state->data[i] = state->data[i] + 1;
        sum += (int)state->data[i];
    }
    return sum;
}

JNIEXPORT jint JNICALL Java_dispose(JNIEnv *env, jobject object) {
    //Makes sure only one thread writes to the stateMap
    std::unique_lock<std::mutex> lck(stateMutex);

    uintptr_t env_addresss = reinterpret_cast<uintptr_t>(env);
    BridgeState *state = reinterpret_cast<BridgeState *>(stateMap[env_addresss]);
    stateMap.erase(env_addresss);

    //cleanup memory
    delete[] state->data;
    delete state;
    return 0;
}
````
</div>

This conceptual code has been lifted from a JNI library doing actual work: The [JGaborator JNI bridge](https://github.com/JorenSix/JGaborator/blob/master/gaborator/jgaborator.cc) . If you need more information on how to compile and use this construct in actual code, please have a look at the [JGaborator GitHub repository](https://github.com/JorenSix/JGaborator)


---

## [JGaborator Updated - Fine grained spectral transforms from Java](https://0110.be/posts/JGaborator_Updated_-_Fine_grained_spectral_transforms_from_Java.md)

- Published: 2021-05-31T00:00:00Z
- Updated: 2021-06-01T10:24:42Z
- Author: Joren
- ID: 480
- Canonical: https://0110.be/posts/JGaborator_Updated_-_Fine_grained_spectral_transforms_from_Java

- Tags: [Code](https://0110.be/tags/Code.md), [UGent](https://0110.be/tags/UGent.md)

I have updated the [JGaborator library](https://github.com/JorenSix/JGaborator). The library calculates fine grained constant-Q spectral representations of audio signals quickly from Java. Such spectral transform can be used for visualisation or as a front-end for audio processing or music information retrieval applications.

The calculation of a [Gabor transform](https://en.wikipedia.org/wiki/Gabor_transform) is done by a C library named [Gaborator](http://gaborator.com). JGaborator provides a Java native interface (JNI) bridge to that library. Thanks to the recent updates, the library is now automatically unpacked which makes it easy to use on supported platforms (intel macOS and x64 Linux).

The new version of JGaborator now also allows multiple Java threads to call the transform. This has the potential to speed up some audio processing chains dramatically.

The visualisation parts of JGaborator also received light touch-ups. Below a number of screenshots can be seen with of spectral representations of several audio files. If you want to try it yourself download the "JGaborator JAR-file":\[JGaborator-0.6.jar\]. Note that it should work only on intel macOS and x64 Linux with [ffmpeg](https://ffmpeg.org/) installed on your path. For other environments, please read and follow the [JGaborator instructions](https://github.com/JorenSix/JGaborator) to get it working.


![Spectral vizualization with JGaborator](https://0110.be/files/photos/480/JGaborator-spectrogram-1.png)

![Spectral vizualization with JGaborator](https://0110.be/files/photos/480/JGaborator-spectrogram-2.png)

![Spectral vizualization with JGaborator](https://0110.be/files/photos/480/JGaborator-spectrogram-3.png)

---

## [Music-based biofeedback to reduce tibial shock in over-ground running: a proof-of-concept study](https://0110.be/posts/Music-based_biofeedback_to_reduce_tibial_shock_in_over-ground_running%3A_a_proof-of-concept_study.md)

- Published: 2021-02-18T00:00:00Z
- Updated: 2021-02-18T14:02:32Z
- Author: Joren
- ID: 479
- Canonical: https://0110.be/posts/Music-based_biofeedback_to_reduce_tibial_shock_in_over-ground_running%3A_a_proof-of-concept_study

- Tags: [UGent](https://0110.be/tags/UGent.md)

<div style="float:right;margin-left:10px;margin-bottom:10px">
<a href="https://www.nature.com/articles/s41598-021-83538-w">\
<img src="https://0110.be/files/attachments/479/2021_Music-based_biofeedback_to_reduce_tibial_shock_in_over-ground_running__a_proof-of-concept_study_Scientific_Reports.png" style="width:160px" ></a>

</div>
For the last couple of years there has been a fruitful collaboration ongoing between the systematic musicology (IPEM) and sports-science departments at Ghent University. IPEM has a rich history of fundamental research on the link between movement and music. In a newly published proof-of-concept study the music-movement link improves running style. The runner is equipped with a musical biofeedback system to lower foot-impact. For more details, see:

<a href="https://www.nature.com/articles/s41598-021-83538-w">Music-based biofeedback to reduce tibial shock in over-ground running: a proof-of-concept study</a>, published in Scientific Reports(2021) by Van den Berghe, P., Lorenzoni, V., Derie, R. et al.

> **Abstract** Methods to reduce impact in distance runners have been proposed based on real-time auditory feedback of tibial acceleration. These methods were developed using treadmill running. In this study, we extend these methods to a more natural environment with a proof-of-concept. We selected ten runners with high tibial shock. They used a music-based biofeedback system with headphones in a running session on an athletic track. The feedback consisted of music superimposed with noise coupled to tibial shock. The music was automatically synchronized to the running cadence. The level of noise could be reduced by reducing the momentary level of tibial shock, thereby providing a more pleasant listening experience. The running speed was controlled between the condition without biofeedback and the condition of biofeedback. The results show that tibial shock decreased by 27% or 2.96 g without guided instructions on gait modification in the biofeedback condition. The reduction in tibial shock did not result in a clear increase in the running cadence. The results indicate that a wearable biofeedback system aids in shock reduction during over-ground running. This paves the way to evaluate and retrain runners in over-ground running programs that target running with less impact through instantaneous auditory feedback on tibial shock.


---

## [ISMIR 2020 - Virtual Conference](https://0110.be/posts/ISMIR_2020_-_Virtual_Conference.md)

- Published: 2020-10-20T00:00:00Z
- Updated: 2025-11-29T14:33:11Z
- Author: Joren
- ID: 477
- Canonical: https://0110.be/posts/ISMIR_2020_-_Virtual_Conference

- Tags: [Command Line Application](https://0110.be/tags/Command%20Line%20Application.md), [Presentation](https://0110.be/tags/Presentation.md), [UGent](https://0110.be/tags/UGent.md)

<img style="float:right;margin:5px" alt="ISMIR 2020 Logo" src="https://0110.be/files/attachments/477/ISMIR2020-logo.png">

From 11-16 October 2020 the latest instalment of the ISMIR conference series was held. Due to the pandemic, the 21st ISMIR conference was the first virtual one. As usual, participants and presenters from around the world joined the conference. For the first time, however, not all participants synchronised their circadian rhythm. By repeating most events with 12h in between, the organisers managed to put together a schedule befitting nearly all participants.

The virtual format had some clear advantages: travel was not needed, so (environmental) cost was low. Attendance fees were lower than usual since no spaces or catering was needed. This democratised the conference experience and attendance reached a record high.

The scientific program of the conference was impressive and varied. It is At the conferences Late Breaking/Demo session I presented [Olaf: Overly Lightweight Acoustic Fingerprinting](https://program.ismir2020.net/lbd_418.html).


<center>
<iframe width="480" height="280" src="https://www.youtube.com/embed/wP29RaQicwE" frameborder="0" allow="picture-in-picture" allowfullscreen>
</iframe>
</center>



---

## [PaPiOM: Patterns in Pitch Organization in Music](https://0110.be/posts/PaPiOM%3A_Patterns_in_Pitch_Organization_in_Music.md)

- Published: 2020-09-30T00:00:00Z
- Updated: 2025-11-29T14:34:31Z
- Author: Joren
- ID: 476
- Canonical: https://0110.be/posts/PaPiOM%3A_Patterns_in_Pitch_Organization_in_Music

- Tags: [Computational ethnomusicology](https://0110.be/tags/Computational%20ethnomusicology.md), [UGent](https://0110.be/tags/UGent.md)

Form the 1st of October 2020 I will start on a new research project. The BOF fund of Ghent University is kind enough to sponsor the project for three years. The abstract is as follows:

> Music is present in every culture in the world. We as a species seem to have an urge to make music. While the diversity of music cultures around the world is phenomenal, they do seem to have patterns in common. Especially for pitch, one of the fundamental building blocks of music, there are strong reasons to believe that there are commonalities amongst cultures on how pitch is organised A better insight in these common patterns may help to answer questions on the definition, origins and evolution of music.

> Common patterns in pitch organisation can be studied from two perspectives. Firstly, the perspective of how humans perceive and make music can be gained from systematic, experimental work. Over the years this has yielded insights in which pitch organisations might be most fit for our perceptual, neurophysiological system. Secondly, these patterns can be observed directly in large-scale, corpus-based, cross-cultural studies which has a potential that is not exploited as of yet.

> During this fellowship a large-scale global corpus with field recordings will be compiled in collaboration. Music Information Retrieval techniques will be employed to describe how pitch is organised in the corpus. More specifically, it will support claims on the use of discrete pitches, octave equivalence, the number of pitch classes in use and the pitch interval structures. The uncovered fundamental properties of pitch will be confronted with findings from experimental work.

Recently I presented the outline of the project with the following slides:

<iframe src="https://0110.be/attachment/presentations/2020.09.PaPiOM-Ghent/index.html" style="width:100%;height:35em">
</iframe>


---

## [Olaf - Acoustic fingerprinting on the ESP32 and in the Browser](https://0110.be/posts/Olaf_-_Acoustic_fingerprinting_on_the_ESP32_and_in_the_Browser.md)

- Published: 2020-08-20T00:00:00Z
- Updated: 2025-11-29T14:35:53Z
- Author: Joren
- ID: 475
- Canonical: https://0110.be/posts/Olaf_-_Acoustic_fingerprinting_on_the_ESP32_and_in_the_Browser

- Tags: [Code](https://0110.be/tags/Code.md), [UGent](https://0110.be/tags/UGent.md)

<img style="width:150px;float:right" alt="Recognition of music." src="https://0110.be/files/attachments/475/web_embedded_olaf.jpg"/> A good year ago I was asked to develop audio recognition technology for an e-costume. The idea was that *lights in the costume would follow a sequence synchronised to a certain song*. Only a single song should trigger the lights, all other music should be ignored. Recognition of music and synchronisation is typically done using audio fingerprinting techniques. The challenge was that the recognition needed to run on a cheap, battery-powered microcontroller with limited CPU and memory. I delivered a prototype but eventually a cheap, battle-tested, off-the-shelf, IP-cleared, alternative was found.

The prototype gathered dust for a while but the idea stuck in my head. With my daughters fourth birthday approaching during the lockdown, I decided to turn the prototype into an over-engineered birthday gift and let an 'Elsa-dress' react to 'Let It Go' from the Frozen soundtrack. With the prototype as a starting point, I ordered an RGB-LED-strip, a beefy Li-Ion Battery, an I2S digital microphone and, of course, an Elsa-dress.

I had an ESP32 microcontroller laying around and used it as the core of the system: it supports [I²S](https://en.wikipedia.org/wiki/I%C2%B2S), has a floating point unit (FPU), is easy to use together with LED strips and has enough memory. The [FPU](https://en.wikipedia.org/wiki/Floating-point_unit) makes it straightforward to use the same code on traditional computers as on embedded devices: fixed-point math can be avoided.

After soldering the components together and with the help from my better half to sew in the LED strip, it all came together. In the video below, the result of our work can be seen. The video first shows a song that should not and is not recognised. Then, "Let It Go" is played and correctly recognised. After the song is stopped, the lights go on for a while and finally stop: this is by design to allow gaps in recognition. Lastly, the song is continued and again correctly recognised.

<center>
<video controls style="width:35%">
<source src="https://0110.be/files/attachments/475/embedded_demo.webm" type="video/webm; codecs=vp9,opus">
<source src="https://0110.be/files/attachments/475/embedded_demo.mp4" type="video/mp4">
</video>
</center>
With my limited C experience the prototype code was not well organised. During my second attempt this improved enough so that I feel comfortable enough to share the code on GitHub: [Olaf - Overly Lightweight Acoustic Fingerprinting](http://github.com/JorenSix/Olaf).

The code went through several iterations and was expanded beyond the original scope and became **a capable general purpose acoustic fingerprinting system** with its [many applications](https://0110.be/publications/Applications_of_Duplicate_Detection_in_Music_Archives%3A_from_Metadata_Comparison_to_Storage_Optimisation). Olaf performs quite well thanks to its resource friendly design and the use of [PFFT](https://bitbucket.org/jpommier/pffft/src/master/) and [LMDB](http://www.lmdb.tech/doc/). Especially LMDB, a fast, B+-tree backed key value store with low storage overhead enables performant storage and lookups.

The GitHub does not contain an example for the ESP32. That code depends on the microcontroller, digital microphone and pins used and Olaf needs to be hacked to exhibit the requested behaviour. All in all that code is much less reusable (and sharable, testable, maintainable). I have, however, included a platformIO project for "Olaf on ESP32":\[ESP32-Olaf.zip\] for reference.

### WASM: Olaf in the browser

Olaf, being written in ANSI C, can run in the browser thanks to the [Emscripten](https://emscripten.org/) compiler. According to its website, Emscripten *'...lets you run C and C on the web at near-native speed without plugins'* Combining the Web Audio API and the WASM version of Olaf makes web-based acoustic fingerprinting applications possible.

Below you can try out Olaf. The *exact same code* is running on your browser as on the ESP32 demonstrated above. This means that Olaf is listening to recognise 'Let It Go' from the Frozen soundtrack. For your convenience the song can be started below on the left. On the right, you can start Olaf by allowing incoming audio to be analysed. The FFT is calculated by Olaf and visualised using [Pixi.js](https://www.pixijs.com/). After a few seconds the red fingerprints should become green, indicating a match. Once you stop the song, the fingerprints will eventually turn red again. As with the video above: going from a match to no match takes a couple of seconds to allow gaps in recognition.

<div style="margin-top:2rem;margin-bottom:2rem;width:100%;display:grid;grid-template-columns: 1fr 1fr;gap: 0.5rem 0.5rem">
<iframe style="width:100%;height:15rem;border:none" src="https://www.youtube.com/embed/moSFlvxnbgk" allow="">
</iframe>
<iframe style="width:100%;height:15rem;border:1px black solid" src="https://0110.be/files/attachments/475/spectrogram.html">
</iframe>
<small>
1. Start the song and play it aloud. Singing along is encouraged.
2. Start the microphone and check whether recognition succeeds.
</small>

</div>
Olaf was featured on [hackaday](https://hackaday.com/2020/08/30/olaf-lets-an-esp32-listen-to-the-music/). There is also a small discussion about Olaf on [Hacker News](https://news.ycombinator.com/item?id=24292817). A write-up of this project also ended up as a contribution to the Late Breaking Demo track of the first virtual ISMIR conference: [Olaf ISMIR 2020 LBD abstract](https://0110.be/files/attachments/475/ISMIR2020_LBD_Olaf.pdf).

<br><br>


- [web\_embedded\_olaf.jpg](https://0110.be/files/attachments/475/web_embedded_olaf.jpg)

- [embedded\_demo.mp4](https://0110.be/files/attachments/475/embedded_demo.mp4)

- [embedded\_demo.webm](https://0110.be/files/attachments/475/embedded_demo.webm)

- [embedded\_use.mp4](https://0110.be/files/attachments/475/embedded_use.mp4)

- [ESP32-Olaf.zip](https://0110.be/files/attachments/475/ESP32-Olaf.zip)

- [ISMIR2020\_LBD\_Olaf.pdf](https://0110.be/files/attachments/475/ISMIR2020_LBD_Olaf.pdf)

---

## [LTC - SMPTE Decoder on Teensy](https://0110.be/posts/LTC_-_SMPTE_Decoder_on_Teensy.md)

- Published: 2019-12-06T00:00:00Z
- Updated: 2025-11-29T14:36:37Z
- Author: Joren
- ID: 474
- Canonical: https://0110.be/posts/LTC_-_SMPTE_Decoder_on_Teensy

- Tags: [Code](https://0110.be/tags/Code.md), [UGent](https://0110.be/tags/UGent.md)

<img style="float:right;margin:5px" width="200" alt="Teensy with audio shield" src="https://0110.be/files/attachments/474/teensy_with_audio_shield.jpg">

For synchronisation between several devices [SMPTE timecode](https://en.wikipedia.org/wiki/SMPTE_timecode) data is often encoded into audio using [LTC](https://en.wikipedia.org/wiki/Linear_timecode) or linear time code.

This blog post presents an LTC decoder for a [Teensy 3.2 microcontroller](https://www.pjrc.com/teensy/teensy31.html) with [audio shield](https://www.pjrc.com/store/teensy3_audio.html).

The audio shield takes care of the line level audio input. This audio input is then decoded. The decoding is done by [libltc](https://github.com/x42/libltc). The library runs as is on a Teensy without modification. The three elements are combined in a relatively simple [teensy patch](https://github.com/ArtScienceLab/ARDUINO_LTC_DECODER/blob/master/LTCDecoder/src/main.cpp)

To use the decoder connect the line level input left channel to an SMPTE source via e.g. an RCA plug.

For code, comments, pull requests please consult the Github repository for the [Teensy SMPTE LTC decoder](https://github.com/ArtScienceLab/ARDUINO_LTC_DECODER)

<div style="clear">
</div>
<figure style="padding:10px">
<center>
<video width="250" controls>
<source src="https://0110.be/files/attachments/474/teensy_ltc_smpte_decoder.mp4" type="video/mp4">
</source>
</video><br>
A teensy decoding an LTC SMPTE signal
</center>
</figure>

<br>


- [teensy\_ltc\_smpte\_decoder.mp4](https://0110.be/files/attachments/474/teensy_ltc_smpte_decoder.mp4)

- [teensy\_with\_audio\_shield.jpg](https://0110.be/files/attachments/474/teensy_with_audio_shield.jpg)

---

## [MIDImorphosis: recording audio and sensor data](https://0110.be/posts/MIDImorphosis%3A_recording_audio_and_sensor_data.md)

- Published: 2019-10-03T00:00:00Z
- Updated: 2025-11-29T14:37:20Z
- Author: Joren
- ID: 470
- Canonical: https://0110.be/posts/MIDImorphosis%3A_recording_audio_and_sensor_data

- Tags: [Code](https://0110.be/tags/Code.md), [Computational ethnomusicology](https://0110.be/tags/Computational%20ethnomusicology.md), [UGent](https://0110.be/tags/UGent.md)

During an experiment which monitors a music performance it might be a requirement to record music, video and sensor data synchronously. Recording analog sensors (balance boards, accelerometers, light sensors, distance sensors) together with audio and video is often problematic. Ideally standard DAW software can be used to record both audio and sensor data. A system is presented here that makes it relatively straightforward to record sensor data together with audio/video.

The basic idea is simple: a microcontroller is programmed to appear as a class compliant MIDI device. Analog measurements on the micro-controller are translated to a specific MIDI protocol. The MIDI data, on the capturing side, can then be converted again into the original sensor data. This setup has several advantages:

-   It makes it easy to record sensor data together with audio data in a standard DAW software package. Recording a recording a midi track and audio track simultaneously in, e.g., Ableton Live, is easy.
-   Communication with the micro-controller is bi-directional. The micro-controller can be programmed to react to certain MIDI messages. A note-on can, for example, be used to start analog sensor recording. These MIDI commands can be send from any possible source that can 'speak' MIDI.
-   Thanks to the Web MIDI API this construct presents an easy way to let analog sensors and websites interact.
-   Real time sonification of the sensor-data is also supported. There are many ready to go options to sonify MIDI. Axoloti, Max/MSP, Zupiter are some of the environments that can be pugged into.

<br>
<center>
<img src="https://0110.be/files/attachments/470/cc_viz_screen.png" alt="screenshot of signal visualization"/><br>
<small style="color:gray">Fig: Visualization in html of analog sensor data, captured as MIDI</small>
</center>
<br>

While the concept is relatively simple, there are many details to get right. Please consult the [MIDImorphosis github](https://github.com/ArtScienceLab/MIDImorphosis) page which details the system that consists of an analog sensor, a MIDI protocol and a clocking infrastructure.

<br>


- [cc\_viz\_screen.png](https://0110.be/files/attachments/470/cc_viz_screen.png)

- [cc\_viz.html](https://0110.be/files/attachments/470/cc_viz.html)

- [cc\_viz.html](https://0110.be/files/attachments/470/cc_viz.html)

---

## [LW Research Day 2019 on Digital Humanities](https://0110.be/posts/LW_Research_Day_2019_on_Digital_Humanities.md)

- Published: 2019-09-10T00:00:00Z
- Updated: 2019-10-03T08:53:08Z
- Author: Joren
- ID: 473
- Canonical: https://0110.be/posts/LW_Research_Day_2019_on_Digital_Humanities

- Tags: [Presentation](https://0110.be/tags/Presentation.md), [Research papers](https://0110.be/tags/Research%20papers.md), [UGent](https://0110.be/tags/UGent.md)

<img src="https://0110.be/files/attachments/473/AfficheLWRD2019_klein-272x300.png" style="width:150;float:right;margin:10px"> On the 9th of September 2019 the second [research day organized by the faculty of Arts and Philosophy of Ghent University](https://www.lwresearchday.ugent.be/) took place. The theme of the day was 'Digital Humanities' and [the program](https://www.lwresearchday.ugent.be/programme/) gave an overview of the breadth of research at our faculty with topics as logic, history, archeology, chemistry, geography

<center>
<img src="https://0110.be/files/attachments/473/balance_board_screenshot.png" style="width:400px"/>

</center>
Together with Jeska, I presented an ongoing study on musical interaction. In the study one of the measurements was the body movement of two participants. This is done with boards that are equipped with weight sensors. The data that comes out of this can be inspected for synchronisation, quality and quantity of movement, movement periodicities.

<center>
<video src="https://0110.be/files/attachments/473/web_2019-09-16_15.30.32.mp4" style="width:400px" controls/>
</center>
<br>

The hardware is the work of Ivan Schepers, the software used to capture and transmit messages is called "the MIDImorphosis" and developed by me. The research is in collaboration with Jeska Buhman, Marc Leman and Alessandro Dell'Anna. An article with detailed findings is forthcoming.


---

## [AAWM/FMA 2019 - Birmingham ](https://0110.be/posts/AAWM%2FFMA_2019_-_Birmingham_.md)

- Published: 2019-07-02T00:00:00Z
- Updated: 2025-11-29T14:37:46Z
- Author: Joren
- ID: 471
- Canonical: https://0110.be/posts/AAWM%2FFMA_2019_-_Birmingham_

- Tags: [Folk Music Analysis (FMA) conference](https://0110.be/tags/Folk%20Music%20Analysis%20%28FMA%29%20conference.md), [Music Information Retrieval](https://0110.be/tags/Music%20Information%20Retrieval.md), [Presentation](https://0110.be/tags/Presentation.md), [Research papers](https://0110.be/tags/Research%20papers.md), [UGent](https://0110.be/tags/UGent.md)

I am currently in Birmingham, UK at the 2019 at the joint [Analytical Approaches to World Music (AAWM) and Folk Music Conference (FMA)](http://fma2019.bcu.ac.uk). The opening concert by the [RBC folk ensemble](https://www.youtube.com/watch?v=TxcMPWTlaaI) already provided the most lively and enthusiastic conference opening probably ever. Especially considering the early morning hour (9.30). At the conference, two studies will be presented on which I collaborated:

### Automatic comparison of human music, speech, and bird song suggests uniqueness of human scales

<a href="https://0110.be/files/attachments/471/FMA2019_paper_12.pdf"><img src="https://0110.be/files/attachments/471/12.png" style="float:right;width:140px"></a> "Automatic comparison of human music, speech, and bird song suggests uniqueness of human scales":\[FMA2019_paper_12.pdf\] by Jiei Kuroyanagi, Shoichiro Sato, Meng-Jou Ho, Gakuto Chiba, Joren Six, Peter Pfordresher, Adam Tierney, Shinya Fujii and Patrick Savage

> The uniqueness of human music relative to speech and animal song has been extensively debated, but rarely directly measured. We applied an automated scale analysis algorithm to a sample of 86 recordings of human music, human speech, and bird songs from around the world. We found that human music throughout the world uniquely emphasized scales with small-integer frequency ratios, particularly a perfect 5th (3:2 ratio), while human speech and bird song showed no clear evidence of consistent scale-like tunings. We speculate that the uniquely human tendency toward scales with small-integer ratios may relate to the evolution of synchronized group performance among humans.

### Automatic comparison of global children's and adult songs

<a href="https://0110.be/files/attachments/471/FMA2019_paper_13.pdf"><img src="https://0110.be/files/attachments/471/13.png" style="float:right;width:140px"></a> "Automatic comparison of global children's and adult songs":\[FMA2019_paper_13.pdf\] by Shoichiro Sato, Joren Six, Peter Pfordresher, Shinya Fujii and Patrick Savage

> Music throughout the world varies greatly, yet some musical features like scale structure display striking crosscultural similarities. Are there musical laws or biological constraints that underlie this diversity? The "vocal mistuning" hypothesis proposes that cross-cultural regularities in musical scales arise from imprecision in vocal tuning, while the integer-ratio hypothesis proposes that they arise from perceptual principles based on psychoacoustic consonance. In order to test these hypotheses, we conducted automatic comparative analysis of 100 children's and adult songs from throughout the world. We found that children's songs tend to have narrower melodic range, fewer scale degrees, and less precise intonation than adult songs, consistent with motor limitations due to their earlier developmental stage. On the other hand, adult and children's songs share some common tuning intervals at small-integer ratios, particularly the perfect 5th (\~3:2 ratio). These results suggest that some widespread aspects of musical scales may be caused by motor constraints, but also suggest that perceptual preferences for simple integer ratios might contribute to cross-cultural regularities in scale structure. We propose a "sensorimotor hypothesis" to unify these competing theories.


---

## [trix: Realtime audio over IP](https://0110.be/posts/trix%3A_Realtime_audio_over_IP.md)

- Published: 2019-03-13T00:00:00Z
- Updated: 2019-07-19T09:57:53Z
- Author: Joren
- ID: 469
- Canonical: https://0110.be/posts/trix%3A_Realtime_audio_over_IP

- Tags: [Code](https://0110.be/tags/Code.md), [UGent](https://0110.be/tags/UGent.md)

At work we have a really nice piano and I wanted to be able to broadcast a live performance over the internet with low latency to potential live listeners. In all honesty, only my significant other gets moderately lukewarm about the idea of hearing me play live. Anyhow:

I did not find any practical tool to easily pump audio over the internet. I did find something that was very close called [trx by Mark Hills](http://www.pogo.org.uk/~mark/trx/): trx is a simple toolset for broadcasting live audio from Linux. It unfortunately only works with the ALSA audio system and is limited to Linux. I decided to extend it to support macOS and [Pulse Audio](https://en.wikipedia.org/wiki/PulseAudio). I also extended its name to form trix.

Audio Transmitter/Receiver over Ip eXchange (trix) is a simple toolset for broadcasting live audio from Linux or macOS. It sends and receives encoded audio over IP networks, via an audio interface. If audio interfaces are properly configured, a low-latency point-to-point or multicast broadband audio connection can be achieved. This could be used for networked music performances. The inclusion of the intermediate [rtAudio](https://www.music.mcgill.ca/~gary/rtaudio/) library provides support for various audio input and outputs.

More information on trix can be found on the [trix](https://github.com/JorenSix/trix) github page.

## Latency

The system can be configured for low latency use. The whole chain is dependent several different components which each add to the total latency: audio input latency, encoder (algorithmic) delay, network latency and finally audio output latency.

Thanks to the use of RtAudio it should be possible to use low latency API's to access audio devices (ASIO on windows or Jack on Unix). This means that audio input and output latencies can be as low as the hardware allows. The [opus](https://en.wikipedia.org/wiki/Opus_(audio_format)) encoder/decoder that is used has a low algorithmic delay. By default it has a 25ms delay but it can be configured to only 2.5ms (see [here](http://www.pogo.org.uk/~mark/trx/latencies.txt)). The network latency (and jitter) is very much dependent on the distance to cover. On a local network this can be kept low, when using wide area networks (the internet) control is lost and latencies can add up depending on the number of hops to take. Jitter can be problematic if the smallest possible buffers are used: then dropouts might occur and this might affect the audio in a noticeable way.


![Piano at the krook](https://0110.be/files/photos/469/2019-03-13_09.19.51.jpg)

---

[Newer posts](https://0110.be/Blog.md?page=1)

[Older posts](https://0110.be/Blog.md?page=3)
