TLDR: there is an extremely high amount of 2-minute tracks, we call it the 2m peak. The 2m peak is half functional music for sleep/relaxation/focus/ambience (think ‘Rainy Nights for Sleep’) and half likely to be traditional music, with a massive surge in releases in 2024. This suggests mainly algorithmically-generated music, including AI.
Here is the distribution of durations over 256M songs:
We see three distinct peaks at 2m, 3m5s and 4m. Here we focus on the 2m peak. Can we explain it?
First, to get a sense of how abnormal it is, let’s zoom in on track counts for durations between 1m50s and 2m30s at a resolution of one second:
Using a linear regression, we see that the 2m peak contains an abnormal excess of about 1.3M songs! Where do they come from?
The most telling hint is to look at the number of releases per year:
This strongly smells like computer-generated, not to say AI-generated, music! And, in fact, 2m tracks hold the record, with 32% of them having been released in 2024:
Interestingly, 2 minutes was the maximum generation length of the AI platform Suno’s V3 (Spring 2024).1
Hypothesis: the 2m peak is fuelled by music that is easy to mass produce at scale, i.e. algorithmically generated music — not necessarily only AI.
Let’s zoom into musical features to see what type of 2m tracks are released and try to identify the proportion of AI-generated tracks in the peak’s excess (spoiler: 50-75%):
We inspect the three following features:
Instrumentalness: a number between 0 and 1 which predicts whether a track contains no vocals (an instrumentalness of 1 means no vocals).
Danceability: a number between 0 and 1 which predicts whether a track is suited for dancing (very suited for dancing is 1).
Valence: a number between 0 and 1 which predicts the positiveness of a track (very positive/happy is 1).
Let’s look at instrumentalness, averaged among all tracks of the same duration:
We see a sharp rise at the 2m mark. For valence and danceability we see a sharp decrease:
This strongly suggests that a differentiating factor of the 2m peak from neighbouring durations is an increased presence of high-instrumentalness, low-valence, low-danceability tracks which typically correspond to functional music (sleep/relaxation) or background/lo-fi.
Let’s look at the distributions of instrumentalness/valence and danceability in the 2m peak:
We notice in each case a ~2% proportion of tracks where the feature is null, which we interpret as a failure of the feature-computing algorithm on those tracks, i.e. ill-defined features for the tracks.
To get a sense of how singular these distributions are, let’s look at them one second later, at 2m1s:
Hence, it is clear that the 2m peak contains an abnormally high amount of super-chill tracks, which we define to be tracks with > 0.5 instrumentalness (or null), < 0.5 valence (or null), and < 0.5 danceability (or null).
The 2m peak contains 1,045,565 super-chill tracks.
Here is the proportion of super-chill tracks per duration:
Hence, we found a good suspect: the 2m peak contains an abnormal proportion of 45% of super-chill tracks!
Let’s look at 10 randomly sampled track names in that set: ‘Raining into the Ocean Waves’, ‘Throat Chakra - 396 Hz’, ‘Blue Dreams, Pt. 15’, ‘Rain Sounds Pt. 15’, ‘Arise’, ‘Dreamy’, ‘Baby Lullaby’, ‘Seraphic Soul Spa’, ‘Fear Removal Tones’, ‘Haunted Smokey Mountain’. It is surprising to us that most of these tunes have valence close to 0, e.g. ‘All Night Pure Calm’, ‘Fine Rain Sound’, ‘Soft Mist Melodies’ have valence equal to 0, whereas they intuitively evoke neutral to mild positiveness.
These track names evoke functional music used for sleep, relaxation, and ambience-setting, but several genres could be represented: from unclassified sounds for sleep to lullabies, lo-fi or space music. The 2m duration suggests they are listened to on repeat.
We can dig deeper with semantic clustering2 of these tracks’ names:
This is a lot, but we see: Spanish relaxation corner: center top left; water-themed (rain, river): center top right; frequency-based (e.g. 396Hz): center bottom right; meditation-based: center bottom right; fire-themed: center bottom right; white noise: center; classical music: center top3 .
Schematically, this gives:
Top left of the space: Spanish-titled functional music
Right part of the space: English-titled functional music (sleep/relaxation)
Center: noise
Bottom left: harder to interpret; listening to a handful of songs indicates cinematic, ambient, lo-fi music
As we can see, contrarily to super-chill tracks, not-super-chill tracks have seen a sudden surge in 2024, which strongly suggests that something changed in 2024, allowing for the mass production of songs — assuming that not-super-chill tracks are mostly like typical songs. The 2m peak seems to be 50% because of super-chill tracks and 50% because of not-super-chill tracks.
Interestingly, for approximately the same amount of music produced, the super-chill tracks are produced by 6x fewer artists.
Let’s explore the assumption that 2m not-super-chill tracks are “like typical songs” by looking at their distribution of instrumentalness, valence, danceability and speechiness:
We see that more than 50% of not-super-chill tracks have low instrumentalness (i.e. they have vocals), and more than 79% have high danceability, and that valence is equi-distributed. Speechiness is very low, which discards easy-to-generate “spoken word” content.
This seems to confirm that most of not-super-chill tracks are “like typical songs”. We note that only about 10% of them have instrumentalness and danceability < 0.5, which would be tracks likely to be similar to super-chill tracks (but with high valence).
If we look at the release dates of not-super-chill tracks with low instrumentalness (< 0.5) and high danceability (> 0.5), which would be good candidates for “typical songs”, we get:
The highish number of such releases <= 2015 does suggest that these tracks sound like typical music, and it is in this bucket that we find the sudden jump in 2024, which suggests that, in 2024, these “typical songs” are mostly AI-generated.
The 2m peak is a significant anomaly due to the recent surge, between 2022 and 2024 of music that is easy to mass-produce. We believe that two relatively independent factors are at play and that they share 50% of the responsibility in 2024:
Super-chill tracks. This category, of high-instrumentness, low-danceability and low-valence music, contains almost only functional music for sleep/relaxation/focus/background. Arguably, one does not need AI to generate the track ‘Harmonic Focus (30 Hz Binaural Frequency)’ or ‘Loopable Rainfall’ and they can be mass-produced using algorithms. However, AI music generation tools are most likely used too, we assume; especially for ambience-setting tracks, e.g. ‘Eerie Laboratory Ambience’ and at the very least for generating track names.
Functional music is appealing for algorithmic generation, as it is not too complex but is consumed a lot; see this post. These low-information-density tracks are most likely meant to be listened to on repeat, which would justify the 2m duration (which is also convenient to the user for offline download).Not super-chill tracks. This category most likely sounds a lot more like music and has seen a dramatic increase in the number of tracks in 2024. Here we hypothesise that AI music generation tools are the culprit, noting that the 2m mark is consistent with limits of models like Suno at the time.4
Hence, our conclusion is that the ~1.3M excess tracks in the 2m peak are most likely the result of programmatic generation, but maybe only 50-75% using AI, at least for now.
The popularity of a track is a value between 0 and 100, with 100 being the most popular. The popularity is calculated by algorithm and is based, in the most part, on the total number of plays the track has had and how recent those plays are.
The 2m mark has the lowest average popularity per song per track duration — ignoring extremely short tracks.
Note: given the abnormally high number of 2m tracks and the fact that track popularity is extremely heavy-tailed, this is not necessarily the best metric to estimate whether it is worth mass-producing songs for streams (i.e. having one hit is enough).













