Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Shazam is one of the very few apps in the past 20 years that STILL "wows" me. I have no idea how the tech works, and I even sort of like not knowing, to be honest. It's one of the very few apps out there that still in exist in a "magical" way to me. I am constantly impressed with how fast/easy it works, even with very obscure music. What an amazing app.

Fun quick related story, about 10 or more years ago there was a back tracking song on a TV show (Scrubs) that I really liked that was only in the Netflix version. It was just an instrumental song with some French sounding words speaking in it so there was no easy way to search for it. However, it was distinct enough that it didn't seem like something made just for the show. It was also pretty quiet and under some talking in the tv show scene. I had posted on reddit asking if anyone knew it, and never got any responses. I searched all over the web, but no source had the track details. It drove me crazy every time I would hear the song in re-watching the show, and I still could not track it down every few years when I tried again. Back then, Shazam had no cataloging of it so it wasn't in there either yet. However, when re-watching it a few years back again, I tried Shazam again and to my surprise it finally worked. I was blown away that Shazam was finally able to solve this 10+ year mystery. It was one of the coolest feelings every to scratch that itch finding this rare French song and hearing it in full. It was truly magical.

EDIT: Oh sorry, I didn't think anyone would actually care about the song itself lol It was called "Sans Hésitation" by the French-Canadian band "Chapeaumelon". https://www.youtube.com/watch?v=Ju4d3YQhByU - It's also interesting cause now the song does in the episode in tv music database sites. Very cool.



Shazam used to wow me, but then as others mentioned in the replies it's essentially matching the signature of the sound to the sounds in the database. If it's one of the song, it gets matched fairly quickly.

Wow blew my mind was when Google introduced 'hum and we'll recognize the song for you' in Google assistant: https://www.google.com/amp/s/blog.google/products/search/hum...

It works so well even with my shitty humming - even my girlfriend can't recognize what the song is but Google can. It doesn't even have the same signature as the original audio file, just similar hums in a noisy environment and it still works. Black magic fuckery.


> it's essentially matching the signature of the sound to the sounds in the database.

You aren't giving it enough credit. The algorithm uses just a few seconds from any part of the song, and has to deal with phone audio quality and often background noise. I mean, you can be in a bar with all that jabber and hold up the phone and it could pick out the song. The app on the phone does the preprocessing to the audio before it is sent to the server that does the matching ... using the comparatively miserable power of a 2001 era cell phone.


Oh that wasn't my intention - Shazam was and is groundbreaking, they did it when no one else could. All I meant was that it seems more "doable and I probably understand how it works" when compared to how Google assistant recognizes songs from my humming.


What is a signature? How is a signature computed from a noisy audio stream, over a mall speaker? How is a signature computed from an arbitrary starting point?


IIRC, it's uses a Fast Fourier Transform of the time delay between high notes in the song to generate a series of "hashes" that are stored a db. Those ids can be calculated locally on the phone and then its a simple db lookup to retrieve potential hits. When Shazam adds a song to the db, they compute a series of "hashes" so you can identify at any point in the tune.


Wow, that's fascinating! I just ended up down the rabbit hole reading Avery Wang's "An Industrial-Strength Audio Search Algorithm" (linked in this thread) - it's such a cool way of "fingerprinting" pieces of music data.


My original comment was from memory of reading a post about how it worked a few years ago. Looking at what you read, I think the gist of what said is right, though it seems they use a different algorithm than FFT.

Totally agree though. It is something that opened my mind to thinking of a way to solve that problem in a way that actually works. Shazam definitely looked like magic the first time I saw it work.



TL;DR (from skimming thru the paper) he figured that a song's spectrogram looks like a starry sky, so matching a song is like finding a constellation on the sky. How do you do it efficiently? By searching for simple features of your constellation, such as pairs or triples of bright stars - those can be pre-hashed to find matches instantaneously. Once a possible match is found, you compare the rest of the constellation. Nothing breathtaking, in other words. However, among all the men who talked, he was the one who both talked and did, and that's his achievement.


Brilliant stuff is easy to understand, a lot harder to come up with. I could do that! (With a little help from wikipedia, audio processing libraries, the answer sheet, and the knowledge that it's possible in the first place)

To me, this highlights how hashing is the closest thing programmers have to magic.


Create a compound signature. You don't just take one measurement but many measurements and then assess the probabilities. You may have people talking in a mall, but they will be in a narrow frequency band. Similarly you can analyze the repeating elements. Keep iterating and adding stuff until f(signal) performs well


The closest to the ideal signature?


> Wow blew my mind was when Google introduced 'hum and we'll recognize the song for you' in Google assistant

Their announcement actually made me roll my eyes a bit, as Soundhound had that functionality nearly a decade before. I had both SH and Shazam installed on my old phone for these usecases - now Shazam is baked into Siri so I don’t even have the app itself installed.


How well does Shazam work for you when you hum or sing a song?


I haven’t tried humming with Shazam recently, but I don’t think it worked well back when I did have the actual app. It works very well for music though. I used it around five times, just this Wednesday night at a concert, and it got every track for me.

Soundhound is what had humming “support” explicitly in its product description, and it worked pretty well from what I remember. It’s been long enough though that I may only be remembering the times it worked.


Doesn't work at all


What do you have to ask Siri to get this to work?


If you prefer to access it via your iPhone's control center, you can configure it that way in the control center settings. It is called "Music Recognition" there.


Nice I will have to check this out, control center is definitely a great little overlay but I haven't reliably figured out how to add things to it. I will investigate further.


For general information have a look here [0]. Also be aware that elements in the control center might even offer additional functionality, e.g. like setting the brightness of the flashlight. In this case instead of just switching the flashlight on by a tap on the button, keep the flashlight button pressed to bring up a slider to set the brightness. Just play around with the other control center elements to find out what is possible.

[0] https://support.apple.com/guide/iphone/iph59095ec58/ios


I usually say "Hey Siri, what song am I listening to?" but it works with a bunch of variations e.g.: "what song is this?"

There's also a bunch of other options to trigger Shazam, main way I use it is from the Control Center: https://support.apple.com/en-us/HT210331


If you have an Apple Watch, you can also set it as a home screen button, which is a lot more discreet in public.


I just say Hey Siri Shazam this.


“Hey Siri, what song is this” works


> Essentially matching the signature of the sound to the sounds in the database.

And Dall-E 2 is just doing fuzzy hashing of images with text keys.

Shazam continues to amaze me because it "just works", and still feels more magical to me than most of the AI out there since it directly solve a major problem I didn't even think was solvable "what is this song!!?"


I enjoy salsa dancing, but I don't know any Spanish, so I use that built-in Google functionality to hum various songs all the time to figure out what they're called.


Dude, spoiler alert. Did you miss the part where OP said they liked not knowing how it works??


xD, spoiling an algorithm


And the downvotes tell me some folks have absolutely no sense of humor


I honestly couldn't tell if you were serious.. the ?? should have given it away but I didn't notice.


There's another where you tap a beat with your space bar, and a website tries to guess the song.


I need to download the Google app (and I presume sign in) to use that feature? Count me out


What really wows me is that Shazam started in 2002. It was a phone number you would call on your cell phone and let it listen to your environment.

Way back then, it was doing everything you describe, but over low quality band limited telephone lines.


As an almost teenager at the time, that (Shazam over the phone with an answer texted back - which I used on a Nokia 3310) was the one thing that convinced me we would soon have pocket devices that really could do anything.

And while it took a few iterations (for me, from palm pilot to blackberry as a teenager, then eventually moving to iPhone after a few too many painful Blackberry upgrades - still missing that unified inbox though, as is everyone else I know who had a BB of that era... and frankly missing a great physical keyboard on a phone, too) I still am impressed on a daily basis that I do indeed have the device in my pocket that 12 year old me dreamed of.


I didn't know it ever worked that way, that's incredible. Reminds me of ChaCha, the texting service where you texted questions and a human would quickly look up the answer and text it back. It's a very cool idea that was quickly outmoded by smart phones and is kind of lost to history now.


Funny enough, even Google used to do that... before smartphones and the Google Assistant, you could text GOOGL (46645, I think) a query and get back a quick answer: https://googleblog.blogspot.com/2004/10/get-411-with-46645.h...

They eventually shut it down :( https://slate.com/technology/2013/05/google-sms-search-shutd...


I remember Sony Ericsson handhelds all came with TrackID back in the day (2007/2008) and I used it to name music I heard in public. It was the same idea. I think it charged £1-2 per track!


Back then phone quality was much better. Cellphones killed that.


What? HD Voice using VoLTE or WiFi Calling is miles better than any land line phone


I'd take a landline over the inconsistency of cell calls any day. Maybe peak/maximum quality is better with today's tech, but reliability definitely isn't. I would bet average quality of calls has dropped too. HD voice is a carrier specific thing, no? Most of the time I can barely even hear the other person, much less have HD anything.


I think you have been misinformed. Voice calls over plain old telephone service were band-limited to <4khz decades before cell phones became popular. This was necessary in order to cram more active calls into our existing telephone infrastructure.

Digital phone calls today are way better.


Today's g.722 (HD Voice) is better, but GSM codecs are also 8kHz sampled, then lossily encoded. If the compression is appropriate, gsm is 13-bit per sample vs 8-bit per sample commonly used for POTS (minus robbed bit signalling in the US), but if the compression isn't appropriate, you get some pretty nasty artifacts. Encode/decode delay can be significant in some applications, but since GSM is TDMA anyway, you're going to have buffering and may as well use that to encode; a T1 PRI multiplexes one sample at a time, so a lot less delay there.


I’m sure that’s all true. I’m also pretty sure I wasn’t regularly shouting into my AT&T landline “I can’t hear you!” Obviously we’ve gained a lot with cell phones and portable music playing but it’s been mostly at the cost of consistent quality.


You can't just post a story like that and not link the song!

Personally, my main usage of Shazam is for identifying vaporwave samples. Often all you have to do is throw the song in Audacity, tweak the speed a bit, and Shazam it.



There are entire albums on Spotify which are full songs of 80s pop classics, played at a slower speed, then uploaded as a new album from another artist.


Interesting, have any of them been hit with copyright violations?

I ask because I like to create bootlegs (basically homebrew remixes, these are substantial re-imaginings of the original track) and would like to put them on Spotify, but am worried about copyright issues and how that might affect posting original music.


haha Sorry, I updated the post. Didn't think anyone would care about that part lol


The first time I heard of Shazam was on a road trip with a friend of mine who had minimal tech skills at best. I was already 10 years into my career as an engineer, and when he told me about it, I honestly didn't believe him; I was positive he was mistaken, and speculated it was a service similar to Aardvark[1], which was a peer-to-peer information engine.

I was wrong, of course, Shazam really did live up to its hype. I think it's interesting that the someone knows about how a technology works the more sceptical they are of what it is capable of.

[1] https://en.wikipedia.org/wiki/Aardvark_(search_engine)


I don't know about Shazam's current algorithm specifically, but years ago I worked at a place with a mathematician that worked on gracenote's algorithms, and asked him for the basics on how it works.

Basically, it records audio chopping it up into small segments and throwing them through a FFT. Then it takes that, and thinking of the data like a greyscale spectrograph image, runs it through a quantization filter that helps reject some noise, then converts that to locality sensitive hashes that are sent to the server. So basically FFT, filter, hash, lookup.


This paper from the Shazam founder describes an approach for doing it: https://www.ee.columbia.edu/~dpwe/papers/Wang03-shazam.pdf


from the Shazam founder

...who by the way holds a PhD from Stanford...


Don't want to spoil it for you if you really don't want to know but I want to share to others in case they do because I found it so interesting when I first learned!

It looks like others shared the paper: https://www.ee.columbia.edu/~dpwe/papers/Wang03-shazam.pdf

It's short but very cool. I read it a while ago and honestly can't pretend I fully grokked everything, but my understanding was that you can't just use a Fourier transformation alone. Noise would basically make this impossible.

So what I'd consider the key insight is that they compressed songs down to "fingerprints". IIRC they noticed that songs, even in noisy environments, preserved certain bits of information. Particularly, they could look at the spectrogram and see peaks of amplitude in the tapestry. They essentially set some radius and scanned the spectrogram. In a given radius, only the largest amplitude value in time and frequency would be preserved. So you've reduce a 3MB song to several bits.

This would be good enough for small databases (I think). But it's intractable for anything practical. So they built hashes out of these fingerprints using pairs of the preserved peak bits. They would choose a certain peak (called the anchor point), record its time offset from the start of the song, and then form pairs with other nearby peaks, saving the pairs of frequencies (but discarding e.g their amplitudes). So for each of these anchor points, you would get a 64 bit value: 32 bits for the time offset and track ID and 32 bits of frequency-pairs.

When you wanted to look up a song, they would fingerprint your snippet into multiple 32bit hashes and compare them against the frequency-pair hashes in the database. If a song was a good match, then you would see that your snippet matched against multiple hashes from that song, and specifically they matched linearly over time (I'm struggling to explain this bit but it's visually obvious if you look at Figure 3 in the paper).

I probably got some of this wrong, but I hope it's a helpful summary of the paper. I remember struggling to understand parts of it, so please let me know if anything I said is egregiously wrong!


I had a similar experience looking for a background track in an episode of This American Life. I couldn't remember which episode it was, and none of the lyrics were in English. Pretty sure I went through backwards through the episodes and listened to all the credited songs to find it. The song was 69 Police by David Holmes, which still feels perfect to me. https://www.youtube.com/watch?v=IWissIWxqKk

On the topic of background music, tons of original background music copies/imitates famous stuff. Sometimes it's "I wanted the sound of X but couldn't afford it", but there are some in-jokes in there too. Wish I could remember some examples.


I think all (so simple) you have to do is parse all the tracks ever made, and say generate a sequence of snapshots of what the tune sounds like and the delta. e.g. if it was notes (for simplicity) E,D,C,D,E,E,E,D,D,D,E,E,E is the start of "Mary had a little Lamb" Millions of tracks contain the note E. Many hundreds of thousands probably have the note D next - and as you work through the sequence, you're pruning down that list until you who what it is. Bit that makes my mind hurt though, is the data-structure you put those sequences into to make it quickly searchable. Users can start recording at any point in the song - so you can't just prune a tree down from a known starting point. There's going be be background nose - so you need some way of "when you have no choice left", I presume sticking wild-cards into the previous decisions, to see if you end up back on a known track.

Yeah - I think it's magic as well.

Other thoughts: I used it back in the UK when it launched, and the first track I ever used it on dialling (2580 - the numbers down the middle of your keypad) was also a French track (MC Solaar – La Vie Est Belle)

I always felt they missed a trick, just identifying music (and then trying to sell you stuff). Surely they could have used the same tech to seamlessly mix all music together. (i.e. take the sequences within tracks they find hard to differentiate, and then use these points to allow two tracks to be mixed together). What's the minimum number of tracks it would say take to seamlessly mix from Megadeth to Mozart?


They used to have a paper on their website describing their algorithm in simplified form but I can't find it any more. Wikipedia has some details: https://en.wikipedia.org/wiki/Acoustic_fingerprint

I believe it's very sensitive to changes in timing, so it doesn't work on live performances etc.

(based on reading I did 13 years ago before an interview at Shazam, which to this day still remains my worst interview performance)



The "doesn't work with live performances" bit is borne out by my consistent experience failing to identify some songs at live performances, but with the "DJ Set" form of live performance, tempo shifting music without pitch shifting it still appears to get the goods more often than not.


I'll also plug AcoustID from MusicBrainz

https://musicbrainz.org/doc/AcoustID


We use AcoustID in MusicBox[0] to identify and deduplicate content, and it works great for us.

What we do is calculate the acoustic fingerprint of every uploaded content and compare/check for duplicates (only authorized staff can upload, but this still helps a bunch with user errors and in cases where you need to reupload a track). Then we compare the fingerprints, using this[1] approach, so we can fine-tune the similarity based on our needs.

In our case it's been very effective. Yes, live versions are treated as different ones (which is exactly what we need in our case, so it's a feature for us), but mechanical differences between tracks (volume, slight distortions from codec, different compression levels or remasters, or track being cut differently) are just ignored.

If you ever want/need audio fingerprinting, I can warmly recommend it.

[0] Music streaming service optimized for cafes, restaurants and other venues - https://musicbox.com.hr/ [1] https://groups.google.com/forum/#!msg/acoustid/Uq_ASjaq3bw/k...


> live versions are treated as different ones

I think you're talking about a live recording vs a studio recording? But what I think zelos was talking about was "someone is currently playing music live, what is it?", which is a lot harder because you need to recognize the essence of a song and not the essence of a recording of a song.


Yeah, agreed, that's way harder and not something AcoustID can do.


>if it was notes (for simplicity) E,D,C,D,E,E,E,D,D,D,E,E,E is the start of "Mary had a little Lamb"

As far as I can tell these operate on audio, not symbolic music.


They do (and I said so) - but I couldn't think of an easy way to write that in a post here.


FFT data tends to get quantized, normalized, and counted for analysis purposes.


My instinct is that it probably isn't as simple as you describe because not only are there multiple notes at a time in a given track (i.e. chords), but there are also several tracks playing at once! It's possible that they're literally generating data like {guitar 1: C chord, guitar 2: single note E, bass: single note E} for every point in time, but even then each instrument isn't playing the exact same rhythm most of the time, so the notes won't exactly line up. I guess I don't think it's completely computationally infeasible to do it this way, but it seems more likely that they're just trying to separate the music from the background noise and then try to find the closest match to the music audio as a whole rather than trying to separate it into component.


Sorry - I wasn't clear. I don't mean they're listening for notes. They're just analyzing the wave-form/fingerprint/whatever-you-want-to-call-it that's being generated at a moment, and then one form the next moment, then the next.

One of these might match random points in many songs, but a far smaller subset of these will have the same three in the same sequence.


Fair enough! I imagine that having many instruments at once would improve the ability to diff the waveform/etc. rather than hindering it then.


> Surely they could have used the same tech to seamlessly mix all music together. (i.e. take the sequences within tracks they find hard to differentiate, and then use these points to allow two tracks to be mixed together). What's the minimum number of tracks it would say take to seamlessly mix from Megadeth to Mozart?

I noodled around with this idea in my free time a few years ago, got absolutely nowhere really usable with it (I probably put in a couple hundred hours).

I knew I was limited by my dataset (small), code quality (terrible) and understanding of musical theory (virtually nil).

Maybe I'll pick up that idea again - even doing beat matching would be kind of neat.


Shazam as a product feels a bit odd. Almost as if they’ve never quite outgrown their slightly sketchy “advertised on MTV2 alongside the Crazy Frog” origins.

They must have loads of data on songs people actually want to know yet never really managed to turn themselves into anything more sophisticated.


> I have no idea how the tech works

It does a Fourier analysis of sections of the song, and puts the results in a database. A Fourier analysis yields what frequencies make up a waveform along with their amplitudes, so it is very compact.


Taking the DTFT of a signal yields exactly the same amount of information, so it's not really more compact. Shazam used a spectrogram (which is more information than the original signal) and searched for peaks to create a finger print.

It's not the analysis that is compact, but the fingerprint derived from it.


You get a spectrogram by applying Fourier transform. Also, getting more information out of something than it contains is literally impossible.


I know it contains the same information, but it makes it easy to discard the low amplitude frequencies, and the frequencies that are not heard by the ears, or are not particularly important to our ears.


And now Chapeaumelon is wondering why the sudden surge of the youtube views. Comments are disabled so we cannot even help them to understand :)


Shazam is great but a similar app that really "wowed" me around 2007 was Midomi - it could recognise humming with good results, even though I'm really bad at hitting right notes and key. It still exist but is not really talked about anymore, Shazam seems to have dominated that market.


Don't skip the credits next time :)


haha wasn't in there! Def lookied :)


Shazam is not particularly complex, however it is a very clever solution and a great example of applying a simple engineering concept broadly. I still hold it as one of the best examples of clever engineering in the app world


It actually wasn't an app in the beginning - it was a phone number that you dialed.


Shazam is probably the only Apple watch app I ever use. Very convenient to have this on the watch.


which episode? and is it still in there? DVD, streaming, and syndication have some different songs because of rights issues.


https://scrubs.fandom.com/wiki/My_Ocardial_Infarction - In the Netflix version. It does now show the song in there, which is cool. But for like 10 years, it was unknown.


What was the song ??


You will get the result once $commenter is online again


Updated the post, sorry! haha


Well?


It’s all just Fourier analysis I’m guessing?

Which I always find to be simultaneously simple and obvious as well as total magic.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: