Small image descriptors for near-copy detection
Introduction
This is quite a long setup towards playing around with small image fingerprints for similarity matching.
If you don't care about the background, feel free to skip straight to the demo!
Many computer vision applications don't process the image directly, but use some kind of representation of the image instead.
Usually this takes the form of some one-way function that calculates a vector of some dimensionality.
In modern machine learning this is called an embedding.
In computer vision we might consider something a feature, a descriptor or even a fingerprint, depending on the application.
The basic ideas are the same: We want to say something about the image, and to have the computer do some calculations with this image it is almost always easier to work with some (often smaller) representation of that image.
Even the biggest representations are usually at least an order of magnitude smaller than the original image, which makes running computations on them much faster.
Different tasks require larger representations of the image.
Object recognition, image classification, or even 3d image reconstruction require more precise input and thus more data to represent the image.
However, some simpler tasks might be accomplished with much less data.
One of these simpler applications is perceptual hashing.
Image compare
Let's imagine that we have two images, and we want to check if these two images are the same image.
Well, you might start by simply comparing data byte-for-byte, which works quite well for direct copies.
GNU diffutils even has a nice cmp tool that will do exactly that.
It will look through two files until it finds a byte that does not match between the two:
$ cp a.jpg a.copy.jpg
$ cmp a.jpg a.copy.jpg && echo "Same" || echo "Not same"
Same
This works for our direct copy.
Now lets see how it handles two entirely different images:
$ cmp a.jpg b.jpg
a.jpg b.jpg differ: byte 14, line 1
The two different images differ on byte 14, which is the byte in the JPEG header which defines the "density", or pixels per length of an image.
In this case, a.jpg does not have anything set, while b.jpg defines its density as pixels per inch.
What about two images that are the same, but reencoded?
$ convert c.png c.jpg && convert c.jpg d.png
$ cmp c.png d.png
c.png d.png differ: byte 36, line 3
In this case we convert a png image into a jpeg and then back into a png.
This introduces jpeg encoding artefacts into the image, which subtly alter the pixel values such that cmp no longer considers them the same image.
This gives us our first hints at why image comparisons are tricky.
Even when both files are jpeg images, the subtle differences in encoding already trip us up before we even get to the actual pixel information.
And even when both files are visually identical (and identically encoded), the subtle differences in pixel values still gets us.
This might all seem a little silly so far.
Of course comparing images in this way does not make sense.
What we are looking for is a method that uses a more complex perception-based model to say if two images are the same.
Even more so, we are looking for a method to not only compare two given images,
but quickly calculate the nearest neighbours of a given image based on their similarity.
Perceptual hashing
Locality-sensitive hashing
I am going to assume you know what a hashing function is.
Different hashing functions have various properties that make them more or less suited to various purposes,
but one common feature is that they try to avoid collisions by having small changes in input correspond to large changes in the output.
When we want to avoid this behaviour we are doing something called locality-sensitive hashing (LSH).
In locality-sensitive hashing we are more likely to put similar data into the same bucket.
In normal hashing we want to avoid the same output values so we call it a collision (bad!).
For LSH we want to encourage the same output values, so we say we "put it in a bucket" (good!).
So what are the requirements of a good perceptual hash?
- Small (i.e. good dimensionality reduction)
- Avoid collisions between different images.
- High similarity between similar images.
We see that we have multiple objectives to optimize for, but we also see that two of our requirements depend on a subjective measure.
Size
I have decided to limit myself to fingerprints that are at most 64 bits (or 8 bytes) in size.
Firstly, this allows me to "compete" in a little happy place of small descriptors.
But more importantly, because modern processors are very good at processing 64 bit integers, this should also be a very performant fingerprint.
The good news is that this takes care of our first requirement.
Subjective image similarity
Broadly speaking, we have to answer the question: when are two images the same?
Lets start by looking at some examples.
These two pictures are, to a human observer, definitely similar, but also not the same.
The framing is similar, the subject is similar, the colors are different.
So we would like a good perceptual hashing algorithm to generate hashes for these images that are close, but not to put them in the same bucket.
These are more similar.
We have the same subject, same coloring, very similar brightness and contrast values.
If we look a bit closer we see that the rotation is a bit different in both images, the left one is a bit sharper, and the right image has more "empty space" around the subject.
All things considered, these image should probably be classified as very similar by most perceptual hashing algorithms.
Resize
The first step is to size the image down. Since we are aiming for a 64-bit descriptor, the logical target size is 8 by 8 since that will give us 64 pixels to work with.
It is obvious, but worth noting, that this action is destructive.
We are throwing away a lot of information here, especially regarding the finer details.
For now, this is a good thing.
After all, we have to describe our image in only 64-bits, so throwing away information was inevitable.
If you squint your eyes the original image can be recognized. Sort of.
DCT
What is DCT?
Frequency domain. Plaatje. Bla bla.
Mauris semper ipsum libero. Praesent laoreet massa sagittis enim consequat malesuada. Nullam viverra nibh sit amet lacus volutpat sollicitudin. Nullam lacus sem, commodo ut vulputate nec, consectetur in erat. Quisque eget nunc ac felis condimentum ultrices eu sit amet arcu. Praesent imperdiet faucibus aliquam. Curabitur sit amet faucibus erat.
Nam pulvinar, mi id sagittis laoreet, risus dolor pellentesque metus, sed mattis nunc erat sed urna. Ut facilisis, velit vel condimentum euismod, ex ante dictum leo, lobortis feugiat eros sem nec velit. Aliquam quis eros lacus. Duis venenatis purus at luctus tempor. Praesent gravida euismod ante, tristique placerat tortor mattis molestie. Duis viverra ex eget lectus tempus consectetur. Aliquam semper, ligula at molestie dapibus, sapien turpis rutrum massa, ac tristique lorem ex eget neque. Donec vel turpis odio.
Mauris semper ipsum libero. Praesent laoreet massa sagittis enim consequat malesuada. Nullam viverra nibh sit amet lacus volutpat sollicitudin. Nullam lacus sem, commodo ut vulputate nec, consectetur in erat. Quisque eget nunc ac felis condimentum ultrices eu sit amet arcu. Praesent imperdiet faucibus aliquam. Curabitur sit amet faucibus erat.
Mauris semper ipsum libero. Praesent laoreet massa sagittis enim consequat malesuada. Nullam viverra nibh sit amet lacus volutpat sollicitudin. Nullam lacus sem, commodo ut vulputate nec, consectetur in erat. Quisque eget nunc ac felis condimentum ultrices eu sit amet arcu. Praesent imperdiet faucibus aliquam. Curabitur sit amet faucibus erat.
Mauris semper ipsum libero. Praesent laoreet massa sagittis enim consequat malesuada. Nullam viverra nibh sit amet lacus volutpat sollicitudin. Nullam lacus sem, commodo ut vulputate nec, consectetur in erat. Quisque eget nunc ac felis condimentum ultrices eu sit amet arcu. Praesent imperdiet faucibus aliquam. Curabitur sit amet faucibus erat.
Mauris semper ipsum libero. Praesent laoreet massa sagittis enim consequat malesuada. Nullam viverra nibh sit amet lacus volutpat sollicitudin. Nullam lacus sem, commodo ut vulputate nec, consectetur in erat. Quisque eget nunc ac felis condimentum ultrices eu sit amet arcu. Praesent imperdiet faucibus aliquam. Curabitur sit amet faucibus erat.
Nam pulvinar, mi id sagittis laoreet, risus dolor pellentesque metus, sed mattis nunc erat sed urna. Ut facilisis, velit vel condimentum euismod, ex ante dictum leo, lobortis feugiat eros sem nec velit. Aliquam quis eros lacus. Duis venenatis purus at luctus tempor. Praesent gravida euismod ante, tristique placerat tortor mattis molestie. Duis viverra ex eget lectus tempus consectetur. Aliquam semper, ligula at molestie dapibus, sapien turpis rutrum massa, ac tristique lorem ex eget neque. Donec vel turpis odio.
Nam pulvinar, mi id sagittis laoreet, risus dolor pellentesque metus, sed mattis nunc erat sed urna. Ut facilisis, velit vel condimentum euismod, ex ante dictum leo, lobortis feugiat eros sem nec velit. Aliquam quis eros lacus. Duis venenatis purus at luctus tempor. Praesent gravida euismod ante, tristique placerat tortor mattis molestie. Duis viverra ex eget lectus tempus consectetur. Aliquam semper, ligula at molestie dapibus, sapien turpis rutrum massa, ac tristique lorem ex eget neque. Donec vel turpis odio.
Nam pulvinar, mi id sagittis laoreet, risus dolor pellentesque metus, sed mattis nunc erat sed urna. Ut facilisis, velit vel condimentum euismod, ex ante dictum leo, lobortis feugiat eros sem nec velit. Aliquam quis eros lacus. Duis venenatis purus at luctus tempor. Praesent gravida euismod ante, tristique placerat tortor mattis molestie. Duis viverra ex eget lectus tempus consectetur. Aliquam semper, ligula at molestie dapibus, sapien turpis rutrum massa, ac tristique lorem ex eget neque. Donec vel turpis odio.
Nam pulvinar, mi id sagittis laoreet, risus dolor pellentesque metus, sed mattis nunc erat sed urna. Ut facilisis, velit vel condimentum euismod, ex ante dictum leo, lobortis feugiat eros sem nec velit. Aliquam quis eros lacus. Duis venenatis purus at luctus tempor. Praesent gravida euismod ante, tristique placerat tortor mattis molestie. Duis viverra ex eget lectus tempus consectetur. Aliquam semper, ligula at molestie dapibus, sapien turpis rutrum massa, ac tristique lorem ex eget neque. Donec vel turpis odio.