Confession: I got nerd-sniped by Sam Henri Goldâs request for my icon galleries:
I'd like to humbly request artwork-level searching in macosicongallery.com
What follows is a train-of-thought blog post as I play with what an implementation might look like.
Iâve actually long-wanted something like this, e.g. let me search for âcoffeeâ and show me all icons that have some depiction of coffee in them.
Similarly, Iâve wanted some kind of ârelatedâ representation for icons. I have this today via existing metadata, e.g. âShow me other icons in the category âProductivityââ or âShow me other icons tagged as âorangeââ. But Iâve wanted a more robust representation of this, so if you were looking at an icon that had a microphone in it, the site would say âHere are other icons that also have microphones in them.â And the relationship would be rich/smart enough to know that âmicrophoneâ was meant broadly, i.e. dynamic mics, condenser mics, ribbon mics, etc.
So how would you do this? I could go through every icon one-by-one and classify/tag any attribute of its design that comes to mind, but that would take ages! Seems like a good use case for a vision model.
Trying CLIP
First, Iâll look at Samâs suggestion: run every image through CLIP.
Iâm not familiar with CLIP so I start with a little research: What is it? How would I use it? And most importantly: is it free/open (because I ainât spending a ton of money to send my thousands of icon PNGs to an AI provider via their API)?
Ok, so CLIP will take an image and spit back an embedding (basically a bunch of numbers representing features of the image). When you do it with multiple images, you can then compare those embeddings to see what the model considers similar (and, if you like, set a threshold for what constitutes a âmatchâ).
After getting a sense of the task in front of me, I work with the LLM to come up with a proof of concept. I donât need to fit this into my existing site. I just want to make one-off HTML pages where I can feel out, âCan this process create anything useful? Whatâs the amount of work required?â
- Write a script that runs a sampling of icons through CLIPâs image encoder
- Read the file locally, e.g.
./ios/256/${icon.id}.png
- 256x256 pixel icons seem to be enough, as the CLIP model Iâm using preprocesses them to ~224px anyway.
- Create a dataset representing the âembeddingsâ (an array of numbers) for each icon that I get from CLIP, e.g.
Array<{ id: String, embedding: Array<number> }>
- Create a dataset representing the top matches between different embeddings, e.g.
{ [id: String]: [id, id, âŚ] }
- Create a
clip.html file has both datasets (plus supplementary icon metadata I already have), render all the sampled icons, and support an onclick for each icon that shows the related[id] icons.
This is enough to create a single HTML file where I can click on an icon and see other icons that look like it.
However, I realize quickly that Iâll need to process my entire icon library to really get a good sense for how well these are matching. So I do that.
[Computer goes brrrrâŚ]
Ok, now when I click on an icon that looks like a camera, I see other icons that look like cameras.
Or if I click on an icon that has a checkmark in it, I see other icons with checkmarks in them â sort-of.
But the results arenât that great unless an icon is visually distinctive. I share some thoughts with Sam. He has a few other suggestions I follow.
DINOv2, SigLIP2, and More
Sam mentions SigLIP2 so I start with that as a keyword. The LLM recommends DINOv2 so I say, âLetâs try itâ.
I give that a try, creating a separate dataset and prototype (e.g. embeddings-dinov2.json and embeddings-dinov2.html) so I can continue to view these different prototypes and compare their outputs.
Itâs fine. Different from CLIP. Honestly not much better.
So I figure letâs try another one. I go with SigLIP2. I ask the LLM to create a page where I can compare the results.
Seems like six of one, half dozen of another. One does better on some kinds of icons, worse on others. The LLM recommends that, at this point, I be done shopping models. Theyâre roughly the same class of tool with different tradeoffs. None are breakthroughs.
So now what?
Try Tagging Icons With Keywords
Sam recommends another approach:
You could also try handing all icons over to a VLM, having it write up a description, and embedding THAT text against what people might search for.
A thoroughly detailed person mightâve done this from the start, e.g. for an icon thatâs a checkmark, add the keyword âcheckmarkâ to its metadata.
That would take me forever to go back through all my icons and do â a perfect task for a computer that never tires.
So I give this a try. First I need a free/open VLM. After a little research I decide to try Moondream via Ollama.
I have the machine go through each image and caption it, then pull out âtagsâ from the caption. For the Clear app icon, I get data like this:
{
"id": "clear-todos-2021-01-10",
"caption": "The image features a red and orange gradient background, with a white checkmark in the center. The checkmark is slightly tilted to the right, giving it a dynamic appearance. The background transitions from red at the top to orange at the bottom, creating a sense of depth and movement. The checkmark is the main object in the image, occupying most of the space and drawing attention to itself.",
"tags": [
"red",
"orange",
"gradient",
"white",
"checkmark",
"slightly",
"tilted",
"dynamic",
"appearance",
"transitions",
"creating",
"sense",
"depth",
"movement",
"object",
"occupying"
]
}
Then the LLM creates a single search.html file where I can test icon matches by searching for tag overlaps (or choosing one of the popular ones).
So, for example, on the search page I can click on âcheckmarkâ and see all the icons with a checkmark.
Or click on âfoxâ and see all the icons with a fox.
Matches are pretty spot to be honest.
But thatâs a different kind of test than what I was doing with CLIP.
- CLIP: click on an icon and see other icons like it.
- Tags: click on a keyword and see other icons with that keyword.
Can I leverage tags for the same kind of ârelated iconsâ work that CLIP is doing?
I get the LLM to cook up a single-page HTML file where I can compare âclick on this icon and find other icons like itâ where Iâm using embeddings from CLIP vs. matching on keywords.
The results seem to fare much better for CLIP. For example, here I matched on what I think of as a âcheckmark iconâ.
You can see the approach that matches on tags didnât work too great. I believe this is because with my simple tag-overlap approach, a distinctive keyword like âcheckmarkâ gets diluted amongst generic tags like âsquareâ, âblueâ, and âsimpleâ.
Whereas with CLIP, if you click on an icon with a checkmark, you get other checkmarks (and not other icons that also have related tags like âsquareâ, âblueâ and âsimpleâ).
Which all makes sense. Pushing on the implementation here could help, but thatâs separate work to do.
So Now What?
Iâm not sure.
While doing all of this was an interesting technical exercise, there are a few important considerations I need to think through before implementing anything, such as:
- What kind of functionality do I actually want?
- A ârelated iconsâ feature? Does it match on keywords or embeddings?
- A âsearchâ feature that matches on keywords?
- Both?
- How do I build these features into my codebase now, given the thousands of icons that already exist?
- How do I maintain this feature in the future?
- e.g. every time I add a new icon to my gallery, is a VLM now a dependency of this project?
- Given all the above, whatâs the time and money cost?
Iâm very picky about adding new dependencies to these icon projects. I like to think thatâs why Iâve been able to maintain and continue contributing to them after so many years â because I make it easy on myself (good job, Past Jim).
So, for example, if I make a VLM a dependency of this project such that every time I add a new icon I have to run it through to create the embeddings, thatâs a big dependency cost IMO. Iâm not sure I want to do that.
That said, Apple now ships foundation models in macOS 27 available through the CLI (go ahead, try typing fm in your Terminal if youâre on Golden Gate). So if my Mac continues to be the primary machine where I add/update metadata for my icon projects, using fm would be a really easy/low-cost way to process each new icon to generate a caption and keywords for matching in search.
But again, I donât know if I want to do that. I wrote this post to try and work through what I want to do, but I am still undecided.
So I guess the only thing for me to do at this point is hit âPublishâ on this post and keep simmering on a decision.
Reply via:
Email
¡ Mastodon ¡
Bluesky