Photo to emoji
Turn a photo into the emoji that fit it, without storing the image.
On this page
Someone shares a photo, and your app suggests the emoji that fit it: 🎂 for a birthday cake, 🐶 for a puppy. Use it for reaction suggestions on images, for alt-text helpers, or to tag uploads.
How it works
You send a small image to POST /v1/classify-image. A vision model on Workers AI looks at the image once. It gives a short caption, the reaction a person would likely have, 3–6 keywords and up to 8 emoji of its own choice. Emojisense then ranks emoji from three signals and fuses them into one list:
| Signal | What it adds |
|---|---|
| Proposed emoji | The emoji the model chose. Each one is checked against the catalog, and unknown ones are dropped. This signal counts the most, because the model sees the image. |
| Keywords | An alias search for each keyword, whole words only. An exact match counts more than a near match. |
| Caption | The emoji nearest to the caption embedding. This signal finds the most candidates, but its far neighbours are weak. |
An emoji with too little evidence is dropped, so the list can be shorter than limit. Gender and direction variants (🚵 🚵♀️ 🚵♂️) appear once.
Call it
Downscale the image in the client first, to about 384 px on the long edge. The API accepts JPEG or WebP up to 256 KB. Smaller images are faster and cost the same.
/** Downscale to about 384 px on the long edge, as a JPEG, before sending. */
async function downscale(file: Blob, maxEdge = 384): Promise<Blob> {
const bitmap = await createImageBitmap(file, { imageOrientation: "from-image" });
const scale = Math.min(1, maxEdge / Math.max(bitmap.width, bitmap.height));
const canvas = document.createElement("canvas");
canvas.width = Math.round(bitmap.width * scale);
canvas.height = Math.round(bitmap.height * scale);
canvas.getContext("2d")?.drawImage(bitmap, 0, 0, canvas.width, canvas.height);
bitmap.close();
return new Promise((resolve, reject) => {
canvas.toBlob(
(blob) => (blob ? resolve(blob) : reject(new Error("encode failed"))),
"image/jpeg",
0.85,
);
});
}
export async function emojiForPhoto(file: File): Promise<string[]> {
const url = "https://api.emojisense.com/v1/classify-image?key=pk_live_…&limit=5";
const response = await fetch(url, {
method: "POST",
headers: { "Content-Type": "image/jpeg" },
body: await downscale(file),
});
if (!response.ok) return [];
const { results } = await response.json();
return results.map((result) => result.emoji);
}curl -X POST "https://api.emojisense.com/v1/classify-image?locale=en&limit=5" \
-H "Authorization: Bearer sk_live_…" \
-H "Content-Type: image/jpeg" \
--data-binary @photo-384.jpg| Part | Value |
|---|---|
| Body | The image bytes, Content-Type: image/jpeg or image/webp, at most 256 KB. |
?locale= | Checked as in search. Default en. The label is in English, so the keywords are always matched with English aliases. |
?limit= | Number of results, 1–50. Default 8. |
X-Image-Hash | Optional. A 64-bit perceptual hash of the image as 16 hex characters. It turns on the label cache. |
Response
The real answer for a photo of a birthday cake with a “3” candle, sent as a cake.jpg (384 px JPEG, 20 KB):
{
"caption": "A birthday cake with a lit number three candle and star decorations",
"reaction": "Happy 3rd birthday!",
"keywords": [
"birthday cake",
"number three candle",
"star decorations",
"celebration",
"party"
],
"results": [
{ "emoji": "🎂", "id": "1F382", "score": 1, "source": "semantic" },
{ "emoji": "🥳", "id": "1F973", "score": 0.721, "source": "semantic" },
{ "emoji": "🕯️", "id": "1F56F", "score": 0.6, "source": "semantic" },
{ "emoji": "🎉", "id": "1F389", "score": 0.571, "source": "semantic" },
{ "emoji": "🎈", "id": "1F388", "score": 0.488, "source": "semantic" }
],
"cached": false,
"degraded": false,
"overLimit": false
}keywords are the model’s search words, most important first. score is the fused evidence, from 0 to 1. 🎂 gets 1 because several signals agree on it. source names the signal that added the most: semantic for a proposed emoji or the caption, alias for a keyword match. Vision models are not fully deterministic: the same photo can get slightly different keywords and emoji.
Cache repeated images
The same meme is often shared many times. Send X-Image-Hash with a perceptual hash (for example a 64-bit difference hash), and the API caches the label for that hash in the data center. Later calls with the same hash skip the vision model. The cache key also holds the vision model and the prompt version, so a new prompt never reads an old label. Every call still counts as one classification, also when it is a cache hit.
Plans and limits
| Plan | Classifications a month |
|---|---|
| Free | 100 |
| Solo | 1k |
| Pro | 10k |
| Scale | 75k |
Over the limit, the API answers 200 with "overLimit": true and no results until the next month. There is no on-device vision model, so show your normal picker in that case. If the vision model is unavailable, the answer has "degraded": true, no results, and is not counted. If only the caption embedding is unavailable, the answer has "degraded": true and fewer results, from the proposed emoji and the keywords.
Errors: 400 for a wrong content type, an unreadable image or a bad hash, and 413 for an image over 256 KB.