← Back to all tools

TOOL · 98 · AI

AI Image Caption Generator

Generate an English caption for any photo using an open-source ViT-GPT2 model — 100% in-browser, no signup. First run ~180 MB.

In short
AI Image Caption Generator describes any photo in plain English using the open-source ViT-GPT2 vision-language model. Runs in your browser via transformers.js — great for accessibility alt text or quick auto-tagging.

Runs in your browser after a one-time model download. Your files stay on your device.

Drop files here or click to browse

JPG, PNG or WebP

No signup · No watermark · Files never leave your device · Up to ~300 MB recommended

How it works

  1. Step · 01

    Drop a photo — JPG, PNG or WebP.

  2. Step · 02

    First run downloads the open-source ViT-GPT2 vision-language model (~180 MB), cached after that.

  3. Step · 03

    The caption is generated locally in your browser — the image never leaves your device.

Related tools

Frequently asked questions

Which model is used?
Xenova/vit-gpt2-image-captioning — a Vision Transformer encoder paired with GPT-2 for caption decoding, ported to run in the browser.
How long is the caption?
One short sentence. It's an image caption model, not a detailed image describer — for longer descriptions, use it as a starting point.
Is the image uploaded?
No. The image is processed entirely on your device after the one-time ~180 MB model download.
Are my files uploaded to a server?
No. Files are processed directly in your browser using JavaScript and WebAssembly. Nothing is uploaded to our servers, so your documents never leave your device.
Does it work on mobile?
Yes. The tools run in any modern browser on desktop, tablet, or mobile (Chrome, Edge, Safari, Firefox). Very large files may be slower on low-end phones.