Upload any image. A pre-trained vision model reads the scene, an attention-based decoder writes the caption underneath — accurate, automatic, and instant.
Configure model settings, upload your images, and get AI-generated captions with attention visualization in real time.
Drag & drop photos, click to browse, or paste from clipboard
JPG, PNG, WEBP — UP TO 10MBUpload, analyze, generate — simple as that.
A pre-trained ResNet50 or VGG16 backbone extracts spatial features from your image — no training data needed from you.
An attention-based decoder focuses on relevant image regions for each word, building a natural sentence one token at a time.
Compare ranked caption candidates, edit any of them inline, visualize attention maps, and export your results as CSV or JSON.
One POST request, ranked captions back. Connect to a CMS, a media library, or your own batch processing workflow.
// POST /api/caption const res = await fetch("/api/caption", { method: "POST", headers: { "Authorization": "Bearer YOUR_KEY", "Content-Type": "application/json" }, body: JSON.stringify({ image_base64: "data:image/jpeg;base64,...", beam_width: 5, num_captions: 3 }) }); const data = await res.json(); // data.captions = [ // { text: "A warm sunset...", score: 0.93 }, // { text: "A golden landscape...", score: 0.84 } // ] // data.analysis = { scene_type, dominant_tone, lighting }