A production REST API that analyzes any image and returns a natural language caption, an array of SEO tags, and a WCAG-compliant alt text string — all in one HTTP call. $0.0004 per image — over 60% cheaper than AWS Rekognition ($0.001) and over 73% cheaper than Google Cloud Vision ($0.0015). The cheapest VLM caption API available. 5,000 free credits on a 24-hour trial, no credit card required.
Sign up, copy your key from the dashboard, and POST your image. Choose a mode and style, then poll the generation endpoint until status=completed to read the caption, tags, and alt text.
curl -X POST https://api.pixelapi.dev/v1/image/caption \
-H "Authorization: Bearer YOUR_API_KEY" \
-F "image=@product.jpg" \
-F "mode=full" \
-F "style=product" \
-F "max_tags=15"
# Response: {"generation_id": "uuid", "status": "queued", "credits_used": 0.0004}
# Poll until completed
curl https://api.pixelapi.dev/v1/image/UUID \
-H "Authorization: Bearer YOUR_API_KEY"
# Response: {
# "status": "completed",
# "caption": "A white ceramic coffee mug with a minimal handle on a natural wood table.",
# "tags": ["coffee mug", "ceramic", "white", "minimalist", "kitchen"],
# "alt_text": "White ceramic coffee mug on wooden table."
# }
pip install pixelapi
---
from pixelapi import PixelAPI
client = PixelAPI(api_key="YOUR_API_KEY")
result = client.caption(
image="product.jpg",
mode="full", # full | caption | tags | alt_text
style="product", # product | creative | seo | technical
max_tags=15,
)
print(result.caption) # natural language description
print(result.tags) # ["coffee mug", "ceramic", ...]
print(result.alt_text) # "White ceramic coffee mug on wooden table."
npm install pixelapi
---
import { PixelAPI } from "pixelapi";
const client = new PixelAPI({ apiKey: process.env.PIXELAPI_KEY });
const result = await client.caption({
image: "./product.jpg",
mode: "full", // full | caption | tags | alt_text
style: "seo", // product | creative | seo | technical
maxTags: 15,
});
console.log(result.caption); // natural language description
console.log(result.tags); // string[]
console.log(result.altText); // WCAG alt text string
composer require pixelapi/pixelapi
---
<?php
use PixelAPI\Client;
$client = new Client(getenv("PIXELAPI_KEY"));
$result = $client->caption([
"image" => "product.jpg",
"mode" => "full",
"style" => "product",
"max_tags" => 15,
]);
echo $result->caption;
print_r($result->tags);
echo $result->alt_text;
gem install pixelapi --- require "pixelapi" client = PixelAPI::Client.new(api_key: ENV["PIXELAPI_KEY"]) result = client.caption( image: "product.jpg", mode: "full", style: "product", max_tags: 15, ) puts result.caption puts result.tags.inspect puts result.alt_text
go get github.com/pixelapi/pixelapi-go
---
import "github.com/pixelapi/pixelapi-go"
client := pixelapi.New("YOUR_API_KEY")
result, err := client.Caption(pixelapi.CaptionParams{
Image: "product.jpg",
Mode: "full",
Style: "product",
MaxTags: 15,
})
if err != nil { panic(err) }
fmt.Println(result.Caption)
fmt.Println(result.Tags)
fmt.Println(result.AltText)
| Provider | Free tier | Per-image cost | Output |
|---|---|---|---|
| PixelAPI | 5,000 credits (24-hour trial), no card | $0.0004 | Caption + Tags + Alt Text (JSON) |
| AWS Rekognition | 1,000 images/month (12 months) | $0.0010 | Label objects (JSON) |
| Google Cloud Vision | 1,000 units/month | $0.0015 | Label objects (JSON) |
| Azure Computer Vision | 5,000 transactions/month | see azure.microsoft.com/en-us/pricing | Captions + Dense Captions |
| OpenAI GPT-4o Vision | trial credits | see openai.com/api/pricing | Free-form text (unstructured) |
AWS Rekognition pricing verified from aws.amazon.com/rekognition/pricing, October 2026. Google Cloud Vision pricing verified from cloud.google.com/vision/pricing, October 2026. Azure and OpenAI pricing pages could not be machine-read at time of publication — visit their respective pricing pages for current rates. PixelAPI's per-image price is set at half the cheapest verifiable mainstream rival per our pricing principle.
The caption API is a pure text API — it returns structured JSON, not images. Each response contains up to three fields depending on the mode you select.
mode=caption or mode=full. A one-to-two sentence description of the image contents. Pass style=product for e-commerce copy, style=creative for marketing, style=seo for keyword-dense descriptions, or style=technical for objective object/scene analysis.
mode=tags or mode=full. A JSON array of search-optimized keywords extracted from the image. Control the count with max_tags (1–50, default 10). Useful for populating Shopify product tags, Etsy keywords, and Pinterest board keywords automatically.
mode=alt_text or mode=full. A concise string (typically 50–125 characters) suitable for the HTML alt attribute, screen readers, and WCAG 2.1 Success Criterion 1.1.1 compliance. Significantly shorter and more prescriptive than the full caption.
mode=fullGet caption + tags + alt_text in a single request at the same $0.0004 price. No extra charge for receiving all three fields. Ideal for pipelines that need to populate alt text, product descriptions, and SEO metadata in one pass.
| Parameter | Value | What it does |
|---|---|---|
mode | full | Returns caption + tags array + alt_text (default) |
mode | caption | Natural language image description only |
mode | tags | SEO/search keyword array only |
mode | alt_text | WCAG accessibility alt text only |
style | product | E-commerce product description phrasing |
style | creative | Marketing/brand copywriting phrasing |
style | seo | Keyword-dense, search-optimized phrasing |
style | technical | Objective, factual object and scene analysis |
max_tags | 1–50 | Max number of tags to return (default 10) |
The caption API is purpose-built for these production automation patterns:
Bulk-generate WCAG alt text for thousands of product images on upload. Use mode=alt_text&style=product. Integrates with Shopify, WooCommerce, and BigCommerce product upload webhooks.
Auto-populate product tags and descriptions from catalog photos. mode=full&style=seo fills title candidates, meta descriptions, and keyword fields in one call per image.
Generate first-draft captions for Instagram, Facebook, and Pinterest image posts. Use style=creative to match brand voice. Pair with your moderation pipeline to filter content before publishing.
Retroactively fill missing alt text across a CMS or document archive. Batch-process image URLs, write alt_text output directly into your database, and pass automated WCAG 2.1 Level AA audits.
Extract structured tags from user-uploaded photos for faceted search and recommendation engines. Store the tags array in Elasticsearch, Typesense, or Algolia alongside the image URL.
Combine caption output with the Moderation API — caption provides context, moderation flags policy violations. Together they power a complete human-review-free content pipeline.
The Caption API is a standard REST endpoint compatible with any HTTP client or automation platform. Official SDKs ship retries, polling, and binary handling out of the box.
pip install pixelapi — includes client.caption(), auto-polling, retry on 429, and typed result objects. Works in Django, FastAPI, and any async or sync context.
npm install pixelapi — fully typed with TypeScript definitions. Works in Next.js API routes, Express, and serverless functions (Vercel, AWS Lambda, Cloudflare Workers).
composer require pixelapi/pixelapi — compatible with Laravel, Symfony, and plain PHP. Handles multipart form uploads and JSON response parsing automatically.
gem install pixelapi — works in Rails and Sinatra. Ships a PixelAPI::Client#caption method that handles file I/O and response wrapping.
go get github.com/pixelapi/pixelapi-go — idiomatic Go with context support. Suitable for high-throughput batch captioning pipelines and CLI tooling.
No SDK needed. Any language that can POST a multipart form (Rust, Java, Swift, Kotlin, .NET) works directly with the /v1/image/caption endpoint. See full API reference.
How PixelAPI's caption API compares to the mainstream vision APIs on price, output format, and developer experience:
AWS Rekognition DetectLabels returns structured label objects with confidence scores — useful for programmatic filtering but not human-readable. PixelAPI returns natural language captions and WCAG alt text alongside tags. At $0.0004 vs $0.001/image, PixelAPI is 60% cheaper with richer text output and no AWS account or IAM setup required.
Google Cloud Vision Label Detection ($0.0015/image) returns label annotations — structured but not sentence-form. PixelAPI returns ready-to-use alt text strings and copy-paste product descriptions alongside keyword tags. Over 73% cheaper, simpler auth (one API key vs service account JSON), and no GCP project setup.
Azure Computer Vision Caption and Dense Caption features require an Azure subscription and Cognitive Services resource. PixelAPI offers equivalent caption + alt text output with a single-key REST API, no Azure Portal required. Pricing on Azure's page is not publicly listed without an account — see azure.microsoft.com/en-us/pricing/details/cognitive-services/computer-vision/.
OpenAI Vision can describe images in rich detail, but billing is per-token (input + output), making per-image costs variable and harder to forecast. PixelAPI's flat $0.0004/image pricing makes costs predictable for bulk pipelines. Output is structured JSON rather than free-form text, so no parsing or prompt engineering is required.
Rate limits apply per API key. Exceeding the limit returns HTTP 429 with a Retry-After header indicating seconds to wait before retrying.
| Plan | Requests / minute | Concurrent jobs |
|---|---|---|
| Free trial | 20 | 3 |
| Starter | 60 | 10 |
| Pro | 120 | 20 |
| Scale | 300 | 50 |
The Python and Node SDKs handle 429 retries with exponential backoff automatically. For manual implementations, start at 2 seconds, double on each retry, and cap at 30 seconds.
# Python SDK — auto-retries 429 with backoff from pixelapi import PixelAPI client = PixelAPI(api_key="...", max_retries=4) result = client.caption(image="product.jpg", mode="full") # Retries automatically on 429; raises after max_retries exceeded
If an image is rejected (corrupt file, unsupported format, NSFW content flagged at the gate), the job returns a 400 or 422 error and credits are not deducted. Failed jobs that pass validation but error during processing are automatically detected and refunded.
PixelAPI Caption API at $0.0004 per image (0.4 credits). That is over 60% cheaper than AWS Rekognition ($0.001/image) and over 73% cheaper than Google Cloud Vision ($0.0015/image). New accounts get 5,000 free credits on a 24-hour trial — no credit card required — enough to caption thousands of product images before paying anything.
POST your image to https://api.pixelapi.dev/v1/image/caption with your API key, mode, and style. The endpoint returns a generation_id; poll GET /v1/image/{id} until status=completed, then read caption, tags, and alt_text from the response. See the Quick Start section above for ready-to-run code in 6 languages.
$0.0004 per image (0.4 credits). On the Starter plan ($10 for 10,000 credits), that is 25,000 caption calls per $10. New accounts receive 5,000 free credits on a 24-hour trial with no credit card required.
Four modes: full (caption + tags + alt_text in one call), caption (natural language description only), tags (keyword array only, up to 50 tags), and alt_text (WCAG-compliant accessibility string only). All modes cost the same $0.0004 per image. Use full to get everything in a single request.
Caption mode generates a natural language description optimized for human readers and content discovery — typically one to two sentences. Alt_text mode generates a shorter, WCAG 2.1-compliant string for screen readers — typically 50–125 characters. For accessibility compliance, use alt_text; for SEO product descriptions, use caption or full.
JSON. In mode=full you get { "caption": "...", "tags": ["tag1", "tag2", ...], "alt_text": "..." }. Single-mode calls return only the requested field. This is a pure text API — no images or binary data are returned, and no output file needs to be downloaded.
Yes. Use mode=full to get caption + tags + alt_text in one call, or mode=tags on its own. Control the number of tags with max_tags (1–50, default 10). Pass style=seo for keyword-dense tags or style=product for e-commerce-style descriptive tags.
JPG, PNG, and WebP. Maximum file size is 20 MB. The API validates the image format on submission and returns a 400 error for corrupt or unsupported files — no credits are deducted on rejected inputs.
Yes — pip install pixelapi. Official SDKs are also available for Node.js (npm install pixelapi), PHP (Composer), Ruby (Gem), and Go (go get github.com/pixelapi/pixelapi-go). All SDKs handle authentication, polling, retry-on-429, and typed response objects automatically.
20 requests/minute on the free trial, 60 on Starter, 120 on Pro, and 300 on Scale, with 3/10/20/50 concurrent jobs respectively. Exceeding the limit returns HTTP 429 with a Retry-After header. The Python and Node SDKs handle exponential backoff automatically. Higher limits for batch pipelines — email support@pixelapi.dev.
PixelAPI Caption API costs $0.0004/image vs AWS Rekognition DetectLabels at $0.001/image for the first 1M images — over 60% cheaper. PixelAPI also returns natural language captions and WCAG alt text alongside tags. No AWS account, IAM policy setup, or region configuration required — just one API key.
Yes. Use mode=alt_text to generate WCAG 2.1 SC 1.1.1-compliant alt text for any image. The output is a concise, screen-reader-optimized string suitable for the HTML alt attribute. Ideal for automating alt text on e-commerce product pages, CMS uploads, and digital documents where manual entry is impractical at scale.