REST API · Text-to-Speech · 33 Languages

TTS API vs OpenAI TTS

Side-by-side comparison of the TTS API vs OpenAI TTS. PixelAPI's text-to-speech API converts text to natural-sounding audio in 33 languages, with voice design (describe a voice in plain text) and voice cloning (upload a reference recording). At $0.0018 flat per request — no per-character metering — it costs a fraction of OpenAI tts-1's $0.015 per 1,000 characters. 5,000 free credits (7-day trial), no credit card.

$0.0018 / request 33 languages Voice design + cloning 5,000 free credits Flat rate, not per-char Async job + polling
Get an API key (free) Quick start See pricing API docs

Quick start — one API call

Sign up, copy your key from the dashboard, and POST your text. The endpoint returns a generation id; poll until status=completed, then download the MP3 from output_url.

# Generate speech (voice design — no reference file needed)
curl -X POST https://api.pixelapi.dev/v1/tts/generate \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -F "text=Hello, welcome to our platform." \
  -F "language=en" \
  -F "voice_description=young woman, warm and friendly"
# {"id": "uuid", "status": "queued", "credits_used": 1.8, ...}

# Poll until completed
curl https://api.pixelapi.dev/v1/tts/status/UUID \
  -H "Authorization: Bearer YOUR_API_KEY"
# {"status": "completed", "output_url": "https://cdn.pixelapi.dev/..."}
pip install pixelapi
# ---
from pixelapi import PixelAPI

client = PixelAPI(api_key="YOUR_API_KEY")
result = client.tts(
    text="Hello, welcome to our platform.",
    language="en",
    voice_description="young woman, warm and friendly"
)
result.save("output.mp3")  # MP3 audio
npm install pixelapi
// ---
import { PixelAPI } from "pixelapi";

const client = new PixelAPI({ apiKey: process.env.PIXELAPI_KEY });
const result = await client.tts({
  text: "Hello, welcome to our platform.",
  language: "en",
  voiceDescription: "young woman, warm and friendly",
});
await result.save("output.mp3"); // MP3 audio
composer require pixelapi/pixelapi
// ---
<?php
use PixelAPI\Client;

$client = new Client(getenv("PIXELAPI_KEY"));
$result = $client->tts([
    "text"             => "Hello, welcome to our platform.",
    "language"         => "en",
    "voice_description"=> "young woman, warm and friendly",
]);
file_put_contents("output.mp3", $result->getBody());
gem install pixelapi
# ---
require "pixelapi"

client = PixelAPI::Client.new(api_key: ENV["PIXELAPI_KEY"])
result = client.tts(
  text: "Hello, welcome to our platform.",
  language: "en",
  voice_description: "young woman, warm and friendly"
)
File.binwrite("output.mp3", result.body)
go get github.com/pixelapi/pixelapi-go
// ---
import "github.com/pixelapi/pixelapi-go"

client := pixelapi.New("YOUR_API_KEY")
result, err := client.TTS(pixelapi.TTSRequest{
    Text:             "Hello, welcome to our platform.",
    Language:         "en",
    VoiceDescription: "young woman, warm and friendly",
})
if err != nil { panic(err) }
result.Save("output.mp3")

TTS API vs OpenAI TTS: Pricing comparison

OpenAI TTS charges per character; PixelAPI charges a flat rate per request — regardless of whether your text is 50 characters or 500. The difference compounds fast at scale.

Provider Free tier Per-unit cost Voice cloning Languages
PixelAPI 5,000 credits (7-day trial) $0.0018 flat/request ✓ $0.10/req 33
OpenAI tts-1 None $0.015 per 1K chars ✗ 57
OpenAI tts-1-hd None $0.030 per 1K chars ✗ 57
ElevenLabs Flash/Turbo 10K chars/mo $0.050 per 1K chars ✓ 32
ElevenLabs Multilingual v2/v3 10K chars/mo $0.100 per 1K chars ✓ 32
PlayHT (PlayAI) 12.5K chars/mo ~$0.040 per 1K chars ✓ 142
Murf.ai — see murf.ai/pricing ✗ 20+

Pricing verified from each rival's public pricing page September 2026. Note: PlayHT was acquired by Meta in July 2025; check play.ht for current availability. PixelAPI's flat-rate pricing means a 100-character text costs the same as a 500-character text — no surprise bills on short utterances.

What you get back

MP3 audio download

Every completed job returns an output_url pointing to an MP3 audio file. Download it directly or stream it from your application. No re-encoding needed.

33-language support

Pass language=auto to detect language automatically, or specify a code: en, zh, hi, es, fr, de, ja, ko, ru, ar, and 23 more. Same endpoint for all languages — no model switching.

Voice design (no reference file)

Pass a natural-language description in voice_description — e.g. "middle-aged man, calm and authoritative" or "energetic young woman, podcast host". Control how closely the voice follows the prompt with cfg_value (0.5–5.0).

Voice cloning (reference audio)

Upload a reference WAV, MP3, or M4A (16 kHz+, up to 10 MB) in the voice_ref field. The API synthesises new speech in the cloned voice. A 5–10 second clean recording is enough. Billed at $0.10 per request.

Common workflows

The TTS API powers these production workflows. Each links to a setup guide:

Audiobook production

Convert manuscript chapters to audio. Voice design for narrators, auto language detection for multi-lingual titles.

Podcast automation

Generate episode intros, ad reads, and show-notes summaries without a recording booth.

IVR & phone systems

Dynamic IVR prompts and hold messages at $0.0018/call — far cheaper than recording new audio.

Language learning apps

Native-sounding pronunciation for 33 languages. Flat per-request pricing keeps costs predictable at learning-app scale.

E-learning narration

Auto-narrate course transcripts and slide decks in any of the 33 supported languages.

Content repurposing

Turn blog posts, newsletters, and articles into audio tracks for distribution on podcast platforms.

Integrations & SDKs

Zapier

No-code TTS trigger — connect any Zapier source to speech output in one step.

Make.com

Drag-and-drop TTS module in Make scenarios — great for newsletter-to-audio pipelines.

Shopify

Auto-generate product description audio for accessibility and listen-while-browse features.

Next.js

Server-side TTS in Next.js API routes — stream audio to the browser or cache to S3.

Webflow

Embed audio players on CMS-backed Webflow pages with auto-generated TTS audio.

WooCommerce

Accessibility-ready audio for WooCommerce product pages, generated on publish.

TTS API vs OpenAI TTS — and other alternatives

vs OpenAI TTS

Flat per-request pricing vs OpenAI's per-character billing. PixelAPI adds voice cloning; OpenAI does not.

vs ElevenLabs

$0.0018/request vs ElevenLabs Flash at $0.05/1K chars. Similar voice cloning quality; simpler pricing model.

vs Murf.ai

Developer REST API vs Murf's studio-first product. PixelAPI is built for programmatic bulk use.

vs PlayHT

PixelAPI charges flat per request; PlayHT billed per character. Note: PlayHT was acquired by Meta in July 2025.

vs Azure TTS

No Azure account or region setup required. Single API key, simpler billing, voice cloning built in.

vs Google TTS

No GCP project needed. Flat per-request pricing vs Google's per-character tiers. Voice design out-of-the-box.

Rate limits & error handling

The TTS endpoint shares the same per-plan rate limits as the rest of the PixelAPI platform: 20 requests/minute on the free trial, 60 on Starter, 120 on Pro, and 300 on Scale — with 3/10/20/50 parallel jobs respectively. Exceeding the limit returns HTTP 429 with a Retry-After header.

# Python SDK handles 429 automatically with exponential backoff
from pixelapi import PixelAPI
client = PixelAPI(api_key="...", max_retries=4)
result = client.tts(text="Hello world", language="en")  # auto-retries on 429

TTS jobs are asynchronous. Common status values:

Transient GPU out-of-memory failures (rare) are auto-retried up to twice server-side before surfacing as failed. Credits are never charged for a failed job.

Frequently asked questions

How does PixelAPI TTS API compare to OpenAI TTS?

PixelAPI charges $0.0018 flat per request with no per-character metering, while OpenAI tts-1 costs $0.015 per 1,000 characters. PixelAPI also supports voice cloning (upload a reference audio file) and voice design (describe a voice in plain text) — neither feature is available in OpenAI's TTS endpoint. Both APIs are asynchronous REST services returning an audio download URL.

What does the TTS API cost?

$0.0018 per text-to-speech request (1.8 credits at $0.001/credit) for voice design or built-in voices. Voice cloning requests (with a reference audio upload) cost $0.10 per request (100 credits). New accounts get 5,000 free credits on a 7-day trial — no credit card required — enough for over 2,700 standard TTS generations to test on real content.

How many languages does the TTS API support?

33 languages: English, Chinese, Hindi, Spanish, French, German, Japanese, Korean, Russian, Arabic, Portuguese, Italian, Dutch, Polish, Turkish, Vietnamese, Thai, Indonesian, Malay, Bengali, Tamil, Telugu, Marathi, Ukrainian, Swedish, Norwegian, Danish, Finnish, Greek, Hebrew, and Swahili — plus language=auto for automatic detection. See the languages endpoint (GET /v1/tts/languages) for the full list.

What is voice design and how does it work?

Voice design lets you describe a voice in plain English — e.g. "energetic young woman, podcast host" or "elderly man, storytelling tone" — and the API synthesises speech in that style without any reference audio. Pass your description in the voice_description field. Tune adherence to the description with cfg_value (0.5–5.0) and quality/speed tradeoff with inference_timesteps (4–20).

What is voice cloning and what audio format do I upload?

Voice cloning lets you upload a reference audio file and the API synthesises new speech in that voice. Supported input formats: WAV, MP3, or M4A at 16 kHz or higher, up to 10 MB. A clean 5–10 second recording delivers the best clone quality; background noise degrades results. Voice cloning is billed at $0.10 per request (100 credits).

What audio format does the TTS API return?

The completed job response includes an output_url pointing to an MP3 audio file. The job is asynchronous: POST to /v1/tts/generate (returns a id), poll GET /v1/tts/status/{id} until status=completed, then fetch the audio from output_url.

Is there a free trial?

Yes. New accounts receive 5,000 free credits on a 7-day trial with no credit card required. At 1.8 credits per standard TTS request, that's over 2,700 generations to test on real content before committing to a paid plan. Each network is eligible for one trial.

What are the rate limits?

20 requests/minute and 3 concurrent jobs on the free trial; 60/10 on Starter; 120/20 on Pro; 300/50 on Scale. Exceeding the limit returns HTTP 429 with a Retry-After header. The Python and Node SDKs retry 429 responses automatically. Need higher limits for batch processing? Email support@pixelapi.dev with your expected volume.

How do I poll for TTS job completion?

After your POST to /v1/tts/generate returns a JSON object with an id field, call GET /v1/tts/status/{id} with your API key. Poll every 1–3 seconds until the status field equals "completed". The response will include output_url for download. Both the Python and Node SDKs handle the polling loop automatically.

Can I use PixelAPI TTS for commercial projects? Is there an invoice?

Yes. All paid plans include commercial usage rights. PixelAPI is a registered Indian business that issues GST invoices — 18% IGST for international clients, CGST/SGST for domestic clients. Invoice download is built into the dashboard. For enterprise pricing and SLA agreements, contact support@pixelapi.dev.

Start free — 5,000 credits, no card Read full API docs Compare all plans