🚀 One API for every frontier model on FastInfra

Browse models

Transparent per-token pricing → compare providers on FastInfra

View pricing

OpenAI-compatible chat completions → start in minutes

Read the docs

⚡ Server Auction → Enterprise bare metal from $78.15/mo (173 in stock)

Browse deals
Developer-first

Build on FastInfra

OpenAI-compatible REST API for chat, LTX-2.5 video, audio, and more. Swap the base URL to https://api.fastinfra.ai/v1 — see the endpoint reference and video API below.

Current rate limits (paid accounts)
Limit Value
Requests per minute (per API key) 600
Requests per minute (per account, all keys combined) 1800
Concurrent requests (per API key) 100
Concurrent requests (server-wide) 1024

Applies to every inference call (chat, video submit, audio, live transcription). There is no IP-based limit. Accounts without a completed top-up use the smaller free tier (60/min per key, 120/min per account, 4 concurrent). See Rate Limits for error handling and retry guidance.

API Overview

FastInfra exposes an OpenAI-compatible REST API at https://api.fastinfra.ai/v1 (same paths as OpenAI: chat, models, audio). Video generation uses a dedicated async job API — it is not available through POST /chat/completions.

Capability Method & path Notes
Chat completions POST https://api.fastinfra.ai/v1/chat/completions OpenAI-compatible; streaming supported
LTX-2.5 video POST https://api.fastinfra.ai/v1/videos/generations HTTP 202 + job id — poll below. Model: lightricks/ltx-2.5
Video job status GET https://api.fastinfra.ai/v1/videos/jobs/{id} Poll until status is completed or failed
Embeddings POST https://api.fastinfra.ai/v1/embeddings OpenAI-compatible; text-embedding-3-* ids accepted as aliases. See Embeddings
Speech-to-text POST https://api.fastinfra.ai/v1/audio/transcriptions Multipart upload; see Audio
Kokoro TTS POST https://api.fastinfra.ai/v1/audio/speech Standalone text→speech; model hexgrad/kokoro-82m (tts-1). See Kokoro TTS
Model catalog GET https://api.fastinfra.ai/v1/models Includes lightricks/ltx-2.5 when video is enabled
Model metadata GET https://api.fastinfra.ai/v1/models/{model_id} Validates a model id before use

Error format

Every error is JSON in the OpenAI shape, plus two fields OpenAI does not send: a stable code you can branch on, and a docs_url that points at the section of this page that explains the fix. Billing and auth errors also carry the URL to act on (topup_url, api_keys_url), and rate-limit errors carry retry_after_seconds.

{
  "error": {
    "message": "'qwen3.6-27b' is a prepaid model. API access to prepaid models unlocks after one completed top-up of at least $5 ...",
    "type": "payment_required",
    "code": "payment_required",
    "model": "qwen3.6-27b",
    "minimum_topup_usd": 5,
    "topup_url": "https://www.fastinfra.ai/Billing",
    "docs_url": "https://www.fastinfra.ai/Docs#credits"
  }
}
HTTPcodeMeaning
400invalid_request, invalid_inputMalformed body or unsupported parameter; the message names the field
401invalid_api_keyMissing, invalid, expired, or deactivated key
401key_not_linkedKey is not owned by a registered account
402payment_requiredPrepaid model before the first $5 top-up
402insufficient_quotaPrepaid model with an empty wallet
404model_not_foundUnknown model id
413payload_too_largeAudio upload over 25 MB
429rate_limit_*, server_capacity, queue_fullSee Rate Limits
502 / 503 / 504upstream_error, upstream_unavailable, upstream_timeoutInference backend failed after failover; retry

How to determine the API contract

LTX video is not part of OpenAI’s public spec — it uses FastInfra-specific paths under https://api.fastinfra.ai/v1. Use any of these sources (they stay in sync):

  1. This page — Video Generations is the canonical human reference: request fields, HTTP 202 submit, poll loop, response JSON, pricing, and error codes.
  2. OpenAPI — /openapi.json documents POST /videos/generations and GET /videos/jobs/{id} for codegen and AI assistants.
  3. Live model catalog — GET https://api.fastinfra.ai/v1/models lists lightricks/ltx-2.5 when video is enabled; GET https://api.fastinfra.ai/v1/models/lightricks/ltx-2.5 validates the id before you integrate.
  4. llms.txt — /llms.txt links back here and to OpenAPI for crawlers and LLM tooling.
Common mistake: sending video prompts to POST https://api.fastinfra.ai/v1/chat/completions. That endpoint is chat-only; video always goes through POST https://api.fastinfra.ai/v1/videos/generations → poll GET https://api.fastinfra.ai/v1/videos/jobs/{id}.

Quick video smoke test

# 1) Submit (returns immediately with HTTP 202)
curl -sS https://api.fastinfra.ai/v1/videos/generations \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"lightricks/ltx-2.5","prompt":"A red balloon floating over a city at sunset.","seconds":5,"size":"1280x704"}'

# 2) Poll (replace JOB_ID from the JSON "id" field; also in Location header)
curl -sS https://api.fastinfra.ai/v1/videos/jobs/JOB_ID \
  -H "Authorization: Bearer YOUR_API_KEY"

Authentication

All API requests require an API key. Create one from the API Keys page.

Pass your key using either header:

Authorization: Bearer YOUR_API_KEY
# or
X-Api-Key: YOUR_API_KEY

Point your OpenAI SDK at base_url="https://api.fastinfra.ai/v1" (include the /v1 suffix).

Credits

Paid chat completions and video generations require a positive prepaid balance. Usage is deducted after a successful response. Video jobs are billed on the first poll that returns status: "completed" — submitting a job (HTTP 202) is not billed. Free-tier models keep working at $0. GET https://api.fastinfra.ai/v1/models is not gated.

When the wallet is empty, paid chat requests return HTTP 402:

HTTP/1.1 402 Payment Required

{
  "error": {
    "message": "Insufficient API credits. Add funds on the Billing page, or use a free-tier model.",
    "type": "insufficient_quota",
    "code": "insufficient_quota"
  }
}

Rate Limits

Inference requests are rate limited to keep latency stable for every customer. The limits below are the complete list: anything that can return 429 is on this page. Limits only go up; they are never lowered under load. Production integrations must handle 429 Too Many Requests and honor the Retry-After response header.

Current limits

Limit Paid accounts Free / unverified accounts Scope
Requests per minute, per API key 600 60 Rolling 60-second window for one key
Requests per minute, per account 1800 120 Rolling 60-second window across every key the account owns
Tokens per minute, per API key 1,000,000 100,000 Chat completions only. Prompt estimate + max_tokens (default 2048) is reserved when the request is admitted and settled to actual usage when it finishes
Tokens per minute, per account 3,000,000 200,000 Across every key the account owns
Concurrent requests, per API key 100 4 In flight at the same time; a streamed response holds its slot until the stream ends
Concurrent requests, server-wide 1024 All API keys combined; a backstop, overflow routing absorbs GPU saturation first
Per IP address none Several sites on one server never share a quota

Paid means the account has completed a top-up on the Billing page. Admin and internal accounts always get the paid limits. Need more? Per-key limits can be raised on request; a human answers within hours.

What is limited

  • Rate limited: POST https://api.fastinfra.ai/v1/chat/completions, POST https://api.fastinfra.ai/v1/videos/generations, POST https://api.fastinfra.ai/v1/audio/transcriptions, POST https://api.fastinfra.ai/v1/audio/translations, POST https://api.fastinfra.ai/v1/audio/speech, and each live transcription WebSocket session (one concurrency slot for its whole duration)
  • Not rate limited: GET https://api.fastinfra.ai/v1/models, GET https://api.fastinfra.ai/v1/videos/jobs/{id}, account pages, and other non-inference routes
  • Under load: paying accounts are admitted ahead of free-tier accounts. Published limits never shrink; free-tier traffic simply queues behind paid traffic when the server is near capacity.

Rate-limit headers

Every rate-limited response, success or 429, carries your current headroom so your client can pace itself instead of discovering the limit by tripping it. The names are the ones the OpenAI and Groq SDKs, LiteLLM, LangChain and the Vercel AI SDK already read; Together AI's names are included as aliases of the same counters. Remaining values are the tighter of your key and your account.

Header Meaning Alias
x-ratelimit-limit-requestsRequests per minute allowed on this keyx-ratelimit-limit
x-ratelimit-remaining-requestsRequests left in the current 60 s windowx-ratelimit-remaining
x-ratelimit-reset-requestsTime until the oldest request leaves the window, e.g. 7.66s or 1m12.4sx-ratelimit-reset (whole seconds)
x-ratelimit-limit-tokensTokens per minute allowed on this keyx-tokenlimit-limit
x-ratelimit-remaining-tokensTokens left in the current window (reservations included)x-tokenlimit-remaining
x-ratelimit-reset-tokensTime until the oldest token reservation leaves the window
x-ratelimit-limit-concurrencyIn-flight requests allowed on this key
x-ratelimit-remaining-concurrencyIn-flight slots currently free
Retry-AfterOn 429 only: real seconds to wait before the request can succeed
HTTP/1.1 200 OK
x-ratelimit-limit-requests: 600
x-ratelimit-remaining-requests: 599
x-ratelimit-reset-requests: 59.8s
x-ratelimit-limit-tokens: 1000000
x-ratelimit-remaining-tokens: 997640
x-ratelimit-reset-tokens: 59.8s
x-ratelimit-limit-concurrency: 100
x-ratelimit-remaining-concurrency: 99
x-ratelimit-limit: 600
x-ratelimit-remaining: 599
x-ratelimit-reset: 60
x-tokenlimit-limit: 1000000
x-tokenlimit-remaining: 997640

429 responses

Every 429 carries type: "rate_limit_error", a code naming the limit that fired, the limit value, and a Retry-After header (also repeated as retry_after_seconds) that is the real number of seconds until the window frees, usually between 1 and 10. It also reports accepted_last_minute and rejected_last_minute for your account, and sets client_retry_loop_detected: true when more requests were rejected in a minute than your entire allowance: that means a retry policy on your side is ignoring Retry-After, and the message says so in plain words.

Possible codes: rate_limit_key, rate_limit_account, rate_limit_key_tokens, rate_limit_account_tokens, rate_limit_concurrency, server_capacity.

rate_limit_key — too many requests on one key within 60 seconds:

HTTP/1.1 429 Too Many Requests
Retry-After: 3

{
  "error": {
    "message": "Rate limit exceeded. Maximum 600 requests per minute per API key.",
    "type": "rate_limit_error",
    "code": "rate_limit_key",
    "limit": 600,
    "retry_after_seconds": 3
  }
}

rate_limit_account — too many requests across all keys on the account within 60 seconds. A request refused here does not count against the key window.

HTTP/1.1 429 Too Many Requests
Retry-After: 2

{
  "error": {
    "message": "Rate limit exceeded. Maximum 1800 requests per minute across all API keys on this account.",
    "type": "rate_limit_error",
    "code": "rate_limit_account",
    "limit": 1800,
    "retry_after_seconds": 2
  }
}

rate_limit_concurrency — the key's in-flight cap stayed full for 30 seconds. Requests queue for a slot before this fires, so it is rare in normal operation:

HTTP/1.1 429 Too Many Requests
Retry-After: 5

{
  "error": {
    "message": "Too many concurrent requests on this API key (limit 100 in flight). Wait for in-flight requests to finish before retrying.",
    "type": "rate_limit_error",
    "code": "rate_limit_concurrency",
    "limit": 100,
    "retry_after_seconds": 5
  }
}

server_capacity — the server-wide backstop was full for 30 seconds. This is our capacity, not your usage; overflow routing normally absorbs GPU saturation long before this fires. Retry after Retry-After:

HTTP/1.1 429 Too Many Requests
Retry-After: 5

{
  "error": {
    "message": "Server-wide capacity is temporarily full. This is not caused by your usage; retry after a few seconds.",
    "type": "rate_limit_error",
    "code": "server_capacity",
    "limit": 1024,
    "retry_after_seconds": 5
  }
}

Video queue full — more than a handful of LTX jobs are already waiting on the GPU. Extra submits should stay queued; this 429 only fires when that queue itself is full. Poll existing job ids; retry the POST after Retry-After.

HTTP/1.1 429 Too Many Requests
Retry-After: 60

{
  "error": {
    "message": "Video generation queue is full. Retry shortly.",
    "type": "rate_limit_error"
  }
}

Integration best practices

  • On 429, wait the number of seconds in Retry-After before retrying, with exponential backoff and jitter if it repeats. Never retry a 429 immediately in a loop: every retry is another request against the same window.
  • Cache failures on your side for a short period. A page that fans out to many sections should not re-fire every failed section for every visitor.
  • Disable your SDK's automatic 429 retry or configure it to honor Retry-After; the OpenAI SDKs retry twice by default.
  • Use "stream": true for long generations — see Streaming.
  • Batch related prompts where your app allows, instead of many tiny back-to-back calls.

Retry example (Python)

import time
from openai import OpenAI

client = OpenAI(api_key="YOUR_API_KEY", base_url="https://api.fastinfra.ai/v1")

def chat_with_retry(messages, max_retries=3):
    for attempt in range(max_retries):
        try:
            return client.chat.completions.create(
                model="llama3.1:8b",
                messages=messages,
            )
        except Exception as exc:
            if getattr(exc, "status_code", None) != 429 or attempt == max_retries - 1:
                raise
            retry_after = int(getattr(exc, "response", {}).headers.get("Retry-After", 60))
            time.sleep(retry_after)

Chat Completions

Create a chat completion using the OpenAI-compatible endpoint. Pass any model ID from GET https://api.fastinfra.ai/v1/models or the pricing catalog.

POST https://api.fastinfra.ai/v1/chat/completions

Request body

{
  "model": "llama3.1:8b",
  "messages": [
    { "role": "system", "content": "You are a helpful assistant." },
    { "role": "user", "content": "Hello!" }
  ],
  "temperature": 0.7
}

Example (local / free-tier model)

curl https://api.fastinfra.ai/v1/chat/completions \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "llama3.1:8b",
    "messages": [{"role": "user", "content": "Summarize quantum computing in one sentence."}]
  }'

Example (catalog model)

Any model ID from the catalog works the same way — routing is handled automatically. See Model Routing.

curl https://api.fastinfra.ai/v1/chat/completions \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "agentica-org/DeepCoder-14B-Preview",
    "messages": [{"role": "user", "content": "Explain model routing in one sentence."}]
  }'

Response

{
  "id": "chatcmpl-...",
  "object": "chat.completion",
  "created": 1234567890,
  "model": "llama3.1:8b",
  "choices": [{
    "index": 0,
    "message": {
      "role": "assistant",
      "content": "Quantum computing uses quantum bits..."
    },
    "finish_reason": "stop"
  }],
  "usage": {
    "prompt_tokens": 12,
    "completion_tokens": 24,
    "total_tokens": 36
  }
}
Instant replies: Self-hosted thinking models (e.g. Qwen3.6-27B, gpt-oss-120b) answer immediately (thinking off / reasoning_effort: low). Pass "enable_thinking": true or a higher "reasoning_effort" to opt in. If you omit max_tokens, these backends cap at 2048.

Vision (image input)

Multimodal models accept images alongside text using the standard OpenAI content-parts format: instead of a string, content is an array of parts, each with a type of text or image_url. The gateway forwards parts arrays to the upstream untouched, so anything the model supports works — including data: URLs for local files.

qwen/qwen3.6-27b is natively multimodal: it reads photos, screenshots, charts, and diagrams (OCR, chart understanding, visual reasoning) on both the dedicated GPU route and its wholesale fallbacks. Plain-string content keeps working exactly as before — the two formats can be mixed across messages in one request.

curl https://api.fastinfra.ai/v1/chat/completions \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen/qwen3.6-27b",
    "messages": [{
      "role": "user",
      "content": [
        {"type": "image_url", "image_url": {"url": "https://example.com/architecture-diagram.png"}},
        {"type": "text", "text": "Explain what this diagram shows."}
      ]
    }]
  }'

For a local file, base64-encode it into a data URL: {"url": "data:image/png;base64,<BASE64>"}. Send one image per request — the dedicated GPU route is tuned for fast text throughput and caps multimodal input at a single image (video input is not enabled).

Text-only models: Free-tier local models served via Ollama ignore image parts — only the text parts reach the model. Send images to a multimodal model such as qwen/qwen3.6-27b.

Streaming (recommended for production)

Streaming is a client choice, not something FastInfra forces. Your app must send "stream": true in the JSON body (or use an SDK with streaming enabled). FastInfra does not turn on streaming for you on non-streaming requests.

  • Your side (API caller): set "stream": true on POST /v1/chat/completions. Tokens arrive as Server-Sent Events (data: … lines).
  • FastInfra side (automatic): when self-hosted models omit max_tokens, the gateway caps at 2048; sets reasoning_effort: low on Qwen3.6 and gpt-oss-120b unless you opt into thinking; keeps the connection alive with SSE pings during slow streams; fails over to wholesale routes if the primary GPU hangs before the first token.
curl https://api.fastinfra.ai/v1/chat/completions \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -N \
  -d '{
    "model": "gpt-oss-120b",
    "stream": true,
    "messages": [{"role": "user", "content": "Explain inference routing briefly."}]
  }'

Use streaming for any response expected to take more than a few seconds — especially gpt-oss-120b and qwen3.6-27b. It reduces silent wait time, keeps proxies from treating the connection as idle, and lets your UI show partial output immediately.

Non-streaming: Still supported. For large models, set a client HTTP timeout of at least 10 minutes. Non-streaming holds a concurrency slot for the entire generation.
Rate limits (paid accounts): This endpoint is limited to 600/min per key, 1800/min per account, 100 concurrent per key, and 1024 server-wide. See Rate Limits for free-tier values and 429 handling.

Video Generations (LTX-2.5)

Public contract: this section, plus /openapi.json and GET https://api.fastinfra.ai/v1/models. See How to determine the API contract.

Text-to-video with synced audio, served on FastInfra GPUs. Use model lightricks/ltx-2.5 (aliases: ltx-2.5, ltx-2.5-distilled, ltx-2-5-fast). This is not a chat-completions model — do not send video prompts to POST https://api.fastinfra.ai/v1/chat/completions. Use the dedicated video endpoints in the API overview. Generation takes minutes, so the API is asynchronous: POST returns HTTP 202 Accepted with a job id in about a second, then poll GET until status is completed. Waiting on POST until the MP4 is ready will 524 through Cloudflare (~100s).

Field Value
Model lightricks/ltx-2.5
Base URL https://api.fastinfra.ai/v1
Submit POST https://api.fastinfra.ai/v1/videos/generations → HTTP 202
Poll GET https://api.fastinfra.ai/v1/videos/jobs/{id}
Auth Authorization: Bearer YOUR_API_KEY
Default size 1280x704 (width × height, divisible by 64)
Duration 5 seconds default, 2–20 seconds

Pricing (input / output tokens)

Billed like every other FastInfra model: USD per token. Video length is mapped to output tokens so the catalog stays consistent with chat SKUs. Competition (fal.ai LTX-2.5 fast) charges per second of video ($0.09/s at 720p, $0.13/s at 1080p). FastInfra’s default 1280x704 clip is priced at $0.10 per second of output.

Direction Price per 1M tokens Video equivalent
Input $0.10 Prompt tokens (≈ 1 token per 4 characters)
Output $100.00 1,000 tokens per second of video → $0.10/s

Example: a 5-second clip is 5,000 output tokens → $0.50 plus a few cents of prompt tokens.

POST https://api.fastinfra.ai/v1/videos/generations
GET https://api.fastinfra.ai/v1/videos/jobs/{id}

Request body (JSON)

FieldRequiredDescription
promptYesText description of the scene (spoken dialogue in the prompt is supported).
modelNoDefaults to lightricks/ltx-2.5. Must be a known LTX alias.
secondsNoClip length, 2–20 (default 5).
sizeNoWIDTHxHEIGHT, both divisible by 64 (default 1280x704).
seedNoInteger for reproducibility when supported by the worker.
{
  "model": "lightricks/ltx-2.5",
  "prompt": "A woman looks at the camera and says, welcome to FastInfra, cinematic lighting.",
  "seconds": 5,
  "size": "1280x704",
  "seed": 42
}

curl

curl https://api.fastinfra.ai/v1/videos/generations \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "lightricks/ltx-2.5",
    "prompt": "A woman looks at the camera and says, welcome to FastInfra, cinematic lighting.",
    "seconds": 5,
    "size": "1280x704"
  }'
# HTTP 202 — copy "id", then poll:
curl https://api.fastinfra.ai/v1/videos/jobs/JOB_ID \
  -H "Authorization: Bearer YOUR_API_KEY"

Python

import base64, json, time, urllib.request

headers = {
    "Authorization": "Bearer YOUR_API_KEY",
    "Content-Type": "application/json",
}
req = urllib.request.Request(
    "https://api.fastinfra.ai/v1/videos/generations",
    data=json.dumps({
        "model": "lightricks/ltx-2.5",
        "prompt": "A golden retriever running through a sunny meadow, cinematic, birdsong.",
        "seconds": 5,
        "size": "1280x704",
    }).encode(),
    headers=headers,
    method="POST",
)
with urllib.request.urlopen(req, timeout=60) as resp:
    job = json.load(resp)
job_id = job["id"]

while True:
    time.sleep(2)
    poll = urllib.request.Request(
        f"https://api.fastinfra.ai/v1/videos/jobs/{job_id}",
        headers={"Authorization": "Bearer YOUR_API_KEY"},
    )
    with urllib.request.urlopen(poll, timeout=60) as resp:
        payload = json.load(resp)
    if payload["status"] == "completed":
        break
    if payload["status"] == "failed":
        raise SystemExit(payload.get("error") or "video job failed")

open("clip.mp4", "wb").write(base64.b64decode(payload["data"][0]["b64_json"]))
print(payload["id"], payload["usage"])

Submit response (HTTP 202)

Response header Location: /v1/videos/jobs/{id} points at the poll URL. Save the JSON id field — that is your gateway job id (not the internal worker id).

HTTP/1.1 202 Accepted
Location: /v1/videos/jobs/video_abc123

{
  "id": "video_abc123",
  "object": "video.job",
  "created": 1234567890,
  "model": "lightricks/ltx-2.5",
  "seconds": 5,
  "size": "1280x704",
  "status": "queued"
}

Failed poll (HTTP 200, job failed)

{
  "id": "video_abc123",
  "object": "video.job",
  "status": "failed",
  "error": "Video generation failed."
}

HTTP errors

  • 400 — missing prompt, unknown model, or seconds out of range.
  • 402 — insufficient prepaid credits (same wallet as chat).
  • 429 — gateway video queue full; honor Retry-After and retry POST.
  • 503 — video backend not configured or unreachable.
  • 404 on GET — unknown or expired job id (jobs expire after ~24h).

Completed poll (HTTP 200)

{
  "id": "video_...",
  "object": "video",
  "created": 1234567890,
  "model": "lightricks/ltx-2.5",
  "seconds": 5,
  "size": "1280x704",
  "status": "completed",
  "data": [{ "b64_json": "<mp4-bytes-as-base64>", "format": "mp4" }],
  "usage": {
    "prompt_tokens": 18,
    "completion_tokens": 5000,
    "total_tokens": 5018
  }
}
Latency: POST returns immediately with HTTP 202 and a job id — the gateway queues your request and a background dispatcher forwards it to the GPU when capacity is available. A 5-second clip typically finishes in 1–2 minutes on a warm GPU — poll every 2 seconds. Status is queued, in_progress, completed, or failed. Do not wait on the POST for the MP4; Cloudflare will 524 around 100 seconds. HTTP 429 with Retry-After means the gateway queue is full — wait and retry the POST.

Tool Calling

Function/tool calling works exactly as with OpenAI: send tools and optionally tool_choice, receive tool_calls on the assistant message with finish_reason: "tool_calls", run the tool, and send the result back as a role: "tool" message carrying the tool_call_id. Streaming delivers delta.tool_calls fragments with an index, which every OpenAI SDK accumulator already stitches together. Supported on the self-hosted flagship models (qwen3.6-27b, gpt-oss-120b) and on wholesale-routed models that support it upstream; see the compatibility matrix.

from openai import OpenAI
import json

client = OpenAI(api_key="YOUR_API_KEY", base_url="https://api.fastinfra.ai/v1")

tools = [{
    "type": "function",
    "function": {
        "name": "get_weather",
        "parameters": {"type": "object", "properties": {"city": {"type": "string"}}, "required": ["city"]},
    },
}]

messages = [{"role": "user", "content": "What's the weather in London?"}]
first = client.chat.completions.create(model="qwen3.6-27b", messages=messages, tools=tools)
call = first.choices[0].message.tool_calls[0]

messages.append(first.choices[0].message)
messages.append({
    "role": "tool",
    "tool_call_id": call.id,
    "content": json.dumps({"temp_c": 18, "sky": "overcast"}),
})
final = client.chat.completions.create(model="qwen3.6-27b", messages=messages, tools=tools)
print(final.choices[0].message.content)
Small free-tier models served by Ollama (llama3.2:3b, phi3:mini, ...) do not return structured tool calls through this gateway; use a flagship model for agents.

Embeddings

POST https://api.fastinfra.ai/v1/embeddings is OpenAI-compatible and served from self-hosted embedding models at $0. The OpenAI model ids are accepted as aliases, so a client that only changes base_url keeps working. encoding_format: "base64" (the OpenAI Python SDK default) and dimensions (Matryoshka truncation with re-normalisation) are supported. Input is a string or an array of up to 2,048 strings; token-id arrays are not accepted.

Model Dimensions Aliases
nomic-embed-text 768 text-embedding-3-small, text-embedding-ada-002, nomic-ai/nomic-embed-text-v1.5
mxbai-embed-large 1024 text-embedding-3-large, mixedbread-ai/mxbai-embed-large-v1
all-minilm 384 sentence-transformers/all-MiniLM-L6-v2
curl https://api.fastinfra.ai/v1/embeddings \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "nomic-embed-text", "input": ["FastInfra runs on its own H200s.", "Embeddings are free."]}'
{
  "object": "list",
  "data": [
    {"object": "embedding", "index": 0, "embedding": [0.0123, -0.0456, ...]},
    {"object": "embedding", "index": 1, "embedding": [0.0789, 0.0012, ...]}
  ],
  "model": "nomic-embed-text",
  "usage": {"prompt_tokens": 17, "total_tokens": 17}
}
from openai import OpenAI
client = OpenAI(api_key="YOUR_API_KEY", base_url="https://api.fastinfra.ai/v1")

# Unchanged from an OpenAI integration: the id resolves to nomic-embed-text.
vectors = client.embeddings.create(model="text-embedding-3-small", input=["hello", "world"])
print(len(vectors.data[0].embedding))  # 768

Rate limited like other inference calls (requests and concurrency; no token budget). Errors: 400 invalid_input / model_not_found for a chat model id, 404 model_not_found when the backend lacks the model, 503 backend_unavailable while the embedding backend is down.

Compatibility Matrix

What each OpenAI request feature does on each class of model, as the gateway actually forwards it. "Upstream" means the field is passed through unchanged and support is whatever that wholesale provider offers. A feature the gateway drops is marked No even where the model could do it natively. Run scripts/compat-check.py against your own key to verify any cell live.

Model family Stream Tools json_object json_schema Vision logprobs seed stop Reasoning
Qwen3.6-27B (self-hosted vLLM, H200)
qwen3.6-27b, qwen/qwen3.6-27b
One image per request, video disabled. Thinking is off by default; set reasoning_effort (medium/high) or enable_thinking: true to turn it on.
Yes Yes Yes Yes Partial Yes Yes Yes Yes
gpt-oss-120b (self-hosted vLLM, H100)
gpt-oss-120b
Text only. reasoning_effort low by default; raise it for harder tasks. Reasoning text arrives in the reasoning field.
Yes Yes Yes Yes No Yes Yes Yes Yes
Free-tier small models (self-hosted Ollama)
llama3.1:8b, llama3.2:3b, phi3:mini, gemma2:2b, qwen2.5:*
$0 models for wiring and smoke tests. temperature, top_p, seed, stop and response_format are forwarded; tools, logprobs and images are not. Use a flagship model for agents.
Yes No Yes Yes No No Yes Yes No
Embedding models (self-hosted Ollama)
nomic-embed-text, mxbai-embed-large, all-minilm
POST /v1/embeddings only. encoding_format float|base64 and dimensions are supported; OpenAI text-embedding-* ids are aliases.
No No No No No No No No No
Wholesale-routed models (400+ ids in GET /v1/models)
anthropic/claude-*, meta-llama/*, mistralai/*, ...
The full OpenAI request body is forwarded unchanged and the reply is passed back with tool_calls, logprobs and extra fields intact. Support is whatever the upstream provider offers for that model.
Yes Upstream Upstream Upstream Upstream Upstream Upstream Upstream Upstream

Not supported anywhere: n > 1 on streaming, token-id arrays as embeddings input, legacy POST /v1/completions. Audio endpoints ignore chat fields such as max_tokens.

Framework recipes

One base_url and one key. Each snippet below is the complete change from an OpenAI integration.

OpenAI SDK (Python)

from openai import OpenAI
client = OpenAI(api_key="YOUR_API_KEY", base_url="https://api.fastinfra.ai/v1", max_retries=2)
# The SDK reads x-ratelimit-* headers and honours Retry-After on 429 automatically.

OpenAI SDK (JavaScript / TypeScript)

import OpenAI from "openai";
const client = new OpenAI({ apiKey: process.env.FASTINFRA_API_KEY, baseURL: "https://api.fastinfra.ai/v1" });

LangChain

from langchain_openai import ChatOpenAI, OpenAIEmbeddings
llm = ChatOpenAI(model="qwen3.6-27b", api_key="YOUR_API_KEY", base_url="https://api.fastinfra.ai/v1")
emb = OpenAIEmbeddings(model="nomic-embed-text", api_key="YOUR_API_KEY", base_url="https://api.fastinfra.ai/v1", check_embedding_ctx_length=False)
# Tool calling: llm.bind_tools([...]) works as with OpenAI on qwen3.6-27b and gpt-oss-120b.

LiteLLM

import litellm
response = litellm.completion(
    model="openai/qwen3.6-27b",           # "openai/" prefix = OpenAI-compatible provider
    api_key="YOUR_API_KEY",
    api_base="https://api.fastinfra.ai/v1",
    messages=[{"role": "user", "content": "hi"}],
)

Vercel AI SDK

import { createOpenAI } from "@ai-sdk/openai";
const fastinfra = createOpenAI({ apiKey: process.env.FASTINFRA_API_KEY, baseURL: "https://api.fastinfra.ai/v1" });
const { text } = await generateText({ model: fastinfra("qwen3.6-27b"), prompt: "hi" });

Continue / Cursor / any "OpenAI-compatible" provider slot

{
  "provider": "openai",
  "model": "qwen3.6-27b",
  "apiBase": "https://api.fastinfra.ai/v1",
  "apiKey": "YOUR_API_KEY"
}

n8n

Credentials → OpenAI → set Base URL to https://api.fastinfra.ai/v1. The OpenAI Chat Model and Embeddings nodes then list our models.

Open WebUI

Settings → Connections → OpenAI API: URL https://api.fastinfra.ai/v1, key YOUR_API_KEY. Models populate from GET /v1/models.

Audio: Speech-to-Text & TTS

OpenAI-compatible audio endpoints for transcription, translation, live streaming, and text-to-speech — all served on FastInfra's own GPU pods (Whisper STT and Kokoro-82M neural TTS on the same worker). Audio is processed in memory and never persisted (see data protection note below).

Whisper and Kokoro are separate calls. You can use POST https://api.fastinfra.ai/v1/audio/speech for text-to-speech without ever uploading audio or calling transcription. Each endpoint has its own model id and billing. Full Kokoro reference: Kokoro TTS.
POST https://api.fastinfra.ai/v1/audio/transcriptions
POST https://api.fastinfra.ai/v1/audio/translations
POST https://api.fastinfra.ai/v1/audio/speech
GET wss://…/v1/realtime/transcription

Models & pricing

Capability Model id Accepted aliases Price
Transcription / translation / live streaming openai/whisper-large-v3-turbo whisper-1, whisper-turbo, whisper-large-v3-turbo, whisper-large-v3, openai/whisper-large-v3 $0.006 / audio minute (OpenAI whisper-1 parity)
Text-to-speech hexgrad/kokoro-82m kokoro, kokoro-82m, tts-1 $15 / 1M input characters (OpenAI tts-1 parity)

The model field is optional on all audio endpoints — omit it and the default model above is used. Live streaming sessions bill wall-clock duration at the same per-minute rate.

Transcription (file upload)

Multipart upload, up to 25 MB. Common formats (wav, mp3, m4a, webm, ogg) are decoded automatically.

curl https://api.fastinfra.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -F file=@answer.wav \
  -F model=whisper-1 \
  -F response_format=json
{ "text": "Newton's second law states that force equals mass times acceleration." }
Form field Notes
file Required. Audio file, max 25 MB.
model Optional. Any alias from the table above.
language Optional ISO code (e.g. en, hi); auto-detected when omitted.
response_format json (default), text, or verbose_json (includes segments, language, duration).
temperature Optional. Use 0 (the default) for maximum accuracy — higher values add sampling randomness and are not recommended for exams.
prompt Optional context hint (names, technical terms) to guide spelling.
Chat-style fields are ignored: max_tokens has no effect on audio endpoints — the full audio is always transcribed and billed by duration, not tokens.

Translation (any language → English)

Same form fields as transcription (except language); the response text is always English.

curl https://api.fastinfra.ai/v1/audio/translations \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -F file=@hindi-answer.wav \
  -F response_format=json
# -> { "text": "My name is Rahul and I am a physics student." }

Live streaming (interactive viva)

Stream microphone audio over WebSocket and receive partial and final transcripts in real time. Send binary frames of PCM16 mono 16 kHz audio; receive JSON events. Browsers cannot set headers on WebSocket upgrades, so the API key may be passed as a query parameter on this endpoint only.

wss://api.fastinfra.ai/v1/realtime/transcription?api_key=YOUR_API_KEY

<- binary frames: PCM16 mono 16 kHz audio chunks
-> { "type": "partial", "text": "newton's second" }
-> { "type": "final",   "text": "Newton's second law states that..." }

A session holds one rate-limit concurrency slot for its whole duration, and billing is per wall-clock minute of the session.

Kokoro text-to-speech

Self-hosted Kokoro-82M neural voices on FastInfra GPUs. OpenAI-compatible POST https://api.fastinfra.ai/v1/audio/speech — JSON in, WAV (or raw PCM) out. No prior transcription step required.

POST https://api.fastinfra.ai/v1/audio/speech
FieldRequiredDescription
inputYesText to speak, max 4,096 characters.
modelNoDefaults to hexgrad/kokoro-82m. Aliases: tts-1, kokoro, kokoro-82m.
voiceNoKokoro voice id (default af_heart). First letter selects language pipeline: a American, b British, h Hindi.
speedNoPlayback speed, 0.5–2.0 (default 1.0).
response_formatNowav (default, 24 kHz mono PCM16) or pcm (raw PCM16 without WAV header).

Quick start (TTS only)

curl https://api.fastinfra.ai/v1/audio/speech \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "tts-1",
    "input": "Can you explain Newton'\''s second law in your own words?",
    "voice": "af_heart"
  }' \
  -o question.wav

Response: HTTP 200, Content-Type: audio/wav, body is the audio file. Billed per input character ($15 / 1M characters, OpenAI tts-1 parity).

Available voices

Voice id Description
af_heartUS female, warm (default)
af_bellaUS female
am_michaelUS male
bf_emmaUK female
bm_georgeUK male
hf_alphaHindi female
hm_omegaHindi male

Python (TTS only)

from openai import OpenAI

client = OpenAI(api_key="YOUR_API_KEY", base_url="https://api.fastinfra.ai/v1")

speech = client.audio.speech.create(
    model="tts-1",
    voice="am_michael",
    input="Welcome to FastInfra. This clip uses Kokoro on our GPU pods.",
    speed=1.0,
)
speech.write_to_file("welcome.wav")

JavaScript (fetch)

const res = await fetch("https://api.fastinfra.ai/v1/audio/speech", {
  method: "POST",
  headers: {
    Authorization: "Bearer YOUR_API_KEY",
    "Content-Type": "application/json",
  },
  body: JSON.stringify({
    model: "tts-1",
    voice: "bf_emma",
    input: "Hello from Kokoro.",
  }),
});
const wav = Buffer.from(await res.arrayBuffer());
// wav is 24 kHz mono PCM16 in a WAV container

Kokoro errors

  • 400 — missing input, unknown model, text over 4,096 chars, or invalid voice id.
  • 402 — insufficient prepaid credits.
  • 503 — Kokoro worker still loading or temporarily unavailable; retry shortly.

Python (full voice loop: STT + TTS)

from openai import OpenAI

client = OpenAI(api_key="YOUR_API_KEY", base_url="https://api.fastinfra.ai/v1")

# Student's answer -> text
with open("answer.wav", "rb") as f:
    transcript = client.audio.transcriptions.create(model="whisper-1", file=f)
print(transcript.text)

# Examiner's question -> natural voice
speech = client.audio.speech.create(
    model="tts-1", voice="af_heart",
    input="Good answer. What happens if the mass doubles?")
speech.write_to_file("question.wav")

Validating a model id

Gateway tools that verify a configured model can call the standard retrieve-model endpoint. Unknown ids return a JSON invalid_request_error with HTTP 404.

curl https://api.fastinfra.ai/v1/models/whisper-large-v3 \
  -H "Authorization: Bearer YOUR_API_KEY"
# -> { "id": "openai/whisper-large-v3-turbo", "object": "model", ... }

Common errors

  • 401 — invalid or missing API key. Create a real key on the API Keys page; placeholder values are rejected.
  • 400 — unknown model id, missing file/input, undecodable audio, or unknown TTS voice. The error message states the exact problem.
  • 402 — insufficient prepaid credits (see Credits).
  • 413 — audio file over 25 MB. Split long recordings or use live streaming.
  • 503 — transcription backends temporarily unavailable; retry shortly.
Data protection: audio is processed in memory on FastInfra's own EU servers and is not persisted. Usage records store durations and token counts only — never audio or transcripts. No third-party transcription API is ever called.

GPU Pods

Rent dedicated GPU machines with any Docker image — training, fine-tuning, ComfyUI, vLLM, Jupyter, or your own stack. Pods are managed from the GPU Cloud dashboard (not the OpenAI API). Billing is per minute from your prepaid credit balance; stop the pod and compute billing stops immediately.

Resource URL
GPU catalog (public) https://gpu.fastinfra.ai
Deploy a pod https://gpu.fastinfra.ai/Deploy (sign in required)
My pods https://gpu.fastinfra.ai/Pods
Add credits Billing
Support Support — choose GPU Pod issue for pod-specific help

Getting started

  1. Sign in and add credits on the Billing page.
  2. Browse the GPU catalog for hourly rates (RTX 4090, A100, H100, H200, and more).
  3. Open Deploy, pick a GPU tier, Docker image, ports, and disk size.
  4. When status is RUNNING, copy SSH or HTTP endpoints from the pod detail page and connect.
  5. Stop when finished — billing for GPU compute ends the moment the pod stops.

Deploy options

Option Details
GPU tier Secure — vetted datacenters. Value — lower price community hosts (when available for that GPU).
GPU count 1–8 GPUs per pod (max depends on GPU type and tier).
Docker image Any public image, e.g. runpod/pytorch:1.0.2-cu1281-torch280-ubuntu2404, ComfyUI, Ollama, or vLLM.
Ports Comma-separated, e.g. 22/tcp, 8888/http. TCP ports get a public host:port; HTTP ports are exposed for web UIs.
Environment One KEY=value per line (up to 50 vars).
Container disk 5–1000 GB ephemeral disk — wiped when the pod stops.
Volume disk 0 (none) or 10–4000 GB persistent volume mounted at /workspace — survives stops.
SSH / Jupyter Enable SSH for shell access; optional JupyterLab (exposes port 8888).

Pricing & billing

GPU pod prices include a 8% platform markup on underlying compute and storage. Catalog hourly rates are rounded up so displayed prices are always what you pay.

  • Compute — billed per minute while the pod is RUNNING (GPU hourly rate × GPU count, plus running disk).
  • Stopped volume — if you attached a persistent volume and stop (but do not terminate) the pod, a lower storage-only rate applies until you terminate or start again.
  • Deploy gate — you need enough credits for at least 1 hour(s) of runtime at the pod's hourly rate before deploy or start.
  • Auto-stop — if your balance hits zero while a pod is running, it is stopped automatically to prevent overdraft. Add credits and start it again.
  • Terminate — permanently deletes the pod and its volume; all billing ends.

Per-minute charges appear on the pod detail page under Billing history. Inference API usage (chat completions) is billed separately — see Credits.

Pod lifecycle

Action Effect
Deploy Provisions a new pod; status moves through PROVISIONING → STARTING → RUNNING (typically 1–2 minutes).
Stop Halts compute billing. Container disk is discarded; volume data is kept if configured.
Start Resumes a stopped pod (requires sufficient credit balance).
Restart Reboots a running pod in place.
Terminate Deletes the pod and volume permanently. Irreversible.

Connecting

When a pod is RUNNING, the detail page lists connection endpoints:

  • SSH — ssh root@HOST -p PORT (when SSH is enabled and port 22 is exposed).
  • HTTP services — Jupyter, Gradio, or custom apps on exposed HTTP ports appear as HOST:PORT.
  • Endpoints refresh — the detail page polls status every 15 seconds while provisioning or running; reload if endpoints are still assigning.
# Example after deploy (values shown on your pod detail page)
ssh root@203.0.113.42 -p 22001

# Jupyter in browser (if enabled)
http://203.0.113.42:8888

Popular Docker images

  • runpod/pytorch:1.0.2-cu1281-torch280-ubuntu2404 — PyTorch + CUDA (default on deploy form)
  • runpod/tensorflow:2.2.0-py3 — TensorFlow
  • Community templates for ComfyUI, Ollama, vLLM — use any public registry image your workload needs
API vs dashboard: GPU pods are provisioned through the web UI today. Chat inference remains on https://api.fastinfra.ai/v1 — a stopped GPU pod does not affect API keys or model routing. For automation, use the inference API; for interactive GPU machines, use GPU pods.

CPU Servers

Rent Linux virtual machines from the CPU Cloud dashboard. Billing is per minute from your prepaid credit balance while the server exists — stopping a server does not stop charges; delete it when finished.

ResourceURL
CPU catalog (public)https://cpu.fastinfra.ai
Deploy a serverhttps://cpu.fastinfra.ai/Cpu/Deploy (sign in)
My servershttps://cpu.fastinfra.ai/Cpu/Servers
Add creditsBilling

Getting started

  1. Sign in and add credits on Billing.
  2. Browse plans on the CPU catalog (vCPU, RAM, region, hourly rate).
  3. Deploy with a name, plan, region, and OS image (Ubuntu or Debian).
  4. When status is RUNNING, copy the public IP and root password from the server detail page (password shown once).
  5. Delete when finished — that is the only way to stop billing.

Prices include a 8% platform markup. Deploy requires ~1 hour(s) of credits at the server's hourly rate. CPU servers are dashboard-only today (no OpenAI-style provisioning API).

Model Routing

FastInfra serves hundreds of models through one OpenAI-compatible API. You send the same request shape for every model; the gateway automatically routes each call to the cheapest healthy capacity for that model, with automatic failover when a route degrades.

Discover models

  • GET https://api.fastinfra.ai/v1/models — full catalog of available model IDs
  • GET https://api.fastinfra.ai/v1/models/count — total model count
  • Pricing page — searchable catalog with per-token pricing
Canonical IDs: GET https://api.fastinfra.ai/v1/models returns clean, stable model IDs (e.g. deepseek/deepseek-v4-flash-0731). Vendor-specific path formats are normalized automatically, so the ID you see in the catalog is the ID you send.

Routing behavior

For every request, the gateway resolves the route in this order:

  1. Free-tier models — served at $0 on self-hosted Ollama capacity (llama3.1:8b, mistral:7b, etc.)
  2. Dedicated GPU primary — e.g. gpt-oss-120b on RunPod vLLM, qwen3.6-27b (multimodal — see Vision) on H200
  3. Fallback chain — wholesale providers tried automatically if the primary fails or is saturated

Routing is automatic — your integration stays identical as capacity changes. Enable wholesale API keys in the admin panel so fallbacks activate when self-hosted GPUs are busy or offline.

Wholesale fallbacks (cheapest first)

When a self-hosted route fails or hits its concurrency cap, these wholesale providers are tried in order (requires API keys in admin):

Model Fallback order Typical wholesale input $/1M tokens
gpt-oss-120b route-21 → route-07 ~$0.15 → ~$0.04
qwen3.6-27b route-21 → route-18 → route-07 ~$0.32 → ~$0.30 → ~$0.60
llama3.1:8b (free tier) route-07 → route-12 → route-03 varies

You are billed per token only when traffic actually routes to a wholesale provider. Self-hosted capacity is used first whenever healthy.

Python example

from openai import OpenAI

client = OpenAI(api_key="YOUR_API_KEY", base_url="https://api.fastinfra.ai/v1")

response = client.chat.completions.create(
    model="agentica-org/DeepCoder-14B-Preview",
    messages=[{"role": "user", "content": "Hello!"}]
)
print(response.choices[0].message.content)

Errors

  • 502 — No inference provider available for the requested model. The model is not in the catalog or is not currently enabled.
  • 503 — Inference provider is not reachable. Upstream capacity is temporarily offline; retry shortly.

If a model does not appear in GET https://api.fastinfra.ai/v1/models, it may not be enabled on this platform yet. Browse the pricing catalog for models available on this deployment.

List Models

Returns the full model catalog available on this deployment. Chat models use POST https://api.fastinfra.ai/v1/chat/completions. The LTX video model (lightricks/ltx-2.5) appears here for discovery but must be called via POST https://api.fastinfra.ai/v1/videos/generations.

GET https://api.fastinfra.ai/v1/models

Response shape

{
  "object": "list",
  "data": [
    { "id": "anthropic/claude-3.5-sonnet", "object": "model" },
    { "id": "deepseek/deepseek-v4-flash-0731", "object": "model" },
    { "id": "llama3.1:8b", "object": "model" }
  ],
  "total_count": 3
}

Use GET https://api.fastinfra.ai/v1/models/count for a lightweight count without downloading the full list.

Currently available (784)

  • agentica-org/DeepCoder-14B-Preview
  • aion-labs/aion-2.0
  • aion-labs/aion-3.0
  • aion-labs/aion-3.0-mini
  • aion-labs/aion-3.5
  • aion-labs/aion-3.5-mini
  • aion-labs/aion-rp-llama-3.1-8b
  • alibaba/happyhorse-1.0-i2v
  • alibaba/happyhorse-1.0-r2v
  • alibaba/happyhorse-1.0-t2v
  • alibaba/happyhorse-1.1-i2v
  • alibaba/happyhorse-1.1-r2v
  • alibaba/happyhorse-1.1-t2v
  • allenai/Molmo-7B-D-0924
  • amazon/nova-2-lite-v1
  • amazon/nova-lite-v1
  • amazon/nova-micro-v1
  • amazon/nova-premier-v1
  • amazon/nova-pro-v1
  • anthracite-org/magnum-v4-72b
  • anthropic/claude-fable-5
  • anthropic/claude-fable-5.1
  • anthropic/claude-fable-5.1:batch
  • anthropic/claude-fable-5:batch
  • anthropic/claude-haiku-4-5
  • anthropic/claude-haiku-4.5:batch
  • anthropic/claude-opus-4-7
  • anthropic/claude-opus-4-8
  • anthropic/claude-opus-4.1
  • anthropic/claude-opus-4.1:batch
  • anthropic/claude-opus-4.5
  • anthropic/claude-opus-4.5:batch
  • anthropic/claude-opus-4.6
  • anthropic/claude-opus-4.6:batch
  • anthropic/claude-opus-4.7:batch
  • anthropic/claude-opus-4.8:batch
  • anthropic/claude-opus-5
  • anthropic/claude-opus-5-5
  • anthropic/claude-opus-5.5
  • anthropic/claude-opus-5.5:batch
  • anthropic/claude-opus-5:batch
  • anthropic/claude-sonnet-4
  • anthropic/claude-sonnet-4-6
  • anthropic/claude-sonnet-4.5
  • anthropic/claude-sonnet-4.5:batch
  • anthropic/claude-sonnet-4.6:batch
  • anthropic/claude-sonnet-5
  • anthropic/claude-sonnet-5-5
  • anthropic/claude-sonnet-5.5
  • anthropic/claude-sonnet-5.5:batch

Showing 50 of 784 models. Call GET https://api.fastinfra.ai/v1/models for the full list.

Code Examples

Chat examples use the OpenAI SDK with base_url="https://api.fastinfra.ai/v1". Video generation uses raw HTTP (async job + poll) — see Video Generations for the full Python/curl flow. Kokoro TTS is standalone — see Kokoro TTS.

Kokoro TTS (Python, minimal)

from openai import OpenAI

client = OpenAI(api_key="YOUR_API_KEY", base_url="https://api.fastinfra.ai/v1")
client.audio.speech.create(
    model="tts-1",
    voice="af_heart",
    input="Hello from Kokoro on FastInfra.",
).write_to_file("hello.wav")

LTX video (Python, minimal)

import json, time, urllib.request

BASE = "https://api.fastinfra.ai/v1"
KEY = "YOUR_API_KEY"
headers = {"Authorization": f"Bearer {KEY}", "Content-Type": "application/json"}

req = urllib.request.Request(
    f"{BASE}/videos/generations",
    data=json.dumps({
        "model": "lightricks/ltx-2.5",
        "prompt": "Ocean waves at golden hour, cinematic.",
        "seconds": 5,
    }).encode(),
    headers=headers,
    method="POST",
)
with urllib.request.urlopen(req, timeout=60) as resp:
    job_id = json.load(resp)["id"]

while True:
    time.sleep(2)
    poll = urllib.request.Request(f"{BASE}/videos/jobs/{job_id}", headers={"Authorization": f"Bearer {KEY}"})
    with urllib.request.urlopen(poll, timeout=60) as resp:
        body = json.load(resp)
    if body["status"] == "completed":
        print(body["usage"])
        break
    if body["status"] == "failed":
        raise SystemExit(body.get("error"))

C# (.NET)

var client = new OpenAIClient(
    new ApiKeyCredential("YOUR_API_KEY"),
    new OpenAIClientOptions { Endpoint = new Uri("https://api.fastinfra.ai/v1") });

var chat = client.GetChatClient("llama3.1:8b");
var response = await chat.CompleteChatAsync("Hello!");
Console.WriteLine(response.Value.Content[0].Text);

Python (free-tier / local model)

from openai import OpenAI

client = OpenAI(
    api_key="YOUR_API_KEY",
    base_url="https://api.fastinfra.ai/v1"
)

response = client.chat.completions.create(
    model="llama3.1:8b",
    messages=[{"role": "user", "content": "Hello!"}]
)
print(response.choices[0].message.content)

Python (provider model)

from openai import OpenAI

client = OpenAI(api_key="YOUR_API_KEY", base_url="https://api.fastinfra.ai/v1")

response = client.chat.completions.create(
    model="agentica-org/DeepCoder-14B-Preview",
    messages=[{"role": "user", "content": "Hello!"}]
)
print(response.choices[0].message.content)

JavaScript

const response = await fetch("https://api.fastinfra.ai/v1/chat/completions", {
  method: "POST",
  headers: {
    "Authorization": "Bearer YOUR_API_KEY",
    "Content-Type": "application/json"
  },
  body: JSON.stringify({
    model: "agentica-org/DeepCoder-14B-Preview",
    messages: [{ role: "user", content: "Hello!" }]
  })
});

const data = await response.json();
console.log(data.choices[0].message.content);

Ready to ship?

Create your free account, generate an API key, and start calling frontier models in minutes.