Glm Flash Latest API
Run ~z-ai/glm-flash-latest through FastInfra's API.
Pay per token. Video is billed as 1,000 output tokens per second of generated video.
Glm Flash Latest pricing
Billed per token. 1,000 output tokens = 1 second of video ($0.00/s at this list price).
| Direction | Price per 1M tokens |
|---|---|
| Input | $0.08 |
| Output | $0.26 |
| Per second of video | $0.00 (1,000 output tokens) |
Served by 1 provider
Requests route to OpenRouter by default (cheapest).
Pin a specific provider with the :provider suffix.
| Provider | Upstream model ID |
|---|---|
| openrouter | ~z-ai/glm-flash-latest |
Call Glm Flash Latest in 30 seconds
Same FastInfra API key and base URL. POST /videos/generations (HTTP 202), then poll GET /videos/jobs/{id} — not chat completions.
Python
import base64, json, time, urllib.request
headers = {"Authorization": "Bearer YOUR_API_KEY", "Content-Type": "application/json"}
req = urllib.request.Request(
"https://api.fastinfra.ai/v1/videos/generations",
data=json.dumps({
"model": "~z-ai/glm-flash-latest",
"prompt": "A woman looks at the camera and says, welcome to FastInfra.",
"seconds": 5,
"size": "1280x704"
}).encode(),
headers=headers,
method="POST",
)
with urllib.request.urlopen(req, timeout=60) as resp:
job = json.load(resp)
while True:
time.sleep(2)
poll = urllib.request.Request(
f"https://api.fastinfra.ai/v1/videos/jobs/{job['id']}",
headers={"Authorization": "Bearer YOUR_API_KEY"},
)
with urllib.request.urlopen(poll, timeout=60) as resp:
payload = json.load(resp)
if payload["status"] == "completed":
break
if payload["status"] == "failed":
raise SystemExit(payload.get("error") or "video job failed")
open("clip.mp4", "wb").write(base64.b64decode(payload["data"][0]["b64_json"]))
curl
curl https://api.fastinfra.ai/v1/videos/generations \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "~z-ai/glm-flash-latest",
"prompt": "A woman looks at the camera and says, welcome to FastInfra.",
"seconds": 5,
"size": "1280x704"
}'
# HTTP 202 — copy "id", then poll:
curl https://api.fastinfra.ai/v1/videos/jobs/JOB_ID \
-H "Authorization: Bearer YOUR_API_KEY"
Glm Flash Latest — common questions
How much does the ~z-ai/glm-flash-latest API cost?
On FastInfra, ~z-ai/glm-flash-latest costs $0.0788/1M input, $0.2625/1M output tokens ($0 per second of video). Billing is per token used, with no subscription.
Is ~z-ai/glm-flash-latest compatible with the OpenAI SDK?
Video models use the same FastInfra API key and base URL (https://api.fastinfra.ai/v1). POST https://api.fastinfra.ai/v1/videos/generations returns HTTP 202 with a job id; poll GET https://api.fastinfra.ai/v1/videos/jobs/{id} until status is completed.
Which providers serve ~z-ai/glm-flash-latest?
~z-ai/glm-flash-latest is available from 1 provider(s): openrouter. Requests route to OpenRouter by default (cheapest); append ":provider" to the model ID to pin one.
Related models
Compare side by side: Glm Flash Latest vs Glm 5.3 Flash · Glm Flash Latest vs Glm 4.7 Flash · Glm Flash Latest vs Glm 5.3 Flash:batch