Llama Nemotron Embed Vl 1b v2 vs Nemotron 3.5 ASR Streaming Multilingual 0.6b
Live API pricing and availability, side by side. Both models run on FastInfra's OpenAI-compatible endpoint — switching between them is a one-line change.
Pricing
Price per 1M tokens
Prices are live and sync automatically from wholesale providers.
| Llama Nemotron Embed Vl 1b v2 | Nemotron 3.5 ASR Streaming Multilingual 0.6b | |
|---|---|---|
| Input | $0.01 | — |
| Output | — | — |
| Default provider | DeepInfra | DeepInfra |
| Providers available | 1 | 1 |
Quickstart
Try both in 30 seconds
Same endpoint, same SDK — only the model string changes.
from openai import OpenAI
client = OpenAI(api_key="YOUR_API_KEY", base_url="https://api.fastinfra.ai/v1")
for model in ["nvidia/llama-nemotron-embed-vl-1b-v2", "nvidia/Nemotron-3.5-ASR-Streaming-Multilingual-0.6b"]:
response = client.chat.completions.create(
model=model,
messages=[{"role": "user", "content": "Hello!"}]
)
print(model, "->", response.choices[0].message.content)