Llama Nemotron Embed Vl 1b v2 vs NVIDIA Nemotron 3 Ultra 550B A55B

Live API pricing and availability, side by side. Both models run on FastInfra's OpenAI-compatible endpoint — switching between them is a one-line change.

Pricing

Price per 1M tokens

Prices are live and sync automatically from wholesale providers.

Llama Nemotron Embed Vl 1b v2 NVIDIA Nemotron 3 Ultra 550B A55B
Input $0.01 $0.53
Output $2.31
Default provider DeepInfra DeepInfra
Providers available 1 1

On input tokens, Llama Nemotron Embed Vl 1b v2 is currently 50× cheaper than NVIDIA Nemotron 3 Ultra 550B A55B on FastInfra. Output-token pricing and quality trade-offs differ per workload — test both with the free tier.

Quickstart

Try both in 30 seconds

Same endpoint, same SDK — only the model string changes.

from openai import OpenAI

client = OpenAI(api_key="YOUR_API_KEY", base_url="https://api.fastinfra.ai/v1")

for model in ["nvidia/llama-nemotron-embed-vl-1b-v2", "nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B"]:
    response = client.chat.completions.create(
        model=model,
        messages=[{"role": "user", "content": "Hello!"}]
    )
    print(model, "->", response.choices[0].message.content)