Llama 3.2 3b Instruct vs Meta Llama 3.1 8B Instruct Turbo
Live API pricing and availability, side by side. Both models run on FastInfra's OpenAI-compatible endpoint — switching between them is a one-line change.
Pricing
Price per 1M tokens
Prices are live and sync automatically from wholesale providers.
| Llama 3.2 3b Instruct | Meta Llama 3.1 8B Instruct Turbo | |
|---|---|---|
| Input | $0.06 | $0.02 |
| Output | $0.06 | $0.04 |
| Default provider | Together AI | DeepInfra |
| Providers available | 2 | 2 |
On input tokens, Meta Llama 3.1 8B Instruct Turbo is currently 3× cheaper than Llama 3.2 3b Instruct on FastInfra. Output-token pricing and quality trade-offs differ per workload — test both with the free tier.
Quickstart
Try both in 30 seconds
Same endpoint, same SDK — only the model string changes.
from openai import OpenAI
client = OpenAI(api_key="YOUR_API_KEY", base_url="https://api.fastinfra.ai/v1")
for model in ["meta-llama/llama-3.2-3b-instruct", "meta-llama/Meta-Llama-3.1-8B-Instruct-Turbo"]:
response = client.chat.completions.create(
model=model,
messages=[{"role": "user", "content": "Hello!"}]
)
print(model, "->", response.choices[0].message.content)