OpenRouter
nvidia/nemotron-3-embed-1b:free
Context window33K tokens
Dimensions2,048
Free budget20 rpm · shared free daily cap
Rate limitsnot published
NVIDIA NIM
nvidia/nemotron-3-embed-1b
Context window4K tokens
Dimensions2,048
Free budget~40 rpm
Rate limitsnot published
Get this model the moment it changes
Limits move, models get replaced, better ones launch. Premium routers see this page's data live. Free routers see last month's.
Live schalten · $19/yr →Use it
FreeLLMAPI is a self-hosted router you run yourself. Install it, paste in your free OpenRouter key, and nvidia/nemotron-3-embed-1b:free answers on an OpenAI-compatible endpoint at http://localhost:3001/v1. No credit card, no hosted middleman: your prompts and your provider keys never leave your machine.
Install the router (macOS, Linux, WSL)
curl -fsSL https://freellmapi.co/install.sh | bashInstall the router (Windows PowerShell)
iwr -useb https://freellmapi.co/install.ps1 | iexCall nvidia/nemotron-3-embed-1b:free with curl
curl http://localhost:3001/v1/embeddings \
-H "Authorization: Bearer freellmapi-your-unified-key" \
-H "Content-Type: application/json" \
-d '{"model": "nvidia/nemotron-3-embed-1b:free", "input": "hello world"}'The same request in Python (openai)
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:3001/v1",
api_key="freellmapi-your-unified-key",
)
resp = client.embeddings.create(
model="nvidia/nemotron-3-embed-1b:free",
input=["the quick brown fox"],
)
print(len(resp.data[0].embedding), "dims")The router answers on /v1/embeddings and every other OpenAI surface, plus the Anthropic Messages API, so existing clients need only a new base_url. Swap the model id for auto and the router picks the best free model that is still under its limits. Full reference: docs/api.md.