FreeLLMAPI is a free, open-source LLM API that routes across every provider with a real free tier. The NVIDIA models below are on the free tier of NVIDIA NIM — reachable through one OpenAI-compatible key, with automatic failover when a provider hits its rate limit.
Free NVIDIA models (18)
| Model | Context | Free limits | Capabilities |
|---|---|---|---|
| nvidia/nemotron-3-ultra-550b-a55b | 1.0M | 40 rpm | tools |
| meta/muse-glimmer-30b | 131K | 40 rpm | tools, vision |
| google/gemma-4-31b-it | 262K | 40 rpm | — |
| poolside/laguna-xs-2.1 | 262K | 40 rpm | tools |
| nvidia/nemotron-3-super-120b-a12b | 262K | 40 rpm | tools |
| nvidia/nemotron-3.5-lightning-30b-a3b | 1M | 40 rpm | tools |
| google/diffusiongemma-26b-a4b-it | 262K | 40 rpm | vision |
| openai/gpt-oss-20b | 131K | 40 rpm | tools |
| nvidia/nemotron-3-nano-omni-30b-a3b-reasoning | 262K | 40 rpm | vision |
| nvidia/ising-calibration-1.5-31b | 131K | 40 rpm | vision |
| meta/llama-3.2-90b-vision-instruct | 131K | 40 rpm | tools, vision |
| meta/llama-3.2-11b-vision-instruct | 131K | 40 rpm | tools, vision |
| nvidia/riva-translate-4b-instruct-v2 | 33K | 40 rpm | — |
| nvidia/riva-translate-4b-instruct-v1.1 | 33K | 40 rpm | — |
| nvidia/nemotron-3.5-content-safety | 131K | 40 rpm | — |
| nvidia/llama-3.1-nemoguard-8b-content-safety | 131K | 40 rpm | — |
| nvidia/llama-3.1-nemoguard-8b-topic-control | 131K | 40 rpm | — |
| nvidia/llama-3.1-nemotron-safety-guard-8b-v3 | 131K | 40 rpm | — |
How to use NVIDIA for free
- Install FreeLLMAPI — the open-source router (GitHub). It runs locally and keeps your keys on your machine.
- Add a free key for NVIDIA NIM on the Keys page — no credit card required.
- Point your OpenAI client at the local endpoint and pick a model:
from openai import OpenAI
# FreeLLMAPI runs locally; grab your unified key + endpoint on the Keys page.
client = OpenAI(base_url="http://localhost:3001/v1", api_key="freellmapi-...")
resp = client.chat.completions.create(
model="google/gemma-4-31b-it",
messages=[{"role": "user", "content": "Hello!"}],
)
print(resp.choices[0].message.content)
Frequently asked questions
Is the NVIDIA API really free?
Yes — these NVIDIA models run on genuine provider free tiers (on the free tier of NVIDIA NIM). Inference costs nothing; you only add a free provider key.
Do I need a credit card?
No. The providers here offer free tiers that work without a card. You add the free key once and FreeLLMAPI routes to it.
How do I call NVIDIA through FreeLLMAPI?
Install the open-source router, add the provider's free key on the Keys page, then point any OpenAI SDK at your local endpoint — see the code sample above.