Provider

Free NVIDIA API

NVIDIA NIM's recurring free tier (build.nvidia.com) — DeepSeek, Llama, Nemotron and Gemma at 40 RPM.

FreeLLMAPI is a free, open-source LLM API that routes across every provider with a real free tier. The NVIDIA models below are on the free tier of NVIDIA NIM — reachable through one OpenAI-compatible key, with automatic failover when a provider hits its rate limit.

Free NVIDIA models (18)

ModelContextFree limitsCapabilities
nvidia/nemotron-3-ultra-550b-a55b1.0M40 rpmtools
meta/muse-glimmer-30b131K40 rpmtools, vision
google/gemma-4-31b-it262K40 rpm—
poolside/laguna-xs-2.1262K40 rpmtools
nvidia/nemotron-3-super-120b-a12b262K40 rpmtools
nvidia/nemotron-3.5-lightning-30b-a3b1M40 rpmtools
google/diffusiongemma-26b-a4b-it262K40 rpmvision
openai/gpt-oss-20b131K40 rpmtools
nvidia/nemotron-3-nano-omni-30b-a3b-reasoning262K40 rpmvision
nvidia/ising-calibration-1.5-31b131K40 rpmvision
meta/llama-3.2-90b-vision-instruct131K40 rpmtools, vision
meta/llama-3.2-11b-vision-instruct131K40 rpmtools, vision
nvidia/riva-translate-4b-instruct-v233K40 rpm—
nvidia/riva-translate-4b-instruct-v1.133K40 rpm—
nvidia/nemotron-3.5-content-safety131K40 rpm—
nvidia/llama-3.1-nemoguard-8b-content-safety131K40 rpm—
nvidia/llama-3.1-nemoguard-8b-topic-control131K40 rpm—
nvidia/llama-3.1-nemotron-safety-guard-8b-v3131K40 rpm—

How to use NVIDIA for free

  1. Install FreeLLMAPI — the open-source router (GitHub). It runs locally and keeps your keys on your machine.
  2. Add a free key for NVIDIA NIM on the Keys page — no credit card required.
  3. Point your OpenAI client at the local endpoint and pick a model:
from openai import OpenAI

# FreeLLMAPI runs locally; grab your unified key + endpoint on the Keys page.
client = OpenAI(base_url="http://localhost:3001/v1", api_key="freellmapi-...")

resp = client.chat.completions.create(
    model="google/gemma-4-31b-it",
    messages=[{"role": "user", "content": "Hello!"}],
)
print(resp.choices[0].message.content)
Use every one of these through one key. FreeLLMAPI is an open-source, self-hosted router that puts all these free tiers behind a single OpenAI-compatible endpoint and fails over when one is rate-limited. Browse the catalog or go live.

Frequently asked questions

Is the NVIDIA API really free?

Yes — these NVIDIA models run on genuine provider free tiers (on the free tier of NVIDIA NIM). Inference costs nothing; you only add a free provider key.

Do I need a credit card?

No. The providers here offer free tiers that work without a card. You add the free key once and FreeLLMAPI routes to it.

How do I call NVIDIA through FreeLLMAPI?

Install the open-source router, add the provider's free key on the Keys page, then point any OpenAI SDK at your local endpoint — see the code sample above.

← Browse all 600+ free models · Best free LLM APIs 2026

Keep reading

All posts →
Provider

Free Gemini API

Google AI Studio's permanent free tier — Gemini and Gemma, no credit card.

Provider

Free DeepSeek API

DeepSeek V3, R1 and V4 served free across multiple providers.

Provider

Free Llama API

Meta's Llama models free across 9 providers, behind one key.