Groq
openai/gpt-oss-20b
Context window131K tokens
Free budget~6M tokens/mo
Requests / min30 RPM
Requests / day1K RPD
Tokens / min8K TPM
Tokens / day200K TPD
Cloudflare Workers AI
@cf/openai/gpt-oss-20b
Context window131K tokens
Free budget~18-45M tokens/mo
Rate limitsnot published
Key is account_id:token info
Cloudflare Workers AI authenticates with a combined credential in the form "account_id:token", not a bare token.
Needs token room info
This route can spend hidden reasoning tokens before visible output. Avoid tiny max_tokens values or the request can finish by length with empty content.
OVH AI Endpoints
gpt-oss-20b
Context window131K tokens
Free budgetfree · 2/min per IP tokens/mo
Requests / min2 RPM
Anonymous tier is 2 req/min warning
OVH AI Endpoints anonymous mode is documented at 2 req/min per IP per model (observed even stricter across models). The 400 req/min authenticated tier requires a Public Cloud project with a payment method, so the catalog ships the keyless path. Treat as a breadth/fallback tier, not a throughput tier.
No API key required info
Routes anonymously — the catalog ships a keyless sentinel row and calls work with no account or key.
Needs token room info
Some free routes spend hidden reasoning tokens before visible output. Avoid tiny max_tokens values or requests can finish by length with empty content.
NVIDIA NIM
openai/gpt-oss-20b
Context window131K tokens
Free budgetfree · 40 RPM tokens/mo
Requests / min40 RPM
Recurring free, 40 RPM, eval-only ToS info
NVIDIA NIM replaced its depleting trial credits with a recurring per-account rate limit (40 RPM default, varies by model), verified June 2026. The trial ToS still scopes usage to evaluation/prototyping, not production.
Ollama Cloud
gpt-oss:20b
Context window131K tokens
Free budget~20-30M tokens/mo
Rate limitsnot published
HuggingFace Router
openai/gpt-oss-20b
Context window131K tokens
Free budget$0.10/mo shared credit tokens/mo
Rate limitsnot published
Small $0.10/month routed credit warning
HuggingFace Inference Providers grants only ~$0.10/month of routed credit on the free tier (PRO is $2/month). Enough for light experimentation; exhausts quickly. Credits apply only to HF-routed requests.
Monthly included credits run out info
HuggingFace meters Inference Providers in dollars, not tokens: free accounts get $0.10 credited every month. When it is spent every route answers 402 "You have depleted your monthly included credits", as happened during the September 18, 2026 audit. This is a spent-wallet state, not a dead route, and it clears when the monthly grant renews. Rows are kept enabled and the router should treat 402 here as a provider-level cooldown to the start of next month rather than a per-model failure.
Get this model the moment it changes
Limits move, models get replaced, better ones launch. Premium routers see this page's data live. Free routers see last month's.
Live schalten · $19/yr →Use it
FreeLLMAPI is a self-hosted router you run yourself. Install it, paste in your free Groq key, and openai/gpt-oss-20b answers on an OpenAI-compatible endpoint at http://localhost:3001/v1. No credit card, no hosted middleman: your prompts and your provider keys never leave your machine.
Install the router (macOS, Linux, WSL)
curl -fsSL https://freellmapi.co/install.sh | bashInstall the router (Windows PowerShell)
iwr -useb https://freellmapi.co/install.ps1 | iexCall openai/gpt-oss-20b with curl
curl http://localhost:3001/v1/chat/completions \
-H "Authorization: Bearer freellmapi-your-unified-key" \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-oss-20b",
"messages": [{"role": "user", "content": "Say hi in five words."}]
}'The same request in Python (openai)
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:3001/v1",
api_key="freellmapi-your-unified-key",
)
resp = client.chat.completions.create(
model="openai/gpt-oss-20b",
messages=[{"role": "user", "content": "Say hi in five words."}],
)
print(resp.choices[0].message.content)The router answers on /v1/chat/completions and every other OpenAI surface, plus the Anthropic Messages API, so existing clients need only a new base_url. Swap the model id for auto and the router picks the best free model that is still under its limits. Full reference: docs/api.md.