What actually counts as free
- Rate-limited free tiers — the real thing. Capped by requests or tokens per day, renewed forever, usually no card. This is most of the catalog.
- One-time signup credits — "$5 free" offers. Useful once, then gone; not a free API.
- Card-first trials — free until the card is charged. Excluded from the catalog.
The strongest free models right now
| Model | Provider | Context | Free limits | Capabilities |
|---|---|---|---|---|
| glm-5.3:free | UnoRouter | 1M | 1 rpm | tools |
| moonshotai/Kimi-K3 | HuggingFace Router | 262K | $0.10/mo credit | tools |
| glm-5.3-flash:free | UnoRouter | 1M | 1 rpm | tools, vision |
| Qwen/Qwen3.8-2.4T-A95B | HuggingFace Router | 1.0M | $0.10/mo shared credit | tools |
| gemini-3.7-flash | Google AI Studio | 1.0M | 10 rpm, 20 rpd | tools, vision |
| dots-studio/dots-3-note-preview:free | OpenRouter | 512K | 20 rpm, 50 rpd | tools, vision |
| dots-studio/dots-3-note-preview:free | Kilo Gateway | 512K | free · 200/hr per IP | tools, vision |
| deepseek-ai/DeepSeek-V4-Pro-0813 | HuggingFace Router | 1.0M | $0.10/mo shared credit | tools |
| deepseek/deepseek-v4-flash-0731free:free | BazaarLink | 1.0M | 10 rpm, 50 rpd | tools |
| deepseek-ai/DeepSeek-V4-Flash-0731 | HuggingFace Router | 1.0M | $0.10/mo shared credit | tools |
| deepseek-ai/DeepSeek-V4-Flash-0731 | Sail Research | 1.0M | $5/month recurring credit · payment method required | tools |
| zai-org/GLM-5.2 | HuggingFace Router | 200K | $0.10/mo credit | tools |
| gemini-3.6-flash | Google AI Studio | 1.0M | 10 rpm, 20 rpd | tools, vision |
| ZhipuAI/GLM-5.2 | ModelScope | 200K | 100 rpd | — |
| zai-org/GLM-5.2-FP8 | HuggingFace Router | 131K | $0.10/mo shared credit | tools |
| @cf/qwen/qwen3.8-27b | Cloudflare Workers AI | 262K | free · shared 10k neurons/day | tools, vision |
| qwen/qwen3.8-27b | Groq | 262K | 30 rpm, 250 rpd | tools, vision |
| gemini-3.5-flash | Google AI Studio | 1.0M | 10 rpm, 20 rpd | tools, vision |
The full, searchable list lives in the free LLM API model catalog. Together the tracked tiers add up to at least 7.4 billion free tokens a month.
Call it like any OpenAI endpoint
FreeLLMAPI is an open-source router you run locally: add each provider's free key once and every model above sits behind one OpenAI-compatible endpoint, with automatic failover when a provider rate-limits.
from openai import OpenAI
# FreeLLMAPI runs locally; grab your unified key + endpoint on the Keys page.
client = OpenAI(base_url="http://localhost:3001/v1", api_key="freellmapi-...")
resp = client.chat.completions.create(
model="gpt-oss-120b",
messages=[{"role": "user", "content": "Hello!"}],
)
print(resp.choices[0].message.content)
Or straight from the shell:
curl http://localhost:3001/v1/chat/completions \
-H "Authorization: Bearer freellmapi-..." \
-H "Content-Type: application/json" \
-d '{"model": "gpt-oss-120b", "messages": [{"role": "user", "content": "Hello!"}]}'
Free tiers by provider
Each guide covers one provider's genuinely free tier — models, limits, quirks, and how to get the key: Gemini · Groq · Mistral · NVIDIA · Cloudflare Workers AI · OpenRouter · Cohere · Cerebras · GLM (Z.ai / Zhipu) · Ollama Cloud · DeepSeek · Llama.
Frequently asked questions
Is there a totally free LLM API?
Yes — 600+ models across 34 providers run on genuine free tiers right now: no card, no trial clock. They are rate-limited (requests or tokens per day), not credit-limited, so they keep working month after month.
Which LLM API is free without a credit card?
Google AI Studio (Gemini), Groq, Cerebras, NVIDIA NIM, Cloudflare Workers AI, Mistral and Cohere all issue keys with no card. FreeLLMAPI puts all of them behind one OpenAI-compatible key.
Is there a free LLM API with no limits?
No — every real free tier has rate limits, and anything claiming otherwise is a trial. The practical fix is combining several free tiers: FreeLLMAPI fails over between providers automatically, so their combined headroom behaves like one bigger limit. How close you can get to unlimited.
Can I use a free LLM API in production?
For low-volume products, yes, with failover across providers. For anything latency- or volume-critical you will outgrow free tiers; until then they are real capacity, not demos.
Get a free LLM API key, step by step → · Best free LLM APIs 2026 · Free LLM API gateways compared · State of Free LLM APIs 2026