The same three complaints, every few months
Threads about OpenRouter's free tier keep returning to the same three points.
- Models vanish without warning. A model in the free pool is there one week and gone the next, and nothing tells you before your requests start failing.
- Free models crawl under load. When the pool is busy, free requests wait behind everyone else's, usually at the hour you need them.
- Free-tier behavior misleads you about the model. A model that is throttled, queued and served on spare capacity is a poor way to judge what that model can actually do.
The first one is the re-shop moment. Here is what a router with more than one provider behind it does with it.
What FreeLLMAPI does when a provider drops a free model
1. The catalog removes it for you
Your router pulls a signed model catalog from freellmapi.co twice a day. When a provider retires a free model, we disable or drop its catalog row, and your next sync applies that: a disabled row is switched off, and a row that vanished from the catalog is deleted locally. You do not edit a config file or pull a new release.
2. Between syncs, the router retires it on its own
If a provider answers a request with a definite end-of-life error, the router disables that model immediately. A merely probable signal, such as a "no longer available" 404, has to be confirmed by a second, separate request within an hour before the model is disabled, so one flaky response from a load balancer cannot kill a healthy model. If a later catalog still lists the model, the retirement is lifted.
3. The request fails over instead of failing
Every request walks your fallback chain: on a 429, a 5xx, a timeout or a retired model it moves to the next candidate, up to 20 hops per request. The same model served by several providers is unified into one entry with failover inside the group, so a request for a model OpenRouter dropped goes to another provider that still serves it, if you hold a key there.
4. Rate limits are tracked, not discovered by failing
The router counts requests and tokens per key against each model's published per-minute and per-day limits and skips a route that is at its cap. A transient 429 benches that key for 90 seconds; a spent daily quota benches it on an escalating ladder that tops out at a day; a provider's Retry-After header is honored. A model that fails three times in fifteen minutes across keys is benched for ten minutes so it stops costing you latency.
What it cannot do
- It cannot serve a model nobody offers free. If only one provider served it and that provider dropped it, it leaves your chain too. Failover only goes to models you can reach.
- Daily quotas still exist. Every free tier is metered, and the router enforces those limits itself. Some providers meter the whole account rather than each model, and the router gates those too: for OpenRouter it allows 1,000 requests a day by default, and
PROVIDER_DAILY_REQUEST_CAP_OPENROUTER=50lowers it if your account is on the smaller pool. Your real ceiling is the sum of the free tiers you hold keys for, not unlimited. - You bring your own keys. One free account per provider, added once on the Keys page. FreeLLMAPI has no shared pool of its own, and your keys never leave your machine.
- Free installs get the catalog on a delay. A free install pulls the monthly snapshot: a new model reaches it 30 days after joining the live feed, and a dropped model leaves at a later sync, normally within a day or two. Premium ($19 a year or $49 once) pulls the live tier.
- It cannot make a free tier faster. The router can rank routes by observed speed and route around a model that keeps timing out, but it cannot change how a provider serves free traffic.
Point your client at it
Run the router, add the free keys you have, and point any OpenAI SDK at it. Each response carries an X-Routed-Via header naming the provider and model that actually answered.
from openai import OpenAI
# FreeLLMAPI runs locally; grab your unified key + endpoint on the Keys page.
client = OpenAI(base_url="http://localhost:3001/v1", api_key="freellmapi-...")
resp = client.chat.completions.create(
model="auto",
messages=[{"role": "user", "content": "Hello!"}],
)
print(resp.choices[0].message.content)
Frequently asked questions
Does FreeLLMAPI include OpenRouter's free models?
Yes. OpenRouter is one of the providers in the catalog. Add your OpenRouter key and its free models join your chain next to every other provider's, with OpenRouter's account-wide daily cap tracked by the router so it stops before the provider does.
How quickly does a dropped model leave my router?
Two ways. The catalog sync runs twice a day: Premium routers pull the live tier, free installs pull the monthly snapshot, and a row we disabled or removed goes with it. Independently of the catalog, a definite end-of-life answer from the provider retires the model on your router at once, and a probable one is confirmed by a second request within an hour.
Do I need Premium for failover?
No. Routing, failover, cooldowns, quota tracking and auto-retirement are all in the free, open-source router. Premium ($19 a year or $49 once) only changes how quickly catalog changes reach you: same day instead of the monthly snapshot.
FreeLLMAPI vs OpenRouter → · OpenRouter free alternative · OpenRouter's free models through one key · Free LLM API model catalog