Google AI Studio

gemma-4-31b-it
Context window33K tokens
Free budget~30M tokens/mo
Requests / min15 RPM
Requests / day1K RPD
Tokens / min250K TPM

Sail Research

google/gemma-4-31B-it
Context window262K tokens
Free budget$5/month recurring credit · payment method required tokens/mo
Rate limitsnot published
Monthly credit requires a payment method warning
Sail grants $5 in free credits each month only while a payment method is attached. Usage becomes pay-as-you-go after that credit is exhausted, so configure Sail spend controls and monitor the account balance. Requests run as background jobs and may remain queued for several minutes.

Get this model the moment it changes

Limits move, models get replaced, better ones launch. Premium routers see this page's data live. Free routers see last month's.

লাইভ করুন · $19/yr →

Use it

FreeLLMAPI is a self-hosted router you run yourself. Install it, paste in your free Google AI Studio key, and gemma-4-31b-it answers on an OpenAI-compatible endpoint at http://localhost:3001/v1. No credit card, no hosted middleman: your prompts and your provider keys never leave your machine.

Install the router (macOS, Linux, WSL)
curl -fsSL https://freellmapi.co/install.sh | bash
Install the router (Windows PowerShell)
iwr -useb https://freellmapi.co/install.ps1 | iex
Call gemma-4-31b-it with curl
curl http://localhost:3001/v1/chat/completions \
  -H "Authorization: Bearer freellmapi-your-unified-key" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "gemma-4-31b-it",
    "messages": [{"role": "user", "content": "Say hi in five words."}]
  }'
The same request in Python (openai)
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:3001/v1",
    api_key="freellmapi-your-unified-key",
)

resp = client.chat.completions.create(
    model="gemma-4-31b-it",
    messages=[{"role": "user", "content": "Say hi in five words."}],
)
print(resp.choices[0].message.content)

The router answers on /v1/chat/completions and every other OpenAI surface, plus the Anthropic Messages API, so existing clients need only a new base_url. Swap the model id for auto and the router picks the best free model that is still under its limits. Full reference: docs/api.md.