LLM API free tiers, observed
Updated · JSON · CC BY 4.0 · 日本語
This page is primary data aggregated from the responses we received while using LLM API free tiers in our own daily work (periods and counts in each section). Key points: Gemini API free-tier limits are per project and per model, and they reset at midnight Pacific time as documented (our first call after the reset succeeded 12 of 17 times). Cloudflare Workers AI documents a reset at 00:00 UTC, but we kept receiving the limit response for a while afterwards: the first success came a median 28.6 minutes (max 82.9) after the reset.
Google Gemini API (free tier)
Observed 2026-08-22 〜 2026-10-03・successful calls 1,597・failed responses (incl. our re-checks while limited) 2,994
| Statement | Source |
|---|---|
| Rate limits are applied per project, not per API key | ai.google.dev/gemini-api/docs/rate-limits (確認 2026-10-02) |
| Requests per day (RPD) quotas reset at midnight Pacific time | ai.google.dev/gemini-api/docs/rate-limits (確認 2026-10-02) |
| The page does not list free-tier numbers; it points to AI Studio for active limits | ai.google.dev/gemini-api/docs/rate-limits (確認 2026-10-02) |
What we observed (differences from the docs)
- The daily-limit 429 body named the quota GenerateRequestsPerDayPerProjectPerModel-FreeTier with value 20 (20 requests per day per project per model). The documentation page lists no free-tier numbers (observed 2026-10-02・n = 1)
| Type | Count (days) | Reset check / time to next success | Response body (start) |
|---|---|---|---|
| Daily limit | 2,528 (18) | After the documented reset our first call succeeded 12 times and was still limited 0 times (of 17). Minutes from reset to first success: median 1.6, max 8.4 (7 days) | { "error": { "code": 429, "message": "You exceeded your current quota, please check your plan and bi |
| Overloaded (503) | 365 (41) | Hours to next success: median 1.5, p90 9.1 (n=365) | { "error": { "code": 503, "message": "This model is currently experiencing high demand. Spikes in de |
| Timeout | 94 (17) | Hours to next success: median 3.5, p90 12.0 (n=94) | The read operation timed out |
| Other | 7 (2) | Hours to next success: median 0.7, p90 1.9 (n=7) | <urlopen error EOF occurred in violation of protocol (_ssl.c:2406)> |
Models that answered: gemini-3.7-flash 950、gemini-3.5-flash 647
Cloudflare Workers AI (free allocation)
Observed 2026-09-22 〜 2026-10-03・successful calls 497・failed responses (incl. our re-checks while limited) 4,025
| Statement | Source |
|---|---|
| 10,000 Neurons per day at no charge | developers.cloudflare.com/workers-ai/platform/pricing/ (確認 2026-10-02) |
| All limits reset daily at 00:00 UTC | developers.cloudflare.com/workers-ai/platform/pricing/ (確認 2026-10-02) |
| Type | Count (days) | Reset check / time to next success | Response body (start) |
|---|---|---|---|
| Daily limit | 4,022 (11) | After the documented reset our first call succeeded 4 times and was still limited 6 times (of 11). Minutes from reset to first success: median 28.6, max 82.9 (6 days) | {"errors":[{"message":"AiError: AiError: you have used up your daily free allocation of 10,000 neuro |
| Timeout | 2 (1) | Hours to next success: median 0.0, p90 0.0 (n=2) | The read operation timed out |
| Other | 1 (1) | Hours to next success: median 9.8, p90 9.8 (n=1) | error code: 502 |
Models that answered: @cf/meta/llama-3.3-70b-instruct-fp8-fast 259、@cf/openai/gpt-oss-120b 124、@cf/mistralai/mistral-small-3.1-24b-instruct 91、@cf/qwen/qwen3.8-27b 23
Alibaba Cloud Model Studio (Qwen, international / Singapore)
Observed 2026-08-13 〜 2026-10-03・successful calls 3,787・failed responses (incl. our re-checks while limited) 56
| Statement | Source |
|---|---|
| Free quota is typically 1,000,000 tokens per model, valid for 90 days, not shared across models | www.alibabacloud.com/help/en/model-studio/new-free-quota (確認 2026-10-02) |
| With 'Free Quota Only' enabled, exhausted quota returns AllocationQuota.FreeTierOnly (disabled by default: usage converts to pay-as-you-go) | www.alibabacloud.com/help/en/model-studio/new-free-quota (確認 2026-10-02) |
| Type | Count (days) | Reset check / time to next success | Response body (start) |
|---|---|---|---|
| Free quota used up | 39 (7) | — | {"error":{"message":"Free quota exhausted. To continue accessing the model on a paid basis, please a |
| Timeout | 10 (2) | Hours to next success: median 2.1, p90 4.3 (n=10) | <urlopen error timed out> |
| Other | 7 (3) | Hours to next success: median 14.9, p90 30.1 (n=7) | {"error":{"message":"Model is not supported in current workspace service site","id":"<id>","type":"M |
Models that answered: qwen-plus 1,669、qwen-max 1,074、qwen-plus-2025-07-28 805、qwen3.5-plus-2026-02-15 146、qwen3-max 59、qwen3.7-plus 34
Method and caveats
- Observed by InfraOne from its own production calls (free tiers, one key per provider); no extra requests were made. Failure counts include our own re-checks while limited, so they depend on our call pattern. Until 2026-10-01 our side paused a provider for hours after a limit, so there are no observations during those pauses.
- application error log (status and response body; bodies truncated to 180 chars before 2026-10-02); scrubbed full response bodies (from 2026-10-02); successful call log (timestamp, provider, model)
- call the provider's free tier from one project/account, log every non-2xx response body and every success timestamp with the model name, and compute the same aggregates over the same window
- Not included: cerebras, chatgpt, claude, codex, mistral, nvidia