LLM API free tiers, observed

Updated · JSON · CC BY 4.0 · 日本語

This page is primary data aggregated from the responses we received while using LLM API free tiers in our own daily work (periods and counts in each section). Key points: Gemini API free-tier limits are per project and per model, and they reset at midnight Pacific time as documented (our first call after the reset succeeded 12 of 17 times). Cloudflare Workers AI documents a reset at 00:00 UTC, but we kept receiving the limit response for a while afterwards: the first success came a median 28.6 minutes (max 82.9) after the reset.

Google Gemini API (free tier)

Observed 2026-08-22 〜 2026-10-03・successful calls 1,597・failed responses (incl. our re-checks while limited) 2,994

What the provider documents
StatementSource
Rate limits are applied per project, not per API keyai.google.dev/gemini-api/docs/rate-limits (確認 2026-10-02)
Requests per day (RPD) quotas reset at midnight Pacific timeai.google.dev/gemini-api/docs/rate-limits (確認 2026-10-02)
The page does not list free-tier numbers; it points to AI Studio for active limitsai.google.dev/gemini-api/docs/rate-limits (確認 2026-10-02)

What we observed (differences from the docs)

  • The daily-limit 429 body named the quota GenerateRequestsPerDayPerProjectPerModel-FreeTier with value 20 (20 requests per day per project per model). The documentation page lists no free-tier numbers (observed 2026-10-02・n = 1)
Error types we received
TypeCount (days)Reset check / time to next successResponse body (start)
Daily limit2,528 (18)After the documented reset our first call succeeded 12 times and was still limited 0 times (of 17). Minutes from reset to first success: median 1.6, max 8.4 (7 days)
{ "error": { "code": 429, "message": "You exceeded your current quota, please check your plan and bi
Overloaded (503)365 (41)Hours to next success: median 1.5, p90 9.1 (n=365)
{ "error": { "code": 503, "message": "This model is currently experiencing high demand. Spikes in de
Timeout94 (17)Hours to next success: median 3.5, p90 12.0 (n=94)
The read operation timed out
Other7 (2)Hours to next success: median 0.7, p90 1.9 (n=7)
<urlopen error EOF occurred in violation of protocol (_ssl.c:2406)>

Models that answered: gemini-3.7-flash 950、gemini-3.5-flash 647

Cloudflare Workers AI (free allocation)

Observed 2026-09-22 〜 2026-10-03・successful calls 497・failed responses (incl. our re-checks while limited) 4,025

What the provider documents
StatementSource
10,000 Neurons per day at no chargedevelopers.cloudflare.com/workers-ai/platform/pricing/ (確認 2026-10-02)
All limits reset daily at 00:00 UTCdevelopers.cloudflare.com/workers-ai/platform/pricing/ (確認 2026-10-02)
Error types we received
TypeCount (days)Reset check / time to next successResponse body (start)
Daily limit4,022 (11)After the documented reset our first call succeeded 4 times and was still limited 6 times (of 11). Minutes from reset to first success: median 28.6, max 82.9 (6 days)
{"errors":[{"message":"AiError: AiError: you have used up your daily free allocation of 10,000 neuro
Timeout2 (1)Hours to next success: median 0.0, p90 0.0 (n=2)
The read operation timed out
Other1 (1)Hours to next success: median 9.8, p90 9.8 (n=1)
error code: 502

Models that answered: @cf/meta/llama-3.3-70b-instruct-fp8-fast 259、@cf/openai/gpt-oss-120b 124、@cf/mistralai/mistral-small-3.1-24b-instruct 91、@cf/qwen/qwen3.8-27b 23

Alibaba Cloud Model Studio (Qwen, international / Singapore)

Observed 2026-08-13 〜 2026-10-03・successful calls 3,787・failed responses (incl. our re-checks while limited) 56

What the provider documents
StatementSource
Free quota is typically 1,000,000 tokens per model, valid for 90 days, not shared across modelswww.alibabacloud.com/help/en/model-studio/new-free-quota (確認 2026-10-02)
With 'Free Quota Only' enabled, exhausted quota returns AllocationQuota.FreeTierOnly (disabled by default: usage converts to pay-as-you-go)www.alibabacloud.com/help/en/model-studio/new-free-quota (確認 2026-10-02)
Error types we received
TypeCount (days)Reset check / time to next successResponse body (start)
Free quota used up39 (7)—
{"error":{"message":"Free quota exhausted. To continue accessing the model on a paid basis, please a
Timeout10 (2)Hours to next success: median 2.1, p90 4.3 (n=10)
<urlopen error timed out>
Other7 (3)Hours to next success: median 14.9, p90 30.1 (n=7)
{"error":{"message":"Model is not supported in current workspace service site","id":"<id>","type":"M

Models that answered: qwen-plus 1,669、qwen-max 1,074、qwen-plus-2025-07-28 805、qwen3.5-plus-2026-02-15 146、qwen3-max 59、qwen3.7-plus 34

Method and caveats

  • Observed by InfraOne from its own production calls (free tiers, one key per provider); no extra requests were made. Failure counts include our own re-checks while limited, so they depend on our call pattern. Until 2026-10-01 our side paused a provider for hours after a limit, so there are no observations during those pauses.
  • application error log (status and response body; bodies truncated to 180 chars before 2026-10-02); scrubbed full response bodies (from 2026-10-02); successful call log (timestamp, provider, model)
  • call the provider's free tier from one project/account, log every non-2xx response body and every success timestamp with the model name, and compute the same aggregates over the same window
  • Not included: cerebras, chatgpt, claude, codex, mistral, nvidia