Skip to main content

GPT Image 2 Usage Limits: API Tiers and What Each 429 Means

The OpenAI API allows gpt-image-2 5 to 250 images per minute, depending on usage tier. A 429 can also be a credit or spend limit that retrying will not clear.

Yingtu AI Editorial
Yingtu AI Editorial
Updated 10 min
GPT Image 2 usage limits: 5 to 250 images per minute by API tier, a per-minute 429 that clears after waiting, and a credit or spend limit 429 that retrying will not fix
yingtu.ai

A GPT Image 2 usage limit on the OpenAI API is a per-minute allowance that OpenAI assigns to your organization according to its usage tier. As of October 1, 2026, gpt-image-2 runs from 5 images per minute on Tier 1 to 250 images per minute on Tier 5, and the Free tier cannot call the model at all.

That allowance is only one of the things an HTTP 429 can report. The same status code is returned when your prepaid credits run out, when a spend limit you set yourself is reached, and when your organization reaches the monthly usage limit OpenAI approved for it. Waiting fixes the first kind. It never fixes the other three. The error.code field in the response body tells you which one you have.

Image generation inside the ChatGPT app is a separate system with its own allowance. A ChatGPT message such as "No image generation quota is currently available." has nothing to do with your API tier, and a paid ChatGPT plan does not add API capacity. For the app side, see How Many Images Can ChatGPT Generate for Free? Limits Explained.

GPT Image 2 API rate limits by usage tier

OpenAI's gpt-image-2 model page publishes two per-minute limits for each tier. TPM is tokens per minute. IPM is images per minute. The rate limits guide adds the conditions for reaching each tier and the monthly usage limit that comes with it.

Usage tierYou qualify afterApproved monthly usage limitTPMIPM
FreeBeing in an allowed geography$100 per monthNot supportedNot supported
Tier 1$5 paid$100 per month100,0005
Tier 2$50 paid$500 per month250,00020
Tier 3$100 paid$1,000 per month800,00050
Tier 4$250 paid$5,000 per month3,000,000150
Tier 5$1,000 paid$200,000 per month8,000,000250

Limits for gpt-image-2 (snapshot gpt-image-2-2026-04-21) on the OpenAI API, as of October 1, 2026.

Four details decide how these numbers behave in practice.

  1. Whichever limit is reached first stops you. A Tier 1 organization that stays under 100,000 tokens per minute is still refused on the sixth image inside one minute. At a steady pace, 5 IPM works out to at most 5 × 60 = 300 images per hour on Tier 1, and fewer if the token limit is reached first.
  2. Limits belong to the organization and the project, not to a user or an API key. Two scripts or two teammates working in the same project draw on the same allowance. A new key in that project adds nothing.
  3. Changing the endpoint does not change the allowance. gpt-image-2 is available through v1/images/generations, v1/images/edits, the Responses API and Batch, and only Batch is counted separately (see the section on getting more capacity). Some model families also share one limit, which the Limits page marks as a shared limit.
  4. The published table is the default, and your account can differ. The Limits page in your organization settings shows the numbers that actually apply to you, per model.

GPT Image 2.5 follows the same pattern: as of September 30, 2026, the model page for gpt-image-2.5-flare lists the same tier table as gpt-image-2.

Which 429 you have, and whether retrying helps

Read error.code before deciding anything. OpenAI's error codes page lists the following cases for image and other API calls, as of October 1, 2026.

What the response saysKind of limitDoes waiting and retrying help?Where it gets fixed
429 "Rate limit reached for requests"Per-minute rate limit (TPM or IPM) for your tierYesIn your code: wait for Retry-After, then send more slowly
429 slow_down (type rate_limit_error)Traffic rose too quickly, even if it is inside your per-minute limitsYes, after you reduce the rateIn your code: drop the rate, then ramp up gradually
429 credit_balance_exhaustedNo prepaid credits leftNoBilling: add credits
429 organization_spend_limit_exceededA hard spend limit set on your organizationNoOrganization settings: raise the hard spend limit
429 project_spend_limit_exceededA hard spend limit set on the projectNoProject settings: raise the hard spend limit
429 organization_usage_limit_exceededThe monthly usage limit OpenAI approved for your organizationNoRequest a higher approved limit or contact OpenAI support
503 server_is_overloaded (type service_unavailable_error)No limit of yours; the model is temporarily overloadedYesNothing to change: retry after a pause
403 "Country, region, or territory not supported", or a refusal that mentions access or verificationNo rate limit at allNoAccount eligibility, not request pacing

Two notes on reading the field. For billing-related errors the broader error.type can still read insufficient_quota, so the code is the part that separates the causes. Older SDK versions and third-party gateways may return only insufficient_quota or rate_limit_exceeded. Treat the first as a billing problem that retrying will not clear, and the second as a per-minute rate limit.

"Rate limit reached for requests" and slow_down

Both of these clear on their own once you send less, and they are the only 429s where a retry is the right response. "Rate limit reached for requests" means you sent more per minute than your tier allows. slow_down is different: it can appear while you are still inside your requests-per-minute and tokens-per-minute limits, because it reflects how fast your traffic grew. OpenAI's rule of thumb for traffic above 1 million input tokens per minute is to increase it by no more than 50% every 15 minutes.

Retrying immediately makes things worse. Failed requests count toward the per-minute limit, so a loop that resends at once keeps the limit exhausted.

credit_balance_exhausted

This code means the organization has no prepaid credits left, and the only fix is adding credits on the billing page. No amount of waiting brings the calls back.

organization_spend_limit_exceeded and project_spend_limit_exceeded

These two codes come from a hard spend limit that someone in your organization configured, at organization or project level. OpenAI offers two controls here. A spend alert only sends a notification, and API traffic continues. A hard spend limit makes the affected requests return 429. If you did not set the limit yourself, ask whoever administers the organization or project. Raising it is a settings change, and it takes no request to OpenAI.

organization_usage_limit_exceeded

This code means your organization reached the approved monthly usage limit in the tier table above, such as $100 per month on Tier 1. OpenAI assigns that limit, and it is separate from the spend limits you configure. You can request a higher approved limit or contact support. Moving up a tier raises it as well.

503 server_is_overloaded

A 503 with this code says the model is temporarily overloaded on OpenAI's side. Your tier, credits and spend limits are not involved. Retry with the same pause logic you use for rate limits.

Access and verification refusals

A refusal that mentions country, access or verification is not a rate limit, and waiting will not clear it. OpenAI may require organization verification before certain models or features can be used. If the message points there, the fix is in your organization's settings and eligibility.

Where to see your own limits

Three places show your real numbers, as opposed to the published defaults.

  • The Limits page in your organization settings lists the per-model limits that apply to your organization.
  • Response headers on each call report the request and token allowance: x-ratelimit-limit-requests, x-ratelimit-remaining-requests, x-ratelimit-reset-requests, the matching -tokens headers, and project-scoped versions such as x-ratelimit-remaining-project-tokens.
  • Retry-After, when it is present on a 429 or 503, gives the minimum number of seconds to wait.

The documented headers cover requests and tokens only. None of them reports how many images you have left this minute, so keep your own count of images sent per minute if you run close to the IPM limit.

Retry-After can appear on a 429 caused by a temporary rate limit and on a 503 caused by overload. OpenAI states that the header does not mean a quota or billing error can be resolved by retrying.

A retry loop that stops on billing errors

The Python sketch below uses the official openai SDK. It retries only the temporary cases, waits at least as long as Retry-After asks, adds a small random delay, and gives up after a fixed number of attempts or a fixed total time. Treat it as a starting point and adapt the limits to your own workload.

hljs python
import random
import time

from openai import OpenAI, InternalServerError, RateLimitError

# The SDK retries eligible 429 and 503 responses by itself.
# Turn that off so this loop is the only one retrying.
client = OpenAI(max_retries=0)

# Codes and types that need an action from you. Never retry these.
NO_RETRY = {
    "credit_balance_exhausted",
    "organization_spend_limit_exceeded",
    "project_spend_limit_exceeded",
    "organization_usage_limit_exceeded",
    "insufficient_quota",
}


def generate_image(prompt, max_attempts=5, max_total_seconds=120):
    started = time.monotonic()
    for attempt in range(max_attempts):
        try:
            return client.images.generate(model="gpt-image-2", prompt=prompt)
        except (RateLimitError, InternalServerError) as e:
            if e.code in NO_RETRY or e.type in NO_RETRY:
                raise  # billing, spend or usage limit: waiting will not help
            if isinstance(e, InternalServerError) and e.status_code != 503:
                raise

            header = e.response.headers.get("retry-after")
            try:
                wait = float(header)
            except (TypeError, ValueError):
                wait = min(2 ** attempt, 30)  # no usable header: exponential backoff
            wait += random.uniform(0, 1)  # small random delay

            out_of_attempts = attempt == max_attempts - 1
            out_of_time = time.monotonic() - started + wait > max_total_seconds
            if out_of_attempts or out_of_time:
                raise
            time.sleep(wait)

The max_retries=0 line matters. The official SDKs already retry eligible 429 and 503 responses automatically. If you leave that on and add your own loop, every outer attempt multiplies into several inner ones, and each failed attempt still counts toward your per-minute limit.

Getting more capacity

Moving up a tier. Tiers rise automatically as your organization's paid total on the API grows, at the amounts in the tier table. Each step raises the per-minute limits and the approved monthly usage limit together. Going from Tier 1 to Tier 2 takes the IPM limit from 5 to 20.

Using Batch for work that can wait. Batch jobs do not consume your synchronous per-minute limits. Batch has its own queue limit, counted in queued input tokens per model. For what that does to the bill, see GPT Image 2 Price Per Image: Official API Costs and Formula.

Spreading requests out. If a job needs 200 images and your tier allows 20 per minute, a queue that releases 20 per minute finishes in about 200 ÷ 20 = 10 minutes without a single 429, provided the token limit is not reached first. Sending all 200 at once produces refusals that then count against the next minute.

Opening a second account or rotating keys to get past a limit is not an option. OpenAI's terms prohibit circumventing rate limits, and limits are counted per organization and project in any case.

Azure OpenAI uses a different table

Azure does not reuse OpenAI's tier numbers. Microsoft counts gpt-image-2 in requests per minute (RPM), not images, and sets its own quota tiers. The figures below come from Microsoft's quotas and limits page, dated August 20, 2026, as of October 1, 2026.

Azure quota tiergpt-image-2 Global Standardgpt-image-2 Data Zone Standardgpt-image-2.5-flare and gpt-image-2.5-sunburst Global Standard
Tier 16 RPM2 RPM5 RPM
Tier 212 RPM4 RPM5 RPM
Tier 318 RPM6 RPM5 RPM
Tier 424 RPM8 RPM5 RPM
Tier 530 RPM10 RPM5 RPM

Azure has seven quota tiers in total (Free and 1 through 6), and tiers rise automatically with usage. Extra quota goes through Microsoft's quota request form. Do not count on extra regions for extra allowance. Microsoft began managing quota at the subscription level after May 7, 2026, starting with a few models and, in its words, soon all models. Under that scheme, Global Standard deployments of one model and version share a single quota pool across all regions in a subscription, and Data Zone Standard deployments share one pool per data zone.

A 429 from a gateway or reseller is a different limit

If your requests go to a third-party endpoint and not to api.openai.com, the 429 comes from that provider's own capacity. A gateway that serves its customers from one shared gpt-image-2 pool returns 429s or timeouts to all of them when that pool is busy. Your OpenAI usage tier plays no part, and the codes in the table above may not appear at all. The provider's own notices and documentation are the places to check.

The same separation applies to per-image services. YingTu's GPT Image 2 page lists the model at $0.03 per image as of October 1, 2026, billed per image through your own LaoZhang API key. That is a separate quota system from your OpenAI organization. It suits someone who wants a handful of images without building up an OpenAI tier, and it does not raise your OpenAI limits. YingTu publishes no images-per-minute figure.

FAQ

Why am I getting a 429 when my monthly budget is not used up?

Because the per-minute rate limit and the monthly limits are counted separately. You can have most of the month's budget left and still send a sixth image inside one minute on Tier 1. If the response is "Rate limit reached for requests" or slow_down, pacing fixes it. If the code is one of the spend or usage codes, check whether a project-level spend limit is lower than the organization budget you were looking at.

Do GPT Image 2 API limits reset every day?

No. The limits published for gpt-image-2 are per minute, so the allowance comes back within a minute once you stop sending. The approved usage limit is monthly. The x-ratelimit-reset-requests and x-ratelimit-reset-tokens headers show the time until the current window resets.

Is there a GPT Image 2 API tier without a per-minute limit?

No. Every paid tier has both a TPM and an IPM limit, and the highest published figure is 250 images per minute on Tier 5 as of October 1, 2026.

Does a ChatGPT Plus subscription raise my API limits?

No. ChatGPT Plus costs $20 per month as of September 30, 2026, and OpenAI states that API usage is separate and billed independently. Your API tier depends only on what your organization has paid for API usage.

Tags

#GPT Image 2#OpenAI API#Rate Limits#429 Error#Usage Tiers

Share this article

XTelegram