AI Gateway REST API Reference
AI Gateway REST API Reference
For sending inference requests, the AI SDK lets you interact with AI Gateway from TypeScript or Python. You can also send requests through the chat completions, responses, Anthropic Messages, or OpenResponses APIs.
AI Gateway exposes a REST API for looking up usage and generations, querying spend reports, and discovering models. This page is the canonical reference for those endpoints. The AI SDK AI Gateway provider also exposes TypeScript APIs for the same data, which you can use as an alternative to calling the REST endpoints directly.
Base URL
https://ai-gateway.vercel.sh/v1
Authentication
Most endpoints require authentication. The exceptions are noted on each endpoint.
API key (Bearer token)
Pass an AI Gateway API key in the Authorization header. AI Gateway infers the team from the key.
HTTP header
Authorization: Bearer <api_key>
Vercel OIDC token
On Vercel deployments, you can pass the Vercel OIDC token in the Authorization header instead. The token is generated automatically per project.
Supported endpoints
| Endpoint | Description |
|---|---|
GET /v1/models |
List all available models |
GET /v1/models/{creator}/{model}/endpoints |
Get every provider endpoint serving a model |
GET /v1/credits |
Check the team's AI Gateway Credits balance and total spend |
GET /v1/generation |
Look up cost, latency, and token usage for a specific generation |
GET /v1/report |
Query aggregated spend reports for the team |
Models
List models
Endpoint
GET /v1/models
Lists every model available through AI Gateway, including reasoning controls, supported parameters, pricing, and modalities. Follows the OpenAI models API format. No authentication required.
For the AI SDK equivalent, see Dynamic Model Discovery.
Example request
list-models.ts
const response = await fetch('https://ai-gateway.vercel.sh/v1/models');
const { data: models }: { data: Array<{ id: string; name: string }> } =
await response.json();
models.forEach((model) => {
console.log(`${model.id}: ${model.name}`);
});
For TypeScript, Python, and cURL examples that filter reasoning models and inspect their controls, see Discover model reasoning support.
Sample response
Response
{
"object": "list",
"data": [
{
"id": "google/gemini-3.1-pro-preview",
"object": "model",
"created": 1755815280,
"released": 1763424000,
"owned_by": "google",
"name": "Gemini 3.1 Pro Preview",
"description": "This model improves upon Gemini 2.5 Pro and is catered towards challenging tasks, especially those involving complex reasoning or agentic workflows.",
"context_window": 1000000,
"max_tokens": 64000,
"type": "language",
"tags": ["file-input", "tool-use", "reasoning", "vision"],
"pricing": {
"input": "0.000002",
"output": "0.000012",
"input_cache_read": "0.0000002",
"input_cache_write": "0.000002"
}
}
]
}
Response fields
| Field | Type | Description |
|---|---|---|
object |
string | Always "list" |
data |
array | Array of available models |
data[].id |
string | Model identifier (for example, openai/gpt-6-astra) |
data[].object |
string | Always "model" |
data[].created |
integer | Unix timestamp when the model was added |
data[].released |
integer | Unix timestamp when the model was released |
data[].owned_by |
string | Model provider or owner |
data[].name |
string | Human-readable model name |
data[].description |
string | Model description |
data[].context_window |
integer | Maximum context length in tokens |
data[].max_tokens |
integer | Maximum output tokens |
data[].type |
string | Model type: language, embedding, reranking, image, or video |
data[].tags |
string[] | Capability tags (for example, reasoning, tool-use, vision) |
data[].pricing |
object | Pricing information (structure varies by model type) |
data[].pricing.input |
string | Base cost per input token |
data[].pricing.output |
string | Base cost per output token (language models only) |
data[].pricing.input_cache_read |
string | Cost per cached input token (read) |
data[].pricing.input_cache_write |
string | Cost per input token (cache write) |
The catalog doesn't expose a structured reasoning default. Missing controls or budget bounds mean unspecified metadata, not unsupported reasoning or an unlimited budget. The controls describe model capabilities; each request format has its own fields and accepted values.
Get model endpoints
Endpoint
GET /v1/models/{creator}/{model}/endpoints
Returns every provider endpoint serving a specific model, along with per-endpoint pricing, capabilities, and supported parameters. No authentication required.
Example request
get-model-endpoints.ts
const response = await fetch(
'https://ai-gateway.vercel.sh/v1/models/google/gemini-3.1-pro-preview/endpoints',
);
const {
data,
}: {
data: {
name: string;
endpoints: Array<{ provider_name: string; context_length: number }>;
};
} = await response.json();
console.log(`Model: ${data.name}`);
data.endpoints.forEach((endpoint) => {
console.log(` ${endpoint.provider_name}: ${endpoint.context_length} tokens`);
});
Sample response
Response
{
"data": {
"id": "google/gemini-3.1-pro-preview",
"name": "Gemini 3.1 Pro Preview",
"created": 1755815280,
"released": 1763424000,
"description": "This model improves upon Gemini 2.5 Pro and is catered towards challenging tasks, especially those involving complex reasoning or agentic workflows.",
"architecture": {
"tokenizer": null,
"instruct_type": null,
"modality": "text+image+file→text",
"input_modalities": ["text", "image", "file"],
"output_modalities": ["text"]
},
"endpoints": [
{
"name": "google | google/gemini-3.1-pro-preview",
"model_name": "Gemini 3.1 Pro Preview",
"context_length": 1000000,
"pricing": {
"prompt": "0.000002",
"completion": "0.000012",
"input_cache_read": "0.0000002",
"input_cache_write": "0.000002"
},
"provider_name": "google",
"max_completion_tokens": 64000,
"supported_parameters": ["max_tokens", "temperature", "tools", "reasoning"],
"status": 0,
"uptime_last_15m": 100,
"uptime_last_1h": 99.8,
"uptime_last_1d": 99.6,
"throughput_last_1h": { "p50": 67, "p95": 69.85 },
"latency_last_1h": { "p50": 2292, "p95": 2685 },
"supports_implicit_caching": false
}
]
}
}
Response fields
| Field | Type | Description |
|---|---|---|
data.id |
string | Model identifier |
data.name |
string | Human-readable model name |
data.created |
integer | Unix timestamp when the model was added |
data.released |
integer | Unix timestamp when the model was released |
data.description |
string | Model description |
data.architecture |
object | Model architecture details |
data.architecture.modality |
string | Input/output modality string |
data.architecture.input_modalities |
string[] | Supported input types |
data.architecture.output_modalities |
string[] | Supported output types |
data.endpoints |
array | Array of provider endpoints |
data.endpoints[].provider_name |
string | Provider name (for example, google, anthropic) |
data.endpoints[].context_length |
integer | Maximum context window in tokens |
data.endpoints[].max_completion_tokens |
integer | Maximum output tokens |
data.endpoints[].pricing.prompt |
string | Cost per prompt token |
data.endpoints[].pricing.completion |
string | Cost per completion token |
data.endpoints[].pricing.input_cache_read |
string | Cost per cached input token (read) |
data.endpoints[].pricing.input_cache_write |
string | Cost per input token (cache write) |
data.endpoints[].supported_parameters |
string[] | API parameters supported by this endpoint |
data.endpoints[].supports_implicit_caching |
boolean | Whether the provider supports automatic caching |
data.endpoints[].status |
integer | Endpoint status: 0 is active |
data.endpoints[].uptime_last_15m |
number | Uptime percentage over the last 15 minutes |
data.endpoints[].uptime_last_1h |
number | Uptime percentage over the last hour |
data.endpoints[].uptime_last_1d |
number | Uptime percentage over the last day |
data.endpoints[].throughput_last_1h |
object | p50 and p95 throughput (tokens/sec) over the last hour |
data.endpoints[].latency_last_1h |
object | p50 and p95 time to first token (ms) over the last hour |
For more on the uptime and metrics fields, see Uptime and Status and Metrics.
Tiered pricing
Some models have tiered pricing based on context size. When tiered pricing applies, the *_tiers arrays contain pricing tiers:
| Field | Type | Description |
|---|---|---|
cost |
string | Cost per token for this tier |
min |
number | Minimum token count (inclusive) |
max |
number | Maximum token count (exclusive), omitted for the highest tier |
Usage and billing
Check credit balance
Endpoint
GET /v1/credits
Returns the team's current AI Gateway Credits balance and lifetime spend.
Example request
credits.ts
const apiKey = process.env.AI_GATEWAY_API_KEY || process.env.VERCEL_OIDC_TOKEN;
const response = await fetch('https://ai-gateway.vercel.sh/v1/credits', {
method: 'GET',
headers: {
Authorization: `Bearer ${apiKey}`,
'Content-Type': 'application/json',
},
});
const credits = await response.json();
console.log(credits);
Sample response
Response
{
"balance": "95.50",
"total_used": "4.50"
}
Response fields
balance: Remaining credit balance, in USDtotal_used: Total credits used to date, in USD
Look up a generation
Endpoint
GET /v1/generation?id={generation_id}
Returns detailed information about a specific generation, including cost, latency, and token usage. Much of this data is also returned in providerMetadata on the chat completion response.
Parameters
id(required): The generation ID to look up. Format:gen_<ulid>.
Example request
generation-lookup.ts
const generationId = 'gen_01ARZ3NDEKTSV4RRFFQ69G5FAV';
const response = await fetch(
`https://ai-gateway.vercel.sh/v1/generation?id=${generationId}`,
{
method: 'GET',
headers: {
Authorization: `Bearer ${process.env.AI_GATEWAY_API_KEY}`,
'Content-Type': 'application/json',
},
},
);
const generation = await response.json();
console.log(generation);
Sample response
Response
{
"data": {
"id": "gen_01ARZ3NDEKTSV4RRFFQ69G5FAV",
"total_cost": 0.00123,
"market_cost": 0.0011,
"surcharge_cost": 0.00013,
"gateway_cost": 0.00123,
"upstream_inference_cost": 0,
"usage": 0.00123,
"created_at": "2026-05-22T00:00:00.000Z",
"model": "anthropic/claude-opus-5",
"is_byok": false,
"provider_name": "anthropic",
"streamed": true,
"finish_reason": "stop",
"latency": 200,
"generation_time": 1500,
"tokens_prompt": 100,
"tokens_completion": 50,
"native_tokens_prompt": 100,
"native_tokens_completion": 50,
"native_tokens_reasoning": 0,
"native_tokens_cached": 0,
"native_tokens_cache_creation": 0,
"billable_web_search_calls": 0
}
}
Response fields
id: The generation IDtotal_cost: Total cost in USD for this generation debited from your gateway balance, including any surcharges (for example, Zero Data Retention or Custom Reporting writes). Does not include cost of BYOK requests.market_cost: Cost of this generation at market list inference rates. Omitted when not recorded for the generation.surcharge_cost: Total surcharges applied to this generation (for example, Zero Data Retention or Custom Reporting writes)gateway_cost: Total amount debited from your AI Gateway balance for this generation. Same astotal_cost.created_at: ISO 8601 timestamp when the generation was createdmodel: Model identifier used for this generationprovider_name: The provider that served this generationlatency: Time to first token in millisecondstokens_prompt: Number of prompt tokenstokens_completion: Number of completion tokens
Custom reporting
Use the Custom Reporting API to break down spend and usage by model, user, tag, provider, or credential type. For concepts, how to attach tags and user IDs to requests, querying with the AI SDK, and the full response field reference, see the Custom Reporting page.
Query spend report
Endpoint
GET /v1/report
Returns aggregated spend over a date range. The team is inferred from the API key or OIDC token. Hobby and Pro-trial plans cannot use this endpoint.
Example request
terminal
curl "https://ai-gateway.vercel.sh/v1/report?start_date=2026-01-01&end_date=2026-01-31&group_by=model" \
-H "Authorization: Bearer $AI_GATEWAY_API_KEY"
Error responses
Errors return a JSON body with an error field. Some endpoints additionally include a type discriminator on the error object.
Error response
{
"error": {
"message": "Invalid request body",
"type": "invalid_request_error"
}
}
| Status | Meaning |
| --- | --- |
| `400` | Invalid request body, missing required fields, or validation error. |
| `401` | Authentication failed (invalid API key or OIDC token). |
| `403` | Endpoint requires a paid plan (for example, Hobby and Pro-trial cannot query `/v1/report`). |
| `404` | Resource not found (for example, a generation that doesn't exist). |
| `500` | Internal server error. |
| `503` | A backing service is misconfigured or unavailable.