Plain-text specification
Plain-text specification
GPT-6 Luna API
GPT-6 Luna is a large language model by OpenAI, available on the Venice API asopenai-gpt-6-luna. Requests are anonymized, so the provider never sees your identity.GPT-6 Luna is OpenAI’s most efficient GPT-6 model for focused, high-volume tasks. It is suited for chat, classification, and lightweight agentic workflows, with a 1.05M token context window (922K input, 128K output), text and image inputs, and capable reasoning for its price tier.GPT-6 Luna API pricing
GPT-6 Luna specifications
How to use the GPT-6 Luna API
Send requests toPOST https://api.venice.ai/api/v1/responses with "model": "openai-gpt-6-luna" and your API key.GPT-6 Luna API FAQ
How much does the GPT-6 Luna API cost?
0.63 per 1M output tokens, with cached input at $0.013 per 1M. Prices are in USD and can be paid in DIEM at parity.What is the GPT-6 Luna model ID?
Useopenai-gpt-6-luna as the model parameter.Is the GPT-6 Luna API private?
GPT-6 Luna is anonymized: Venice forwards requests to the provider without your identity, but the provider may retain prompt data, so use a private model for sensitive work.What is the context window of GPT-6 Luna?
1.05M tokens of context, with up to 125K output tokens per response.What does GPT-6 Luna support?
GPT-6 Luna supports function calling, structured outputs, reasoning, image input, web search and prompt caching. Reasoning effort is adjustable withreasoning_effort: none, low, medium, high, xhigh and max (default high).Which endpoint does the GPT-6 Luna API use?
CallPOST /responses. /chat/completions is also supported. OpenAI reasoning models are designed around the Responses API: reasoning, tool calls and messages come back as typed output items. Venice’s /responses endpoint is in alpha.Related models
- GPT-5.6 Luna API: 1.50 output per 1M tokens
- GLM 5.3 Flash API: 0.50 output per 1M tokens
- Qwen 3.8 Flash API: 0.49 output per 1M tokens
- MiMo-V2.6-Flash API: 0.35 output per 1M tokens
- DeepSeek V4 Flash 0731 API: 0.35 output per 1M tokens