Plain-text specification
Plain-text specification
Kimi K2.5 API
Kimi K2.5 is a large language model by Moonshot AI, available on the Venice API askimi-k2-5. It runs privately, with zero data retention.Kimi K2.5 is Moonshot AIs most advanced open reasoning model, featuring trillion-parameter Mixture-of-Experts architecture with 32B active parameters and 256K context windows.Kimi K2.5 API pricing
Kimi K2.5 specifications
How to use the Kimi K2.5 API
Send requests toPOST https://api.venice.ai/api/v1/chat/completions with "model": "kimi-k2-5" and your API key.Kimi K2.5 API FAQ
How much does the Kimi K2.5 API cost?
3.50 per 1M output tokens, with cached input at $0.22 per 1M. Prices are in USD and can be paid in DIEM at parity.What is the Kimi K2.5 model ID?
Usekimi-k2-5 as the model parameter.Is the Kimi K2.5 API private?
Kimi K2.5 is private: requests run on infrastructure Venice controls with zero data retention, and prompts and outputs are never stored or used for training.What is the context window of Kimi K2.5?
250K tokens of context, with up to 64K output tokens per response.What does Kimi K2.5 support?
Kimi K2.5 supports function calling, structured outputs, reasoning, image input, web search and prompt caching.Which endpoint does the Kimi K2.5 API use?
CallPOST /chat/completions. /responses (Alpha) is also supported.Related models
- Kimi K3 API: 18.75 output per 1M tokens
- Kimi K2.6 API: 3.50 output per 1M tokens
- Qwen 3.6 Plus Uncensored API: 3.75 output per 1M tokens
- Gemini 3 Flash Preview API: 3.75 output per 1M tokens
- Grok Build 0.1 API: 2.00 output per 1M tokens
- GLM 5 API: 3.20 output per 1M tokens