Plain-text specification
Plain-text specification
GLM 5.3 Flash API
GLM 5.3 Flash is a large language model by Z.ai, available on the Venice API asz-ai-glm-5-3-flash. Private and end-to-end encrypted variants are available.GLM-5.3 Flash is a reasoning model designed for coding, sustained agentic work, and production workloads. It is suited for long-horizon software engineering, complex reasoning, and workflows that combine text with visual context.GLM 5.3 Flash API pricing
GLM 5.3 Flash specifications
How to use the GLM 5.3 Flash API
Send requests toPOST https://api.venice.ai/api/v1/chat/completions with "model": "z-ai-glm-5-3-flash" and your API key.GLM 5.3 Flash API FAQ
How much does the GLM 5.3 Flash API cost?
0.50 per 1M output tokens, with cached input at 0.16 input and $0.54 output. Prices are in USD and can be paid in DIEM at parity.What is the GLM 5.3 Flash model ID?
Usez-ai-glm-5-3-flash as the model parameter. Other variants: e2ee-glm-5-3-flash (E2EE).Is the GLM 5.3 Flash API private?
The Standard variant is private: requests run on infrastructure Venice controls with zero data retention, and prompts and outputs are never stored or used for training. The E2EE variant is end-to-end encrypted: prompts are encrypted on your device and decrypted only inside an attested hardware enclave, so neither Venice nor the GPU provider can read them.What is the context window of GLM 5.3 Flash?
1M tokens of context, with up to 128K output tokens per response.What does GLM 5.3 Flash support?
GLM 5.3 Flash supports function calling, structured outputs, reasoning, image input, web search and prompt caching. Reasoning effort is adjustable withreasoning_effort: none, low, medium and high (default high).Which endpoint does the GLM 5.3 Flash API use?
CallPOST /chat/completions. /responses (Alpha) is also supported.Related models
- GLM 4.7 Flash API: 0.40 output per 1M tokens
- Qwen 3.8 Flash API: 0.49 output per 1M tokens
- GPT-6 Luna API: 0.63 output per 1M tokens
- DeepSeek V4 Flash 0731 API: 0.35 output per 1M tokens
- MiMo-V2.6-Flash API: 0.35 output per 1M tokens