Skip to main content

GLM 5 Turbo API

GLM 5 Turbo is a large language model by Z.ai, available on the Venice API as z-ai-glm-5-turbo. Requests are anonymized, so the provider never sees your identity.GLM-5 Turbo is a fast inference model from Z.ai tuned for strong performance in agent-driven environments and production coding workflows.

GLM 5 Turbo API pricing

GLM 5 Turbo specifications

How to use the GLM 5 Turbo API

Send requests to POST https://api.venice.ai/api/v1/chat/completions with "model": "z-ai-glm-5-turbo" and your API key.

GLM 5 Turbo API FAQ

How much does the GLM 5 Turbo API cost?

1.20per1Minputtokensand1.20 per 1M input tokens and 4.00 per 1M output tokens, with cached input at $0.24 per 1M. Prices are in USD and can be paid in DIEM at parity.

What is the GLM 5 Turbo model ID?

Use z-ai-glm-5-turbo as the model parameter.

Is the GLM 5 Turbo API private?

GLM 5 Turbo is anonymized: Venice forwards requests to the provider without your identity, but the provider may retain prompt data, so use a private model for sensitive work.

What is the context window of GLM 5 Turbo?

200K tokens of context, with up to 32K output tokens per response.

What does GLM 5 Turbo support?

GLM 5 Turbo supports function calling, structured outputs, reasoning, web search and prompt caching. Reasoning effort is adjustable with reasoning_effort: none, low, medium and high (default high).

Which endpoint does the GLM 5 Turbo API use?

Call POST /chat/completions. /responses (Alpha) is also supported.

Related models