Skip to main content

Qwen 3.8 Max API

Qwen 3.8 Max is a large language model by Alibaba Qwen, available on the Venice API as qwen-3-8-max. Requests are anonymized, so the provider never sees your identity.Qwen 3.8 Max is Alibaba’s flagship 2.4-trillion-parameter MoE model, with major gains over Qwen 3.7 Max in software engineering and office-productivity workflows and strong long-horizon, multi-agent performance. It accepts both text and vision-language input (images and video), operates in thinking mode only, and supports a 1M-token context window.

Qwen 3.8 Max API pricing

Qwen 3.8 Max specifications

How to use the Qwen 3.8 Max API

Send requests to POST https://api.venice.ai/api/v1/chat/completions with "model": "qwen-3-8-max" and your API key.

Qwen 3.8 Max API FAQ

How much does the Qwen 3.8 Max API cost?

2.50per1Minputtokensand2.50 per 1M input tokens and 7.50 per 1M output tokens, with cached input at $0.31 per 1M. Prices are in USD and can be paid in DIEM at parity.

What is the Qwen 3.8 Max model ID?

Use qwen-3-8-max as the model parameter.

Is the Qwen 3.8 Max API private?

Qwen 3.8 Max is anonymized: Venice forwards requests to the provider without your identity, but the provider may retain prompt data, so use a private model for sensitive work.

What is the context window of Qwen 3.8 Max?

1M tokens of context, with up to 128K output tokens per response.

What does Qwen 3.8 Max support?

Qwen 3.8 Max supports function calling, reasoning, image input, web search and prompt caching.

Which endpoint does the Qwen 3.8 Max API use?

Call POST /chat/completions. /responses (Alpha) is also supported.

Related models

  • Qwen 3.7 Max API: 2.70input/2.70 input / 8.05 output per 1M tokens
  • Grok 4.5 API: 2.27input/2.27 input / 6.80 output per 1M tokens
  • Grok 4.6 API: 2.27input/2.27 input / 6.80 output per 1M tokens
  • Gemini 3.5 Flash API: 1.55input/1.55 input / 9.45 output per 1M tokens
  • Grok 4.7 API: 2.27input/2.27 input / 6.80 output per 1M tokens