> ## Documentation Index
> Fetch the complete documentation index at: https://veniceai-feat-models-redesign.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# GLM 5.1 API

> GLM 5.1 API on Venice: 200K context, $1.54 input and $4.84 output per 1M tokens. Private, with zero data retention. Supports reasoning and function calling.

export const HubMount = ({view, children, ...props}) => {
  const [hub, setHub] = useState(null);
  const [failed, setFailed] = useState(false);
  useEffect(() => {
    let alive = true;
    const w = window;
    if (!w.__veniceModelHub) {
      const urls = w.location.hostname === 'localhost' ? ['http://localhost:3333/data/model-hub.bundle.json', '/data/model-hub.bundle.json'] : ['/data/model-hub.bundle.json'];
      const load = i => fetch(urls[i], {
        cache: 'no-cache'
      }).then(res => {
        if (!res.ok) throw new Error(`bundle ${res.status}`);
        return res.json();
      }).catch(err => i + 1 < urls.length ? load(i + 1) : Promise.reject(err));
      const Frag = <></>.type;
      const h = (type, props, ...kids) => {
        const T = type;
        const {key, ...rest} = props || ({});
        if (!kids.length) return <T key={key} {...rest} />;
        if (kids.length === 1) return <T key={key} {...rest}>{kids[0]}</T>;
        return <T key={key} {...rest}>{kids.map((kid, i) => <Frag key={i}>{kid}</Frag>)}</T>;
      };
      w.__veniceModelHub = load(0).then(bundle => new Function(`return (${bundle.code})`)()({
        h,
        Fragment: Frag,
        useState,
        useEffect,
        useRef,
        useMemo,
        useCallback
      }));
    }
    w.__veniceModelHub.then(instance => {
      if (alive) setHub(instance);
    }).catch(() => {
      w.__veniceModelHub = null;
      if (alive) setFailed(true);
    });
    return () => {
      alive = false;
    };
  }, []);
  const View = hub ? hub[view] : null;
  if (View) return <View {...props}>{children}</View>;
  if (failed) {
    return <div className="vx-mount is-failed">
        <p className="vx-mount-note">The interactive model catalog could not load. The full data is below.</p>
        {children}
      </div>;
  }
  return <div className="vx-mount" aria-busy="true">
      <div className="vx-mount-skeleton" aria-hidden="true"><span /><span /><span /></div>
      <div className="vx-mount-source">{children}</div>
    </div>;
};

<HubMount view="ModelPage" data={{"family":{"slug":"glm-5-1","name":"GLM 5.1","modality":"text","task":"chat","provider":"zai","description":"GLM-5.1 is the next-generation large language model developed by Zhiyuan AI, featuring significantly enhanced reasoning capabilities, improved instruction following, and support for multiple languages. Supports large context windows for processing extensive text and detailed analysis with fast inference speed.","primary":"zai-org-glm-5-1","variants":["zai-org-glm-5-1"],"created":1775520000,"updated":1775520000,"privacy":["private"],"openWeights":true},"models":[{"id":"zai-org-glm-5-1","name":"GLM 5.1","type":"text","modality":"text","task":"chat","variant":"standard","provider":"zai","created":1775520000,"description":"GLM-5.1 is the next-generation large language model developed by Zhiyuan AI, featuring significantly enhanced reasoning capabilities, improved instruction following, and support for multiple languages. Supports large context windows for processing extensive text and detailed analysis with fast inference speed.","source":"https://huggingface.co/zai-org/GLM-5.1","privacy":"private","openWeights":true,"text":{"context":200000,"maxOutput":80000,"quantization":"fp8","reasoning":{"supported":true,"effort":["none","low","medium","high"],"defaultEffort":"low"},"caps":{"tools":true,"structured":true,"webSearch":true}},"pricing":{"input":1.54,"output":4.84,"cacheRead":0.286,"blended":2.365},"headline":{"value":2.365,"unit":"per 1M tokens","basis":"blended"},"endpoints":[{"id":"chat","method":"POST","path":"/chat/completions","name":"Chat Completions","status":"stable","recommended":true},{"id":"responses","method":"POST","path":"/responses","name":"Responses","status":"alpha"}],"family":"glm-5-1"}],"related":{"similar":[{"slug":"gpt-5-4-mini","name":"GPT-5.4 Mini","provider":"openai","modality":"text","privacy":["anonymized"],"created":1774569600,"headline":{"value":2.109375,"unit":"per 1M tokens","basis":"blended"},"variants":1},{"slug":"deepseek-v4-pro","name":"DeepSeek V4 Pro","provider":"deepseek","modality":"text","privacy":["private"],"created":1776988800,"headline":{"value":2.06275,"unit":"per 1M tokens","basis":"blended"},"variants":1},{"slug":"inkling","name":"Inkling","provider":"inkling","modality":"text","privacy":["private"],"created":1784160000,"headline":{"value":2.203125,"unit":"per 1M tokens","basis":"blended"},"variants":1},{"slug":"deepseek-v4-pro-0813","name":"DeepSeek V4 Pro 0813","provider":"deepseek","modality":"text","privacy":["private"],"created":1786665600,"headline":{"value":2.475,"unit":"per 1M tokens","basis":"blended"},"variants":1}],"versions":[{"slug":"glm-5-3","name":"GLM 5.3","provider":"zai","modality":"text","privacy":["private","e2ee"],"created":1787011200,"headline":{"value":2.6875,"unit":"per 1M tokens","basis":"blended"},"variants":2},{"slug":"glm-5-2","name":"GLM 5.2","provider":"zai","modality":"text","privacy":["private","e2ee"],"created":1781568000,"headline":{"value":2.15,"unit":"per 1M tokens","basis":"blended"},"variants":2},{"slug":"glm-5","name":"GLM 5","provider":"zai","modality":"text","privacy":["private"],"created":1770768000,"headline":{"value":1.55,"unit":"per 1M tokens","basis":"blended"},"variants":1},{"slug":"glm-4-7","name":"GLM 4.7","provider":"zai","modality":"text","privacy":["private"],"created":1766534400,"headline":{"value":1.075,"unit":"per 1M tokens","basis":"blended"},"variants":1},{"slug":"glm-4-6","name":"GLM 4.6","provider":"zai","modality":"text","privacy":["private"],"created":1711929600,"headline":{"value":0.76,"unit":"per 1M tokens","basis":"blended"},"variants":1}]},"providers":{"zai":{"slug":"zai","name":"Z.ai","logo":"/images/icons/models/Zhipu.svg"},"openai":{"slug":"openai","name":"OpenAI","logo":"/images/icons/models/openai.svg"},"deepseek":{"slug":"deepseek","name":"DeepSeek","logo":"/images/icons/models/deepseek.svg"},"inkling":{"slug":"inkling","name":"Inkling","logo":"/images/icons/models/text.svg","inferred":true}},"faq":[{"q":"How much does the GLM 5.1 API cost?","a":"$1.54 per 1M input tokens and $4.84 per 1M output tokens, with cached input at $0.29 per 1M. Prices are in USD and can be paid in DIEM at parity."},{"q":"What is the GLM 5.1 model ID?","a":"Use `zai-org-glm-5-1` as the `model` parameter."},{"q":"Is the GLM 5.1 API private?","a":"GLM 5.1 is private: requests run on infrastructure Venice controls with zero data retention, and prompts and outputs are never stored or used for training."},{"q":"What is the context window of GLM 5.1?","a":"200K tokens of context, with up to 80K output tokens per response."},{"q":"What does GLM 5.1 support?","a":"GLM 5.1 supports function calling, structured outputs, reasoning, web search and prompt caching. Reasoning effort is adjustable with `reasoning_effort`: none, low, medium and high (default low)."},{"q":"Which endpoint does the GLM 5.1 API use?","a":"Call `POST /chat/completions`. `/responses` (Alpha) is also supported."}]}} />

<div className="vx-static">
  <Accordion title="Plain-text specification">
    # GLM 5.1 API

    GLM 5.1 is a large language model by Z.ai, available on the Venice API as `zai-org-glm-5-1`. It runs privately, with zero data retention.

    GLM-5.1 is the next-generation large language model developed by Zhiyuan AI, featuring significantly enhanced reasoning capabilities, improved instruction following, and support for multiple languages. Supports large context windows for processing extensive text and detailed analysis with fast inference speed.

    ## GLM 5.1 API pricing

    | Model ID | Variant | Privacy | Price |
    | - | - | - | - |
    | `zai-org-glm-5-1` | Standard | Private | $1.54 input / $4.84 output per 1M tokens |

    ## GLM 5.1 specifications

    | Spec | Value |
    | - | - |
    | Provider | Z.ai |
    | Released | Apr 7, 2026 |
    | Privacy | Private |
    | Open weights | Yes |
    | Context window | 200K tokens |
    | Max output | 80K tokens |
    | Reasoning effort | none, low, medium, high |
    | Served precision | FP8 |

    ## How to use the GLM 5.1 API

    Send requests to `POST https://api.venice.ai/api/v1/chat/completions` with `"model": "zai-org-glm-5-1"` and your API key.

    ```bash theme={null}
    curl https://api.venice.ai/api/v1/chat/completions \
      -H "Authorization: Bearer $VENICE_API_KEY" \
      -H "Content-Type: application/json" \
      -d '{
        "model": "zai-org-glm-5-1",
        "messages": [{ "role": "user", "content": "Explain TEE attestation in two sentences." }],
        "reasoning_effort": "low"
      }'
    ```

    ## GLM 5.1 API FAQ

    ### How much does the GLM 5.1 API cost?

    $1.54 per 1M input tokens and $4.84 per 1M output tokens, with cached input at \$0.29 per 1M. Prices are in USD and can be paid in DIEM at parity.

    ### What is the GLM 5.1 model ID?

    Use `zai-org-glm-5-1` as the `model` parameter.

    ### Is the GLM 5.1 API private?

    GLM 5.1 is private: requests run on infrastructure Venice controls with zero data retention, and prompts and outputs are never stored or used for training.

    ### What is the context window of GLM 5.1?

    200K tokens of context, with up to 80K output tokens per response.

    ### What does GLM 5.1 support?

    GLM 5.1 supports function calling, structured outputs, reasoning, web search and prompt caching. Reasoning effort is adjustable with `reasoning_effort`: none, low, medium and high (default low).

    ### Which endpoint does the GLM 5.1 API use?

    Call `POST /chat/completions`. `/responses` (Alpha) is also supported.

    ## Related models

    * [GLM 5.3 API](/models/glm-5-3): $1.75 input / $5.50 output per 1M tokens
    * [GLM 5.2 API](/models/glm-5-2): $1.40 input / $4.40 output per 1M tokens
    * [GLM 5 API](/models/glm-5): $1.00 input / $3.20 output per 1M tokens
    * [GLM 4.7 API](/models/glm-4-7): $0.55 input / $2.65 output per 1M tokens
    * [GLM 4.6 API](/models/glm-4-6): $0.43 input / $1.75 output per 1M tokens
    * [GPT-5.4 Mini API](/models/gpt-5-4-mini): $0.94 input / $5.63 output per 1M tokens
    * [DeepSeek V4 Pro API](/models/deepseek-v4-pro): $1.65 input / $3.30 output per 1M tokens
    * [Inkling API](/models/inkling): $1.25 input / $5.06 output per 1M tokens
    * [DeepSeek V4 Pro 0813 API](/models/deepseek-v4-pro-0813): $1.65 input / $4.95 output per 1M tokens
  </Accordion>
</div>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.