> ## Documentation Index
> Fetch the complete documentation index at: https://veniceai-feat-models-redesign.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Kokoro Text to Speech API

> Kokoro Text to Speech API on Venice: text to speech with 54 voices at $3.50 per 1M characters (about $0.0029 per minute). Private, with zero data retention.

export const HubMount = ({view, children, ...props}) => {
  const [hub, setHub] = useState(null);
  const [failed, setFailed] = useState(false);
  useEffect(() => {
    let alive = true;
    const w = window;
    if (!w.__veniceModelHub) {
      const urls = w.location.hostname === 'localhost' ? ['http://localhost:3333/data/model-hub.bundle.json', '/data/model-hub.bundle.json'] : ['/data/model-hub.bundle.json'];
      const load = i => fetch(urls[i], {
        cache: 'no-cache'
      }).then(res => {
        if (!res.ok) throw new Error(`bundle ${res.status}`);
        return res.json();
      }).catch(err => i + 1 < urls.length ? load(i + 1) : Promise.reject(err));
      const Frag = <></>.type;
      const h = (type, props, ...kids) => {
        const T = type;
        const {key, ...rest} = props || ({});
        if (!kids.length) return <T key={key} {...rest} />;
        if (kids.length === 1) return <T key={key} {...rest}>{kids[0]}</T>;
        return <T key={key} {...rest}>{kids.map((kid, i) => <Frag key={i}>{kid}</Frag>)}</T>;
      };
      w.__veniceModelHub = load(0).then(bundle => new Function(`return (${bundle.code})`)()({
        h,
        Fragment: Frag,
        useState,
        useEffect,
        useRef,
        useMemo,
        useCallback
      }));
    }
    w.__veniceModelHub.then(instance => {
      if (alive) setHub(instance);
    }).catch(() => {
      w.__veniceModelHub = null;
      if (alive) setFailed(true);
    });
    return () => {
      alive = false;
    };
  }, []);
  const View = hub ? hub[view] : null;
  if (View) return <View {...props}>{children}</View>;
  if (failed) {
    return <div className="vx-mount is-failed">
        <p className="vx-mount-note">The interactive model catalog could not load. The full data is below.</p>
        {children}
      </div>;
  }
  return <div className="vx-mount" aria-busy="true">
      <div className="vx-mount-skeleton" aria-hidden="true"><span /><span /><span /></div>
      <div className="vx-mount-source">{children}</div>
    </div>;
};

<HubMount view="ModelPage" data={{"family":{"slug":"kokoro-text-to-speech","name":"Kokoro Text to Speech","modality":"audio","task":"tts","provider":"hexgrad","primary":"tts-kokoro","variants":["tts-kokoro"],"created":1742418046,"updated":1742418046,"privacy":["private"],"openWeights":true},"models":[{"id":"tts-kokoro","name":"Kokoro Text to Speech","type":"tts","modality":"audio","task":"tts","variant":"standard","provider":"hexgrad","created":1742418046,"source":"https://huggingface.co/hexgrad/Kokoro-82M","privacy":"private","openWeights":true,"audio":{"voices":["af_alloy","af_aoede","af_bella","af_heart","af_jadzia","af_jessica","af_kore","af_nicole","af_nova","af_river","af_sarah","af_sky","am_adam","am_echo","am_eric","am_fenrir","am_liam","am_michael","am_onyx","am_puck","am_santa","bf_alice","bf_emma","bf_lily","bm_daniel","bm_fable","bm_george","bm_lewis","ef_dora","em_alex","em_santa","ff_siwis","hf_alpha","hf_beta","hm_omega","hm_psi","if_sara","im_nicola","jf_alpha","jf_gongitsune","jf_nezumi","jf_tebukuro","jm_kumo","pf_dora","pm_alex","pm_santa","zf_xiaobei","zf_xiaoni","zf_xiaoxiao","zf_xiaoyi","zm_yunjian","zm_yunxi","zm_yunxia","zm_yunyang"],"formats":["mp3","opus","aac","flac","wav","pcm"],"defaultFormat":"mp3"},"pricing":{"per1MChars":3.5,"perMinute":0.002888,"perHour":0.17328},"headline":{"value":3.5,"unit":"per 1M characters","basis":"characters"},"endpoints":[{"id":"speech","method":"POST","path":"/audio/speech","name":"Create speech","status":"stable","recommended":true}],"family":"kokoro-text-to-speech"}],"related":{"similar":[{"slug":"inworld-tts-1-5-max","name":"Inworld TTS-1.5 Max","provider":"inworld","modality":"audio","privacy":["anonymized"],"created":1776384000,"headline":{"value":12.5,"unit":"per 1M characters","basis":"characters"},"variants":1},{"slug":"xai-tts-v1","name":"xAI TTS v1","provider":"xai","modality":"audio","privacy":["anonymized"],"created":1776384000,"headline":{"value":18.75,"unit":"per 1M characters","basis":"characters"},"variants":1},{"slug":"chatterbox-hd","name":"Chatterbox HD (Resemble AI)","provider":"resemble","modality":"audio","privacy":["private"],"created":1776384000,"headline":{"value":50,"unit":"per 1M characters","basis":"characters"},"variants":1},{"slug":"gradium-tts","name":"Gradium TTS","provider":"gradium","modality":"audio","privacy":["anonymized"],"created":1780617600,"headline":{"value":47.5,"unit":"per 1M characters","basis":"characters"},"variants":1}],"versions":[]},"providers":{"hexgrad":{"slug":"hexgrad","name":"hexgrad","logo":"/images/icons/models/music.svg"},"inworld":{"slug":"inworld","name":"Inworld","logo":"/images/icons/models/music.svg"},"xai":{"slug":"xai","name":"xAI","logo":"/images/icons/models/grok.svg"},"resemble":{"slug":"resemble","name":"Resemble AI","logo":"/images/icons/models/music.svg"},"gradium":{"slug":"gradium","name":"Gradium","logo":"/images/icons/models/music.svg"}},"faq":[{"q":"How much does the Kokoro Text to Speech API cost?","a":"$3.50 per 1M characters of input text, about $0.0029 per minute of generated speech. Prices are in USD and can be paid in DIEM at parity."},{"q":"What is the Kokoro Text to Speech model ID?","a":"Use `tts-kokoro` as the `model` parameter."},{"q":"Is the Kokoro Text to Speech API private?","a":"Kokoro Text to Speech is private: requests run on infrastructure Venice controls with zero data retention, and prompts and outputs are never stored or used for training."},{"q":"How many voices does Kokoro Text to Speech have?","a":"54 voices, for example `af_alloy`, `af_aoede`, `af_bella`. Output formats: mp3, opus, aac, flac, wav, pcm."},{"q":"Which endpoint does the Kokoro Text to Speech API use?","a":"Call `POST /audio/speech`."}]}} />

<div className="vx-static">
  <Accordion title="Plain-text specification">
    # Kokoro Text to Speech API

    Kokoro Text to Speech is a text-to-speech model by hexgrad, available on the Venice API as `tts-kokoro`. It runs privately, with zero data retention.

    ## Kokoro Text to Speech API pricing

    | Model ID | Variant | Privacy | Price |
    | - | - | - | - |
    | `tts-kokoro` | Standard | Private | \$3.50 per 1M characters |

    ## Kokoro Text to Speech specifications

    | Spec | Value |
    | - | - |
    | Provider | hexgrad |
    | Released | Mar 19, 2025 |
    | Privacy | Private |
    | Open weights | Yes |
    | Voices | 54 |
    | Formats | mp3, opus, aac, flac, wav, pcm |

    ## How to use the Kokoro Text to Speech API

    Send requests to `POST https://api.venice.ai/api/v1/audio/speech` with `"model": "tts-kokoro"` and your API key.

    ```bash theme={null}
    curl https://api.venice.ai/api/v1/audio/speech \
      -H "Authorization: Bearer $VENICE_API_KEY" \
      -H "Content-Type: application/json" \
      -d '{
        "model": "tts-kokoro",
        "input": "Private AI for everyone.",
        "voice": "af_alloy",
        "response_format": "mp3"
      }' --output speech.mp3
    ```

    ## Kokoro Text to Speech API FAQ

    ### How much does the Kokoro Text to Speech API cost?

    $3.50 per 1M characters of input text, about $0.0029 per minute of generated speech. Prices are in USD and can be paid in DIEM at parity.

    ### What is the Kokoro Text to Speech model ID?

    Use `tts-kokoro` as the `model` parameter.

    ### Is the Kokoro Text to Speech API private?

    Kokoro Text to Speech is private: requests run on infrastructure Venice controls with zero data retention, and prompts and outputs are never stored or used for training.

    ### How many voices does Kokoro Text to Speech have?

    54 voices, for example `af_alloy`, `af_aoede`, `af_bella`. Output formats: mp3, opus, aac, flac, wav, pcm.

    ### Which endpoint does the Kokoro Text to Speech API use?

    Call `POST /audio/speech`.

    ## Related models

    * [Inworld TTS-1.5 Max API](/models/inworld-tts-1-5-max): \$12.50 per 1M characters
    * [xAI TTS v1 API](/models/xai-tts-v1): \$18.75 per 1M characters
    * [Chatterbox HD (Resemble AI) API](/models/chatterbox-hd): \$50.00 per 1M characters
    * [Gradium TTS API](/models/gradium-tts): \$47.50 per 1M characters
  </Accordion>
</div>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.