Skip to main content
The official Ruby SDK for the Typecast API. Convert text to lifelike speech using AI-powered voices, generate timestamps, list voices, and create custom voices. The Ruby SDK uses only the Ruby standard library at runtime and supports Ruby 2.6+.

RubyGems

Typecast Ruby SDK

Source Code

Typecast Ruby SDK Source Code

Installation

Install from RubyGems:
Latest registered version: 0.1.6 on RubyGems.
Or add it to your Gemfile:
Requires Ruby 2.6 or higher. Check your version with ruby -v.

Quick Start

Features

  • Multiple Voice Models: Support for ssfm-v30 and ssfm-v21 AI voice models
  • Multi-language Support: 37 languages including English, Korean, Japanese, Chinese, Spanish, and more
  • Emotion Control: Preset emotions or smart context-aware inference
  • Audio Customization: Control loudness, pitch, tempo, and output format
  • Voice Discovery: V2 Voices API with filtering by model, gender, age, and use cases
  • Streaming Endpoint: Access streaming TTS responses from Ruby
  • Timestamp TTS: Word- and character-level alignment data with SRT/VTT helpers
  • Instant Voice Cloning: Upload a WAV sample and create a custom voice ID
  • No Runtime Dependencies: Built on Ruby standard library net/http

Voice Recommendations

Use recommend_voices when you know the desired style but not the exact voice_id.
Recommendation results contain only voice_id, voice_name, and score. Use get_voice_v2 or get_voices_v2 when you need detailed metadata such as supported models, emotions, gender, age, or use cases.

Configuration

Set your API key via environment variable or constructor:
When requests go through your own proxy, set base_url to the proxy endpoint and omit api_key. The SDK will not send the X-API-KEY header for empty or missing keys. Requests to the default Typecast host still require an API key.
Proxy without API key
You can also override the API host and HTTP timeouts:

Advanced Usage

Emotion Control (ssfm-v30)

ssfm-v30 offers two emotion control modes: Preset and Smart.
Let the AI infer emotion from context:

Audio Customization

Control loudness, pitch, tempo, and output format:

Generate audio to a file

Use generate_to_file when you want the SDK to synthesize speech and write the audio bytes directly to a local file. The model defaults to ssfm-v30, and .mp3 / .wav extensions infer the output format when no output format is set. Browse available voice IDs on the Voices page.

Text pauses

Use text pause markup when you only need silent gaps inside one composed text segment. Put <|5s|>, <|1s|>, <|0.3s|>, or <|0.34413s|> directly in the text. The value is interpreted as seconds and must end with s. This keeps the pause expression visible in plain text without adding separate pause calls.

Multi-speaker composition

Use the composer chaining API when one output file needs different voices or per-segment options such as pitch, tempo, prompt, or seed. The composer generates each segment as WAV, trims leading/trailing silent PCM samples, and concatenates the result. If you need MP3, generate WAV first and convert it in your app or server pipeline.

Voice Discovery (V2 API)

List and filter available voices with enhanced metadata:

Streaming

Use text_to_speech_stream() to call the streaming endpoint:
WAV streaming format: 32000 Hz, 16-bit, mono PCM. The first chunk includes a 44-byte WAV header (size = 0xFFFFFFFF); subsequent chunks are raw PCM only.

Timestamp TTS

text_to_speech_with_timestamps() wraps POST /v1/text-to-speech/with-timestamps and returns audio together with word- or character-level alignment data.

Granularity

Pass granularity: "word" (default) or granularity: "char" to control the alignment unit.

Subtitle Export

Instant Voice Cloning

Upload a short WAV sample to create a custom voice:
Voice cloning audio must be 25 MB or smaller, and the custom voice name must be 1-30 characters.