Skip to content
GitHub

Multimedia Tts (lexigram-multimedia-tts)

Text-to-speech generation for the Lexigram Framework — local and API-based backends (local-http, elevenlabs, openai, chatterbox, kokoro, f5-tts, piper).


lexigram-multimedia-tts synthesizes speech from text. The default backend calls a local HTTP reference server (http://localhost:5002) so the package works out of the box with no API keys; hosted backends (ElevenLabs, OpenAI) and in-process reference servers (Chatterbox, Kokoro, F5-TTS, Piper) are selectable via config.

Full documentation: docs.lexigram.dev

Terminal window
uv add lexigram-multimedia-tts
# Optional extras
uv add "lexigram-multimedia-tts[elevenlabs]" # ElevenLabs API
uv add "lexigram-multimedia-tts[openai]" # OpenAI TTS API
uv add "lexigram-multimedia-tts[chatterbox-server]" # local Chatterbox server (torch)
uv add "lexigram-multimedia-tts[kokoro-server]" # local Kokoro server
uv add "lexigram-multimedia-tts[f5-tts-server]" # local F5-TTS server (torch)
uv add "lexigram-multimedia-tts[piper-server]" # local Piper server
from lexigram import Application
from lexigram.di.module import Module, module
from lexigram.multimedia.tts import AudioTTSModule
from lexigram.contracts.multimedia import TTSProvider, TTSRequest
@module(imports=[AudioTTSModule.configure()])
class AppModule(Module):
pass
async def main() -> None:
async with Application.boot(modules=[AppModule]) as app:
tts = await app.container.resolve(TTSProvider)
result = await tts.generate(TTSRequest(text="Hello from Lexigram", voice="alloy"))
if result.is_ok():
asset = result.unwrap() # MediaAsset — audio bytes or URI
if __name__ == "__main__":
import asyncio
asyncio.run(main())

Zero-config usage: Call AudioTTSModule.configure() with no arguments to use the local-http backend at http://localhost:5002.

application.yaml
multimedia:
tts:
backend: "elevenlabs"
elevenlabs_voice_id: "21m00Tcm4TlvDq8ikWAM"

Option 2 — Profiles + Environment Variables

Section titled “Option 2 — Profiles + Environment Variables”
Terminal window
export LEX_PROFILE=production
export LEX_MULTIMEDIA__TTS__BACKEND=openai
export LEX_MULTIMEDIA__TTS__OPENAI_VOICE=echo
from lexigram.multimedia.tts import AudioTTSModule
from lexigram.multimedia.tts.config import TTSConfig
AudioTTSModule.configure(
config=TTSConfig(backend="elevenlabs", elevenlabs_voice_id="21m00Tcm4TlvDq8ikWAM")
)
FieldDefaultEnv varDescription
backend"local-http"LEX_MULTIMEDIA__TTS__BACKENDlocal-http, elevenlabs, openai, chatterbox, kokoro, f5-tts, piper
local_http_base_url"http://localhost:5002"LEX_MULTIMEDIA__TTS__LOCAL_HTTP_BASE_URLLocal reference server URL
elevenlabs_voice_idNoneLEX_MULTIMEDIA__TTS__ELEVENLABS_VOICE_IDElevenLabs voice ID (required for elevenlabs)
elevenlabs_api_key_secret_name"elevenlabs_api_key"LEX_MULTIMEDIA__TTS__ELEVENLABS_API_KEY_SECRET_NAMESecret name for the ElevenLabs API key
openai_api_key_secret_name"openai_api_key"LEX_MULTIMEDIA__TTS__OPENAI_API_KEY_SECRET_NAMESecret name for the OpenAI API key
openai_voice"alloy"LEX_MULTIMEDIA__TTS__OPENAI_VOICEOpenAI voice (alloy, echo, fable, onyx, nova, shimmer)
openai_model"tts-1"LEX_MULTIMEDIA__TTS__OPENAI_MODELOpenAI TTS model
openai_base_url"https://api.openai.com"LEX_MULTIMEDIA__TTS__OPENAI_BASE_URLOpenAI-compatible base URL
chatterbox_base_url"http://localhost:5100"LEX_MULTIMEDIA__TTS__CHATTERBOX_BASE_URLChatterbox server URL
chatterbox_exaggeration0.5LEX_MULTIMEDIA__TTS__CHATTERBOX_EXAGGERATIONChatterbox exaggeration factor
chatterbox_cfg_weight0.5LEX_MULTIMEDIA__TTS__CHATTERBOX_CFG_WEIGHTChatterbox classifier-free guidance weight
chatterbox_temperature0.85LEX_MULTIMEDIA__TTS__CHATTERBOX_TEMPERATUREChatterbox sampling temperature
kokoro_base_url"http://localhost:5101"LEX_MULTIMEDIA__TTS__KOKORO_BASE_URLKokoro server URL
kokoro_default_voice"af_heart"LEX_MULTIMEDIA__TTS__KOKORO_DEFAULT_VOICEDefault Kokoro voice
f5_tts_base_url"http://localhost:5102"LEX_MULTIMEDIA__TTS__F5_TTS_BASE_URLF5-TTS server URL
piper_base_url"http://localhost:5103"LEX_MULTIMEDIA__TTS__PIPER_BASE_URLPiper server URL
piper_default_voice"en_US-lessac-medium"LEX_MULTIMEDIA__TTS__PIPER_DEFAULT_VOICEDefault Piper voice
timeout60.0LEX_MULTIMEDIA__TTS__TIMEOUTRequest timeout in seconds
MethodDescription
AudioTTSModule.configure(config)Configure with explicit TTS config
AudioTTSModule.stub()Real module pinned to the default local-http backend for tests
  • Seven backendslocal-http, elevenlabs, openai, chatterbox, kokoro, f5-tts, piper
  • Reference serverslexigram-tts-*-serve console scripts run each local model server
  • Secret-managed API keys — provider keys resolved by name through the secrets backend
  • Result-basedgenerate() -> Result[MediaAsset, MultimediaError]; errors are domain values, not exceptions
from lexigram import Application
from lexigram.multimedia.tts import AudioTTSModule
async def test_boot():
async with Application.boot(modules=[AudioTTSModule.stub()]) as app:
assert app.container is not None
FileWhat it contains
src/lexigram/multimedia/tts/module.pyAudioTTSModule.configure() and .stub()
src/lexigram/multimedia/tts/config.pyTTSConfig
src/lexigram/multimedia/tts/di/provider.pyAudioTTSProvider — registers TTSProvider, wires task handlers
src/lexigram/multimedia/tts/providers/Backend implementations (local_http, elevenlabs, openai, chatterbox, kokoro, f5_tts, piper)
src/lexigram/multimedia/tts/servers/Reference-server entry points (lexigram-tts-*-serve)
src/lexigram/multimedia/tts/tasks.pyBackground generation task handlers
src/lexigram/multimedia/tts/exceptions.pyTTSError hierarchy