OpenAI Speech-to-Text
Provides speech-to-text transcription capabilities using OpenAI's Whisper API with configurable language settings and optional file saving
Directory
Provides speech-to-text transcription capabilities using OpenAI's Whisper API with configurable language settings and optional file saving
frontend for generic MCP server based chatbot
Access 47 AI models and 30 services via MCP. LLM, image gen, video, speech, embeddings, reranking, PDF parsing & more. Pay-per-use with x402 (USDC) or API key.
Hosted MCP server for generating AI images, videos, speech, and music from agents and IDEs. Supports OAuth and Clipia API keys.
A JavaScript/TypeScript server for MiniMax MCP, offering image/video generation, text-to-speech, and voice cloning.
Give your AI agents the ability to listen. Microphone capture and speech-to-text tools for MCP-compatible agents.
30+ pay-per-use API tools for AI agents via MCP: image generation (Flux 1.1 Pro), web scraping, browser automation (Playwright), TTS, PDF, QR code, weather…
Text-to-Speech protocol server that synthesizes text from LLMs and plays audio natively through the host system's desk speakers.
Quick example of building a speaker agent with Google ADK and ElevenLabs' MCP server
An MCP server for GPT-SoVITS, providing text-to-speech synthesis, voice cloning, and multi-language support.
Model Context Protocol (MCP) server for Kokoro text-to-speech with female voice. 100% local, no Python required. Supports SSE and stdio transports.
Add one endpoint and Claude, Cursor, or any MCP client can make video, images, music, and speech in chat. 100+ models, one API key.
Remote MCP for AI video, image, music and speech generation in Claude, Cursor and ChatGPT.
Integrates with ClickSend's API to enable sending SMS messages and initiating Text-to-Speech calls for automated communication workflows.
Provides speech-to-text, diarization, translation, and text summarization via the Whissle AI API.
Hosted MCP server for AudioPod's audio AI: text-to-speech, voice cloning, music generation, stem and speaker separation, transcription, denoise, and voice…
Integrates with ElevenLabs to provide high-quality text-to-speech, voice cloning, and conversational capabilities with customizable voice profiles and audio…
Access Whissle API for speech-to-text, diarization, translation, and text summarization.
Connects AI systems to VOICEVOX text-to-speech engine for Japanese voice synthesis, supporting both default transport and Server-Sent Events with configurable…
One API key gives agents access to 80+ tools: web search, deep search, browser automation, screenshots, 400+ LLM models, image generation, text-to-speech…
Provides a full suite of AI tools via DeepInfra’s OpenAI-compatible API, including image generation, text processing, embeddings, and speech recognition.
Generate images, video, and audio directly in Claude Code, Cursor, Windsurf, or any MCP-compatible AI agent.
MCP server for AI image generation, video generation, music creation, text-to-speech, and LLM chat. 130+ models from Flux, Kling, Seedance, Veo, Suno…
A Chat Client for LLMs, written in Compose Multiplatform.