Verbatik Launches First MCP Text-to-Speech Server: 2,700+ Neural Voices Now a Native Capability for Claude, Codex, and Other AI Assistants
Technology📅 July 16, 2026👤 FreeReadText Team

Verbatik Launches First MCP Text-to-Speech Server: 2,700+ Neural Voices Now a Native Capability for Claude, Codex, and Other AI Assistants

London-based Verbatik Technologies releases the first text-to-speech server for the Model Context Protocol (MCP), connecting 2,700+ neural voices across 50+ languages directly to AI assistants — turning voice generation into a native agent capability rather than an external API call.

On July 16, 2026, London-based Verbatik Technologies Limited launched a text-to-speech server for the Model Context Protocol (MCP), the open standard developed by Anthropic for connecting AI assistants to external tools. The server exposes nine MCP-accessible tools that let compatible AI assistants — including Claude Desktop, Claude Code, Cursor, Windsurf, Kiro, Codex, OpenAI, and Lovable — generate, clone, and customize speech within a single workflow without leaving the agent environment. Co-founder Lucian Deleu framed the launch as making 'voice a native capability of the agent itself.'

The server provides access to 2,700+ neural voices spanning more than 50 languages, with text-to-speech generation supporting up to 50,000 characters per call and SSML (Speech Synthesis Markup Language) for fine-grained control. Voice cloning is available from a 10-second audio sample, with seven adjustable emotional settings and support for natural interjections such as laughs, sighs, and custom pauses. Pricing is granular: pre-trained voice TTS costs $0.002 per 1,000 characters, cloned-voice TTS costs $0.10 per 1,000 characters, and each voice clone costs $3.00. A built-in cost estimation tool lets agents preview expenses before processing large jobs — for example, narrating a batch of 50 articles.

The MCP ecosystem has grown rapidly since Anthropic open-sourced the protocol in late 2024, with servers emerging for databases, file systems, web search, and code execution. Verbatik's server is the first dedicated TTS integration for the protocol, opening up multi-step workflows where an AI assistant can read a document, select an appropriate voice, generate narration, and return audio — all within one continuous agent session. The server also supports multilingual pipelines where scripts are first translated and then synthesized in multiple languages, collapsing what previously required multiple services and manual coordination into a single agent-driven workflow.

Setup is designed for speed: the company claims connection takes under one minute, with free initial credits and no credit card required. Authentication uses per-client API keys and OAuth 2.1 with PKCE and token rotation over Streamable HTTP. The launch arrives at a moment when voice is rapidly becoming a primary AI interface — OpenAI's GPT-Live, xAI's Grok Voice, and Google's Gemini voice capabilities all launched or received major updates within the same two-week window. Verbatik's bet is that as AI assistants evolve from text-based chatbots to multimodal agents, voice generation will become as essential a tool as web search or code execution — and that MCP-native integration, rather than bolt-on API calls, is the right architecture for that future.

VerbatikMCPModel Context ProtocolText-to-SpeechAI AssistantsVoice CloningClaude

출처

← Back to News