أخبار الصناعة

ابقَ على اطلاع بآخر التطورات في تقنية الصوت بالذكاء الاصطناعي وتركيب الكلام والمشهد التنظيمي المتطور

OpenAI Pauses Astra Development Over Critical Cyber Risk Threshold
Regulation

OpenAI Pauses Astra Development Over Critical Cyber Risk Threshold

OpenAI has halted internal development of its Astra model after evaluations indicated potential capabilities for autonomous zero-day exploit development, marking the first time a model has triggered the 'Critical' threshold under the company's Preparedness Framework.

👤 FreeReadText Team📅 August 8, 2026
OpenAIAstraAI SafetyCybersecurityPreparedness Framework
اقرأ المزيد
ElevenLabs Releases v3: Its Most Expressive AI Voice Model with Audio Tags and Text-to-Dialogue
Technology

ElevenLabs Releases v3: Its Most Expressive AI Voice Model with Audio Tags and Text-to-Dialogue

ElevenLabs has officially launched Eleven v3, expanding support to over 70 languages and introducing inline audio tags for precise emotional control, alongside a new Text-to-Dialogue feature for seamless multi-speaker generation.

👤 FreeReadText Team📅 August 8, 2026
ElevenLabsEleven v3Text-to-SpeechAI AudioText-to-DialogueVoice Synthesis
اقرأ المزيد
Microsoft Unveils MAI-Voice-1: Hyper-Realistic Speech Generation from Just One Minute of Audio
Technology

Microsoft Unveils MAI-Voice-1: Hyper-Realistic Speech Generation from Just One Minute of Audio

Microsoft launches three new foundational AI models including MAI-Voice-1, which delivers hyper-realistic voice synthesis and custom brand voice creation, marking a major leap in enterprise TTS capabilities.

👤 FreeReadText Team📅 April 2, 2026
MicrosoftMAI-Voice-1Enterprise TTSVoice SynthesisFoundry
اقرأ المزيد
ElevenLabs Reaches $11 Billion Valuation, Eyes IPO as Voice AI Becomes Enterprise Standard
Business

ElevenLabs Reaches $11 Billion Valuation, Eyes IPO as Voice AI Becomes Enterprise Standard

AI voice startup ElevenLabs raises $500 million at an $11 billion valuation, tripling its worth in just over a year while forging major partnerships with IBM and planning a potential IPO.

👤 FreeReadText Team📅 February 4, 2026
ElevenLabsFundingIPOIBM PartnershipEnterprise AI
اقرأ المزيد
Global AI Voice Regulation Tightens: EU AI Act Deepfake Rules Take Effect as Voice Cloning Crosses 'Indistinguishable Threshold'
Regulation

Global AI Voice Regulation Tightens: EU AI Act Deepfake Rules Take Effect as Voice Cloning Crosses 'Indistinguishable Threshold'

As voice cloning technology reaches human-level quality, regulators worldwide respond with new laws — the EU AI Act's deepfake labeling rules, the US ELVIS Act, and emerging biometric voice data protections reshape the industry landscape.

👤 FreeReadText Team📅 March 15, 2026
EU AI ActELVIS ActVoice CloningDeepfakeBiometric DataCompliance
اقرأ المزيد
OpenAI Launches Voice Engine to the Public: Real-Time Conversational TTS Now Available to All Developers
Technology

OpenAI Launches Voice Engine to the Public: Real-Time Conversational TTS Now Available to All Developers

After over a year of limited preview, OpenAI opens Voice Engine to all API developers, introducing real-time streaming TTS with emotional awareness and 40+ language support at significantly reduced pricing.

👤 FreeReadText Team📅 March 28, 2026
OpenAIVoice EngineReal-Time TTSAPIConversational AI
اقرأ المزيد
Google DeepMind Brings Studio-Quality TTS to Smartphones with SoundStorm 2 Edge — No Internet Required
Technology

Google DeepMind Brings Studio-Quality TTS to Smartphones with SoundStorm 2 Edge — No Internet Required

Google DeepMind announces SoundStorm 2 Edge, a compact on-device TTS model that runs entirely on mobile hardware, delivering studio-quality voice synthesis without cloud connectivity and opening new possibilities for offline accessibility.

👤 FreeReadText Team📅 March 20, 2026
Google DeepMindSoundStorm 2On-Device AIMobile TTSAccessibility
اقرأ المزيد
AI Dubbing Market Surges Past $2 Billion as Hollywood, Streaming Giants, and Game Studios Embrace Automated Localization
Business

AI Dubbing Market Surges Past $2 Billion as Hollywood, Streaming Giants, and Game Studios Embrace Automated Localization

The AI-powered dubbing and localization market crosses the $2 billion mark in Q1 2026, driven by adoption from Netflix, Disney+, and major game publishers seeking to reach global audiences at a fraction of traditional costs.

👤 FreeReadText Team📅 April 5, 2026
AI DubbingLocalizationNetflixGamingVoice ActingStreaming
اقرأ المزيد
Apple Unveils 'Personal Voice 2.0' in iOS 20: On-Device Voice Cloning Creates Your Digital Twin in 3 Minutes
Technology

Apple Unveils 'Personal Voice 2.0' in iOS 20: On-Device Voice Cloning Creates Your Digital Twin in 3 Minutes

Apple announces Personal Voice 2.0 at its spring event, allowing users to create a highly realistic clone of their own voice in just 3 minutes of recording — all processed entirely on-device with Apple Silicon, positioning it as the privacy-first alternative to cloud-based voice AI.

👤 FreeReadText Team📅 April 8, 2026
AppleiOS 20Personal VoiceOn-Device AIPrivacyVoice Cloning
اقرأ المزيد
Spotify Rolls Out AI Voice Translation for Podcasts Globally: Your Favorite Hosts Now Speak 40 Languages in Their Own Voice
Business

Spotify Rolls Out AI Voice Translation for Podcasts Globally: Your Favorite Hosts Now Speak 40 Languages in Their Own Voice

Spotify launches its AI-powered podcast translation feature worldwide, using voice cloning technology to automatically dub podcasts into 40 languages while preserving each host's unique voice characteristics — opening 100,000+ shows to global audiences overnight.

👤 FreeReadText Team📅 April 7, 2026
SpotifyPodcast TranslationVoice CloningLocalizationStreaming AudioCreator Economy
اقرأ المزيد
FDA Clears First AI Voice Assistant for Clinical Use: Voice-Based Patient Screening Enters the Hospital
Technology

FDA Clears First AI Voice Assistant for Clinical Use: Voice-Based Patient Screening Enters the Hospital

The FDA grants its first clearance for an AI voice assistant designed for clinical patient interaction, allowing automated voice-based symptom screening and triage in emergency departments — marking a historic milestone for voice AI in healthcare.

👤 FreeReadText Team📅 April 10, 2026
Healthcare AIFDAVoice AssistantClinical AIPatient ScreeningHippocratic AI
اقرأ المزيد
Meta Releases Llama-Voice: First Fully Open-Source TTS Model to Match Commercial Giants in 50+ Languages
Technology

Meta Releases Llama-Voice: First Fully Open-Source TTS Model to Match Commercial Giants in 50+ Languages

Meta drops Llama-Voice under an Apache 2.0 license, delivering near state-of-the-art voice synthesis, zero-shot voice cloning from 10 seconds of audio, and 52-language coverage — all runnable on a single consumer GPU.

👤 FreeReadText Team📅 April 12, 2026
MetaLlama-VoiceOpen SourceMultilingual TTSVoice CloningHugging Face
اقرأ المزيد
NVIDIA Launches Voice Foundry NIM: Blackwell-Optimized Microservices Cut Real-Time TTS Costs by 70%
Technology

NVIDIA Launches Voice Foundry NIM: Blackwell-Optimized Microservices Cut Real-Time TTS Costs by 70%

NVIDIA unveils Voice Foundry, a dedicated suite of NIM inference microservices for TTS and STT optimized for Blackwell GB200 hardware, promising sub-80ms first-token latency and 70% lower per-character costs for enterprise voice applications.

👤 FreeReadText Team📅 April 15, 2026
NVIDIAVoice FoundryNIMBlackwellEnterprise InfrastructureTensorRT
اقرأ المزيد
Audible Opens AI-Narrated Audiobook Catalog to 400,000 Backlist Titles — Narrators Split on Landmark Royalty Model
Business

Audible Opens AI-Narrated Audiobook Catalog to 400,000 Backlist Titles — Narrators Split on Landmark Royalty Model

Amazon's Audible launches the industry's largest AI-narrated audiobook catalog, adding 400,000 previously unnarrated titles using voice clones of consenting narrators, with a first-of-its-kind per-listen residual model that splits the narration community.

👤 FreeReadText Team📅 April 17, 2026
AudibleAmazonAI NarrationAudiobooksVoice ActingRoyaltiesSAG-AFTRA
اقرأ المزيد
Google Launches Gemini 3.1 Flash TTS: 70+ Languages, Multi-Speaker Dialogue, and a Top Spot on the Artificial Analysis Leaderboard
Technology

Google Launches Gemini 3.1 Flash TTS: 70+ Languages, Multi-Speaker Dialogue, and a Top Spot on the Artificial Analysis Leaderboard

Google introduces Gemini 3.1 Flash TTS, a new text-to-speech model with audio tags for fine-grained vocal control, native multi-speaker dialogue, and 70+ language support — landing in the 'most attractive quadrant' of the Artificial Analysis TTS leaderboard with an Elo of 1,211.

👤 FreeReadText Team📅 April 15, 2026
GoogleGeminiFlash TTSMulti-SpeakerSynthIDVertex AI
اقرأ المزيد
OpenAI Launches GPT-Realtime-2: Voice Models with GPT-5-Class Reasoning, Live Translation, and Streaming Transcription
Technology

OpenAI Launches GPT-Realtime-2: Voice Models with GPT-5-Class Reasoning, Live Translation, and Streaming Transcription

OpenAI introduces three new Realtime API voice models — GPT-Realtime-2 with GPT-5-class reasoning, GPT-Realtime-Translate covering 70+ input languages, and GPT-Realtime-Whisper for live transcription — quadrupling the context window to 128K tokens and bringing voice agents closer to production-ready workflows.

👤 FreeReadText Team📅 May 7, 2026
OpenAIGPT-Realtime-2Realtime APILive TranslationSpeech-to-TextVoice Agents
اقرأ المزيد
Microsoft Launches MAI-Voice-2 at Build 2026: Expressive Speech and Zero-Shot Voice Cloning Across 15 Languages
Technology

Microsoft Launches MAI-Voice-2 at Build 2026: Expressive Speech and Zero-Shot Voice Cloning Across 15 Languages

Microsoft unveils MAI-Voice-2, calling it the most expressive and natural-sounding text-to-speech model it has built, expanding from English-only to 15 languages with granular emotion control, code-switching, and zero-shot voice prompting from a few seconds of audio.

👤 FreeReadText Team📅 June 2, 2026
MicrosoftMAI-Voice-2Multilingual TTSVoice CloningBuild 2026Azure
اقرأ المزيد
Wispr Hits ~$2 Billion Valuation as AI Voice Dictation Becomes a Workplace Standard
Business

Wispr Hits ~$2 Billion Valuation as AI Voice Dictation Becomes a Workplace Standard

Wispr, the startup behind the AI dictation tool Wispr Flow, is raising roughly $260 million at a near-$2 billion valuation led by Menlo Ventures — nearly tripling its worth in six months as voice-to-text moves from novelty to everyday workplace productivity tool.

👤 FreeReadText Team📅 May 13, 2026
WisprWispr FlowFundingMenlo VenturesVoice DictationSpeech-to-Text
اقرأ المزيد
FTC Begins Enforcing the TAKE IT DOWN Act: Platforms Face $53,088-Per-Violation Penalties for AI Deepfakes
Regulation

FTC Begins Enforcing the TAKE IT DOWN Act: Platforms Face $53,088-Per-Violation Penalties for AI Deepfakes

The FTC's civil enforcement of the TAKE IT DOWN Act took effect on May 19, 2026, requiring platforms to remove nonconsensual intimate imagery — including AI-generated deepfakes — within 48 hours, with penalties of $53,088 per violation. The agency promptly sent warning letters to major platforms and 'nudify' websites.

👤 FreeReadText Team📅 May 19, 2026
TAKE IT DOWN ActFTCDeepfakeVoice CloningSynthetic MediaCompliance
اقرأ المزيد
Poland Government Takes Stake in ElevenLabs, Launches AI Lab to Build Voice AI from Europe
Business

Poland Government Takes Stake in ElevenLabs, Launches AI Lab to Build Voice AI from Europe

The Government of Poland invests in ElevenLabs through its Vinci/BGK Group, joining Andreessen Horowitz and Sequoia as a strategic backer, while launching AI Lab Poland to nurture the next generation of voice AI companies with global ambition.

👤 FreeReadText Team📅 June 18, 2026
ElevenLabsPolandGovernment InvestmentAI LabEuropeVoice AI
اقرأ المزيد
ElevenLabs Launches Dubbing v2: Emotion-Preserving AI Dubbing Across 90+ Languages
Technology

ElevenLabs Launches Dubbing v2: Emotion-Preserving AI Dubbing Across 90+ Languages

ElevenLabs releases Dubbing v2, a breakthrough AI dubbing model that preserves the original speaker's emotion, tone, and pacing across 90+ languages by conditioning directly on the performance rather than just transcripts.

👤 FreeReadText Team📅 May 28, 2026
ElevenLabsDubbing v2AI DubbingTranslationLocalizationAudio AI
اقرأ المزيد
ElevenLabs Partners with UK Government to Bring Voice AI to Public Services, Doubles London Headquarters
Business

ElevenLabs Partners with UK Government to Bring Voice AI to Public Services, Doubles London Headquarters

ElevenLabs signs a Memorandum of Understanding with the UK's Department for Science, Innovation and Technology to deploy voice AI in public services, focusing on accessibility for the visually impaired, elderly, and linguistically diverse communities.

👤 FreeReadText Team📅 June 8, 2026
ElevenLabsUK GovernmentPublic ServicesAccessibilityAI PolicyDSIT
اقرأ المزيد
Rumik Launches Silk Mulberry 1.5: 'Describe a Voice Into Existence' with Plain-Language Prompts, Matching Commercial TTS Giants at 95% Lower Cost
Technology

Rumik Launches Silk Mulberry 1.5: 'Describe a Voice Into Existence' with Plain-Language Prompts, Matching Commercial TTS Giants at 95% Lower Cost

Indian AI startup Rumik releases Silk Mulberry 1.5, a text-to-speech model that replaces preset voice menus with plain-language voice descriptions, achieving MOS scores competitive with ElevenLabs and Google at roughly $0.0046 per minute.

👤 FreeReadText Team📅 June 19, 2026
RumikSilk MulberryIndian AIText-to-SpeechVoice DesignCode-Switching
اقرأ المزيد
Michael Caine's AI Voice Narrates 13-Hour 'The Odyssey' Audiobook — 20 AI Characters, Original Score, Built by 4 Producers in 6 Weeks
Business

Michael Caine's AI Voice Narrates 13-Hour 'The Odyssey' Audiobook — 20 AI Characters, Original Score, Built by 4 Producers in 6 Weeks

ElevenLabs releases a cinematic audiobook of Homer's The Odyssey narrated by an authorized AI replica of Sir Michael Caine's voice, featuring ~20 AI-generated character voices, original music, and sound design — all produced by a four-person team in six weeks.

👤 FreeReadText Team📅 June 23, 2026
ElevenLabsMichael CaineAI AudiobookVoice CloningIconic MarketplaceHollywood
اقرأ المزيد
Five9 Launches Voice AI Agents with ElevenLabs, Deepgram, and OpenAI Under the Hood — Targeting Legacy IVR Replacement
Technology

Five9 Launches Voice AI Agents with ElevenLabs, Deepgram, and OpenAI Under the Hood — Targeting Legacy IVR Replacement

Five9 unveils Voice AI Agents at Customer Contact Week 2026, combining ElevenLabs TTS, Deepgram ASR, and OpenAI reasoning in a proprietary three-model architecture built to replace scripted IVR systems with natural, human-like voice self-service.

👤 FreeReadText Team📅 June 23, 2026
Five9Voice AI AgentsContact CenterAgentic AIEnterpriseCustomer Experience
اقرأ المزيد
xAI Launches Voice Agent Builder: No-Code Platform Harnesses Grok Voice to Beat GPT and Gemini in Telephony Benchmarks
Technology

xAI Launches Voice Agent Builder: No-Code Platform Harnesses Grok Voice to Beat GPT and Gemini in Telephony Benchmarks

Elon Musk's xAI enters the voice AI market with Voice Agent Builder, a no-code platform powered by Grok Voice Think Fast 1.0 that scores 67.3% on the τ-voice Bench — far outpacing Google Gemini 3.1 Flash Live (43.8%) and OpenAI GPT Realtime 1.5 (35.3%) — with pricing starting at $0.05 per minute.

👤 FreeReadText Team📅 July 1, 2026
xAIGrok VoiceVoice Agent BuilderNo-CodeSpeech-to-SpeechCall Center AI
اقرأ المزيد
Bland.ai Raises $50M Series C After 180 Investor Rejections, Now Powers 3.5 Million Voice Calls Per Week
Business

Bland.ai Raises $50M Series C After 180 Investor Rejections, Now Powers 3.5 Million Voice Calls Per Week

San Francisco voice AI startup Bland.ai closes a $50 million Series C led by Dell Technologies Capital, bringing total funding past $100 million — after founders were rejected by 180 investors who told them 'phone calls won't exist in a year.'

👤 FreeReadText Team📅 June 16, 2026
Bland.aiSeries CFundingDell Technologies CapitalVoice AgentsEnterprise AI
اقرأ المزيد
NetEase Youdao Releases Confucius4-TTS: Open-Source 14-Language Voice Cloning from Just 3 Seconds of Audio
Technology

NetEase Youdao Releases Confucius4-TTS: Open-Source 14-Language Voice Cloning from Just 3 Seconds of Audio

Chinese edtech giant NetEase Youdao open-sources Confucius4-TTS under Apache 2.0, a 1.3B-parameter voice cloning model achieving 85%+ voice similarity from 3 seconds of audio across 14 languages — with no reference text needed for cross-lingual cloning.

👤 FreeReadText Team📅 June 23, 2026
NetEase YoudaoConfucius4-TTSOpen SourceVoice CloningMultilingual TTSFlow Matching
اقرأ المزيد
NO FAKES Act Unanimously Passes Senate Judiciary Committee, Creating Federal Voice and Likeness Protection
Regulation

NO FAKES Act Unanimously Passes Senate Judiciary Committee, Creating Federal Voice and Likeness Protection

The bipartisan NO FAKES Act clears the Senate Judiciary Committee by unanimous voice vote, creating a federal intellectual property right over AI-generated digital replicas of voice and visual likeness — with platform liability, 70-year post-mortem protections, and DMCA-style takedown provisions.

👤 FreeReadText Team📅 June 18, 2026
NO FAKES ActSenate Judiciary CommitteeDigital ReplicaVoice RightsDeepfake RegulationFederal IP Law
اقرأ المزيد
Kotoba Technologies Raises $10 Million to Bring Real-Time Voice AI to East Asian Languages
Business

Kotoba Technologies Raises $10 Million to Bring Real-Time Voice AI to East Asian Languages

San Francisco and Tokyo-based Kotoba Technologies raises an additional $10 million in seed funding led by Kindred Ventures, with Salesforce Ventures and Sony Innovation Fund participating, to expand its Koto voice AI model optimized for Japanese, Korean, and Chinese — languages spoken by roughly 1.6 billion people.

👤 FreeReadText Team📅 June 24, 2026
Kotoba TechnologiesEast Asian AISpeech-to-SpeechVoice AISeed FundingOn-Device AIMultilingual
اقرأ المزيد
ViiTorVoice-NAR Goes Open Source: First TTS Model That Edits Single Words Inside Finished Audio
Technology

ViiTorVoice-NAR Goes Open Source: First TTS Model That Edits Single Words Inside Finished Audio

Chinese startup Yunshang Qulv releases ViiTorVoice-NAR under Apache 2.0, introducing word-level audio editing that replaces individual words without regenerating surrounding content — alongside sub-60ms latency and benchmark-leading accuracy on both English and Chinese.

👤 FreeReadText Team📅 July 1, 2026
ViiTorVoiceOpen Source TTSWord-Level EditingChinese AINAR ArchitectureApache 2.0
اقرأ المزيد
OpenAI Launches GPT-Live: Full-Duplex Voice Model Lets ChatGPT Listen and Speak Simultaneously
Technology

OpenAI Launches GPT-Live: Full-Duplex Voice Model Lets ChatGPT Listen and Speak Simultaneously

OpenAI rolls out GPT-Live-1 and GPT-Live-1 mini globally, introducing full-duplex architecture that enables ChatGPT to listen and speak at the same time — with background task delegation to GPT-5.5 for complex reasoning, marking voice AI's shift from turn-based chat to continuous conversation.

👤 FreeReadText Team📅 July 8, 2026
OpenAIGPT-LiveFull-DuplexChatGPT VoiceReal-Time AIGPT-5.5
اقرأ المزيد
Gradium Raises $100M Seed Round Backed by Nvidia to Build Ultra-Low-Latency Voice AI
Business

Gradium Raises $100M Seed Round Backed by Nvidia to Build Ultra-Low-Latency Voice AI

Paris-based voice AI startup Gradium, spun out of French research lab Kyutai, extends its seed round to over $100 million with Nvidia joining as a strategic investor — signaling that the race to eliminate latency in AI voice conversations is attracting infrastructure-level capital.

👤 FreeReadText Team📅 July 8, 2026
GradiumNvidiaSeed FundingKyutaiVoice AIParisUltra-Low Latency
اقرأ المزيد
Tencent Cloud Partners with Inworld AI to Deliver One-Stop Real-Time Voice AI with Sub-130ms Latency Across 100+ Languages
Business

Tencent Cloud Partners with Inworld AI to Deliver One-Stop Real-Time Voice AI with Sub-130ms Latency Across 100+ Languages

Tencent Cloud and Inworld AI announce a strategic partnership integrating Inworld's top-ranked TTS models into Tencent RTC's global infrastructure, creating a production-grade voice AI solution with sub-130ms first-chunk latency, 100+ language support, and voice cloning — backed by 3,200+ global edge nodes.

👤 FreeReadText Team📅 June 16, 2026
Tencent CloudInworld AIReal-Time VoiceTTSPartnershipEnterprise AITencent RTC
اقرأ المزيد
Rime Raises $24 Million Series A to Build Enterprise-Ready Speech-to-Speech Voice AI
Business

Rime Raises $24 Million Series A to Build Enterprise-Ready Speech-to-Speech Voice AI

San Francisco-based Rime raises $24 million led by M13 to scale its linguistics-first speech-to-speech voice AI platform, already handling nearly 100 million calls monthly for Mayo Clinic and Dialpad — and hires ex-Meta audio research lead Rafael Valle as Chief Science Officer.

👤 FreeReadText Team📅 July 15, 2026
RimeSeries AM13Speech-to-SpeechEnterprise Voice AIContact Center
اقرأ المزيد
Omilia Launches Lexis: First Native Generative TTS Built Into an Enterprise Contact Center Platform
Technology

Omilia Launches Lexis: First Native Generative TTS Built Into an Enterprise Contact Center Platform

Omilia releases Lexis, a generative text-to-speech model built natively into its Cloud Platform — not via third-party API — delivering sub-45ms latency, 25+ languages, instant voice cloning, and PCI-DSS, HIPAA, and GDPR compliance for regulated enterprise contact centers.

👤 FreeReadText Team📅 July 8, 2026
OmiliaLexis TTSContact CenterEnterprise VoiceVoice CloningCompliance
اقرأ المزيد
Ex-Waymo Engineer Raises $28 Million to Bring Autonomous-Vehicle-Grade Testing to Voice AI Agents
Business

Ex-Waymo Engineer Raises $28 Million to Bring Autonomous-Vehicle-Grade Testing to Voice AI Agents

Coval, founded by a former Waymo evaluation infrastructure lead, raises $28 million from Norwest, Base10 Partners, and Twilio Ventures to build simulation-based testing for voice AI agents — applying the same safety discipline that made self-driving cars possible to autonomous voice systems.

👤 FreeReadText Team📅 June 24, 2026
CovalSeries ANorwestVoice AI TestingSimulationEnterpriseWaymo
اقرأ المزيد
Alibaba's Qwen-Audio-3.0-TTS Tops Global Leaderboard: 16 Languages, Voice Cloning, and Sub-200ms Latency
Technology

Alibaba's Qwen-Audio-3.0-TTS Tops Global Leaderboard: 16 Languages, Voice Cloning, and Sub-200ms Latency

Alibaba releases Qwen-Audio-3.0-TTS in Flash and Plus variants, claiming the #1 spot on the Artificial Analysis Speech Arena with an Elo score of 1,234 — surpassing Google, ElevenLabs, and Cartesia in blind listening tests while adding support for 16 languages and 20 Chinese dialects.

👤 FreeReadText Team📅 July 20, 2026
AlibabaQwen-AudioTTSArtificial AnalysisVoice CloningMultilingual TTSChinese AI
اقرأ المزيد
ByteDance Unveils Seed Audio 1.0: A Single Prompt Generates Dialogue, Sound Effects, and Ambient Audio in One Pass
Technology

ByteDance Unveils Seed Audio 1.0: A Single Prompt Generates Dialogue, Sound Effects, and Ambient Audio in One Pass

ByteDance's Seed team launches Seed Audio 1.0, an 'audio creation model' that jointly generates speech, sound effects, music, and environmental audio within a unified framework — eliminating the multi-tool pipeline that traditional audio production requires and delivering 90%+ usable audio in a single inference step.

👤 FreeReadText Team📅 July 20, 2026
ByteDanceSeed AudioAudio GenerationTTSSound EffectsChinese AIVolcano Ark
اقرأ المزيد
Japan Moves to Legally Protect Voices from Unauthorized AI Use — First-of-Its-Kind Framework in Asia
Regulation

Japan Moves to Legally Protect Voices from Unauthorized AI Use — First-of-Its-Kind Framework in Asia

Japan's Ministry of Justice unveils draft guidelines extending right-of-publicity protections to voices, proposing that unauthorized AI-generated voice content causing harm be treated as a civil violation — responding to over 43,000 suspected unauthorized uploads and an estimated ¥4.5 billion in economic losses to rights holders.

👤 FreeReadText Team📅 July 14, 2026
JapanVoice RightsRight of PublicityAI RegulationVoice CloningAnime IndustryMinistry of Justice
اقرأ المزيد
Speechify's Simba 3.2 Tops Global TTS Leaderboard: Consumer-First Voice Model Beats OpenAI, Google, and ElevenLabs in Blind Tests
Technology

Speechify's Simba 3.2 Tops Global TTS Leaderboard: Consumer-First Voice Model Beats OpenAI, Google, and ElevenLabs in Blind Tests

Speechify's Simba 3.2 reaches #1 on the Artificial Analysis TTS leaderboard, beating ElevenLabs, OpenAI, Google DeepMind, and Cartesia in independent blind listening tests — while priced at just $6–10 per million characters, roughly 15x cheaper than rivals in the top ten.

👤 FreeReadText Team📅 July 10, 2026
SpeechifySimba 3.2TTS BenchmarkArtificial AnalysisVoice AIConsumer TTS
اقرأ المزيد
Verbatik Launches First MCP Text-to-Speech Server: 2,700+ Neural Voices Now a Native Capability for Claude, Codex, and Other AI Assistants
Technology

Verbatik Launches First MCP Text-to-Speech Server: 2,700+ Neural Voices Now a Native Capability for Claude, Codex, and Other AI Assistants

London-based Verbatik Technologies releases the first text-to-speech server for the Model Context Protocol (MCP), connecting 2,700+ neural voices across 50+ languages directly to AI assistants — turning voice generation into a native agent capability rather than an external API call.

👤 FreeReadText Team📅 July 16, 2026
VerbatikMCPModel Context ProtocolText-to-SpeechAI AssistantsVoice CloningClaude
اقرأ المزيد
AMD Lemonade 11.0 Brings Text-to-Speech and Voice Cloning to Open-Source Local AI Server
Technology

AMD Lemonade 11.0 Brings Text-to-Speech and Voice Cloning to Open-Source Local AI Server

AMD releases Lemonade 11.0, a major update to its open-source local AI server, adding text-to-speech with voice cloning and voice design via the OpenMOSS backend — enabling entirely offline, privacy-preserving speech synthesis on consumer AMD hardware.

👤 FreeReadText Team📅 July 15, 2026
AMDLemonade 11.0Local AIOpen Source TTSVoice CloningOpenMOSSPrivacy
اقرأ المزيد
Deepgram Brings Enterprise Voice AI to Snapdragon PCs: Nova-3 Speech Recognition Runs Fully On-Device with 6.89% Word Error Rate
Technology

Deepgram Brings Enterprise Voice AI to Snapdragon PCs: Nova-3 Speech Recognition Runs Fully On-Device with 6.89% Word Error Rate

Deepgram partners with Qualcomm to optimize its Nova-3 speech-to-text model for Snapdragon X Series processors, enabling real-time enterprise voice recognition that runs entirely on-device via the Hexagon NPU — no cloud required — targeting AI PCs, automotive, XR, and industrial edge deployments.

👤 FreeReadText Team📅 July 21, 2026
DeepgramQualcommSnapdragonNova-3On-Device AISpeech-to-TextEdge Computing
اقرأ المزيد
China's National Security Ministry Warns AI Voice Cloning Now Takes Just 3 Seconds — Outlines Three Major Risk Categories
Regulation

China's National Security Ministry Warns AI Voice Cloning Now Takes Just 3 Seconds — Outlines Three Major Risk Categories

China's Ministry of State Security issues a landmark public advisory warning that voice cloning platforms now require as little as 3 seconds of audio to create convincing fakes, outlining risks of impersonation fraud, voice rights infringement, and public opinion manipulation — while detailing an existing three-pillar legal framework for enforcement.

👤 FreeReadText Team📅 July 22, 2026
ChinaNational SecurityVoice CloningDeepfake RegulationAI SafetyTelecom FraudCivil Code
اقرأ المزيد
ElevenLabs Explores $22 Billion Tender Offer, Doubling Valuation in Six Months as Voice AI Revenue Surges Past $5 Billion ARR
Business

ElevenLabs Explores $22 Billion Tender Offer, Doubling Valuation in Six Months as Voice AI Revenue Surges Past $5 Billion ARR

ElevenLabs enters early-stage negotiations for a secondary share sale at a $22 billion valuation — doubling its February 2026 price tag — backed by surging enterprise adoption, $5 billion-plus ARR, and 33 million AI-powered conversations handled by its agents platform in the first half of 2026.

👤 FreeReadText Team📅 July 2, 2026
ElevenLabsTender OfferValuationSecondary SaleVoice AIEnterprise AIIPO
اقرأ المزيد
Cartesia Launches Sonic 3.5 and Ink 2: SSM Architecture Makes It the First Provider to Top Both TTS and STT Leaderboards
Technology

Cartesia Launches Sonic 3.5 and Ink 2: SSM Architecture Makes It the First Provider to Top Both TTS and STT Leaderboards

Cartesia releases Sonic 3.5 (TTS) and Ink 2 (STT) built on State Space Model architecture rather than Transformers, achieving sub-90ms latency and 42-language support — becoming the only provider to simultaneously hold the #1 ranking for both speaking and listening on Artificial Analysis benchmarks.

👤 FreeReadText Team📅 June 17, 2026
CartesiaSonic 3.5Ink 2State Space ModelSSMReal-Time TTSSpeech-to-Text
اقرأ المزيد
Vapi Raises $50M at $500M Valuation After Amazon Ring Picks Its Voice AI Platform Over 40 Rivals
Business

Vapi Raises $50M at $500M Valuation After Amazon Ring Picks Its Voice AI Platform Over 40 Rivals

Voice AI infrastructure startup Vapi closes a $50 million Series B led by Peak XV Partners at a $500 million valuation, after Amazon Ring selected its platform from over 40 vendors to handle 100% of inbound customer support calls — with the company now processing over 1 billion total calls.

👤 FreeReadText Team📅 May 12, 2026
VapiSeries BPeak XVAmazon RingVoice AgentsEnterprise AIContact Center
اقرأ المزيد
Lingraphica Launches AI-Powered 'Conversations' AAC Tool: Voice AI Enables Spontaneous Dialogue for People With Speech Disabilities
Technology

Lingraphica Launches AI-Powered 'Conversations' AAC Tool: Voice AI Enables Spontaneous Dialogue for People With Speech Disabilities

Lingraphica releases Conversations, an AI-powered communication tool for its AAC devices that transcribes a conversation partner's speech and suggests contextually relevant responses — cutting conversation time by 80% and achieving a 100% user test completion rate in clinical trials.

👤 FreeReadText Team📅 July 8, 2026
LingraphicaAACVoice AIAccessibilitySpeech DisabilitiesHealthcare AIAssistive Technology
اقرأ المزيد
Fish Audio Raises $52 Million Seed Round After Turning a Bedroom GPU Project Into One of Voice AI's Fastest-Growing Companies
Business

Fish Audio Raises $52 Million Seed Round After Turning a Bedroom GPU Project Into One of Voice AI's Fastest-Growing Companies

Palo Alto-based Fish Audio raises a $52 million seed round co-led by Coreline Ventures and Capital Today, scaling from a solo developer's open-source side project into an 8-million-user platform with $21 million in annual recurring revenue — all within its first year.

👤 FreeReadText Team📅 July 28, 2026
Fish AudioSeed FundingOpen Source TTSVoice CloningCoreline VenturesCapital Today$52M
اقرأ المزيد
xAI Launches Grok Voice Think Fast 2.0: Tops Agentic Benchmark and Crushes Transcription Rivals by Up to 10x in Noisy Environments
Technology

xAI Launches Grok Voice Think Fast 2.0: Tops Agentic Benchmark and Crushes Transcription Rivals by Up to 10x in Noisy Environments

Elon Musk's xAI releases Grok Voice Think Fast 2.0, a speech-to-speech model that ranks #1 on the τ-voice Bench for agentic performance while transcribing 1.5–2x more accurately than dedicated STT services — and up to 10x better in noisy conditions — at $0.08 per minute.

👤 FreeReadText Team📅 July 29, 2026
xAIGrok VoiceThink Fast 2.0Speech-to-SpeechTranscriptionBenchmarksElon Musk
اقرأ المزيد
EU AI Act Article 50 Takes Effect: Synthetic Audio Marking and AI Voice Disclosure Become Law on August 2, With Penalties Up to €15 Million
Regulation

EU AI Act Article 50 Takes Effect: Synthetic Audio Marking and AI Voice Disclosure Become Law on August 2, With Penalties Up to €15 Million

The EU AI Act's transparency obligations become enforceable on August 2, 2026, requiring all AI-generated synthetic audio to carry machine-readable watermarks and all AI voice interactions to disclose their artificial nature — with fines of up to €15 million or 3% of global turnover for non-compliance.

👤 FreeReadText Team📅 July 30, 2026
EU AI ActArticle 50Synthetic AudioWatermarkingTransparencyComplianceGDPR
اقرأ المزيد
Smallest.ai Raises $13M Series A, Unveils Voice 4.0 and Hydra: The First Async Speech-to-Speech Model That Listens, Reasons, and Responds in Parallel
Technology

Smallest.ai Raises $13M Series A, Unveils Voice 4.0 and Hydra: The First Async Speech-to-Speech Model That Listens, Reasons, and Responds in Parallel

San Francisco-based Smallest.ai raises a $13 million Series A led by Seligman Ventures, bringing total funding past $21 million, alongside the launch of Voice 4.0 — a paradigm shift that processes listening, reasoning, and speech in parallel through its Hydra speech-to-speech model, enabling AI to begin responding while conversations are still unfolding.

👤 FreeReadText Team📅 July 30, 2026
Smallest.aiVoice 4.0HydraSpeech-to-SpeechSeries AAsync AISeligman Ventures
اقرأ المزيد
PolyAI Debuts Dialog-RSN-1: An Audio-Native Voice Model That Hears Calls the Way Humans Do, With Sub-300ms Latency
Technology

PolyAI Debuts Dialog-RSN-1: An Audio-Native Voice Model That Hears Calls the Way Humans Do, With Sub-300ms Latency

PolyAI launches Dialog-RSN-1, an audio-native dialog model that reasons directly over raw call audio instead of text transcripts — fusing turn-taking, speech recognition, function calling, and response generation into a single LLM that responds in under 300 milliseconds, outperforming cascaded and speech-to-speech architectures on call center benchmarks.

👤 FreeReadText Team📅 July 30, 2026
PolyAIDialog-RSN-1Audio-Native AICall CenterVoice AIContact CenterLow Latency
اقرأ المزيد
DXC Technology Partners With ElevenLabs, Participates in $500M Series D as Enterprise IT Giants Bet on Voice AI
Business

DXC Technology Partners With ElevenLabs, Participates in $500M Series D as Enterprise IT Giants Bet on Voice AI

DXC Technology announces a strategic partnership with ElevenLabs to embed voice AI across its enterprise operations and customer solutions, while revealing its participation in ElevenLabs' $500 million Series D round — marking one of the first major commitments by a traditional IT services giant to make voice AI a core enterprise platform.

👤 FreeReadText Team📅 July 28, 2026
DXC TechnologyElevenLabsStrategic PartnershipSeries DEnterprise ITVoice AISystem Integrator
اقرأ المزيد
China Enacts World's First AI Companion Regulation: Voice-Based Emotional AI Faces Mandatory Addiction Warnings, Minor Protections, and Crisis Intervention Rules
Regulation

China Enacts World's First AI Companion Regulation: Voice-Based Emotional AI Faces Mandatory Addiction Warnings, Minor Protections, and Crisis Intervention Rules

China's Interim Measures for the Administration of Anthropomorphic AI Interaction Services took effect on July 15, 2026 — the first comprehensive national regulation targeting AI emotional companions — requiring mandatory overuse reminders every two hours, banning AI romantic partners for minors, and forcing major platforms to suspend companion features overnight.

👤 FreeReadText Team📅 July 15, 2026
ChinaAI Companion RegulationVoice AIMinor ProtectionCACEmotional AIAnthropomorphic AI
اقرأ المزيد
Resemble AI Releases Chatterbox Nano and Flash: Open-Source TTS Goes Edge-Native with 110M-Parameter Model That Runs on CPU
Technology

Resemble AI Releases Chatterbox Nano and Flash: Open-Source TTS Goes Edge-Native with 110M-Parameter Model That Runs on CPU

Resemble AI ships Chatterbox Nano (110M parameters, runs at 3x realtime on CPU) and Chatterbox Flash (diffusion-LLM architecture for high-throughput production) as MIT-licensed open-source TTS models, alongside multilingual V3 supporting 23+ languages — bringing enterprise-grade voice synthesis to edge devices and on-premise deployments.

👤 FreeReadText Team📅 July 6, 2026
Resemble AIChatterboxOpen Source TTSEdge AIMIT LicenseVoice CloningWatermarking
اقرأ المزيد
WellSpan Health Expands Hippocratic AI Partnership: Voice Agent 'Ana' Now Handles 160,000 Patient Calls a Month as Health Systems Bet on AI at Scale
Business

WellSpan Health Expands Hippocratic AI Partnership: Voice Agent 'Ana' Now Handles 160,000 Patient Calls a Month as Health Systems Bet on AI at Scale

WellSpan Health announces a platform-wide expansion of its Hippocratic AI partnership, with voice agent 'Ana' now managing over 160,000 patient calls and 7,000 hours of conversation monthly — expanding from appointment scheduling into post-discharge follow-ups, chronic disease management, and clinical triage co-development.

👤 FreeReadText Team📅 July 30, 2026
WellSpan HealthHippocratic AIVoice AIHealthcare AIPatient EngagementClinical AIAna Voice Agent
اقرأ المزيد
Hugging Face and Cerebras Bring Gemma 4 to Real-Time Voice AI: Open-Source Pipeline Solves the Latency Problem That Breaks Conversational Flow
Technology

Hugging Face and Cerebras Bring Gemma 4 to Real-Time Voice AI: Open-Source Pipeline Solves the Latency Problem That Breaks Conversational Flow

Hugging Face and Cerebras Systems partner to deploy Google DeepMind's Gemma 4 in a fully open, modular speech-to-speech pipeline — pairing Cerebras' wafer-scale inference hardware with best-in-class ASR and TTS to eliminate the P95 tail latency spikes that have long frustrated real-world voice AI deployments.

👤 FreeReadText Team📅 July 3, 2026
Hugging FaceCerebrasGemma 4Real-Time Voice AIOpen SourceSpeech-to-SpeechGoogle DeepMind
اقرأ المزيد
Microsoft Releases VibeVoice-ASR-BitNet: 7B Speech Recognition Model Runs Real-Time on CPU, Beating Whisper.cpp by Up to 2.3x
Technology

Microsoft Releases VibeVoice-ASR-BitNet: 7B Speech Recognition Model Runs Real-Time on CPU, Beating Whisper.cpp by Up to 2.3x

Microsoft open-sources VibeVoice-ASR-BitNet, an edge-optimized inference engine that compresses a 7-billion-parameter speech recognition model from 4.62 GB to 1.58 GB using heterogeneous quantization — achieving real-time transcription on just 3 CPU threads with no GPU required.

👤 FreeReadText Team📅 July 23, 2026
MicrosoftVibeVoiceASREdge AIOpen SourceSpeech RecognitionBitNet
اقرأ المزيد
Krafton Open-Sources 21B Speech-to-Speech Voice AI Model, Bringing Real-Time Conversational Game Characters to Life
Technology

Krafton Open-Sources 21B Speech-to-Speech Voice AI Model, Bringing Real-Time Conversational Game Characters to Life

South Korean gaming giant Krafton releases A.X K2 Raon-Speech on Hugging Face, a 21-billion-parameter end-to-end speech-to-speech model that ranks #1 in Korean and #3 in English among open-source voice AI — designed to power game characters that understand player voice and intent in real time.

👤 FreeReadText Team📅 July 29, 2026
KraftonRaon-SpeechGaming AISpeech-to-SpeechOpen SourceKorean AIGame Characters
اقرأ المزيد
NVIDIA Releases NemotronLabs VoiceChat: First Open Full-Duplex Speech-to-Speech Model with Live Tool Calling
Technology

NVIDIA Releases NemotronLabs VoiceChat: First Open Full-Duplex Speech-to-Speech Model with Live Tool Calling

NVIDIA open-sources an 11-billion-parameter full-duplex speech-to-speech model on Hugging Face — the first open-weight model to support live tool calling mid-conversation, collapsing the traditional ASR-LLM-TTS pipeline into a single unified architecture with ~448ms turn-taking latency.

👤 FreeReadText Team📅 August 3, 2026
NVIDIANemotronLabsVoiceChatFull-DuplexSpeech-to-SpeechOpen SourceTool Calling
اقرأ المزيد
Microsoft Begins Testing MAI Realtime: Its First Self-Developed Full-Duplex AI Voice Model Speaks 16 Languages
Technology

Microsoft Begins Testing MAI Realtime: Its First Self-Developed Full-Duplex AI Voice Model Speaks 16 Languages

Microsoft begins partner testing of MAI Realtime, its first self-developed native full-duplex AI voice model, supporting 16 languages with automatic detection and mid-conversation switching — signaling a strategic move to reduce Azure's dependence on OpenAI for voice infrastructure.

👤 FreeReadText Team📅 August 4, 2026
MicrosoftMAI RealtimeFull-DuplexVoice AIAzureMultilingualEnterprise
اقرأ المزيد
Kakao Upgrades Kanana-o AI Voice: Natural-Language Prompts Now Control Emotion, Accent, and Delivery — Outperforms OpenAI on Korean TTS Benchmark
Technology

Kakao Upgrades Kanana-o AI Voice: Natural-Language Prompts Now Control Emotion, Accent, and Delivery — Outperforms OpenAI on Korean TTS Benchmark

South Korean internet giant Kakao announces a major upgrade to its Kanana-o omni AI model, enabling users to control voice emotion, tone, accent, and delivery style through plain-language text prompts — scoring 94.50 on the Korean InstructTTSEval benchmark, ahead of OpenAI's GPT-4o-mini-tts at 91.10.

👤 FreeReadText Team📅 August 4, 2026
KakaoKanana-oKorean AIVoice AIEmotion ControlText-to-SpeechLM-SPT
اقرأ المزيد