ElevenLabs Launches V4 Speech Models With 90-Language Support
The v4 and v4 Turbo models add expanded expression controls, faster voice generation and voice cloning from 10 seconds of audio.

ElevenLabs launched its v4 and v4 Turbo speech models Monday with expanded expression controls, lower latency and support for more than 90 languages, targeting developers building voice agents and other audio applications.
The new models allow users to clone a voice with 10 seconds of audio. They are also designed to preserve a speaker’s identity across longer passages and adjust delivery based on the context of the text being read.
ElevenLabs introduced inline expression tags with its v3 model and has expanded the system in v4. Users can stack multiple tags, allowing the model to follow a sequence of directions during a passage.
Language support has increased from 70 languages in the previous version to more than 90. The company identified the largest quality improvements in Japanese, Brazilian Portuguese, Mandarin and Cantonese.
The v4 generation is also aimed at voice agents, where response time affects the flow of a conversation. The model can begin producing audio as soon as the language model behind it starts generating a response. It can also handle confrontations, escalations and hold periods differently to support issue resolution.
ElevenLabs has expanded its enterprise calling business, with large companies accounting for more than 55% of its business. The startup raised $500 million in a funding round led by Sequoia that valued it at $11 billion.
Its annualized revenue run rate increased from roughly $330 million at the start of the year to more than $600 million. The company has also expanded hiring across markets including India, Europe and Brazil, bringing its headcount above 800.


