Open-Source VieNeu-TTS v3 Turbo Runs 48 kHz Vietnamese Speech on CPUs
Takeaways − VieNeu-TTS v3 Turbo ships 48 kHz Vietnamese TTS, Apache-2.0, trained from scratch on ~10k hours. 16 concurrent real-time streams on a single RTX 3060 with ~115 ms first-audio latency.

- Takeaways − VieNeu-TTS v3 Turbo ships 48 kHz Vietnamese TTS, Apache-2.0, trained from scratch on ~10k hours.
- 23 preset voices across Northern, Central, and Southern dialects, plus instant voice cloning.
- OpenAI-compatible /v1/audio/speech endpoint drops into Pipecat, LiveKit, and OpenAI SDK clients.
Takeaways − VieNeu-TTS v3 Turbo ships 48 kHz Vietnamese TTS, Apache-2.0, trained from scratch on ~10k hours. 16 concurrent real-time streams on a single RTX 3060 with ~115 ms first-audio latency. 23 preset voices across Northern, Central, and Southern dialects, plus instant voice cloning. OpenAI-compatible /v1/audio/speech endpoint drops into Pipecat, LiveKit, and OpenAI SDK clients. CPU path is torch-free via ONNX Runtime; GPU path auto-switches to PyTorch with batching. Supports English-Vietnamese code-switching, inline emotion cues, and single-GPU LoRA fine-tuning. VieNeu-TTS v3 Turbo offers 48 kHz Vietnamese speech synthesis on CPUs VieNeu-TTS v3 Turbo is an open-source text-to-speech model with 48 kHz output, voice cloning, Vietnamese-English code-switching, inline emotion cues, and an OpenAI-compatible streaming API. Its Hugging Face page reports nearly 780,000 downloads. The project describes v3 Turbo as an original architecture trained from scratch on approximately 10,000 hours of English and Vietnamese speech. At roughly 100 million parameters, the model supports CPU inference while providing a separate CUDA path for higher concurrency. The weights use the Apache-2.0 license. The release publishes a single checkpoint in ONNX, Safetensors, and GGUF formats, along with a Python SDK named vieneu . CPU inference runs through ONNX Runtime without importing PyTorch. On CUDA systems, the SDK selects a PyTorch engine with automatic batching and a continuous-batching scheduler. Applications use the same API on either backend. 48 kHz audio: The model uses the MOSS Audio Tokenizer neural codec, doubling the 24 kHz output rate of the v2 series. 23 preset voices: The 23 preset voices cover Northern, Central, and Southern Vietnamese regions with several reading characters. You've reached the end of the free preview. Upgrade to AlphaSignal Pro to read the full article - and everything else behind the paywall.
Sources
Related stories

Google froze its open source bug bounty program due to a ‘significant rise’ in AI submissions
Google froze its open source bug bounty program due to a significant rise in AI submissions | TechCrunch Last day to exhibit your breakthrough to 10,000+ tech leaders at Disrupt is on Oct 2 . Book Exhibit Table Now.

NASA and IBM's open source lunar model turns 17 years of orbiter data into a foundation for lunar science
NASA and IBM's open source lunar model turns 17 years of orbiter data into a foundation for lunar science The NASA-IBM Lunar Foundation Model makes decades of lunar observation data usable for machine learning. It's especially strong at predicting ice deposits at the poles and detecting craters.

FlashML Runs MiniMax H3 Video AI on 8 GB Consumer GPUs
Takeaways − FlashML-org released FreeVideo , a local inference engine for MiniMax H3 video generation. Runs in 8 GB VRAM and 16 GB RAM via aggressive weight offloading and streaming.

The Agent Said It Was Done. The Database Disagreed.
The Agent Said It Was Done. The Database Disagreed. The Agent Said It Was Done. The Database Disagreed. Microsoft ThinkingBox grades AI agents on the records they leave behind, not the sentences they generate, and then asks whether they can do it twenty times in a row.