phoneme free download

Showing 20 open source projects for "phoneme"

View related business solutions

AI-powered service management for IT and enterprise teams
Enterprise-grade ITSM, for every business

Give your IT, operations, and business teams the ability to deliver exceptional services—without the complexity. Maximize operational efficiency with refreshingly simple, AI-powered Freshservice.

Try it Free
Try Google Cloud Risk-Free With $300 in Credit
No hidden charges. No surprise bills. Cancel anytime.

Use your credit across every product. Compute, storage, AI, analytics. When it runs out, 20+ products stay free. You only pay when you choose to.

Start Free
1

GLM-TTS

Controllable & emotion-expressive zero-shot TTS

...The system introduces a multi-reward reinforcement learning framework that jointly optimizes for voice similarity, emotional expressiveness, pronunciation, and intelligibility, yielding output that can rival commercial options in naturalness and expressiveness. GLM-TTS also supports phoneme-level control and hybrid text + phoneme input, giving developers precise control over pronunciation critical for multilingual or polyphone-rich languages.

Downloads: 1 This Week

Last Update: 2026-04-10
See Project
2

Fish Speech

SOTA Open Source TTS

...Fish Speech emphasizes expressive and controllable voices: it supports a long list of emotion tags, tone markers, and special audio effect markers that can be embedded in the text to drive prosody and vocal style, from basic emotions to nuanced states like sarcastic, conciliative, or hysterical. The system is multilingual and cross-lingual, handling multiple languages in a single input without explicit phoneme markup, and is trained on large-scale datasets.

Downloads: 11 This Week

Last Update: 2025-11-28
See Project
3

FastKoko

Dockerized FastAPI wrapper for Kokoro-82M text-to-speech model

...The project exposes an OpenAI-compatible speech endpoint, which means existing code that talks to the OpenAI audio API can often be pointed at a Kokoro-FastAPI instance with minimal changes. It supports multiple languages and voicepacks and allows phoneme based generation for more accurate pronunciation and prosody. The server also offers per-word timestamped captions, which makes it useful for creating subtitles or aligning audio with text. A built in web UI, API documentation, and debug endpoints for monitoring system status help users explore voices, test requests, and integrate the service into larger systems.

Downloads: 3 This Week

Last Update: 2025-12-13
See Project
4

Matcha-TTS

A fast TTS architecture with conditional flow matching

...The repository provides an end-to-end TTS pipeline: a PyTorch/Lightning training stack, configuration files, pre-trained checkpoints, a command-line interface, and a Gradio app for interactive testing. Users can train on standard datasets like LJSpeech or plug in their own corpora, with helper tools for computing dataset statistics, extracting phoneme durations, and running multi-GPU training.

Downloads: 0 This Week

Last Update: 2025-11-28
See Project
MongoDB Atlas runs apps anywhere
Deploy in 115+ regions with the modern database for every enterprise.

MongoDB Atlas gives you the freedom to build and run modern applications anywhere—across AWS, Azure, and Google Cloud. With global availability in over 115 regions, Atlas lets you deploy close to your users, meet compliance needs, and scale with confidence across any geography.

Start Free
5

PaddleSpeech

Easy-to-use Speech Toolkit including Self-Supervised Learning model

...We provide high-speed and ultra-lightweight models, and also cutting-edge technology. We provide production ready streaming asr and streaming tts system. Our frontend contains Text Normalization and Grapheme-to-Phoneme (G2P, including Polyphone and Tone Sandhi). Moreover, we use self-defined linguistic rules to adapt Chinese context.

Downloads: 0 This Week

Last Update: 2025-03-04
See Project
6

Infinite Monkeys 5.0

...(A complementary Script Forge, allowing users to design their own script-templates from scratch, is currently under development.) IM5.26 also incorporates robust features such as macro definitions, conditional logic, loops, string and array manipulation, phoneme-based operations, dictionary filtering, and semantics-aware word selection. Designed for writers, digital humanists, and e-literature creators, Infinite Monkeys 5.26 functions entirely client-side and requires no installation. Whe

Downloads: 1 This Week

Last Update: 2025-12-17
See Project
7

Pen Possible

scans a given textual string in 146 pen on paper possible combinations

Application scans a given textual string in 146 pen on paper possible combinations- horizontal, vertical, diagonal, reverse, join top, join bottom, groups(2/3/4..), edges & in quadrant dimensions of your choice

Downloads: 0 This Week

Last Update: 2024-12-01
See Project
8

EmotiVoice

Multi-Voice and Prompt-Controlled TTS Engine

EmotiVoice is a multi-voice, prompt-controlled text-to-speech engine designed to generate highly expressive speech across thousands of voices. It supports both English and Chinese and ships with over 2,000 preset voices, making it suitable for everything from characters and virtual anchors to narration and dialogue. The core idea is prompt-based emotional and style control: you can ask the engine to speak “happy,” “sad,” “excited,” or with other high-level style prompts that shape prosody,...

Downloads: 4 This Week

Last Update: 2025-11-30
See Project
9

VALL-E

PyTorch implementation of VALL-E (Zero-Shot Text-To-Speech)

We introduce a language modeling approach for text to speech synthesis (TTS). Specifically, we train a neural codec language model (called VALL-E) using discrete codes derived from an off-the-shelf neural audio codec model, and regard TTS as a conditional language modeling task rather than continuous signal regression as in previous work. During the pre-training stage, we scale up the TTS training data to 60K hours of English speech which is hundreds of times larger than existing systems....

Downloads: 0 This Week

Last Update: 2023-04-14
See Project
Go From AI Idea to AI App Fast
One platform to build, fine-tune, and deploy ML models. No MLOps team required.

Access Gemini 3 and 200+ models. Build chatbots, agents, or custom models with built-in monitoring and scaling.

Try Free
10

VITS

Conditional Variational Autoencoder with Adversarial Learning

VITS is a foundational research implementation of “VITS: Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech,” a well-known neural TTS architecture. Unlike traditional two-stage systems that separately train an acoustic model and a vocoder, VITS trains an end-to-end model that maps text directly to waveform using a conditional variational autoencoder combined with normalizing flows and adversarial training. This architecture enables parallel generation...

Downloads: 0 This Week

Last Update: 2025-11-28
See Project
11

Transformer TTS

Implementation of a Transformer based neural network

...It takes inspiration from architectures like FastSpeech, FastSpeech 2, FastPitch, and Transformer TTS, and extends them with its own aligner and forward models. The system separates alignment learning and acoustic modeling: an autoregressive Transformer is used as an aligner to extract phoneme-to-frame durations, while a non-autoregressive “ForwardTransformer” generates mel-spectrograms conditioned on text and durations. This design addresses common autoregressive issues such as repetition, skipped words, and unstable attention, and results in robust, fast synthesis where all frames are predicted in parallel. The repository ships with tooling to build datasets (especially LJSpeech) and create training data, plus scripts to train both the aligner and the TTS model, monitor training with TensorBoard, and resume or reset training runs.

Downloads: 0 This Week

Last Update: 2025-11-28
See Project
12

PORORO

Platform of neural models for natural language processing

pororo performs Natural Language Processing and Speech-related tasks. It is easy to solve various subtasks in the natural language and speech processing field by simply passing the task name. Recognized speech sentences using the trained model. Currently English, Korean and Chinese support. Get vector or find similar words and entities from pretrained model using Wikipedia.

Downloads: 0 This Week

Last Update: 2022-08-19
See Project
13

Vietnamese Grapheme to Phoneme

Converting any Vietnamese word in grapheme to phoneme

Vietnamese is a language that any Vietnamese word can be correctly pronounced even if the speaker does not know its meaning and has never seen it. This tool , a grapheme-to-phoneme method, converts any Vietnamese word from grapheme-based into a phoneme-based pronunciation that integrates tone information. It is usefull to create a lexicon for deverloping a Vi LVCSR system. Detail in: http://ieeexplore.ieee.org/xpl/login.jsp?tp=&arnumber=7357876&url=http%3A%2F%2Fieeexplore.ieee.org%2Fxpls%2Fabs_all.jsp%3Farnumber%3D7357876 ---------- Van Huy Nguyen, TNUT, huynguyen@tnut.edu.vn

Downloads: 0 This Week

Last Update: 2016-08-11
See Project
14

ProseVis

ProseVis is a visualization tool for analyzing the sound of text.

...Mellon Foundation through a grant titled "SEASR Services," in which we seek to identify other features than the "word" to analyze texts. These features comprise sound including parts-of-speech, accent, phoneme, stress, tone, break index. ProseVis allows a reader to map the features extracted from OpenMary (http://mary.dfki.de/) Text-to-speech System and predictive classification data to the "original" text. We developed this project with the ultimate goal of facilitating a reader's ability to analyze and disseminate the results in human readable form. ...

Downloads: 0 This Week

Last Update: 2014-08-11
See Project
15

OpenWTK

This project aims to provide a completely open source alternative to Sun/Oracle Java WTK. It uses open technology from OpenJDK, PhoneME, Ant and microemulator projects and gives you the power of easy J2ME project management with command line tools.

Downloads: 0 This Week

Last Update: 2016-02-17
See Project
16

speakalator

A phrase to phoneme code converter for the SpeakJet chip by Magnevation. Speakalator runs on Unix type operating systems.

Downloads: 0 This Week

Last Update: 2016-06-14
See Project
17

Grapheme to Phoneme Forge

Use our tools to hand edit phonetic word dictionaries for speech recognition engines. The new G2P4J format supporting SAMPA and Kirshenbaum IPA is portable to Sphinx, Julius and others. Demo medical, legal and technical dictionaries are featured.

Downloads: 0 This Week

Last Update: 2013-04-03
See Project
18

pymbrola: a python phonemiser for MBROLA

pymbrola aims to be a universal text-to-phoneme engine which supports and promotes the use of the MBROLA TTS synthesizer.

Downloads: 0 This Week

Last Update: 2013-04-23
See Project
19

SAPI Lipsync

SAPI Lipsync (phoneme alignment) C++ Software

Downloads: 0 This Week

Last Update: 2015-07-03
See Project
20

mms-300m-1130-forced-aligner

CTC-based forced aligner for audio-text in 158 languages

mms-300m-1130-forced-aligner is a multilingual forced alignment model based on Meta’s MMS-300M wav2vec2 checkpoint, adapted for Hugging Face’s Transformers library. It supports forced alignment between audio and corresponding text across 158 languages, offering broad multilingual coverage. The model enables accurate word- or phoneme-level timestamping using Connectionist Temporal Classification (CTC) emissions. Unlike other tools, it provides significant memory efficiency compared to the TorchAudio forced alignment API. Users can integrate it easily through the Python package ctc-forced-aligner, and it supports GPU acceleration via PyTorch. The alignment pipeline includes audio processing, emission generation, tokenization, and span detection, making it suitable for speech analysis, transcription syncing, and dataset creation. ...

Downloads: 0 This Week

Last Update: 2025-07-02
See Project