const-set-1 free download

Whisper

Robust Speech Recognition via Large-Scale Weak Supervision

...These tasks are jointly represented as a sequence of tokens to be predicted by the decoder, allowing a single model to replace many stages of a traditional speech-processing pipeline. The multitask training format uses a set of special tokens that serve as task specifiers or classification targets.

Downloads: 65 This Week

Last Update: 2025-06-26

See Project

Diffgram

Training data (data labeling, annotation, workflow) for all data types

From ingesting data to exploring it, annotating it, and managing workflows. Diffgram is a single application that will improve your data labeling and bring all aspects of training data under a single roof. Diffgram is world’s first truly open source training data platform that focuses on giving its users an unlimited experience. This is aimed to reduce your data labeling bills and increase your Training Data Quality. Training Data is the art of supervising machines through data. This...

Downloads: 4 This Week

Last Update: 2024-10-14

See Project

Kaldi

kaldi-asr/kaldi is the official location of the Kaldi project

...It includes extensive tools for data preparation, feature extraction, acoustic and language modeling, decoding, and evaluation. With its modular design, Kaldi allows users to adapt the system to a wide range of languages and domains. As one of the most influential projects in speech recognition, it has become a foundation for much of the modern work in ASR.

Downloads: 3 This Week

Last Update: 15 hours ago

See Project

The SpeechBrain Toolkit

A PyTorch-based Speech Toolkit

SpeechBrain is an open-source and all-in-one conversational AI toolkit. It is designed to be simple, extremely flexible, and user-friendly. Competitive or state-of-the-art performance is obtained in various domains. SpeechBrain supports state-of-the-art methods for end-to-end speech recognition, including models based on CTC, CTC+attention, transducers, transformers, and neural language models relying on recurrent neural networks and transformers.

Downloads: 0 This Week

Last Update: 2025-04-07

See Project

WhisperJAV

A subtitle generator for Japanese Adult Videos.

A subtitle generator for Japanese Adult Videos. Transformer-based ASR architectures like Whisper suffer significant performance degradation when applied to the spontaneous and noisy domain of JAV. This degradation is driven by specific acoustic and temporal characteristics that defy the statistical distributions of standard training data.

1 Review

Downloads: 79 This Week

Last Update: 4 days ago

See Project

Mice MX OS speech to text Voice Control

Mice speech to text with MX Cinnamon OS ISO

...However, only German settings are currently implemented. category: System commands comment: Screen grid trigger: Display screen (Ras.*|Grid)* terminal_command: /opt/micesttm/read-aloud/screen_grid.py & sleep 1 && xdotool search --name "screen grid" windowactivate intern_command: tts: Screen grid for the mouse click was selected.

Downloads: 0 This Week

Last Update: 2025-05-14

See Project

Tensor2Tensor

Library of deep learning models and datasets

Deep Learning (DL) has enabled the rapid advancement of many useful technologies, such as machine translation, speech recognition and object detection. In the research community, one can find code open-sourced by the authors to help in replicating their results and further advancing deep learning. However, most of these DL systems use unique setups that require significant engineering effort and may only work for a specific problem or architecture, making it hard to run new experiments and compare the results. ...

Downloads: 0 This Week

Last Update: 2021-05-24

See Project

Lip Reading

Cross Audio-Visual Recognition using 3D Architectures

...Audio-visual recognition (AVR) has been considered as a solution for speech recognition tasks when the audio is corrupted, as well as a visual recognition method used for speaker verification in multi-speaker scenarios. The approach of AVR systems is to leverage the extracted information from one modality to improve the recognition ability of the other modality by complementing the missing information. The essential problem is to find the correspondence between the audio and visual streams, which is the goal of this work. We proposed the utilization of a coupled 3D Convolutional Neural Network (CNN) architecture that can map both modalities into a representation space to evaluate the correspondence of audio-visual streams using the learned multimodal features.

Downloads: 3 This Week

Last Update: 2022-08-11

See Project

Search Results for "const-set-1"

Showing 8 open source projects for "const-set-1"

Whisper

Diffgram

Kaldi

The SpeechBrain Toolkit

WhisperJAV

Mice MX OS speech to text Voice Control

Tensor2Tensor

Lip Reading

Search Results for "const-set-1"

Showing 8 open source projects for "const-set-1"

Whisper

Diffgram

Kaldi

The SpeechBrain Toolkit

WhisperJAV

Mice MX OS speech to text Voice Control

Tensor2Tensor

Lip Reading

Related Searches

Related Categories