Showing 23 open source projects for "audio transcript"

View related business solutions
  • Custom VMs From 1 to 96 vCPUs With 99.95% Uptime Icon
    Custom VMs From 1 to 96 vCPUs With 99.95% Uptime

    General-purpose, compute-optimized, or GPU/TPU-accelerated. Built to your exact specs.

    Live migration and automatic failover keep workloads online through maintenance. One free e2-micro VM every month.
    Start Free
  • $300 Free Credits to Build on Google Cloud Icon
    $300 Free Credits to Build on Google Cloud

    New customers can spin up VMs, build with AI, and query data at no cost.

    Put your $300 in credit toward real workloads, then keep building with free monthly usage for 20+ products. No commitment and no charge until you upgrade.
    Start Free
  • 1
    Podcastfy.ai

    Podcastfy.ai

    Transforming Multimodal Content into Captivating Multilingual Audio

    Podcastfy is an open-source Python package that transforms multi-modal content (text, images) into engaging, multi-lingual audio conversations using GenAI. Input content includes websites, PDFs, youtube videos as well as images. Unlike UI-based tools focused primarily on note-taking or research synthesis (e.g. NotebookLM), Podcastfy focuses on the programmatic and bespoke generation of engaging, conversational transcripts and audio from a multitude of multi-modal sources enabling...
    Downloads: 10 This Week
    Last Update:
    See Project
  • 2
    quill macOS

    quill macOS

    Ultraminimalist macOS recording + transcription

    Quill is a minimalist macOS meeting recorder and transcription utility that runs from the menu bar. It captures microphone input and system audio as separate tracks, which naturally distinguishes the user from other speakers. When recording stops, both tracks are transcribed locally and merged into a timestamped, speaker-labeled transcript. The app uses an on-device Parakeet model, so recordings and text never need to leave the Mac. Sessions are stored as audio, metadata, transcript, and log files in organized folders. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 3
    Buzz

    Buzz

    Buzz transcribes and translates audio offline

    Buzz is a desktop application for transcribing and translating audio locally with speech recognition models based on Whisper. It can process audio files, video files, and YouTube links without requiring cloud transcription. Live microphone transcription supports real-time captions and a presentation view for accessible events. Speech separation can improve results on noisy recordings, while speaker identification distinguishes voices within transcribed media. Multiple Whisper backends,...
    Downloads: 671 This Week
    Last Update:
    See Project
  • 4
    GladiaFlow

    GladiaFlow

    A desktop app for real-time voice dictation

    GladiaFlow is an open-source desktop app for turning spoken language into text in virtually any text field. It captures microphone audio through a global hotkey and streams it to Gladia’s Live Transcription API. Partial and final results arrive in real time, then the app cleans punctuation, capitalization, and duplicate text before pasting the transcript into the active application. Users can choose push-to-talk or toggle activation and configure languages, code switching, vocabulary, and pronunciations. ...
    Downloads: 7 This Week
    Last Update:
    See Project
  • Ship Agents Faster Icon
    Ship Agents Faster

    Transform your applications and workflows into powerful agentic systems at global scale.

    Gemini Enterprise Agent Platform lets you rapidly build, scale, govern and optimize production-ready agents grounded in your organization's data. The platform enables developers to build custom or pre-built agents for virtually any use case. New customers get $300 in free credits.
    Start Free
  • 5
    claude-video

    claude-video

    Give Claude the ability to watch any video

    Claude Video is an agent skill that gives Claude and compatible coding assistants the ability to analyze video content. It accepts public video URLs or local video files, then extracts the information needed to answer user questions about what happened on screen and in the audio. The workflow checks captions first, downloads only what is necessary, extracts timestamped frames, and produces a transcript through native captions or Whisper fallback. It supports different detail levels so users can trade speed, token cost, and visual coverage depending on the task. The skill is useful for summarizing videos, reviewing screen recordings, analyzing content structure, diagnosing visual bugs, and turning course material into notes. ...
    Downloads: 7 This Week
    Last Update:
    See Project
  • 6
    claude-real-video

    claude-real-video

    Let Claude (or any LLM) actually watch a video

    claude-real-video is a local video processing tool that prepares videos for analysis by Claude or other LLMs. It accepts public video URLs or local files and extracts the visual and audio evidence an AI model needs. Instead of sampling at a fixed interval, it detects scene changes and removes near-duplicate frames to reduce unnecessary token usage. It can transcribe audio, generate a clean output folder, and create a local viewer with the video, keyframe grid, and transcript. The tool can also be installed as a Claude Code skill so an agent can process videos more directly. ...
    Downloads: 1 This Week
    Last Update:
    See Project
  • 7
    Meetily

    Meetily

    Privacy first, AI meeting assistant with 4x faster Parakeet/Whisper

    This project is a privacy-first AI meeting assistant that captures meeting audio, produces real-time transcripts, and generates summaries while keeping processing entirely on your own machine or infrastructure. It’s built for organizations that want meeting intelligence without sending recordings or transcripts to third-party cloud services, which helps address compliance and data sovereignty requirements. The app supports live transcription with local model options (including Whisper- and Parakeet-based workflows) and presents the transcript as the meeting happens, making it useful both for note-taking and accessibility. ...
    Downloads: 29 This Week
    Last Update:
    See Project
  • 8
    video-use

    video-use

    Edit videos with Claude Code

    Video Use is an open-source AI-powered video editing tool that allows users to transform raw footage into polished videos using natural language commands. Designed to work with Claude Code, it automates the entire editing process—from cutting clips to rendering the final output—without requiring manual timelines or complex software interfaces. The system intelligently analyzes audio transcripts and visual cues to make precise, context-aware editing decisions. It supports a wide range of...
    Downloads: 12 This Week
    Last Update:
    See Project
  • 9
    LARA is software for musical analysis using (new) scientific methods for analysis and visualization. LARA is part of the core research: “Interpretation and performance” of the HSLU – Musik (University of Applied Sciences Luzern – Music depart
    Downloads: 0 This Week
    Last Update:
    See Project
  • MongoDB Atlas runs apps anywhere Icon
    MongoDB Atlas runs apps anywhere

    Deploy in 115+ regions with the modern database for every enterprise.

    MongoDB Atlas gives you the freedom to build and run modern applications anywhere—across AWS, Azure, and Google Cloud. With global availability in over 115 regions, Atlas lets you deploy close to your users, meet compliance needs, and scale with confidence across any geography.
    Start Free
  • 10
    HearWrite PDF

    HearWrite PDF

    Local audio-to-PDF transcription for Windows and Debian.

    HearWrite PDF is a privacy-focused desktop transcription application created by Don Fritz. It converts a single audio recording or an entire folder of recordings into polished PDF and editable TXT transcripts. For folders containing multiple recordings, users can arrange files by date and time or filename, create one combined transcript, individual transcripts, or both, and optionally organize the combined document into calendar-date sections.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 11
    CC2.TV / CC2 - Audio- und TV-Datenbank

    CC2.TV / CC2 - Audio- und TV-Datenbank

    Meta-Datenbank-Anwendung für die Audio- und TV-Sendungen des CC2.TV

    Dieses Programm stellt eine Meta-Datenbank-Anwendung für die Audio- und Video-Sendungen des CC2.TV für GNU/Linux Systeme zur Verfügung. Es ermöglicht das Durchsuchen, Verwalten und Abspielen der umfangreichen Inhalte des CC2.TV-Audiocasts und -Videocasts. Ziel ist es, die über 3000 Audiocast-Themen und über 1000 Videocast-Themen, die sich auf Computerthemen, Technik und gesellschaftliche Aspekte konzentrieren, komfortabel zugänglich zu machen. Für die volle Funktionalität,...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 12
    AudioEnhancerMAX

    AudioEnhancerMAX

    Local-first AI audio processing, transcription and mastering

    AudioEnhancerMAX 3.7 is an open-source, local-first audio production suite for podcasters, creators, journalists, educators, researchers, and developers. It combines deterministic Smart Enhance diagnosis, non-destructive transcript-linked speech editing, audio cleanup and mastering, local Faster-Whisper transcription, delivery profiles, review tools, TTS, system monitoring, and optional trusted-LAN Android workers.
    Leader badge
    Downloads: 37 This Week
    Last Update:
    See Project
  • 13
    TranscribeGeek

    TranscribeGeek

    Free offline Whisper transcription for Windows - no upload, no limit

    TranscribeGeek turns audio and video recordings into text on your own machine. It runs OpenAI Whisper models locally, writes a transcript and an optional SubRip .srt subtitle file, and has no account, no upload and no per-minute limit. Windows 10 1809 or later, 64-bit. GPL-3.0.
    Downloads: 1 This Week
    Last Update:
    See Project
  • 14
    Grabbit

    Grabbit

    Free Windows app to download videos from YouTube, TikTok, Instagram an

    Grabbit is a desktop video downloader for Windows. Download videos and audio from YouTube, TikTok, Instagram, Facebook, Twitter and more — up to 4K/8K quality. Includes a Chrome Extension to grab videos directly from your browser. No cloud, no account needed. Free tier available. Pro plan from $8.99/month.
    Downloads: 11 This Week
    Last Update:
    See Project
  • 15
    MARS5

    MARS5

    MARS5 speech model (TTS) from CAMB.AI

    ...The model is built to handle prosodically challenging content such as sports commentary, anime dialogue, and other high-energy or highly varied speech patterns with realistic rhythm and intonation. To control speaker identity, MARS5 uses a short reference audio clip, typically between 2 and 12 seconds, from which it learns the voice characteristics. It supports two main inference modes: shallow clone, which is faster and only needs the reference audio, and deep clone, which additionally uses the transcript of the reference audio to increase similarity and naturalness at the cost of more computation.
    Downloads: 1 This Week
    Last Update:
    See Project
  • 16
    Meeting AI Analyser

    Meeting AI Analyser

    Live AI co-pilot for Windows meetings: local Whisper + Claude AI.

    Meeting AI Analyser is a Windows desktop app that acts as a live AI co-pilot during any meeting. It captures system audio and microphone on your PC and transcribes everything in real time using OpenAI Whisper locally, so your audio never leaves your machine. Only the resulting text transcript is sent to Claude AI via the Claude Code CLI to generate structured summaries every 60 seconds with decisions, action items, participants and next steps. Claude can also explain unfamiliar jargon and translate in real time. ...
    Downloads: 3 This Week
    Last Update:
    See Project
  • 17
    Ainee

    Ainee

    Ainee - AI Notetaking and Learning Companion

    Ainee is your ultimate AI-powered notetaking and learning companion. Capture lecture notes in real-time and effortlessly transform audio, text, files, and YouTube videos into formatted notes, mindmaps, quizzes, flashcards, podcasts, and more. Explore our AI meeting note taker, AI notes, video transcript generator, PDF to AI converter, and AI flashcard maker. Enhance your learning with our AI voice recorder, article summarizer AI, and AI quiz generator. Additionally, share your knowledge base with others to foster the flow of information and help new users benefit from collective insights. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 18
    footswitch2basic

    footswitch2basic

    Audio Transcription software for Linux (Vlc) with a foot pedal

    Footswitch 2 (Basic) is a media player for transcribers on Linux. This version is a stripped down version of Footswitch2, containing only the absolute essentials for transcription. Written in python and using the python bindings for VLC it allows a transcriber to control the audio or video with a footpedal, and includes a set of macros that integrate into LibreOffice. This allows the transcriber to control the media player from within Libreoffice as well, making it useful for those who do...
    Downloads: 1 This Week
    Last Update:
    See Project
  • 19
    Transcription Aid

    Transcription Aid

    Transcription Aid helps you type text from recordings.

    This software is to help type in text from speech recordings. It has several functions proven to help this type of work. However it is fully manual (aside from auto-completion), so no speech recognition if you are looking for that, but it is a great tool to do the job.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 20

    transqript

    a program to transcript audio files

    transqript can be used to transcribe audio files of interviews etc. to text files.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 21
    A Sermon Content Management System based on asp.net. It allows for collection of: speaker, summary, topics, keywords, date, title and transcript information. Creates podcast of mp3 and allow for embedding of video from popular online video providers.
    Downloads: 1 This Week
    Last Update:
    See Project
  • 22
    Speech Made Visible
    Speech Made Visible is an experiment in showing some of the qualities of speech in printed text. Analyze a recording for attributes like pitch, intensity (loudness), and speed; then style the words in a transcript to suggest those characteristics.
    Downloads: 1 This Week
    Last Update:
    See Project
  • 23
    Youtube-Transcript-API

    Youtube-Transcript-API

    Extract & Download Youtube Video Transcripts

    Go beyond basic YouTube captions. Extract transcripts from any youtube video using audio-based transcription when captions aren't available. Then do more with your transcripts — generate mindmaps, create summaries, or chat with the content to find exactly what you need.
    Downloads: 0 This Week
    Last Update:
    See Project
  • Previous
  • You're on page 1
  • Next