Showing 9 open source projects for "audio transcript"

View related business solutions
  • Build Agents and Models on One Platform Icon
    Build Agents and Models on One Platform

    Everything you need to build production-ready agents and models. Access 200+ Google and third-party AI models and tools.

    Gemini Enterprise Agent Platform is Google Cloud's comprehensive platform for developers to build, scale, govern, and optimize agents and models. Choose from Google's most advanced models and third-party models like Anthropic's Claude Model Family.
    Try It Free
  • MongoDB Atlas runs apps anywhere Icon
    MongoDB Atlas runs apps anywhere

    Deploy in 115+ regions with the modern database for every enterprise.

    MongoDB Atlas gives you the freedom to build and run modern applications anywhere—across AWS, Azure, and Google Cloud. With global availability in over 115 regions, Atlas lets you deploy close to your users, meet compliance needs, and scale with confidence across any geography.
    Start Free
  • 1
    Podcastfy.ai

    Podcastfy.ai

    Transforming Multimodal Content into Captivating Multilingual Audio

    Podcastfy is an open-source Python package that transforms multi-modal content (text, images) into engaging, multi-lingual audio conversations using GenAI. Input content includes websites, PDFs, youtube videos as well as images. Unlike UI-based tools focused primarily on note-taking or research synthesis (e.g. NotebookLM), Podcastfy focuses on the programmatic and bespoke generation of engaging, conversational transcripts and audio from a multitude of multi-modal sources enabling...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 2
    claude-video

    claude-video

    Give Claude the ability to watch any video

    Claude Video is an agent skill that gives Claude and compatible coding assistants the ability to analyze video content. It accepts public video URLs or local video files, then extracts the information needed to answer user questions about what happened on screen and in the audio. The workflow checks captions first, downloads only what is necessary, extracts timestamped frames, and produces a transcript through native captions or Whisper fallback. It supports different detail levels so users can trade speed, token cost, and visual coverage depending on the task. The skill is useful for summarizing videos, reviewing screen recordings, analyzing content structure, diagnosing visual bugs, and turning course material into notes. ...
    Downloads: 3 This Week
    Last Update:
    See Project
  • 3
    claude-real-video

    claude-real-video

    Let Claude (or any LLM) actually watch a video

    claude-real-video is a local video processing tool that prepares videos for analysis by Claude or other LLMs. It accepts public video URLs or local files and extracts the visual and audio evidence an AI model needs. Instead of sampling at a fixed interval, it detects scene changes and removes near-duplicate frames to reduce unnecessary token usage. It can transcribe audio, generate a clean output folder, and create a local viewer with the video, keyframe grid, and transcript. The tool can also be installed as a Claude Code skill so an agent can process videos more directly. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 4
    video-use

    video-use

    Edit videos with Claude Code

    Video Use is an open-source AI-powered video editing tool that allows users to transform raw footage into polished videos using natural language commands. Designed to work with Claude Code, it automates the entire editing process—from cutting clips to rendering the final output—without requiring manual timelines or complex software interfaces. The system intelligently analyzes audio transcripts and visual cues to make precise, context-aware editing decisions. It supports a wide range of...
    Downloads: 20 This Week
    Last Update:
    See Project
  • Save Up to 91% on Cloud Compute With Spot VMs Icon
    Save Up to 91% on Cloud Compute With Spot VMs

    Automatic sustained-use discounts. One free VM per month. No negotiation needed.

    Run batch jobs at 60-91% off with Spot VMs. Long-running workloads get automatic discounts with sustained use.
    Try Free
  • 5
    Meetily

    Meetily

    Privacy first, AI meeting assistant with 4x faster Parakeet/Whisper

    This project is a privacy-first AI meeting assistant that captures meeting audio, produces real-time transcripts, and generates summaries while keeping processing entirely on your own machine or infrastructure. It’s built for organizations that want meeting intelligence without sending recordings or transcripts to third-party cloud services, which helps address compliance and data sovereignty requirements. The app supports live transcription with local model options (including Whisper- and Parakeet-based workflows) and presents the transcript as the meeting happens, making it useful both for note-taking and accessibility. ...
    Downloads: 26 This Week
    Last Update:
    See Project
  • 6
    LARA is software for musical analysis using (new) scientific methods for analysis and visualization. LARA is part of the core research: “Interpretation and performance” of the HSLU – Musik (University of Applied Sciences Luzern – Music depart
    Downloads: 0 This Week
    Last Update:
    See Project
  • 7
    MARS5

    MARS5

    MARS5 speech model (TTS) from CAMB.AI

    ...The model is built to handle prosodically challenging content such as sports commentary, anime dialogue, and other high-energy or highly varied speech patterns with realistic rhythm and intonation. To control speaker identity, MARS5 uses a short reference audio clip, typically between 2 and 12 seconds, from which it learns the voice characteristics. It supports two main inference modes: shallow clone, which is faster and only needs the reference audio, and deep clone, which additionally uses the transcript of the reference audio to increase similarity and naturalness at the cost of more computation.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 8
    AudioEnhancerMAX

    AudioEnhancerMAX

    Local-first AI audio processing, transcription and mastering

    AudioEnhancerMAX is an open-source, local-first audio processing suite for podcasters, creators, and audio professionals. It combines enhancement, mastering, transcription, TTS, monitoring, and trusted-LAN edge workers in one FastAPI app. The core pipeline runs on your machine: uploads, DSP processing, Whisper transcription, local Ollama/Gemma smart mode, Kokoro TTS, timing history, and reports stay under user control when local dependencies are installed. Optional Edge Neural TTS and...
    Leader badge
    Downloads: 18 This Week
    Last Update:
    See Project
  • 9
    Speech Made Visible
    Speech Made Visible is an experiment in showing some of the qualities of speech in printed text. Analyze a recording for attributes like pitch, intensity (loudness), and speed; then style the words in a transcript to suggest those characteristics.
    Downloads: 1 This Week
    Last Update:
    See Project
  • Our Free Plans just got better! | Auth0 Icon
    Our Free Plans just got better! | Auth0

    With up to 25k MAUs and unlimited Okta connections, our Free Plan lets you focus on what you do best—building great apps.

    You asked, we delivered! Auth0 is excited to expand our Free and Paid plans to include more options so you can focus on building, deploying, and scaling applications without having to worry about your security. Auth0 now, thank yourself later.
    Try free now
  • Previous
  • You're on page 1
  • Next
Auth0 Logo