Showing 54 open source projects for "sentence"

View related business solutions
  • Demo Series - Small Business Backup By Veeam Icon
    Demo Series - Small Business Backup By Veeam

    Learn how to protect your Microsoft 365 data, with simple, actionable tips today.

    Watch this on-demand demo series and learn how to protect your Microsoft 365 data with clear, simple, actionable steps that are easy to implement for businesses of all sizes.
    Watch Demo Series
  • Custom VMs From 1 to 96 vCPUs With 99.95% Uptime Icon
    Custom VMs From 1 to 96 vCPUs With 99.95% Uptime

    General-purpose, compute-optimized, or GPU/TPU-accelerated. Built to your exact specs.

    Live migration and automatic failover keep workloads online through maintenance. One free e2-micro VM every month.
    Start Free
  • 1
    SentenceTransformers

    SentenceTransformers

    Multilingual sentence & image embeddings with BERT

    SentenceTransformers is a Python framework for state-of-the-art sentence, text and image embeddings. The initial work is described in our paper Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. You can use this framework to compute sentence / text embeddings for more than 100 languages. These embeddings can then be compared e.g. with cosine-similarity to find sentences with a similar meaning.
    Downloads: 1 This Week
    Last Update:
    See Project
  • 2
    amrlib

    amrlib

    A python library that makes AMR parsing, generation and visualization

    A python library that makes AMR parsing, generation and visualization simple. amrlib is a python module designed to make processing for Abstract Meaning Representation (AMR) simple by providing the following functions. Sentence to Graph (StoG) parsing to create AMR graphs from English sentences. Graph to Sentence (GtoS) generation for turning AMR graphs into English sentences. A QT-based GUI to facilitate the conversion of sentences to graphs and back to sentences. Methods to plot AMR graphs in both the GUI and as library functions. Training and test code for both the StoG and GtoS models. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 3
    SimpleEnglish

    SimpleEnglish

    Agent skill: make LLMs write docs in ASD-STE100

    ...It replaces vague, promotional, and overly complex AI prose with short, direct, testable statements. The instructions enforce controlled practices such as active voice, simple tenses, consistent terminology, limited sentence length, and one instruction per sentence. It is designed for documentation, error messages, runbooks, incident reports, release notes, prompts, and translation preparation rather than marketing copy. The repository includes a complete skill, reusable system prompts, examples, evaluations, and a compact version for limited context budgets. ...
    Downloads: 13 This Week
    Last Update:
    See Project
  • 4
    SetFit

    SetFit

    Efficient few-shot learning with Sentence Transformers

    SetFit is an efficient and prompt-free framework for few-shot fine-tuning of Sentence Transformers. It achieves high accuracy with little labeled data - for instance, with only 8 labeled examples per class on the Customer Reviews sentiment dataset, SetFit is competitive with fine-tuning RoBERTa Large on the full training set of 3k examples.
    Downloads: 0 This Week
    Last Update:
    See Project
  • Paessler - Monitor Your Whole Network in Minutes Icon
    Paessler - Monitor Your Whole Network in Minutes

    Auto-discovery finds your devices and deploys pre-configured sensors instantly. No project plan required, just visibility from day one.

    Waiting weeks for a monitoring rollout isn't an option when infrastructure doesn't stop running. PRTG's auto-discovery scans your network and suggests from over 200 pre-configured sensor types, so you're watching servers, applications and devices within minutes, not after a multi-week deployment. Enterprise-strength monitoring, without the enterprise complexity. Start your free trial today.
    Start Free 30-Day Trial
  • 5
    BudouX

    BudouX

    Standalone, small, language-neutral

    Standalone. Small. Language-neutral. BudouX is the successor to Budou, the machine learning-powered line break organizer tool. It is standalone. It works with no dependency on third-party word segmenters such as Google cloud natural language API. It is small. It takes only around 15 KB including its machine learning model. It's reasonable to use it even on the client-side. It is language-neutral. You can train a model for any language by feeding a dataset to BudouX’s training...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 6
    BERTopic

    BERTopic

    Leveraging BERT and c-TF-IDF to create easily interpretable topics

    ...However, that takes quite some time and lacks a global representation. Instead, we can visualize the topics that were generated in a way very similar to LDAvis. By default, the main steps for topic modeling with BERTopic are sentence-transformers, UMAP, HDBSCAN, and c-TF-IDF run in sequence.
    Downloads: 2 This Week
    Last Update:
    See Project
  • 7
    No AI slop

    No AI slop

    Removes 20+ patterns of AI slop from any piece of writing

    ...It targets formulaic contrasts, throat-clearing openings, false insight setups, dramatic fragments, vague attribution, inflated importance, and other recognizable habits. The skill improves clarity, grammar, sentence structure, active voice, and the use of concrete details without flattening the writer’s personal voice. It can also analyze a passage for these patterns without claiming to determine whether AI created it. When editing, it makes the minimum effective changes and returns the revised draft with a brief explanation. A built-in evaluation file provides pass-or-fail checks that the skill applies to its own work. ...
    Downloads: 20 This Week
    Last Update:
    See Project
  • 8
    Hazm

    Hazm

    Persian NLP Toolkit

    Hazm is a natural language processing (NLP) library for Persian text, offering various tools for text preprocessing, tokenization, part-of-speech tagging, and more.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 9
    gTTS

    gTTS

    Python library and CLI tool to interface with Google Translate

    ...It lets you send text to the Google Translate TTS endpoint and receive spoken audio back as MP3 data, either written to a file, a file-like object, or standard output. The library is designed to handle long texts, using a speech-specific sentence tokenizer that keeps intonation and punctuation natural while splitting requests into acceptable chunks. It supports customizable text pre-processors, which can correct pronunciations, tweak formatting, or handle domain-specific vocabulary before sending it to the API. gTTS is primarily aimed at developers who want a quick way to add cloud-backed speech to scripts, apps, or pipelines without managing any model weights locally. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • $300 Free Credits to Build on Google Cloud Icon
    $300 Free Credits to Build on Google Cloud

    New customers can spin up VMs, build with AI, and query data at no cost.

    Put your $300 in credit toward real workloads, then keep building with free monthly usage for 20+ products. No commitment and no charge until you upgrade.
    Start Free
  • 10
    VideoCaptioner

    VideoCaptioner

    AI-powered tool for generating, optimizing, and translating subtitles

    VideoCaptioner is an open source AI-powered subtitle processing tool designed to simplify the workflow of creating subtitles for videos. It integrates speech recognition, language processing, and translation technologies to automatically generate and refine subtitles from video or audio sources. VideoCaptioner uses speech-to-text engines such as Whisper variants to transcribe spoken content and convert it into subtitle text with accurate timestamps. After transcription, large language models...
    Downloads: 19 This Week
    Last Update:
    See Project
  • 11
    LocalGPT

    LocalGPT

    Chat with your documents on your local device using GPT models

    ...Its data remains on the user’s machine, making it suitable for confidential or offline workflows. The retrieval system combines semantic similarity, keyword matching, late chunking, contextual enrichment, and sentence-level pruning. A smart router chooses between retrieval-augmented generation and direct model responses for each query. An independent verification pass is intended to improve answer reliability before results are returned. The platform works with Ollama models, offers a browser interface and API, and can run through local or Docker-based setups. ...
    Downloads: 3 This Week
    Last Update:
    See Project
  • 12
    ContextGem

    ContextGem

    ContextGem: Effortless LLM extraction from documents

    ContextGem is an open-source framework designed to simplify the extraction of structured data and insights from documents using large language models (LLMs). It provides a flexible, intuitive API that minimizes boilerplate code, enabling developers to build complex extraction workflows efficiently. ContextGem supports various document formats and integrates with multiple LLM providers, making it a versatile tool for tasks like contract analysis, anomaly detection, and information retrieval.​
    Downloads: 0 This Week
    Last Update:
    See Project
  • 13
    spaCy models

    spaCy models

    Models for the spaCy Natural Language Processing (NLP) library

    spaCy is designed to help you do real work, to build real products, or gather real insights. The library respects your time, and tries to avoid wasting it. It's easy to install, and its API is simple and productive. spaCy excels at large-scale information extraction tasks. It's written from the ground up in carefully memory-managed Cython. If your application needs to process entire web dumps, spaCy is the library you want to be using. Since its release in 2015, spaCy has become an industry...
    Downloads: 3 This Week
    Last Update:
    See Project
  • 14
    spaCy

    spaCy

    Industrial-strength Natural Language Processing (NLP)

    spaCy is a library built on the very latest research for advanced Natural Language Processing (NLP) in Python and Cython. Since its inception it was designed to be used for real world applications-- for building real products and gathering real insights. It comes with pretrained statistical models and word vectors, convolutional neural network models, easy deep learning integration and so much more. spaCy is the fastest syntactic parser in the world according to independent benchmarks, with...
    Downloads: 3 This Week
    Last Update:
    See Project
  • 15
    Sepia

    Sepia

    De-AI writing skill for any Agent Skills-compatible agent

    ...Rather than focusing only on vocabulary substitutions, it targets higher-level narrative structure and document-specific writing habits. Fiction processing emphasizes narrative architecture before sentence-level polishing. Professional modes apply different rules to release notes, technical articles, pull request replies, postmortems, tickets, and other workplace formats. Four workflows are provided: writing from scratch, diagnostic review, minimal refactoring, and complete recreation. The skill follows the standard Agent Skills format and can therefore be installed across many compatible assistants. ...
    Downloads: 1 This Week
    Last Update:
    See Project
  • 16
    model2Vec

    model2Vec

    Fast State-of-the-Art Static Embeddings

    model2vec is an innovative embedding framework that converts large sentence transformer models into compact, high-speed static embedding models while preserving much of their semantic performance. The project focuses on dramatically reducing the computational cost of generating embeddings, achieving significant improvements in speed and model size without requiring large datasets for retraining. By using a distillation-based approach, it can produce lightweight models that run efficiently on CPUs, making it suitable for edge applications and large-scale processing pipelines. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 17
    AutoTrain Advanced

    AutoTrain Advanced

    Faster and easier training and deployments

    ...The project provides a no-code and low-code interface that allows users to train models using custom datasets without needing extensive expertise in machine learning engineering. It supports a wide range of tasks including text classification, sequence-to-sequence modeling, token classification, sentence embedding training, and large language model fine-tuning. The system integrates closely with the Hugging Face ecosystem and allows developers to train models using datasets hosted on the Hugging Face Hub. AutoTrain Advanced can run locally or in cloud environments, making it adaptable to different computational setups. By automating tasks such as model configuration, hyperparameter selection, and training pipelines, the project significantly reduces the technical barrier to building AI systems.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 18
    Infinity

    Infinity

    Low-latency REST API for serving text-embeddings

    Infinity is a high-throughput, low-latency REST API for serving vector embeddings, supporting all sentence-transformer models and frameworks. Infinity is developed under MIT License. Infinity powers inference behind Gradient.ai and other Embedding API providers.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 19
    Large Concept Model

    Large Concept Model

    Language modeling in a sentence representation space

    Large Concept Model is a research codebase centered on concept-centric representation learning at scale, aiming to capture shared structure across many categories and modalities. It organizes training around concepts (rather than just raw labels), encouraging models to understand attributes, relations, and compositional structure that transfer across tasks. The repository provides training loops, data tooling, and evaluation routines to learn and probe these concept embeddings, typically...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 20
    RealtimeTTS

    RealtimeTTS

    Converts text to speech in realtime

    RealtimeTTS is a low-latency text-to-speech library built for real-time applications such as voice chat with LLMs, assistants, and interactive tools. It is designed around a streaming model: you can feed it text incrementally (for example, as an LLM responds) and get audio output almost immediately, which keeps end-to-end latency very low. The library is engine-agnostic and plugs into a wide range of cloud and local TTS systems, including OpenAI, ElevenLabs, Azure, Coqui, Piper, StyleTTS2,...
    Downloads: 1 This Week
    Last Update:
    See Project
  • 21
    TextSight AI Detector API Client

    TextSight AI Detector API Client

    AI detector and AI humanizer API client for Python and JavaScript

    Free, open-source Python and JavaScript client for the TextSight AI detector and AI humanizer API. Detect if text was written by ChatGPT, Claude or Gemini with a sentence-by-sentence AI score, and rewrite AI-sounding text so it reads naturally, in 3 lines of code. A developer-friendly alternative to the GPTZero, Originality.ai and ZeroGPT APIs. Zero dependencies, TypeScript types, automatic retries and a command line tool. Requires a TextSight API key.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 22
    Whisper Batch Transcriber

    Whisper Batch Transcriber

    Unlimited, private and free Speech-To-Text program

    ...Works 100% offline on your computer, privately and locally. ## Usecases: Convert speeches, podcasts, webinars, monologues, storytellings and other audio speech into a formatted .txt file. One sentence per new line. ## Notes: - Its 2GB in size and requires 2-6GB of GPU VRAM too. (basically you need atleast a mid-range gaming PC to use this.) - Its fairly slow to start (10min) and transcribe, this is normal behavior. - Includes a python installer to install Python on your computer so you can directly run the 'whisper_transcriber.py' file like you would an .exe by double-clicking it. ...
    Downloads: 7 This Week
    Last Update:
    See Project
  • 23
    Universal-TTS-Pro

    Universal-TTS-Pro

    Portable, fully offline Text-to-Speech (TTS) supporting 30+ languages

    ... • Max Chunk Length control, Volume settings, and Supertonic 3 phonetic corrections. • Clean, smaller executable build with improved audio export quality (no clipping). • Interactive playback: Click any sentence during speech to instantly skip or resume. • Export formats: WAV or space-efficient, high-fidelity OPUS files.
    Downloads: 6 This Week
    Last Update:
    See Project
  • 24
    CLIP-as-service

    CLIP-as-service

    Embed images and sentences into fixed-length vectors

    ...Horizontally scale up and down multiple CLIP models on single GPU, with automatic load balancing. Easy-to-use. No learning curve, minimalist design on client and server. Intuitive and consistent API for image and sentence embedding. Async client support. Easily switch between gRPC, HTTP, WebSocket protocols with TLS and compression. Smooth integration with neural search ecosystem including Jina and DocArray. Build cross-modal and multi-modal solutions in no time.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 25
    Coqui TTS

    Coqui TTS

    A deep learning toolkit for Text-to-Speech, battle-tested in research

    TTS is a library for advanced Text-to-Speech generation. It's built on the latest research, was designed to achieve the best trade-off among ease-of-training, speed and quality. TTS comes with pre-trained models, tools for measuring dataset quality and is already used in 20+ languages for products and research projects. High-performance Deep Learning models for Text2Speech tasks. Text2Spec models (Tacotron, Tacotron2, Glow-TTS, SpeedySpeech). Speaker Encoder to compute speaker embeddings...
    Downloads: 14 This Week
    Last Update:
    See Project
  • Previous
  • You're on page 1
  • 2
  • 3
  • Next