Showing 2019 open source projects for "language"

View related business solutions
  • Veeam Data Platform v13.1 - Get Your Free Trial Icon
    Veeam Data Platform v13.1 - Get Your Free Trial

    Secure by design, portable by default. Recover clean, fast, anywhere. Start a free trial.

    Try Veeam Data Platform today. Experience the unified platform that's secure by design, portable by default, and proven to recover clean, fast, and anywhere.
    Try it Free
  • Custom VMs From 1 to 96 vCPUs With 99.95% Uptime Icon
    Custom VMs From 1 to 96 vCPUs With 99.95% Uptime

    General-purpose, compute-optimized, or GPU/TPU-accelerated. Built to your exact specs.

    Live migration and automatic failover keep workloads online through maintenance. One free e2-micro VM every month.
    Start Free
  • 1
    Playwright for Python

    Playwright for Python

    Python version of the Playwright testing and automation library

    Playwright enables reliable end-to-end testing for modern web apps. Single API to automate Chromium, Firefox and WebKit. Capable automation for single page apps that rely on the modern web platform. Use the Playwright API in JavaScript & TypeScript, Python, .NET and, Java. With Playwright, test how your app behaves in Apple Safari with WebKit builds for Windows, Linux and macOS. Test locally and on CI. Use device emulation to test your responsive web apps in mobile web browsers. Playwright...
    Downloads: 1 This Week
    Last Update:
    See Project
  • 2
    RAPTOR

    RAPTOR

    The official implementation of RAPTOR

    RAPTOR is a retrieval architecture designed to improve retrieval-augmented generation systems by organizing documents into hierarchical structures that enable more effective context retrieval. Traditional RAG systems typically retrieve small text chunks independently, which can limit a model’s ability to understand broader document context. RAPTOR addresses this limitation by recursively embedding, clustering, and summarizing documents to create a tree-structured hierarchy of information....
    Downloads: 17 This Week
    Last Update:
    See Project
  • 3
    UniMate

    UniMate

    UniMate: One Unified Model to Animate Diverse Skeletons

    UniMate is a generative motion model that animates rigged 3D assets with different skeletal structures from natural-language prompts. It is designed to handle bipeds, quadrupeds, birds, marine creatures, insects, serpentine rigs, and articulated rigid objects with one unified model. Its topology-aware diffusion transformer conditions generation on skeleton structure rather than requiring a fixed character template. The accompanying UniML3D dataset contains 13,006 text-paired motion sequences with unified skeletal canonicalization. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 4
    AI Infra

    AI Infra

    Understanding of AI Infra: Quantitative Analysis and System Design

    AI Infra Book is an open-source technical book and companion repository focused on the infrastructure behind modern large language models. It approaches inference and training through quantitative analysis of hardware limits, data movement, model architecture, and distributed systems. The book contains twelve chapters supported by formulas, diagrams, experiments, and case studies. Companion tools help readers reproduce resource calculations and inspect the assumptions behind system-design decisions. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • MongoDB Atlas runs apps anywhere Icon
    MongoDB Atlas runs apps anywhere

    Deploy in 115+ regions with the modern database for every enterprise.

    MongoDB Atlas gives you the freedom to build and run modern applications anywhere—across AWS, Azure, and Google Cloud. With global availability in over 115 regions, Atlas lets you deploy close to your users, meet compliance needs, and scale with confidence across any geography.
    Start Free
  • 5
    MSA: Memory Sparse Attention

    MSA: Memory Sparse Attention

    Trainable latent-memory framework for 100M-token contexts

    MSA, or Memory Sparse Attention, is a research framework for scaling language-model memory to extremely long contexts. It replaces full attention over all tokens with sparse selection of compressed latent memory states. Document-wise rotary position encoding and top-k routing keep training and inference close to linear complexity. A tiered KV-cache design stores routing keys on GPU while larger content states can remain on CPU.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 6
    US Job Market Visualizer

    US Job Market Visualizer

    A research tool for visually exploring Bureau of Labor Statistics

    ...The included AI exposure layer is intended as an exploratory estimate rather than a prediction of job elimination. A generated prompt file also packages the complete dataset for data-grounded analysis with language models.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 7
    Human Writing

    Human Writing

    Make the Chinese written by AI read like a specific person speaking

    Human Writing is an AI writing skill designed to make Chinese prose sound like it was written by a specific person rather than a generic language model. Before drafting, it checks whether enough factual or fictional material exists to support meaningful writing. For nonfiction, it emphasizes verified facts, numbers, quotations, and lived details, while fiction must still contain purposeful actions and causal progression. Each paragraph is expected to add new information instead of repeating or paraphrasing earlier points. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 8
    FlashKDA

    FlashKDA

    High-performance Kimi Delta Attention kernels

    FlashKDA is an open-source library of high-performance CUDA kernels for Kimi Delta Attention, implemented on NVIDIA CUTLASS. It is intended to accelerate the forward pass used by KDA-based language models on modern NVIDIA GPUs. The package integrates with flash-linear-attention and can be selected automatically as the backend for chunk_kda during inference. It supports recurrent state input and output, variable-length batches, internal gating, query-key normalization, and beta activation. Builds can target the detected GPU architecture or multiple supported architectures for wheels and CI pipelines. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 9
    jlens

    jlens

    Companion code for the global workspace interpretability paper

    Jacobian Lens is Anthropic’s reference implementation for examining what a language model’s internal activations are inclined to produce as text. It transports residual-stream vectors from selected layers and positions into the final-layer basis using an averaged input-output Jacobian. The transformed vectors are decoded through the model’s own unembedding into ranked vocabulary predictions. The package can fit new lenses, load saved ones, apply them to prompts, and merge results from parallel fitting jobs. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • $300 Free Credits to Build on Google Cloud Icon
    $300 Free Credits to Build on Google Cloud

    New customers can spin up VMs, build with AI, and query data at no cost.

    Put your $300 in credit toward real workloads, then keep building with free monthly usage for 20+ products. No commitment and no charge until you upgrade.
    Start Free
  • 10
    DeepSpec

    DeepSpec

    A full-stack codebase for training and evaluating speculative decoding

    ...The evaluation pipeline measures speculative decoding performance across benchmark tasks such as math, coding, instruction-following, and chat-style datasets. Overall, it is useful for researchers and engineers studying faster language model inference through speculative decoding methods.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 11
    Pycorrector

    Pycorrector

    Pycorrector is a toolkit for text error correction

    ...The repository includes usage examples, evaluation materials, datasets, documentation, and model references. It is useful for NLP engineers, researchers, and application developers building Chinese language quality tools.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 12
    HRM-Text

    HRM-Text

    1B text generation model based on the HRM architecture

    ...The repository supports reference pretraining runs for smaller and larger configurations, with Hopper-class GPUs expected for the attention path. It is useful for researchers and engineers exploring efficient language model pretraining, reasoning-focused architectures, and reproducible foundation model experiments.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 13
    GitHubPoster

    GitHubPoster

    Make everything a GitHub svg poster and Skyline

    ...The project is built around loaders that import data and render it as visually recognizable contribution-style graphics. It is useful for people who want to display habits, reading, coding, health, language learning, or other quantified-life records in a GitHub-inspired format. It can be run locally from the command line and can also be automated through GitHub Actions. Its modular approach makes it possible for contributors to add new data sources over time.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 14
    GEO Content Writer

    GEO Content Writer

    Backlog-row-first content production system for teams

    ...It focuses on producing articles, pages, and structured content that align with both traditional SEO requirements and emerging AI search patterns. The system leverages language models to generate content that is context-aware, location-specific, and optimized for discoverability. It supports automated workflows for generating large volumes of content while maintaining consistency and relevance. The tool is particularly useful for businesses targeting local markets or region-specific audiences. It integrates into broader SEO pipelines, allowing content generation to be part of a continuous optimization process. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 15
    CLI-Anything

    CLI-Anything

    Making ALL Software Agent-Native

    ...The system provides a methodology and tooling for generating CLI wrappers around existing applications, allowing them to be controlled programmatically using natural language instructions interpreted by AI agents. It integrates with multiple AI platforms such as Claude Code, OpenClaw, Codex, and GitHub Copilot CLI, enabling cross-platform compatibility and flexibility. CLI-Anything emphasizes structured outputs such as JSON to reduce parsing complexity and improve reliability in automation scenarios.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 16
    machine_learning_examples

    machine_learning_examples

    A collection of machine learning examples and tutorials

    ...It includes implementations of many machine learning algorithms and neural network architectures using Python and popular libraries such as TensorFlow and NumPy. The repository covers a wide range of topics including supervised learning, unsupervised learning, reinforcement learning, and natural language processing. Many of the examples are accompanied by tutorials and educational materials that explain how the algorithms work and how they can be applied in real-world projects. The code is organized into small independent experiments so that learners can explore specific algorithms or techniques without needing to understand the entire codebase.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 17
    Machine Learning Engineering Open Book

    Machine Learning Engineering Open Book

    Machine Learning Engineering Open Book

    Machine Learning Engineering Open Book is an open “living book” that captures practical methodologies, tooling advice, and operational knowledge for successfully training and deploying large language models and multimodal systems. The repository functions as a field guide compiled from real-world experience, particularly from work on large-scale models such as BLOOM-176B and IDEFICS-80B. It is heavily oriented toward practitioners who need hands-on solutions, including copy-paste commands, infrastructure comparisons, and performance tuning strategies. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 18
    List of independent blogs in Chinese

    List of independent blogs in Chinese

    List of independent blogs in Chinese

    List of independent blogs in Chinese is a curated open repository that aggregates and maintains a large list of independent Chinese-language blogs across technology, design, and personal knowledge domains. The project aims to promote the independent blogging ecosystem by making it easier for readers to discover high-quality personal sites outside major content platforms. It is community-driven, allowing contributors to submit and update blog entries so the directory remains current and diverse. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 19
    Linux insides

    Linux insides

    A book-in-progress about the Linux kernel and its insides

    ...The project’s stated goal is to share knowledge about Linux kernel internals and related low-level concepts in an accessible narrative format. It is written for readers who already have some familiarity with C and assembly language and want to understand what happens under the hood of Linux. The material is continuously updated as the kernel evolves, reflecting changes in modern kernel versions. Overall, linux-insides is widely regarded as a deep technical learning resource for systems programmers and advanced Linux enthusiasts.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 20
    Claude Code Hooks Mastery

    Claude Code Hooks Mastery

    Master Claude Code Hooks

    ...Although the project itself doesn’t include a single coherent application, it functions as a curated collection of advanced hook examples, best practices, and coding patterns that show how to tailor Claude Code to specific use cases such as automated CI workflows, custom command triggers, and integrations with external tools. The repository is part of a larger ecosystem of Claude Code tooling that enables natural-language-driven coding tasks, and the hooks contained here help users go beyond default behaviors to solve real problems efficiently.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 21
    Auto-Deep-Research

    Auto-Deep-Research

    Your Fully-Automated Personal AI Assistant

    Auto-Deep-Research is a system designed to fully automate deep research workflows using language models, retrieval, planning, and multi-stage reasoning to produce structured research artifacts such as surveys, benchmarks, reports, and even prototypes without heavy human intervention. Users provide a research topic or multifaceted goal, and the system autonomously breaks the objective down into subtasks like literature collection, critical summarization, cross-comparison, citation extraction, metric evaluation, and structured writing. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 22
    bu-agent-sdk

    bu-agent-sdk

    An agent is just a for-loop

    The bu-agent-sdk from the Browser Use project is a minimalistic Python framework that defines an AI agent as a simple loop of tool calls, aiming to keep abstractions low so developers can build autonomous agents without unnecessary complexity. At its core, the agent loop repeatedly queries a large language model, interprets its output, and executes defined “tools” — functions annotated with task names — to perform actions, allowing the agent to complete tasks like arithmetic, decision-making, or domain-specific work. The SDK emphasizes simplicity and control, avoiding heavy orchestration frameworks and instead letting developers specify exactly what tools an agent can employ and how it should signal task completion. ...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 23
    Moondream

    Moondream

    Tiny vision language model

    Moondream is a creative code project and visual experimentation repository that explores generative graphics, aesthetic patterns, and interactive art through code. The project typically showcases procedural visualizations, algorithmic designs, and artistic experiments that push the boundaries of what can be expressed with programming languages and rendering frameworks. While the exact nature can vary by commit or branch, Moondream’s work often blends geometry, color theory, and motion to...
    Downloads: 0 This Week
    Last Update:
    See Project
  • 24
    ValueCell

    ValueCell

    Community-driven, multi-agent platform for financial applications

    ...The system brings together a suite of collaborative agents—such as research agents that gather and interpret fundamentals, strategy agents that implement trading logic, and news agents that deliver personalized updates—to help users make more informed financial decisions across stocks, crypto, and other markets. ValueCell supports integrations with multiple language model providers and market data sources, giving developers flexibility in customizing agents and incorporating external APIs to enhance insights. Sensitive user data is stored locally, a design choice that prioritizes privacy and security while still enabling rich analytic workflows.
    Downloads: 0 This Week
    Last Update:
    See Project
  • 25
    TruLens

    TruLens

    Evaluation and Tracking for LLM Experiments

    TruLens is an open-source Python library designed to systematically evaluate and track Large Language Model (LLM) applications. It provides fine-grained instrumentation, feedback functions, and a user interface to compare and iterate on app versions, facilitating rapid development and improvement of LLM-based applications. Programmatic tools that assess the quality of inputs, outputs, and intermediate results from LLM applications, enabling scalable evaluation.
    Downloads: 0 This Week
    Last Update:
    See Project