kernel-ntfs free download

FlashInfer

FlashInfer: Kernel Library for LLM Serving

FlashInfer is a kernel library designed to enhance the serving of Large Language Models (LLMs) by optimizing inference performance. It provides a high-performance framework that integrates seamlessly with existing systems, aiming to reduce latency and improve efficiency in LLM deployments. FlashInfer supports various hardware architectures and is built to scale with the demands of production environments.

Downloads: 17 This Week

Last Update: 2 days ago

See Project

RWKV Runner

A RWKV management and startup tool, full automation, only 8MB

...So it's combining the best of RNN and transformer - great performance, fast inference, fast training, saves VRAM, "infinite" ctxlen, and free text embedding. Moreover it's 100% attention-free. Default configs has enabled custom CUDA kernel acceleration, which is much faster and consumes much less VRAM. If you encounter possible compatibility issues, go to the Configs page and turn off Use Custom CUDA kernel to Accelerate.

Downloads: 5 This Week

Last Update: 2026-02-01

See Project

TensorRT

C++ library for high performance inference on NVIDIA GPUs

NVIDIA® TensorRT™ is an SDK for high-performance deep learning inference. It includes a deep learning inference optimizer and runtime that delivers low latency and high throughput for deep learning inference applications. TensorRT-based applications perform up to 40X faster than CPU-only platforms during inference. With TensorRT, you can optimize neural network models trained in all major frameworks, calibrate for lower precision with high accuracy, and deploy to hyperscale data centers,...

Downloads: 17 This Week

Last Update: 2026-03-25

See Project

MACE

Deep learning inference framework optimized for mobile platforms

...Chip-dependent power options like big.LITTLE scheduling, Adreno GPU hints are included as advanced APIs. UI responsiveness guarantee is sometimes obligatory when running a model. Mechanism like automatically breaking OpenCL kernel into small units is introduced to allow better preemption for the UI rendering task. Graph level memory allocation optimization and buffer reuse are supported. The core library tries to keep minimum external dependencies to keep the library footprint small.

Downloads: 0 This Week

Last Update: 2022-01-13

See Project

TurboTransformers

Fast and user-friendly runtime for transformer inference

TurboTransformers is a high-performance inference framework optimized for running Transformer models efficiently on CPUs and GPUs. It improves latency and throughput for NLP applications.

Downloads: 0 This Week

Last Update: 2025-01-24

See Project

Search Results for "kernel-ntfs"

Showing 5 open source projects for "kernel-ntfs"

FlashInfer

RWKV Runner

TensorRT

MACE

TurboTransformers

Search Results for "kernel-ntfs"

Showing 5 open source projects for "kernel-ntfs"

FlashInfer

RWKV Runner

TensorRT

MACE

TurboTransformers

Related Searches

Related Categories