inference engine free download

Transformer Engine

A library for accelerating Transformer models on NVIDIA GPUs

Transformer Engine (TE) is a library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit floating point (FP8) precision on Hopper GPUs, to provide better performance with lower memory utilization in both training and inference. TE provides a collection of highly optimized building blocks for popular Transformer architectures and an automatic mixed precision-like API that can be used seamlessly with your framework-specific code.

Downloads: 30 This Week

Last Update: 2026-03-31

See Project

X-AnyLabeling

Effortless data labeling with AI support from Segment Anything

X-AnyLabeling is an open-source data annotation platform designed to streamline the process of labeling datasets for computer vision and multimodal AI applications. The software integrates an AI-powered labeling engine that allows users to generate annotations automatically with the assistance of modern vision models such as Segment Anything and various object detection frameworks. It supports labeling tasks across images and videos and enables developers to prepare training datasets for...

Downloads: 58 This Week

Last Update: 2026-03-26

See Project

FlexLLMGen

Running large language models on a single GPU

FlexLLMGen is an open-source inference engine designed to run large language models efficiently on limited hardware resources such as a single GPU. The system focuses on high-throughput generation workloads where large batches of text must be processed quickly, such as large-scale data extraction or document analysis tasks. Instead of requiring expensive multi-GPU systems, the framework uses techniques such as memory offloading, compression, and optimized batching to run large models on commodity hardware. ...

Downloads: 0 This Week

Last Update: 2026-03-10

See Project

Lightning Bolts

Toolbox of models, callbacks, and datasets for AI/ML researchers

Bolts package provides a variety of components to extend PyTorch Lightning, such as callbacks & datasets, for applied research and production. Torch ORT converts your model into an optimized ONNX graph, speeding up training & inference when using NVIDIA or AMD GPUs. We can introduce sparsity during fine-tuning with SparseML, which ultimately allows us to leverage the DeepSparse engine to see performance improvements at inference time.

Downloads: 0 This Week

Last Update: 2024-08-15

See Project

Minkowski Engine

Auto-diff neural network library for high-dimensional sparse tensors

The Minkowski Engine is an auto-differentiation library for sparse tensors. It supports all standard neural network layers such as convolution, pooling, unspooling, and broadcasting operations for sparse tensors. The Minkowski Engine supports various functions that can be built on a sparse tensor. We list a few popular network architectures and applications here.

Downloads: 0 This Week

Last Update: 2022-08-11

See Project

Search Results for "inference engine"

Showing 5 open source projects for "inference engine"

Transformer Engine

X-AnyLabeling

FlexLLMGen

Lightning Bolts

Minkowski Engine

Search Results for "inference engine"

Showing 5 open source projects for "inference engine"

Transformer Engine

X-AnyLabeling

FlexLLMGen

Lightning Bolts

Minkowski Engine

Related Searches

Related Categories