novel. free download - SourceForge

Kimi K2

Kimi K2 is the large language model series developed by Moonshot AI

...It was trained on an enormous corpus of over 15.5 trillion tokens to push frontier capabilities in coding, reasoning, and general agentic tasks while addressing training stability through novel optimizer and architecture design strategies. The model family includes variants like a foundational base model that researchers can fine-tune for specific use cases and an instruct-optimized variant primed for general-purpose chat and agent-style interactions, offering flexibility for both experimentation and deployment. With its high-dimensional attention mechanisms and expert routing, Kimi-K2 excels across benchmarks in live coding, math reasoning, and problem solving.

Downloads: 55 This Week

Last Update: 2026-01-27

See Project

SimpleLLM

950 line, minimal, extensible LLM inference engine built from scratch

SimpleLLM is a minimal, extensible large language model inference engine implemented in roughly 950 lines of code, built from scratch to serve both as a learning tool and a research platform for novel inference techniques. It provides the core components of an LLM runtime—such as tokenization, batching, and asynchronous execution—without the abstraction overhead of more complex engines, making it easier for developers and researchers to understand and modify. Designed to run efficiently on high-end GPUs like NVIDIA H100 with support for models such as OpenAI/gpt-oss-120b, Simple-LLM implements continuous batching and event-driven inference loops to maximize hardware utilization and throughput. ...

Downloads: 2 This Week

Last Update: 2026-01-28

See Project

Coconut

Training Large Language Model to Reason in a Continuous Latent Space

Coconut is the official PyTorch implementation of the research paper “Training Large Language Models to Reason in a Continuous Latent Space.” The framework introduces a novel method for enhancing large language models (LLMs) with continuous latent reasoning steps, enabling them to generate and refine reasoning chains within a learned latent space rather than relying solely on discrete symbolic reasoning. It supports training across multiple reasoning paradigms—including standard Chain-of-Thought (CoT), no-thought, and hybrid configurations—using configurable training stages and latent representations. ...

Downloads: 2 This Week

Last Update: 11 hours ago

See Project

Qwen-Image

Qwen-Image is a powerful image generation foundation model

...Qwen-Image supports sophisticated editing tasks such as style transfer, object insertion and removal, detail enhancement, and even human pose manipulation, making it suitable for both professional and casual users. It also includes advanced image understanding capabilities like object detection, semantic segmentation, depth and edge estimation, and novel view synthesis.

1 Review

Downloads: 11 This Week

Last Update: 2026-02-10

See Project

rwkv.cpp

INT4/INT5/INT8 and FP16 inference on CPU for RWKV language model

Besides the usual FP32, it supports FP16, quantized INT4, INT5 and INT8 inference. This project is focused on CPU, but cuBLAS is also supported. RWKV is a novel large language model architecture, with the largest model in the family having 14B parameters. In contrast to Transformer with O(n^2) attention, RWKV requires only state from the previous step to calculate logits. This makes RWKV very CPU-friendly on large context lengths.

Downloads: 0 This Week

Last Update: 2025-03-23

See Project

MoBA

MoBA: Mixture of Block Attention for Long-Context LLMs

MoBA, short for Mixture of Block Attention, is an open-source research implementation of a novel attention mechanism designed to improve the efficiency of large language models processing extremely long contexts. The architecture adapts ideas from Mixture-of-Experts networks and applies them directly to the attention mechanism of transformer models. Instead of forcing each token to attend to every other token in the sequence, MoBA divides the context into blocks and dynamically routes queries to only the most relevant segments of information. ...

Downloads: 1 This Week

Last Update: 2026-03-06

See Project

Magicoder

Empowering Code Generation with OSS-Instruct

Magicoder is an open-source family of large language models designed specifically for code generation and software development tasks. The project focuses on improving the quality and diversity of code generation by training models with a novel dataset construction approach known as OSS-Instruct. This technique uses open-source code repositories as a foundation for generating more realistic and diverse instruction datasets for training language models. By grounding training data in real open-source examples, Magicoder aims to reduce bias and improve the reliability of code generation results compared to models trained solely on synthetic instructions. ...

Downloads: 1 This Week

Last Update: 2026-03-06

See Project

Qwen2.5-Omni

Capable of understanding text, audio, vision, video

Qwen2.5-Omni is an end-to-end multimodal flagship model in the Qwen series by Alibaba Cloud, designed to process multiple modalities (text, images, audio, video) and generate responses both as text and natural speech in streaming real-time. It supports “Thinker-Talker” architecture, and introduces innovations for aligning modalities over time (for example synchronizing video/audio), robust speech generation, and low-VRAM/quantized versions to make usage more accessible. It holds...

Downloads: 0 This Week

Last Update: 2025-09-23

See Project

Graph of Thoughts

Official Implementation of "Graph of Thoughts

Graph of Thoughts is an open-source framework that implements a novel reasoning paradigm for large language models by organizing reasoning steps as a structured graph instead of a simple linear chain. Traditional reasoning methods such as chain-of-thought generate sequential reasoning steps, but Graph of Thoughts introduces a more flexible structure where multiple reasoning paths can be explored and evaluated simultaneously.

Downloads: 0 This Week

Last Update: 2026-03-05

See Project

GPT-NeoX

Implementation of model parallel autoregressive transformers on GPUs

This repository records EleutherAI's library for training large-scale language models on GPUs. Our current framework is based on NVIDIA's Megatron Language Model and has been augmented with techniques from DeepSpeed as well as some novel optimizations. We aim to make this repo a centralized and accessible place to gather techniques for training large-scale autoregressive language models, and accelerate research into large-scale training. For those looking for a TPU-centric codebase, we recommend Mesh Transformer JAX. If you are not looking to train models with billions of parameters from scratch, this is likely the wrong library to use. ...

Downloads: 3 This Week

Last Update: 2023-03-23

See Project

Emb-GAM

An interpretable and efficient predictor using pre-trained models

...The final model (which we call Emb-GAM) is a transparent, linear function of its input features and feature interactions. Leveraging the language model allows Emb-GAM to learn far fewer linear coefficients, model larger interactions, and generalize well to novel inputs. Across a variety of natural-language-processing datasets, Emb-GAM achieves strong prediction performance without sacrificing interpretability.

Downloads: 0 This Week

Last Update: 2023-03-22

See Project

Search Results for "novel."

Showing 11 open source projects for "novel."

Kimi K2

SimpleLLM

Coconut

Qwen-Image

rwkv.cpp

MoBA

Magicoder

Qwen2.5-Omni

Graph of Thoughts

GPT-NeoX

Emb-GAM

Search Results for "novel."

Showing 11 open source projects for "novel."

Kimi K2

SimpleLLM

Coconut

Qwen-Image

rwkv.cpp

MoBA

Magicoder

Qwen2.5-Omni

Graph of Thoughts

GPT-NeoX

Emb-GAM

Related Searches

Related Categories