Page 3 | context-shredder free download

MAE (Masked Autoencoders)

PyTorch implementation of MAE

...It trains a Vision Transformer (ViT) by randomly masking a high percentage of image patches (typically 75%) and reconstructing the missing content from the remaining visible patches. This forces the model to learn semantic structure and global context without supervision. The encoder processes only the visible patches, while a lightweight decoder reconstructs the full image—making pretraining computationally efficient. After pretraining, the encoder serves as a powerful backbone for downstream tasks like image classification, segmentation, and detection, achieving top performance with minimal fine-tuning. ...

Downloads: 0 This Week

Last Update: 2025-10-06

See Project

Retrieval-Based Conversational Model

Dual LSTM Encoder for Dialog Response Generation

Retrieval-Based Conversational Model in Tensorflow is a project implementing a retrieval-based conversational model using a dual LSTM encoder architecture in TensorFlow, illustrating how neural networks can be trained to select appropriate responses from a fixed set of candidate replies rather than generate them from scratch. The core idea is to embed both the conversation context and potential replies into vector representations, then score how well each candidate fits the current dialogue, choosing the best match accordingly. Designed to work with datasets like the Ubuntu Dialogue Corpus, this codebase includes data preparation, model training, and evaluation components for building and assessing dialog models that can handle multi-turn conversations.

Downloads: 0 This Week

Last Update: 2026-02-13

See Project

Nemotron 3 Super

Open language model developed by NVIDIA as part of Nemotron-3 family

...The model contains approximately 120 billion parameters, but employs a Mixture-of-Experts architecture that activates only a smaller subset of parameters during inference, improving computational efficiency while maintaining high capability. Its architecture combines Transformer attention layers with Mamba state-space components to balance long-context reasoning, memory efficiency, and high-quality language generation. The model is optimized for building AI agents that must perform complex tasks such as planning, tool usage, coding assistance, and multi-step reasoning.

Downloads: 0 This Week

Last Update: 2026-03-13

See Project

Mistral Small 4

Model that fuses instruct, reasoning and agentic skills

The Mistral Small 4 collection is a set of open-weight large language models developed by Mistral AI that aim to unify multiple capabilities, including instruction following, reasoning, and coding, within a single efficient architecture. These models are part of the broader Mistral Small family, which is designed to deliver strong performance across a wide range of everyday AI tasks while maintaining relatively low latency and efficient deployment requirements. The collection reflects an...

Downloads: 0 This Week

Last Update: 2026-03-17

See Project

Nemotron 3 Nano

LL model providing reasoning and conversational capabilities

...This architecture allows the system to maintain strong reasoning capabilities while improving throughput and reducing the computational cost associated with large context processing. The model is designed as a general-purpose language system capable of handling tasks such as chat interaction, coding assistance, document analysis, and instruction following.

Downloads: 0 This Week

Last Update: 2026-03-13

See Project

DeepSeek-V3.2-Speciale

High-compute ultra-reasoning model surpassing model surpassing GPT-5

DeepSeek-V3.2-Speciale is the high-compute, ultra-reasoning variant of DeepSeek-V3.2, designed specifically to push the boundaries of mathematical, logical, and algorithmic intelligence. It builds on the DeepSeek Sparse Attention (DSA) framework, delivering dramatically improved long-context efficiency while preserving full model quality. Unlike the standard version, Speciale is tuned exclusively for deep reasoning and therefore does not support tool-calling, focusing its full capacity on pure cognitive performance. The model uses a scaled reinforcement learning framework that allows it to surpass GPT-5 in several evaluations and reach reasoning performance comparable to Gemini-3.0-Pro. ...

Downloads: 0 This Week

Last Update: 2025-12-01

See Project

DeepSeek-V3.2

High-efficiency reasoning and agentic intelligence model

DeepSeek-V3.2 is a cutting-edge large language model developed by DeepSeek-AI, focused on achieving high reasoning accuracy and computational efficiency for agentic tasks. It introduces DeepSeek Sparse Attention (DSA), a new attention mechanism that dramatically reduces computational overhead while maintaining strong long-context performance. Built with a scalable reinforcement learning framework, it reaches near-GPT-5 levels of reasoning and outperforms comparable models like DeepSeek-V3.1 and Gemini-3.0-Pro in advanced benchmarks. The model was notably used in competitive AI challenges such as the 2025 International Mathematical Olympiad (IMO) and IOI, achieving top-tier results. ...

Downloads: 0 This Week

Last Update: 2025-12-01

See Project

Mellum-4b-base

JetBrains’ 4B parameter code model for completions

...Built with 4 billion parameters and a LLaMA-style architecture, it was trained on over 4.2 trillion tokens across multiple programming languages, including datasets such as The Stack, StarCoder, and CommitPack. With a context window of 8,192 tokens, it excels at code completion, fill-in-the-middle tasks, and intelligent code suggestions for professional developer tools and IDEs. The model is efficient for both cloud inference with vLLM and local deployment using llama.cpp or Ollama, thanks to its bf16 precision and AMP training. While the base model is not fine-tuned for downstream tasks, it is designed to be easily adapted through supervised fine-tuning (SFT) or reinforcement learning (RL). ...

Downloads: 0 This Week

Last Update: 2025-09-11

See Project

Search Results for "context-shredder" - Page 3

Showing 58 open source projects for "context-shredder"

MAE (Masked Autoencoders)

Retrieval-Based Conversational Model

Nemotron 3 Super

Mistral Small 4

Nemotron 3 Nano

DeepSeek-V3.2-Speciale

DeepSeek-V3.2

Mellum-4b-base

Search Results for "context-shredder" - Page 3

Showing 58 open source projects for "context-shredder"

MAE (Masked Autoencoders)

Retrieval-Based Conversational Model

Nemotron 3 Super

Mistral Small 4

Nemotron 3 Nano

DeepSeek-V3.2-Speciale

DeepSeek-V3.2

Mellum-4b-base

Related Categories