A powerful tool for creating datasets for LLM fine-tuning
Framework to easily create LLM powered bots over any dataset
A dataset consists of 15,140 ChatGPT prompts from Reddit
Experimental notebook pipeline for creating language models
Synthetic data curation for post-training and data extraction
Driving with Graph Visual Question Answering
Open-source evaluation toolkit of large multi-modality models (LMMs)
NBA sports betting using machine learning
Code for the paper "Evaluating Large Language Models Trained on Code"
Scalable data pre processing and curation toolkit for LLMs
Unleashing 10,000+ Word Generation from Long Context LLMs
Empowering Code Generation with OSS-Instruct
Retrieval and Retrieval-augmented LLMs
Unifying 3D Mesh Generation with Language Models
Learn to build your Second Brain AI assistant with LLMs
Data Lake for Deep Learning. Build, manage, and query datasets
Training Large Language Model to Reason in a Continuous Latent Space
A straightforward method for training your LLM
Mainly record the knowledge and interview questions
State-of-the-art Parameter-Efficient Fine-Tuning
Framework and no-code GUI for fine-tuning LLMs
Pre & Post-training & Dataset & Evaluation & Depoly & RAG
One-stop solution for creating your digital avatar from chat history
Chat with your SQL database
Curated list of datasets and tools for post-training