Arxiv Sanity Preserver is a web application for discovering, organizing, and filtering machine-learning research papers from arXiv. It was created to make the large daily flow of new submissions easier for researchers to manage. The backend downloads papers, extracts their text, creates thumbnails, and computes TF-IDF representations for similarity analysis. Users can search papers, find work similar to a selected paper, browse popular submissions, and maintain a personal library. Personalized recommendations are generated through models trained from user preferences. The web interface uses Flask, Tornado, and SQLite, while the indexing pipeline relies on scientific Python libraries. A newer rewrite called arxiv-sanity-lite later became the recommended successor.

Features

  • Recent arXiv paper discovery
  • Full-text paper search
  • Paper-to-paper similarity ranking
  • Personal research libraries
  • Personalized paper recommendations
  • Automated PDF indexing and TF-IDF analysis

Project Samples

Project Activity

See All Activity >

License

MIT License

Follow Arxiv Sanity Preserver

Arxiv Sanity Preserver Web Site

Other Useful Business Software
Ship Agents Faster Icon
Ship Agents Faster

Transform your applications and workflows into powerful agentic systems at global scale.

Gemini Enterprise Agent Platform lets you rapidly build, scale, govern and optimize production-ready agents grounded in your organization's data. The platform enables developers to build custom or pre-built agents for virtually any use case. New customers get $300 in free credits.
Start Free
Rate This Project
Login To Rate This Project

User Reviews

Be the first to post a review of Arxiv Sanity Preserver!

Additional Project Details

Programming Language

Python

Related Categories

Python User Interface (UI) Software

Registered

3 hours ago