Redundancy due to cut-paste operations in text creates bias in machine learning for NLP.
This module takes a directory and produces a subset of the files in that directory (in a list) with an upper bound on similarity between two files.
Features
- Identify copy paste redundancy in a document corpus
- Input: a folder with text documents and similarity threshold
- Output (a) a list of non-redundant documents (a non-redundant subset of the corpus)
- Output (b) list of document pairs found to be redundant with the amount of redundancy for the pair
- Python script (2.6) - tested on various Linux flavours + Windows XP/7
License
GNU General Public License version 3.0 (GPLv3)Follow Corpus redundancy manager
Other Useful Business Software
$300 Free Credits to Build on Google Cloud
Start your next project with $300 in free Google Cloud credit. Spin up VMs, run containers, query petabytes in BigQuery, or build agents with Gemini Enterprise Agent Platform. Once your credits are used, keep building with 20+ always-free tier products including Compute Engine, Cloud Storage, GKE, and Cloud Run functions. No commitment required—just sign up and start building.
Rate This Project
Login To Rate This Project
User Reviews
Be the first to post a review of Corpus redundancy manager!