Redundancy due to cut-paste operations in text creates bias in machine learning for NLP.
This module takes a directory and produces a subset of the files in that directory (in a list) with an upper bound on similarity between two files.
Features
- Identify copy paste redundancy in a document corpus
- Input: a folder with text documents and similarity threshold
- Output (a) a list of non-redundant documents (a non-redundant subset of the corpus)
- Output (b) list of document pairs found to be redundant with the amount of redundancy for the pair
- Python script (2.6) - tested on various Linux flavours + Windows XP/7
License
GNU General Public License version 3.0 (GPLv3)Follow Corpus redundancy manager
Other Useful Business Software
Stop Storing Third-Party Tokens in Your Database
Rolling your own OAuth token storage can be a security liability. Token Vault securely stores access and refresh tokens from federated providers and handles exchange and renewal automatically. Connected accounts, refresh exchange, and privileged worker flows included.
Rate This Project
Login To Rate This Project
User Reviews
Be the first to post a review of Corpus redundancy manager!