Redundancy due to cut-paste operations in text creates bias in machine learning for NLP.
This module takes a directory and produces a subset of the files in that directory (in a list) with an upper bound on similarity between two files.

Features

  • Identify copy paste redundancy in a document corpus
  • Input: a folder with text documents and similarity threshold
  • Output (a) a list of non-redundant documents (a non-redundant subset of the corpus)
  • Output (b) list of document pairs found to be redundant with the amount of redundancy for the pair
  • Python script (2.6) - tested on various Linux flavours + Windows XP/7

Project Activity

See All Activity >

License

GNU General Public License version 3.0 (GPLv3)

Follow Corpus redundancy manager

Corpus redundancy manager Web Site

Other Useful Business Software
Sales CRM and Pipeline Management Software | Pipedrive Icon
Sales CRM and Pipeline Management Software | Pipedrive

The easy and effective CRM for closing deals

Pipedrive’s simple interface empowers salespeople to streamline workflows and unite sales tasks in one workspace. Unlock instant sales insights with Pipedrive’s visual sales pipeline and fine-tune your strategy with robust reporting features and a personalized AI Sales Assistant.
Try it for free
Rate This Project
Login To Rate This Project

User Reviews

Be the first to post a review of Corpus redundancy manager!

Additional Project Details

Intended Audience

Science/Research

User Interface

Console/Terminal

Programming Language

Python

Related Categories

Python Linguistics Software, Python Natural Language Processing (NLP) Tool

Registered

2011-05-09