WebCorpus is a Hadoop-based framework that enables you to calculate statistics on large web corpora extracted from web crawls.
Features
- linguistic processing of text corpora with multiple GB or TB in size using Apache Hadoop
- extracts and counts sentences, word n-grams (with or without POS-tags) and cooccurrences
- reads popular web crawl formats (ARC and WARC)
- filters input data by language, duplicate URL, duplicate content and encoding errors
- can be extended by further linguistic counts based on custom UIMA annotations
License
Apache License V2.0Follow WebCorpus
Other Useful Business Software
Build Data Resilience - Take the Assessment Today
Is your recovery strategy as strong as you think? Take this quick self-assessment to check your recovery readiness and gain tailored insights. In only 2 minutes, you'll learn where you fall on the recovery readiness scale.
Rate This Project
Login To Rate This Project
User Reviews
Be the first to post a review of WebCorpus!