cc_net provides tools to download, segment, clean, and filter Common Crawl to build large-scale text corpora, including monolingual datasets and the multilingual CC-100 collection introduced in the associated paper. It includes pipelines to fetch snapshots, extract text, de-duplicate, identify language, and apply quality filtering based on heuristics and language models. The outputs are intended for pretraining language models and for creating standardized corpora that can be reproduced or updated with new crawls. The repository documents practical concerns like HTTP failures, snapshot differences, and stats JSONs, reflecting community use across many languages. While powerful, the repo has been archived and is read-only, so users should expect to run it as-is or fork for maintenance. Even in archived state, issues and releases pages remain useful references for implementation details and dataset lineage.

Features

  • End-to-end Common Crawl download and extraction
  • Language identification and monolingual segmentation
  • Quality filtering and de-duplication pipelines
  • Support for building multilingual datasets like CC-100
  • Reproducible statistics and corpus metadata outputs
  • Scripts and configs for snapshot-by-snapshot processing

Project Samples

Project Activity

See All Activity >

License

MIT License

Follow CC-Net

CC-Net Web Site

Other Useful Business Software
Keep company data safe with Chrome Enterprise Icon
Keep company data safe with Chrome Enterprise

Protect your business with AI policies and data loss prevention in the browser

Make AI work your way with Chrome Enterprise. Block unapproved sites and set custom data controls that align with your company's policies.
Download Chrome
Rate This Project
Login To Rate This Project

User Reviews

Be the first to post a review of CC-Net!

Additional Project Details

Operating Systems

Linux, Mac

Programming Language

Python

Related Categories

Python Natural Language Processing (NLP) Tool

Registered

4 days ago