Dawarich is a command-line tool (likely Ruby-based) for transforming and analyzing Arabic text data with normalization, diacritic handling, segmentation, and morphological tokenization. Designed for text mining and NLP workflows in Arabic-language contexts.
Features
- Normalizes Arabic script variants and punctuation
- Removes or processes diacritics for text standardization
- Tokenization and segmentation suited to Arabic morphology
- Supports stop word removal and light stemming
- Command‑line interface for batch NLP preprocessing
- Output formats compatibility: plain text, CSV/JSON
Categories
MappingLicense
Affero GNU Public LicenseFollow Dawarich
Other Useful Business Software
Build Securely on Azure with Proven Frameworks
Moving to the cloud brings new challenges. How can you manage a larger attack surface while ensuring great network performance? Turn to Fortinet’s Tested Reference Architectures, blueprints for designing and securing cloud environments built by cybersecurity experts. Learn more and explore use cases in this white paper.
Rate This Project
Login To Rate This Project
User Reviews
Be the first to post a review of Dawarich!