Stream-oriented Java library and a set of command line tools for high quality sentence boundary detection. (Sentence segmentation / splitting / disambiguation). Currently has one model for German (trained on general text and Wikipedia lynx dumps).
Features
- model for German (trained on general text and wikipedia lynx dumps)
- highly accurate
- handles a wide range of potential boundaries
- can cope with headlines, lists, tables
- preserves whitespace
- stream-oriented
License
GNU General Public License version 3.0 (GPLv3)Follow Sentrick
Other Useful Business Software
Stop Cyber Threats with VM-Series Next-Gen Firewall on Azure
Gain integrated visibility across all traffic in a single pass. Deploy Palo Alto Networks VM-Series to determine application identity and content while automating security policy updates via rich APIs.
Rate This Project
Login To Rate This Project
User Reviews
Be the first to post a review of Sentrick!