Java Suffix array library for phrase discovery. Inspired initially by the classic paper of Yamamoto & Church, with newer ideas from Abouelhoda et al and Kim et al. Adapted for large alphabet so that words can be tokenized as alphabet characters.

Features

  • Adapted to large alphabet for NLP
  • Includes tokenizers, normalizers and symbol table.
  • Calculates term and document frequency; full distribution across texts
  • Modular design, user can apply various statistics to phrases
  • Includes Aho Corasick automaton, also for large alphabet
  • Needs: Better sorting, though radix qsort usually works okay
  • Needs: improvement to the symbol table. This is the slowest part.
  • Needs: Tokenizers, normalizers for more languages.
  • Needs: Some links to foma (foma.sourceforge.net/)

Project Activity

See All Activity >

License

Apache Software License

Follow suffix arrays for phrase extraction

suffix arrays for phrase extraction Web Site

You Might Also Like
Top-Rated Free CRM Software Icon
Top-Rated Free CRM Software

216,000+ customers in over 135 countries grow their businesses with HubSpot

HubSpot is an AI-powered customer platform with all the software, integrations, and resources you need to connect your marketing, sales, and customer service. HubSpot's connected platform enables you to grow your business faster by focusing on what matters most: your customers.
Rate This Project
Login To Rate This Project

User Reviews

Be the first to post a review of suffix arrays for phrase extraction!

Additional Project Details

Intended Audience

Information Technology

Programming Language

Java

Related Categories

Java Linguistics Software, Java Natural Language Processing (NLP) Tool

Registered

2010-05-02