Similarity is a Java toolkit for calculating similarity scores between text strings. It provides a collection of algorithms for word similarity, phrase similarity, sentence similarity, paragraph similarity, semantic comparison, sentiment tendency, and approximate word discovery. The project is designed to teach and apply natural language similarity methods while keeping the architecture practical and customizable. It includes approaches such as edit distance, cosine similarity, Euclidean distance, Jaccard similarity, Jaro distance, Jaro-Winkler distance, Manhattan distance, SimHash with Hamming distance, and Sørensen-Dice coefficient. It also supports Java dependency integration through Maven or Gradle workflows. It is useful for Chinese NLP projects, search features, duplicate detection, recommendation systems, and text analysis experiments.

Features

  • Java text similarity toolkit
  • Word, phrase, sentence, and paragraph comparison
  • Multiple distance algorithms
  • Sentiment tendency analysis
  • Approximate word discovery
  • Maven and Gradle integration

Project Samples

Project Activity

See All Activity >

Categories

Libraries

License

Apache License V2.0

Follow Similarity

Similarity Web Site

Other Useful Business Software
Build Data Resilience - Take the Assessment Today Icon
Build Data Resilience - Take the Assessment Today

Can you recover when it matters most? Take this quick assessment to identify gaps and build greater recovery confidence.

Is your recovery strategy as strong as you think? Take this quick self-assessment to check your recovery readiness and gain tailored insights. In only 2 minutes, you'll learn where you fall on the recovery readiness scale.
Take the Assessment
Rate This Project
Login To Rate This Project

User Reviews

Be the first to post a review of Similarity!

Additional Project Details

Programming Language

Java

Related Categories

Java Libraries

Registered

2026-06-18