data processing free download

Chinese-LLaMA-Alpaca 2

Chinese LLaMA-2 & Alpaca-2 Large Model Phase II Project

This project is developed based on the commercially available large model Llama-2 released by Meta. It is the second phase of the Chinese LLaMA&Alpaca large model project. The Chinese LLaMA-2 base model and the Alpaca-2 instruction fine-tuning large model are open-sourced. These models expand and optimize the Chinese vocabulary on the basis of the original Llama-2, use large-scale Chinese data for incremental pre-training, and further improve the basic semantics and command understanding of...

Downloads: 0 This Week

Last Update: 2024-01-23

See Project

funNLP

Resources, corpora, and tools for Chinese natural language processing

...The project is highly community-oriented, frequently updated with contributions and new resources, and it’s widely used in both academic and applied NLP research. Its value lies in providing not just tools but also curated, domain-specific data, which can be hard to find elsewhere.

Downloads: 0 This Week

Last Update: 2025-10-01

See Project

Common Resource Grep - crgrep

Common Resource Grep

CRGREP searches for matching text in databases, various document formats, archives and other difficult to access resources. A command line tool for name and content text matching in database tables, plain files, MS Office documents, PDF, archives, MP3 audio, image meta-data, scanned documents, maven dependencies and web resources. CRGREP will search resources within resources of any arbitrary combination or depth, so text within a document within a zip archive, and so on. Here you...

3 Reviews

Downloads: 6 This Week

Last Update: 2023-04-23

See Project

XLM (Cross-lingual Language Model)

PyTorch original implementation of Cross-lingual Language Model

XLM (Cross-lingual Language Model) is a family of multilingual pretraining methods that align representations across languages to enable strong zero-shot transfer. It popularized objectives like Masked Language Modeling (MLM) across many languages and Translation Language Modeling (TLM) that jointly trains on parallel sentence pairs to tighten cross-lingual alignment. Using a shared subword vocabulary, XLM learns language-agnostic features that work well for classification and sequence...

Downloads: 0 This Week

Last Update: 2025-10-07

See Project

cocoNLP

A Chinese information extraction tool

...Its API is intentionally simple, so you can drop it into scripts, ETL jobs, or dashboards without deep ML expertise. Because it aims at utility over complexity, it’s useful for prototyping data products or building lightweight text analytics where large models would be overkill. The repository also includes examples and test snippets to help you understand expected inputs and typical outputs, which shortens the learning curve for newcomers.

Downloads: 0 This Week

Last Update: 2025-11-05

See Project

Hermes Natural Language Processing

A repository of software, documentation and data for NLP

Hermes is a repository of software, documentation and data for NLP. I am currently adding corpora extracted from Wikipedia (mostrly in Romance languages).

Downloads: 3 This Week

Last Update: 2013-04-26

See Project

CRFSharp

CRFSharp is a .NET(C#) implementation of Conditional Random Field

CRFSharp(aka CRF#) is a .NET(C#) implementation of Conditional Random Fields, an machine learning algorithm for learning from labeled sequences of examples. It is widely used in Natural Language Process (NLP) tasks, for example: word breaker, postagging, named entity recognized, query chunking and so on. CRF#'s mainly algorithm is the same as CRF++ written by Taku Kudo. It encodes model parameters by L-BFGS. Moreover, it has many significant improvement than CRF++, such as totally...

Downloads: 0 This Week

Last Update: 2015-08-03

See Project

Sanchay

Sanchay is a collection of tools and APIs for language researchers. It has some implementations of NLP algorithms, some flexible APIs, several user friendly annotation interfaces and Sanchay Query Language for language resources.

Downloads: 0 This Week

Last Update: 2013-04-11

See Project

Semantic Weblog Monitoring Framework

Facilitates data mining/natural language processing experiments to be executed on weblogs, such as classification, clustering and rating. As part of these experiments, it is possible to apply Latent Semantic Analysis.

Downloads: 0 This Week

Last Update: 2014-03-29

See Project

MutationFinder

MutationFinder is a biomedical natural language processing (NLP) system for extracting mentions of point mutations from free text. MutationFinder achieves high performance (99% precision, 81% recall on blind test data) as an information extraction system

Downloads: 1 This Week

Last Update: 2013-03-22

See Project

Suffix Trees for NLP

A Java API for using suffix trees with natural language and an Eclipse/SWT-based GUI for suffix tree visualization using Graphviz.

Downloads: 0 This Week

Last Update: 2013-05-02

See Project

Search Results for "data processing"

11 projects for "data processing" with 2 filters applied:

Chinese-LLaMA-Alpaca 2

funNLP

Common Resource Grep - crgrep

XLM (Cross-lingual Language Model)

cocoNLP

Hermes Natural Language Processing

CRFSharp

Sanchay

Semantic Weblog Monitoring Framework

MutationFinder

Suffix Trees for NLP

Search Results for "data processing"

11 projects for "data processing" with 2 filters applied:

Chinese-LLaMA-Alpaca 2

funNLP

Common Resource Grep - crgrep

XLM (Cross-lingual Language Model)

cocoNLP

Hermes Natural Language Processing

CRFSharp

Sanchay

Semantic Weblog Monitoring Framework

MutationFinder

Suffix Trees for NLP

Related Searches

Related Categories