...It organizes data around criminal charges, sentencing cases, legal question-and-answer pairs, and related legal information. A multiclass model predicts likely offense categories from written case descriptions using document embeddings and a support vector machine. Separate classifiers sort consultation questions into predefined legal categories before retrieving or generating relevant responses from the prepared knowledge base. The repository includes training scripts, inference programs, dictionaries, models, and utilities for building the question-answer database. Its published experiments use millions of case records and hundreds of thousands of consultation pairs. ...