pc benchmark test free download

BIG-bench

Beyond the Imitation Game collaborative benchmark for measuring

BIG-bench (Beyond the Imitation Game Benchmark) is a large, collaborative benchmark suite designed to probe the capabilities and limitations of large language models across hundreds of diverse tasks. Rather than focusing on a single metric or domain, it aggregates many hand-authored tasks that test reasoning, commonsense, math, linguistics, ethics, and creativity. Tasks are intentionally heterogeneous: some are multiple-choice with exact scoring, others are free-form generation judged by model-based or human evaluation. ...

Downloads: 1 This Week

Last Update: 2025-10-09

See Project

MiniMax-M1

Open-weight, large-scale hybrid-attention reasoning model

MiniMax-M1 is presented as the world’s first open-weight, large-scale hybrid-attention reasoning model, designed to push the frontier of long-context, tool-using, and deeply “thinking” language models. It is built on the MiniMax-Text-01 foundation and keeps the same massive parameter budget, but reworks the attention and training setup for better reasoning and test-time compute scaling. Architecturally, it combines Mixture-of-Experts layers with lightning attention, enabling the model to...

Downloads: 0 This Week

Last Update: 2025-12-01

See Project

$Grade School Math$

Grade School Math

8.5K high quality grade school math problems

...JSONL) for problem + answer pairs, and is used broadly in research to benchmark model performance under “word problem” settings. Issues are tracked (people report incorrect problems, ambiguous statements), and contributions are possible for cleaning or expanding the set.

Downloads: 0 This Week

Last Update: 2025-10-03

See Project

Search Results for "pc benchmark test"

Showing 3 open source projects for "pc benchmark test"

BIG-bench

MiniMax-M1

Grade School Math

Search Results for "pc benchmark test"

Showing 3 open source projects for "pc benchmark test"

BIG-bench

MiniMax-M1

Grade School Math

Related Categories