Download Latest Version maxtext-v0.2.4 source code.zip (43.5 MB) Google Add to Preferred Sources
Home / maxtext-v0.2.4
Name Modified Size InfoDownloads / Week
Parent folder
maxtext-v0.2.4 source code.tar.gz 2026-08-21 42.3 MB
maxtext-v0.2.4 source code.zip 2026-08-21 43.5 MB
README.md 2026-08-21 9.6 kB
Totals: 3 Items   85.8 MB 0

Changes

  • Flax NNX Migration: Enabled pure_nnx, enable_nnx, and pure_nnx_decoder configurations by default (PR [#3526]), migrating MaxText primarily on Flax NNX (PR [#2885]).

  • Dependency Upgrades: Upgraded JAX to version 0.10.2 for pre-training and 0.11.0 for post-training.

  • Model Support & Architecture:

  • DeepSeek-V4: Full model integration, decoders, and configuration stack (PR [#4153]), added HyperHead, aligned Sinkhorn implementation (PR [#4337]), and added checkpoint conversion support (PR [#4336]). See the user guide for more details.

  • Qwen3-VL: Added support for Qwen3-VL models (PR [#4293], PR [#4517]) and Qwen3-VL-4B (PR [#4263]).
  • Apple Envy MoE: Added model configurations and support for Apple Envy Switch architectures.
  • Chunked MoE: Added chunked MoE support via num_moe_token_chunks to reduce memory footprint (PR [#4499]).
  • Block Diffusion: Added block-diffusion pre-training support (PR [#4776]), model-independent block corruption utilities (PR [#4737]), and causal-block attention across Dense, Splash, and Tokamax kernels (PR [#4743]).

  • LoRA & QLoRA: Added native LoRA and QLoRA support for Gemma4, Gemma3, Qwen3, and Llama3, along with interactive tutorials (PR [#3969], PR [#4265], PR [#4068], PR [#3968], PR [#3970], PR [#4417]).

  • Context Parallelism (CP), Ring Attention:

  • Added Ulysses and USP CP strategy and packing (PR [#4687], PR [#4825], PR [#4836]), Tokamax load-balanced Ring Attention (PR [#4266], PR [#4537], PR [#4622]), and sequence packing for USP and All-Gather CP (PR [#4230], PR [#4887]).

  • DeepSeek MoE & MLA: Added Ring Attention with DSA Sparse Indexer PR [#4767], auxiliary loss-free and sequence-wise load balancing PR [#4753], MLA QK head chunking PR [#4564], optimized generate_mask PR [#4437], and Approximate Top-K PR [#4243].
  • Positional Embeddings: Added YaRN RoPE config PR [#4238], standardized MRoPE to BS3 convention for multimodal training PR [#4709], and fixed Qwen3.5 partial rotary factor handling.
  • Kernels & Megacore: Added configurable attention_for_vit kernels PR [#4232] and enabled Megacore for Splash Attention dkv backward PR [#4755].

  • Quantization & Performance: Added FP4 [E2M1] (PR [#4495]) and experimental attention quantization (PR [#4487]); enabled TE Collective GEMMs (PR [#4470]) and overlap (PR [#4307]), MoE comms with collective matmul (PR [#4295]), Tokamax GMM v2 (MoE configuration guide), and double-buffered inner scans during gradient accumulation (PR [#4316]).

  • Checkpointing: Added support for Multi-tier checkpointing in Pathways.

  • Goodput & Elasticity:

  • Added Goodput support for Pathways Elasticity & Slice Efficiency, including record_slice_state() to query live slice counts (PR [#4840]).

  • Implemented checkpoint-based elasticity using set-based slice tracking (PR [#4245]).

  • Post Training:

  • Added reward_functions_path and reward_functions CLI knobs for custom rewards (PR [#4149]) to RL training.

  • Updated tutorials with AgenticGRPOLearner for async RL training (PR [#4181]) and added GRPO Gemma4-e4b tutorial (PR [#4427]).
  • Added RL support for Qwen3 30B and GPT-OSS 20B. See the Qwen3 30B RL tutorial and GPT-OSS 20B RL tutorial for recipes.
  • Added support for DPO along with tutorials (PR [#4362]).

  • Usability & Infrastructure:

  • Added wandb logging support (PR [#3053]).

  • Added Hugging Face Grain streaming integration and onboarding guide (PR [#4486]).
  • Added Simple-evals runner support for gpt-oss model family (PR [#4644]).
  • Added scripts to run vanilla DiLoCo on MaxText (PR [#4095]).
  • Added option to enable on-demand profiling server in ML Diagnostics (PR [#4131]).

Bug Fixes

  • Post-Training:

  • Resolved Gemma 3/4 RL rollout gibberish issue by unrolling scanned weights for vLLM adapter (PR [#4536], PR [#4519], PR [#4404]).

  • Fixed RL LR schedule defaults (PR [#4225]), added drop_remainder=True to prevent shape mismatches on tail batches during GRPO training (PR [#4252]) and resolved Qwen3.5 MRoPE/Kv-cache rollout issues (PR [#4177]).

  • Compilation:

  • Fixed double-compilation in train_step by matching input sharding (PR [#4174]).

  • Truncated out_sharding on extra pspec dimensions (PR [#4769]) and restricted GMM quantization to fp8_full (PR [#4842]).

  • Model-Specific Fixes:

  • Qwen3.5: Applied partial MRoPE for Qwen3.5 (PR [#4764]).

  • Mixtral: Fixed EP throughput via configurable expert-axis batch sharding (PR [#4179]).

  • NNX, MoE & MTP:

  • Resolved silent zero-loss (PR [#4525]) and targets_segmentation bugs (PR [#4756]) in Multi-Token Prediction (MTP).

  • Preserved scanned layer intermediates for MoE load-balancing loss in NNX (PR [#4829]).
  • Relanded Qwix quantization on NNX (PR [#4198]) and fixed Qwix LoRA mesh sharding (PR [#4866]).

Deprecations

  • Tensor Transpose Parallelism Removed: Completely removed the tensor_transpose physical mesh axis and deleted ici_tensor_transpose_parallelism and dcn_tensor_transpose_parallelism configuration options.
  • Flax Linen Deprecation Warning: Flax Linen is now deprecated in favor of Flax NNX; running with pure_nnx=False or enable_nnx=False will issue a deprecation warning.
Source: README.md, updated 2026-08-21