Hy4 preview
770B MoE model for coding, research, reasoning, and long-context work
...It contains 770B backbone parameters while activating 49B per token across 78 layers, with 256 routed experts and one shared expert in each MoE layer. Its architecture uses Gated DeepSeek Sparse Attention with IndexCache for cross-layer sparse-index reuse and identity Hyper-Connections to improve information flow. A native 10B-parameter Multi-Token Prediction layer enables speculative decoding for faster inference. Hy4-preview supports a native 1M-token context window, allowing it to process large codebases, numerous files, and complex extended workflows. Tencent specifically optimized the model for software engineering, office and financial analysis, game development, and scientific research.