K-EXAONE 2.0 Brings 262K Context to Frontier AI
Overview
K-EXAONE-2.0-750B-A37B is a frontier-scale multilingual language model developed by LGAI-EXAONE featuring 750 billion total parameters with 37 billion active parameters during inference. The model uses a Mixture-of-Experts architecture with 256 total experts and 8 activated experts per token, supports a context window of 262,144 tokens, and covers ten languages: Korean, English, Spanish, German, Japanese, Vietnamese, French, Italian, Polish, and Portuguese. Built through upcycling its predecessor and scaling both depth and width, the model incorporates dual-mode functionality with reasoning and non-reasoning modes, making it suitable for both high-accuracy tasks and latency-sensitive applications. It implements advanced attention mechanisms including global attention without positional embeddings, sliding window attention, and a Multi-Token Prediction layer for speculative decoding. The knowledge cutoff is 2025 Q2. You can serve it using SGLang or vLLM with custom forks that include K-EXAONE-specific optimizations, and it is released under the Apache 2.0 license for broad ecosystem deployment.
Best use...
Copyright of this story solely belongs to hackernoon.com. To see the full text click HERE