DeepSeek V4.1-Flash Cuts Memory and API Costs

https://assets.techrepublic.com/uploads/2026/09/solen-feyissa-sdreva1fKbc-unsplash-1.jpg?f=jpeg

DeepSeek launches V4.1-Flash. Image: Solen Feyissa/Unsplash

DeepSeek V4.1-Flash promises lower memory use and API costs, but buyers should test its performance, compatibility and total deployment expenses.

Sep 11, 2026

DeepSeek says its latest model lowers memory requirements and API costs while outperforming V4 Pro on several internal benchmarks.

Chinese AI company DeepSeek launched V4.1-Flash on Thursday as the smallest model in its new V4.1 architecture family, combining native visual understanding with an architecture designed to improve speed, throughput and serving costs.

The model has 552 billion total parameters in a mixture-of-experts (MoE) system, but DeepSeek says it activates approximately 8 billion parameters per input token and 16 billion per output token. DeepSeek says its new Causal Encoder-Decoder architecture, combined with new pretraining methods and larger-scale reinforcement learning, allows the model to deliver stronger results without using the full model for every request.

DeepSeekalso gives V4.1-Flash a one-million-token context window,...

Copyright of this story solely belongs to www.techrepublic.com. To see the full text click HERE

Read more