DeepSeek's new model sets a template for powerful LLMs that run lean

https://image.theregister.com/262727.jpg?imageId=262727&x=0&y=0&cropw=100&croph=100&panox=0&panoy=0&panow=100&panoh=100&width=1200&height=683

Chinese AI darling DeepSeek unveiled an updated version of its cost-and-latency-optimized Flash model on Thursday, with a new version 4.1 that includes architectural improvements more significant than you would expect in a point release because the changes might open the door to larger, smarter, and less resource-intensive models.

At 763 billion parameters, the point release is more than 2.5x the size of the model it replaces. In fact, the model is larger than the V3 and R1 models that put DeepSeek on the map back in early 2025.

Despite its ginormous parameter count, DeepSeek V4.1 Flash’s memory requirements aren’t nearly as high as you’d expect for a model of its size. Under the hood, DeepSeek's devs have made numerous architectural changes that see the LLM become smarter while dramatically reducing the resources necessary to serve it.

DeepSeek has managed this through two key improvements. First, it made significant changes to...

Copyright of this story solely belongs to www.theregister.com. To see the full text click HERE

Read more