The next age of LLMs? Dev gets a small LLM running at 10 tokens a second locally on a $10 microcontroller
- A developer has a 28.9-million-parameter model generating TinyStories-style text at 9.88 tokens per second, fully offline, on a $10 microcontroller
- It was achieved by fitting a language model on a chip with 512KB of RAM by leaving most of it in flash storage
- The project is available on GitHub under an MIT license
A developer going by slvDev has a 28.9-million-parameter language model generating text on an ESP32-S3, a microcontroller built for sensor nodes and smart plugs, at 9.88 tokens per second, with nothing leaving the chip.
The project, esp32-ai, went up on GitHub under an MIT license in late July 2026, was showcased on the Better Stack YouTube channel, and has since collected over 3,600 stars and has more than 470 forks of the underlying code.
Given the underlying hardware's limitations, the attempt meant the model, which was trained on the TinyStories dataset by MicrosoftResearch, had to be...
Copyright of this story solely belongs to techradar.com. To see the full text click HERE