Forget Expensive GPUs, This DIY AI Chatbot Cluster Runs On E-Waste
How much hardware do you need to run a large language model (LLM) locally? You need a massive GPU, or a powerful AI-targeted machine like the HP Z2 Mini G1a, right? That's certainly one way to go. Alternatively you could grab a handful of e-waste motherboards, cheap USB Ethernet adapters, and hook them all up in an unsteady 3D printed frame with terrifying hand-wired power delivery and commodity DDR4 memory. That's exactly what JoeC-J on YouTube did, and his local AI cluster can run an 80B sparse model at around 4 tokens per second.
Now, as Joe says, "is 4 tokens per second going to replace ChatGPT? No. As a proof of concept? I'd call it a win." Indeed, the fact that it works at all is quite fascinating. The motherboards he's using are laptop motherboards from Lenovo Thinkpad L380 thin and light systems from 2018. They are not...
Copyright of this story solely belongs to hothardware.com. To see the full text click HERE