GLM-4.7-Flash on 2x RTX 3090: My Hands-On Experience
Last time I wrote about GLM-4.7-Flash, I was annoyed. It ran at a third of the speed I expected on my mini PC, fell off a cliff as context grew, and the fix turned out to be a runtime that actually implements Multi-head Latent Attention. If you missed it, the whole diagnosis is in the mini-PC writeup. I ended that one with a promise: fixing a model is not the same as it being worth running. The real question is whether GLM-4.7-Flash earns its place against a boring, plain grouped-query model of the same size that is already fast. So I moved it onto a proper pair of NVIDIA cards and measured everything. This is the hands-on writeup, on 2x RTX 3090.
A quick note on why I care beyond curiosity. I build ModelDirectory, and a real chunk of its backend runs on local hardware rather than rented GPUs....
Copyright of this story solely belongs to hackernoon.com. To see the full text click HERE