Idk about you guys but I'd find 0.5t/s useless. Even for long tasks. I'd rather just shell out the money to offload as much as possible to say 2x 4060ti 16gb with tensor parallelisation. Anything but that low token rate. This is the sort of thing I'd expect in 20 years for some cyberpunk esque "turtlebot" that thinks at 0.5t/s, is solar powered and performs some menial civic maintenance background task like cutting g…
It could be handy if it's some batch work you don't need quickly. But that's pretty rare, tbh. It's way too slow for interactive stuff. With this kind of model I like to spar, fire ideas at and get feedback. Having hours of delay every time doesn't work for my ADHD brain.