This is super cool, but unfortunately it also seems super impractical. Models tend to be quite large, so even if a browser can run them, getting them to the browser involves either: 1. Large downloads on every visit to a website. 2. Large downloads and high storage consumption for each website using large models. (150 websites x 800 MB models => 120 GB of storage used) Both of those options seem terrible. I think it…
It's an inherent problem with on-device AI processing, not just in the browser. I think this will only get better when operating systems start to preinstall models and provide an API that browser vendors can use as well. Even then I think cloud hosted models will probably always be far better for most tasks.
Separately, having to rely on preinstallation very likely means stagnating on overly sanitized poorly done official instruction-tunes. With the exception of mixtral7x8, the trend has been the community overtime arrives at finetunes which far eclipse official ones.