Bagel: Open-source unified multimodal model
21–30 of 35 posts
Re: Bagel: Open-source unified multimodal model
#22Re: Bagel: Open-source unified multimodal model
#23Is it from ByteDance Team, right? The team behind TikTok, CapCut, BuzzVideo and more. Any thoughts on that?
Re: Bagel: Open-source unified multimodal model
#24Re: Bagel: Open-source unified multimodal model
#25Earlier quoted context omitted.
As someone who used to be in the academia, I think is isn't bad in itself , I just worry that by comparison it raises the burden of effort that one has to make in order to get their work noticed.
Compared to the effort required to play in that field at all , making a video is almost negligible.
Re: Bagel: Open-source unified multimodal model
#26It works following readme instructions at least on Ubuntu, on my RTX 3090 GPU with 24 gigs of memory, just barely. Have to close most other windows and lower screen resolution to be able to load the model. Then it generates or edits images in 2-3 minutes. I only have this one GPU and am using Chrome to use the browser interface on the same machine.
The original release won't run on this hardware, but the compressed one is supposed to give identical results.
Re: Bagel: Open-source unified multimodal model
#27I found a losslessly compressed version: https://github.com/LeanModels/Bagel-DFloat11 It works following readme instructions at least on Ubuntu, on my RTX 3090 GPU with 24 gigs of memory, just barely. Have to close most other windows and lower screen resolution to be able to load the model. Then it generates or edits images in 2-3 minutes. I only have this one GPU and am using Chrome to use the browser interface on t…
Re: Bagel: Open-source unified multimodal model
#28Re: Bagel: Open-source unified multimodal model
#29Has anyone here experimented with fine-tuning this for domain-specific applications?