Viewing profile — memossy
memossy
HN member- Joined
- Tue, Oct 08, 2013, 11:11 AM UTC
- HN karma
- 341
- Public activity
- 133 items
- HN profile
- View on Hacker News ↗
About memossy
No profile information was provided.
Recent public activity
-
comment
Comment #39670417
It was interesting Aurora used GPU Max & definitely looking forward to Falcon Shores. I think Gaudi2 was bad timed & they had to build stack, Gaudi3 is where I think we will see ma…
-
comment
Comment #39670400
It's a great company & will do well, plenty of demand & B100s/BH200s etc coming The Hopper stuff is particulalry interesting
-
comment
Comment #39670390
Falcon shores next year will be crazy with 300gb VRAM & new lith
-
comment
Comment #39670376
Gaudi2s started coming out in 2022 ( https://huggingface.co/blog/habana-gaudi-2-benchmark ) but didn't hit mass scale. I think Gaudi3 will & others have seen similar performance fo…
-
comment
Comment #39670337
We use v4s, v5es & v5ps. Mostly v5ps, very stable int8 training (versus the horror that is fp8 stability)
-
comment
Comment #39670329
I mean they work well, here is another blog by Databricks: https://www.databricks.com/blog/llm-training-and-inference-i...
-
comment
Comment #39669694
I think as we go to enterprise workloads the total cost of ownership becomes important. NVIDIA is still the best for research given ecosystem but once the models are standardised a…
-
comment
Comment #39669678
It took less than a day to port our code over, we do custom CUDA across modalities. Gaudi2 was actually announced 2 years ago and is 7nm like the A100 80Gb it was meant to be compe…
-
comment
Comment #39669658
The v5es and v5ps are pretty amazing at running SD, giving code for SD3 now to optimise it on those. v5es are particularly interesting given the millions that will land and the lar…
-
comment
Comment #39669648
Think it'll probably crack on with Gaudi3 at 4x performance, twice VRAM etc later this year. We found cuda sycl conversion surprisingly good https://www.intel.com/content/www/us/en…
-
comment
Comment #39669009
"For Stable Diffusion 3, we measured the training throughput for the 2B Multimodal Diffusion Transformer (MMDiT) architecture model. Gaudi 2 trained images 1.5x faster than the H10…
- story
-
comment
Comment #39468583
I could but I won't as legal stuff :)
-
comment
Comment #39468578
I mean open models yo
-
comment
Comment #39468067
If we trained it with videos yes but need more GPUs for that.
-
comment
Comment #39468053
It'll be out soon, doing benchmark tests etc
-
comment
Comment #39467874
We have highly efficient models for inference and a quantization team. Need moar GPUs to do a video version of this model similar to Sora now they have proved that Diffusion Transf…
-
comment
Comment #39467786
As the leader in open image models it is incumbent upon us as the models get to this level of quality to take seriously how we can release open and safe models from a legal, societ…
-
comment
Comment #39467741
800m is good for mobile, 8b for graphics cards. Bigger than that is also possible, not saturated yet but need more GPUs.
-
comment
Comment #39466979
We did this for every stable diffusion release, you get the feedback data to improve it continuously ahead of open release.
-
comment
Comment #39454823
Training on 4096 v5es how did you handle crazy batch size :o
-
comment
Comment #29913457
You should go web3. There are incredibly few lawyers who know anything about crypto and huge demand. You can also join DAOs etc and earn significant amounts of money for your effor…
-
comment
Comment #28479908
"Payment can be accepted in any currency, including cryptocurrencies as we outrun the end of civilisation."
-
comment
Comment #26356036
I was helped with my sinus pressure and dizziness by 2 x 1g N Acetyl Cysteine per day. Very short half life, very quick response (if any), available on Amazon.
-
comment
Comment #26356025
You may wish to look into Clemastine/Tavegil antihistamine, easy to get over the counter.