Viewing profile — adefa
adefa
HN member- Joined
- Fri, Aug 17, 2012, 5:58 PM UTC
- HN karma
- 622
- Public activity
- 119 items
- HN profile
- View on Hacker News ↗
About adefa
No profile information was provided.
Recent public activity
-
comment
Comment #48452734
I built a tmux clone in Rust: https://github.com/TrevorS/rmux
-
comment
Comment #47651217
Released uncensored versions of all four Gemma 4 models. bf16 + GGUF for each. Collection: https://huggingface.co/collections/TrevorJS/gemma-4-uncensor... Code: https://github.com/…
- story
-
comment
Comment #46983678
True :) After some performance improvements, it is realtime on my DGX Spark with an RTF of .416 -- now getting ~19.5 tokens per second. Check it out, see if it's better for you.
-
comment
Comment #46983656
I'm curious to see if you are able to run the model now from the CLI?
-
comment
Comment #46983646
The cubecl-wgpu were only needed to reduce the number of kernel workgroups, otherwise I was getting errors in WASM.
-
comment
Comment #46983630
This should be fixed now. There were a number of bugs that kept the model from working correctly in different environments. Please let me know if you test again. :)
-
comment
Comment #46983622
Please try again. The model weights are unchanged, but the inference code is improved.
-
comment
Comment #46983617
this should be fixed
-
comment
Comment #46983613
Hello everyone, thanks for the interest. I merged a number of significant performance improvements that increase speed and accuracy across CUDA, Metal, and WASM as well as improve …
-
comment
Comment #46983602
Hello, I pushed up and merged a PR that greatly improves performance on CUDA, Metal, and in WASM. Depending on your hardware, the model is definitely real time (able to transcribe …
-
story
Show HN: Voxtral Mini 4B Realtime running in the browser
Hello! Earlier this week Mistral released: https://huggingface.co/mistralai/Voxtral-Mini-4B-Realtime-26... Last time I ported a TTS model to Rust using candle, this time I ported a…
-
comment
Comment #46878131
Benchmarks using DGX Spark on vLLM 0.15.1.dev0+gf17644344 FP8: https://huggingface.co/Qwen/Qwen3-Coder-Next-FP8 Sequential (single request) Prompt Gen Prompt Processing Token Gen T…
-
story
Show HN: Qwen 3 TTS ported to Rust
I love pushing these coding platforms to their (my? our?) limits! This time I ported the new Qwen 3 TTS model to Rust using Candle: https://github.com/TrevorS/qwen3-tts-rs It took …
-
comment
Comment #46672959
Absolutely -- it's perfectly understandable. I wanted to be completely upfront about AI usage and while I was willing and did start to break the PR down into parts, it's totally OK…
-
comment
Comment #46671396
I ran a similar experiment last month and ported Qwen 3 Omni to llama cpp. I was able to get GGUF conversion, quantization, and all input and output modalities working in less than…
- story
-
comment
Comment #45058859
Thank you, I posted this earlier and it was flagged: https://news.ycombinator.com/item?id=45058762
- story
-
comment
Comment #44373888
Here is the article as a PDF with some screen shots in it: https://gofile.io/d/4aahPJ
-
comment
Comment #44373827
This looks like a leak or very early post, date reads the 25th.
-
comment
Comment #44036474
I have been using Claude Code a lot since the Max plan change and I've never hit the limits myself.
-
comment
Comment #43965384
Here’s a CLI I’m experimenting with https://github.com/TrevorS/rhizome that indexes local repos with Tree‑sitter, stores ONNX embeddings in SQLite, and answers semantic queries off…
- story
-
comment
Comment #43790441
If you missed it, check out this MusicFX DJ: https://labs.google/fx/tools/music-fx-dj It's pretty fun :) https://imgur.com/a/ohTZXZ0