Viewing profile — kamranjon
kamranjon
HN member- Joined
- Wed, Feb 08, 2017, 2:01 PM UTC
- HN karma
- 2,375
- Public activity
- 607 items
- HN profile
- View on Hacker News ↗
About kamranjon
Recent public activity
-
comment
Comment #49249340
I honestly think of the T5 model family to sort of be the real beginning of this open model craze - I know BERT was already popular for classification etc, but T5 was the first sor…
-
comment
Comment #49232056
1 to 200 is a pretty big spread to between losing 18k dollars and making 3.6 milllion - do you have any actual numbers on the value produced from this 18k investment?
-
comment
Comment #49216390
This is pretty interesting, I've never heard of this approach before - do you know if there is a research paper that covers how this was achieved?
-
comment
Comment #49216009
The sort of sad but interesting thing here is that while this claims to be a more direct translation, Claude models were certainly trained on all of the pre-existing translations a…
-
comment
Comment #49212210
This wasn't written with AI... obviously... I feel like there might need to be a new definition for whatever this paranoia is called because it's getting a bit out of hand.
-
comment
Comment #49181359
Because street photography is very spontaneous it’s pretty common practice to set an aperture of 8 and just snap away - it’s a helpful trick for rangefinder cameras that often take…
-
comment
Comment #49178081
“Current approaches to low precision primarily focus on converting models trained in full precision to lower bitwidths. We view this as fundamentally the wrong approach…” Very exci…
-
comment
Comment #49168592
“While technically true the hallucination rates on modern models is low…” Isn’t this entirely context dependent? Where did you get the information that modern models have low hallu…
-
comment
Comment #49167688
This seems really interesting - I was curious about this line from the website. “The whole post-training stack in one CLI. Soup doctors your data pre-flight, picks the method, writ…
-
comment
Comment #49138030
Do you have an example? Would love to read one.
-
comment
Comment #49137956
I love this analogy and think it possibly also applies to search which just doesn’t work anymore and is full of AI generated garbage. What the hell happened to stack overflow? I do…
-
comment
Comment #49137350
In what way is it similar?
-
comment
Comment #49136453
This is amazing and I think will probably end up being a pretty important development. I was just reading this great breakdown of how diffusion Gemma works: https://newsletter.maar…
-
comment
Comment #49123811
I actually run it as a server - so most of the time I don't have to listen to it right next to me - it's just sitting in another room in my house - but I often am traveling with it…
-
comment
Comment #49123790
The really interesting thing about this is how big of a jump was achieved with just extra fine-tuning here. No structural changes to the model, just more data, compute and time. It…
-
comment
Comment #49123711
Generally get 20-25tps - prefill is pretty good around 400-450tps. I have been using compaction at around 100k tokens but mostly just cause it was the default in pi coding agent - …
-
comment
Comment #49122982
Can't wait for the DwarfStar quants - I have been using DeepSeek v4 flash (preview) as my main coding agent for months now (running on my 128gb mbp) - it seems this model outperfor…
-
comment
Comment #49119510
I know many companies are spending quite a bit of money, I don’t know if it bears out that the increased spend has resulted in increased profits, even if there has been some increa…
-
comment
Comment #49089968
I haven't spun up thinkingcap yet but I'm aware of it and am intending to try it out soon. How did you find it?
-
comment
Comment #49087072
Would you happen to have a link to that interview? Sounds like an interesting read.
-
comment
Comment #49082106
Sorry I should have clarified - I meant that a ternary 27b model would outperform a non-quantized or 8 bit quantized 9 or 12b model - which it is generally close to (or much smalle…
-
comment
Comment #49081221
I’ve actually been really impressed with the 27b model they recently released - amazing performance approaching 40 tok/s on m4 max and I didn’t run into any quality issues in the s…
-
comment
Comment #49080083
Unfortunately Fermion Research appears to entirely AI generate all of their content here, even for the research section: https://www.fermionresearch.com/research/neutrino-8b/ "Neut…
-
comment
Comment #49079919
There's a really interesting trend of labs using "proprietary" methods to convert existing models to a compressed ternary format. PrismML actually targeted the same Qwen 8b model a…
-
comment
Comment #49070333
There is a really interesting startup in Prague that is doing just that. They fine-tuned Qwen 3.6 27b to have 46% fewer reasoning tokens while maintaining most of the performance c…