Live data from Hacker News

Viewing profile — kamranjon

kamranjon

HN member
Joined
Wed, Feb 08, 2017, 2:01 PM UTC
HN karma
2,375
Public activity
607 items

About kamranjon

kamranjon.com

Recent public activity

  1. comment
    Comment #49249340

    I honestly think of the T5 model family to sort of be the real beginning of this open model craze - I know BERT was already popular for classification etc, but T5 was the first sor…

  2. comment
    Comment #49232056

    1 to 200 is a pretty big spread to between losing 18k dollars and making 3.6 milllion - do you have any actual numbers on the value produced from this 18k investment?

  3. comment
    Comment #49216390

    This is pretty interesting, I've never heard of this approach before - do you know if there is a research paper that covers how this was achieved?

  4. comment
    Comment #49216009

    The sort of sad but interesting thing here is that while this claims to be a more direct translation, Claude models were certainly trained on all of the pre-existing translations a…

  5. comment
    Comment #49212210

    This wasn't written with AI... obviously... I feel like there might need to be a new definition for whatever this paranoia is called because it's getting a bit out of hand.

  6. comment
    Comment #49181359

    Because street photography is very spontaneous it’s pretty common practice to set an aperture of 8 and just snap away - it’s a helpful trick for rangefinder cameras that often take…

  7. comment
    Comment #49178081

    “Current approaches to low precision primarily focus on converting models trained in full precision to lower bitwidths. We view this as fundamentally the wrong approach…” Very exci…

  8. comment
    Comment #49168592

    “While technically true the hallucination rates on modern models is low…” Isn’t this entirely context dependent? Where did you get the information that modern models have low hallu…

  9. comment
    Comment #49167688

    This seems really interesting - I was curious about this line from the website. “The whole post-training stack in one CLI. Soup doctors your data pre-flight, picks the method, writ…

  10. comment
    Comment #49138030

    Do you have an example? Would love to read one.

  11. comment
    Comment #49137956

    I love this analogy and think it possibly also applies to search which just doesn’t work anymore and is full of AI generated garbage. What the hell happened to stack overflow? I do…

  12. comment
    Comment #49137350

    In what way is it similar?

  13. comment
    Comment #49136453

    This is amazing and I think will probably end up being a pretty important development. I was just reading this great breakdown of how diffusion Gemma works: https://newsletter.maar…

  14. comment
    Comment #49123811

    I actually run it as a server - so most of the time I don't have to listen to it right next to me - it's just sitting in another room in my house - but I often am traveling with it…

  15. comment
    Comment #49123790

    The really interesting thing about this is how big of a jump was achieved with just extra fine-tuning here. No structural changes to the model, just more data, compute and time. It…

  16. comment
    Comment #49123711

    Generally get 20-25tps - prefill is pretty good around 400-450tps. I have been using compaction at around 100k tokens but mostly just cause it was the default in pi coding agent - …

  17. comment
    Comment #49122982

    Can't wait for the DwarfStar quants - I have been using DeepSeek v4 flash (preview) as my main coding agent for months now (running on my 128gb mbp) - it seems this model outperfor…

  18. comment
    Comment #49119510

    I know many companies are spending quite a bit of money, I don’t know if it bears out that the increased spend has resulted in increased profits, even if there has been some increa…

  19. comment
    Comment #49089968

    I haven't spun up thinkingcap yet but I'm aware of it and am intending to try it out soon. How did you find it?

  20. comment
    Comment #49087072

    Would you happen to have a link to that interview? Sounds like an interesting read.

  21. comment
    Comment #49082106

    Sorry I should have clarified - I meant that a ternary 27b model would outperform a non-quantized or 8 bit quantized 9 or 12b model - which it is generally close to (or much smalle…

  22. comment
    Comment #49081221

    I’ve actually been really impressed with the 27b model they recently released - amazing performance approaching 40 tok/s on m4 max and I didn’t run into any quality issues in the s…

  23. comment
    Comment #49080083

    Unfortunately Fermion Research appears to entirely AI generate all of their content here, even for the research section: https://www.fermionresearch.com/research/neutrino-8b/ "Neut…

  24. comment
    Comment #49079919

    There's a really interesting trend of labs using "proprietary" methods to convert existing models to a compressed ternary format. PrismML actually targeted the same Qwen 8b model a…

  25. comment
    Comment #49070333

    There is a really interesting startup in Prague that is doing just that. They fine-tuned Qwen 3.6 27b to have 46% fewer reasoning tokens while maintaining most of the performance c…