Live data from Hacker News

Viewing profile — jumpCastle

jumpCastle

HN member
Joined
Fri, Dec 30, 2016, 5:06 PM UTC
HN karma
42
Public activity
66 items

About jumpCastle

I design channel codes for flash memories.

Recent public activity

  1. comment
    Comment #43709763

    It was a good benchmark until it entered the training set.

  2. comment
    Comment #43709734

    But open source like aider

  3. comment
    Comment #43598729

    Deepseek v3 was FP8

  4. comment
    Comment #43598708

    It's quite similar to muP https://github.com/microsoft/mup

  5. comment
    Comment #42167971

    Stock prices cannot go up without if the planet is destroyed

  6. comment
    Comment #41154448

    Attention was invented because Bengio lab had to be disciplined about a black box (google had more compute)

  7. comment
    Comment #40977064

    Aren't services like runpod solve half of these concerns?

  8. comment
    Comment #40967668

    Are the updates really slow?

  9. comment
    Comment #40362132

    Spend more time with my side projects.

  10. comment
    Comment #39995573

    You can try his book. https://www.math.ias.edu/avi/book

  11. comment
    Comment #39991442

    You can use older models with the api.

  12. comment
    Comment #39790830

    Also the parameters are optimized also with loss of future tokens in the sequence.

  13. comment
    Comment #39581312

    Work pretty well for all classes of problems?

  14. comment
    Comment #39365915

    Rsu exist before ipo?

  15. comment
    Comment #39364495

    Use api, embed history and retrieve.

  16. comment
    Comment #38664239

    Use model output as training data. For better performance you can get some top log probs and minimize kl divergence.

  17. comment
    Comment #38601286

    Without open weights why would anyone care about them? By the time they could compete with gpt4 there's probably be gpt5 already.

  18. comment
    Comment #37681858

    The future is the AI also writes the high level descriptions of stuff.

  19. comment
    Comment #37497900

    But you can fine tune gpt 3.5 turbo, so your comparison is not clear.

  20. comment
    Comment #37127838

    I still use infinity for reddit though.

  21. comment
    Comment #36709386

    Gzip every query with all training data can get more expensive.

  22. comment
    Comment #36622345

    Silly is a charitable interpretation.

  23. comment
    Comment #36621172

    Title with a 10 digits number, meaningless first page figure and no experiments related to the main claim. Did a rogue author posted it without permission again?

  24. comment
    Comment #36611264

    They no longer require sign-in. Got scared of the reaction.

  25. comment
    Comment #36565303

    Is Azure as simple to use as OpenAI? Their documentation are much less simple.