Live data from Hacker News

Viewing profile — gdiamos

gdiamos

HN member
Joined
Thu, Dec 18, 2014, 9:52 PM UTC
HN karma
861
Public activity
368 items

About gdiamos

No profile information was provided.

Recent public activity

  1. comment
    Comment #49248374

    How big is the open model? 30B?

  2. comment
    Comment #49215749

    Christopher Nolan beat you to it

  3. comment
    Comment #49205405

    vLLM is originally marketed as paged attention, but in hindsight, separating the web server and GPU process, continuous batching, kv caching / chunking, and a huge model library in…

  4. comment
    Comment #49093835

    I wish I could get a model to state its assumptions.

  5. comment
    Comment #49023539

    How do you ban melted sand?

  6. comment
    Comment #49023475

    I think we should shut it off. It would force US companies to build open models.

  7. comment
    Comment #49009544

    I want a hosted paper to be archived. That means that 10 years from now I don’t want think about making sure the hosting server is up. I also want it to have a standard format for …

  8. comment
    Comment #49001122

    thank god, these parameters are so confusing

  9. comment
    Comment #48943306

    as soon as you release a way of measuring it, you give LLMs a signal to optimize

  10. comment
    Comment #48913792

    Being on the review board comes with a promise to not be evil right?

  11. comment
    Comment #48897350

    No, I want arxiv to host the paper, not to review the paper. I wouldn't want my google drive to start telling me my paper was too sloppy. I just want a link.

  12. comment
    Comment #48780501

    I wonder if Amazon eventually gets cut out by 3D printing/replicators for imitable objects.

  13. comment
    Comment #48742554

    Scaling laws assume the error metric and data distribution. There is a lot of follow on work that explains what happens as you change them, e.g. Scaling Laws for Transfer - https:/…

  14. comment
    Comment #48741817

    When I first saw scaling laws in that deep speech experiment notebook, I didn’t believe it could be real. I was worried for months that we made a mistake, or that it only worked fo…

  15. comment
    Comment #48695102

    He started with tinkercad and thingiverse. I tried basic elegoo and bambu printers. He can’t read very well but he likes dragging shapes around on a tablet. He would ask me to find…

  16. comment
    Comment #48694707

    My kindergartner has a 3D printer. I got a call from the school principal. She said “another parent called and said your son 3D printed a gun and brought it to school”. I looked at…

  17. comment
    Comment #48664296

    Usually breakthroughs in computing lead to more usage of computing, not less.

  18. comment
    Comment #48537851

    This is why I use a router to send my own IP to my own models, and general information to Claude. https://split-brain-ui.scalarxlm.com/docs/clients I expect Claude to train on my g…

  19. comment
    Comment #48537809

    This weekend I was reading this paper on programming the Cerebras wafer scale engine, https://arxiv.org/html/2405.07898v1 . Data movement is the expensive part of computing, and so…

  20. story
  21. comment
    Comment #48475250

    What I do is route general data to Mythos, and my own IP to a local model. I expect them to train on their traffic, and I train on mine.

  22. comment
    Comment #48434804

    I don’t get it. That’s what I am using.

  23. comment
    Comment #48434787

    Think about the worst enterprise SaaS apps you have used…

  24. comment
    Comment #48434552

    What I tell my team to do is to drop using so many cloud saas apps, and build more themselves using LLMs. I’m not planning on firing people, but I am planning on building more, usi…

  25. comment
    Comment #48390948

    It’s about enterprises who care about supply chain risk and having a throat to choke if they have a problem. Here’s a real example. I’m in a design meeting talking about a model us…