Live data from Hacker News

Viewing profile — sacred_numbers

sacred_numbers

HN member
Joined
Fri, Feb 07, 2020, 7:14 PM UTC
HN karma
744
Public activity
153 items

About sacred_numbers

No profile information was provided.

Recent public activity

  1. comment
    Comment #37644764

    GPT-4 is not the same product. I know it seems like it due to the way they position 3.5 and 4 on the same page, but they are really quite separate things. When I signed up for Chat…

  2. comment
    Comment #37487337

    Based on my research, GPT-3.5 is likely significantly smaller than 70B parameters, so it would make sense that it's cheaper to run. My guess is that OpenAI significantly overtraine…

  3. comment
    Comment #37153732

    The reason cement is a major contributor to CO2 emissions is because of how much cement we produce. I don't know the lifetime or effectiveness of this catalyst, but typically you o…

  4. comment
    Comment #37059513

    Yes, with 4 bit quantization.

  5. comment
    Comment #36823694

    We could not do that inadvertently. To block even 1% of light using Starlink sized satellites (~30 m^2 with solar panels deployed) would require tens of billions of satellites. We …

  6. comment
    Comment #36526662

    The quality difference is substantial. I don't care if it's wasteful to use something that has many uses for a supposedly narrow task (although I don't see translation as a particu…

  7. comment
    Comment #36035730

    It's worse on English and a lot of other common languages (see Appendix C of the paper). It does better on less common languages like Latvian or Tajik, though.

  8. comment
    Comment #35857763

    There will always be overhead, but that doesn't mean it will always be a huge amount of overhead. I believe the state of the art is 97% efficiency ( https://www.osti.gov/biblio/149…

  9. comment
    Comment #35819927

    I did my own calculations based on plotting loss on benchmarks compared to models with known parameters and training data, as well as using a quote from Sam Altman that said that G…

  10. comment
    Comment #35817624

    I would bet money against that. Replicating GPT-4 pre-training with current hardware would cost about 40-50m in compute. Compute will continue to decrease in cost and algorithmic i…

  11. comment
    Comment #35562565

    They do update the model in the background, although I'm not sure how often or how much they update it. To avoid issues with this practice they offer gpt-4-0314 which says this in …

  12. comment
    Comment #35559792

    It's unlikely that OSS LLMs will ever be able to compete with corporate LLMs. I can only think of a few scenarios where this could work: 1. Someone develops a procedure for trainin…

  13. comment
    Comment #35474014

    I can think of a few ways: 1. The ChatGPT web search plugin becomes standard protocol for every prompt. If you ask a factual question ChatGPT will first look up an answer with a se…

  14. comment
    Comment #35391976

    If you bought an 8xA100 machine for $140k you would have to run it continuously for over 10,000 hours (about 14 months) to train the 7B model. By that time the value of the A100s y…

  15. comment
    Comment #35363846

    When I checked yesterday I believe the signature said OpenAI CEO Sam Altman, so it was definitely a joke signature, not a case of two people having the same name.

  16. comment
    Comment #35335890

    The Reflexion paper ( https://arxiv.org/abs/2303.11366 ) that came out recently shows how this kind of mistake might be overcome. Asking the model to think about the answer after i…

  17. comment
    Comment #35255894

    Theoretically it should be way less energy intensive as well, since there won't be an animal expending energy to live for months before slaughter. Nor will there be a need to grow …

  18. comment
    Comment #35202485

    When you are speaking to a person, they have inner thoughts and outer actions/words. If a person sees a chess board they will either consciously or unconsciously evaluate all the l…

  19. comment
    Comment #35111685

    Alternatively: 1. Quickly reduce costs by increasing model and computation efficiency. 2. Massively reduce prices while still maintaining some gross margin. 3. Massively increase m…

  20. comment
    Comment #34988052

    It could be even smaller than a Chinchilla optimal model. The Chinchilla paper was about training the most capable models with the least training compute. If you are optimizing for…

  21. comment
    Comment #34876365

    We have, but it's not a single process. We can convert light to electricity quite cheaply and efficiently with solar PV panels and then use that electricity to electrolyze hydrogen…

  22. comment
    Comment #34843539

    Vinyl chloride, when burned, can create poisonous byproducts such as phosgene and carbon monoxide. Vinyl chloride that leaks into the environment is a carcinogen that can cause dam…

  23. comment
    Comment #34805932

    I think the biggest reason for Tesla's gross margins is that millions of people want EVs for various reasons (gas prices, environmental concerns, fun, status) and Tesla is one of t…

  24. comment
    Comment #34799119

    Surprisingly it appears not to be too far off standard solar panel efficiencies. According to this source[0], five nines silicon (5N) is called Upgraded Mettalurgical-grade (UMG) s…

  25. comment
    Comment #34795482

    Unfortunately I think you're off by an order of magnitude. I think it would be 810 Kilojoules, which is approximately equivalent to a 1kg lithium-ion battery. Of course, you could …