Viewing profile — sacred_numbers
sacred_numbers
HN member- Joined
- Fri, Feb 07, 2020, 7:14 PM UTC
- HN karma
- 744
- Public activity
- 153 items
- HN profile
- View on Hacker News ↗
About sacred_numbers
No profile information was provided.
Recent public activity
-
comment
Comment #37644764
GPT-4 is not the same product. I know it seems like it due to the way they position 3.5 and 4 on the same page, but they are really quite separate things. When I signed up for Chat…
-
comment
Comment #37487337
Based on my research, GPT-3.5 is likely significantly smaller than 70B parameters, so it would make sense that it's cheaper to run. My guess is that OpenAI significantly overtraine…
-
comment
Comment #37153732
The reason cement is a major contributor to CO2 emissions is because of how much cement we produce. I don't know the lifetime or effectiveness of this catalyst, but typically you o…
-
comment
Comment #37059513
Yes, with 4 bit quantization.
-
comment
Comment #36823694
We could not do that inadvertently. To block even 1% of light using Starlink sized satellites (~30 m^2 with solar panels deployed) would require tens of billions of satellites. We …
-
comment
Comment #36526662
The quality difference is substantial. I don't care if it's wasteful to use something that has many uses for a supposedly narrow task (although I don't see translation as a particu…
-
comment
Comment #36035730
It's worse on English and a lot of other common languages (see Appendix C of the paper). It does better on less common languages like Latvian or Tajik, though.
-
comment
Comment #35857763
There will always be overhead, but that doesn't mean it will always be a huge amount of overhead. I believe the state of the art is 97% efficiency ( https://www.osti.gov/biblio/149…
-
comment
Comment #35819927
I did my own calculations based on plotting loss on benchmarks compared to models with known parameters and training data, as well as using a quote from Sam Altman that said that G…
-
comment
Comment #35817624
I would bet money against that. Replicating GPT-4 pre-training with current hardware would cost about 40-50m in compute. Compute will continue to decrease in cost and algorithmic i…
-
comment
Comment #35562565
They do update the model in the background, although I'm not sure how often or how much they update it. To avoid issues with this practice they offer gpt-4-0314 which says this in …
-
comment
Comment #35559792
It's unlikely that OSS LLMs will ever be able to compete with corporate LLMs. I can only think of a few scenarios where this could work: 1. Someone develops a procedure for trainin…
-
comment
Comment #35474014
I can think of a few ways: 1. The ChatGPT web search plugin becomes standard protocol for every prompt. If you ask a factual question ChatGPT will first look up an answer with a se…
-
comment
Comment #35391976
If you bought an 8xA100 machine for $140k you would have to run it continuously for over 10,000 hours (about 14 months) to train the 7B model. By that time the value of the A100s y…
-
comment
Comment #35363846
When I checked yesterday I believe the signature said OpenAI CEO Sam Altman, so it was definitely a joke signature, not a case of two people having the same name.
-
comment
Comment #35335890
The Reflexion paper ( https://arxiv.org/abs/2303.11366 ) that came out recently shows how this kind of mistake might be overcome. Asking the model to think about the answer after i…
-
comment
Comment #35255894
Theoretically it should be way less energy intensive as well, since there won't be an animal expending energy to live for months before slaughter. Nor will there be a need to grow …
-
comment
Comment #35202485
When you are speaking to a person, they have inner thoughts and outer actions/words. If a person sees a chess board they will either consciously or unconsciously evaluate all the l…
-
comment
Comment #35111685
Alternatively: 1. Quickly reduce costs by increasing model and computation efficiency. 2. Massively reduce prices while still maintaining some gross margin. 3. Massively increase m…
-
comment
Comment #34988052
It could be even smaller than a Chinchilla optimal model. The Chinchilla paper was about training the most capable models with the least training compute. If you are optimizing for…
-
comment
Comment #34876365
We have, but it's not a single process. We can convert light to electricity quite cheaply and efficiently with solar PV panels and then use that electricity to electrolyze hydrogen…
-
comment
Comment #34843539
Vinyl chloride, when burned, can create poisonous byproducts such as phosgene and carbon monoxide. Vinyl chloride that leaks into the environment is a carcinogen that can cause dam…
-
comment
Comment #34805932
I think the biggest reason for Tesla's gross margins is that millions of people want EVs for various reasons (gas prices, environmental concerns, fun, status) and Tesla is one of t…
-
comment
Comment #34799119
Surprisingly it appears not to be too far off standard solar panel efficiencies. According to this source[0], five nines silicon (5N) is called Upgraded Mettalurgical-grade (UMG) s…
-
comment
Comment #34795482
Unfortunately I think you're off by an order of magnitude. I think it would be 810 Kilojoules, which is approximately equivalent to a 1kg lithium-ion battery. Of course, you could …