I'm confused. Is this the "GPT-5" that was coming in summer, just with a different name? Or is this more like a parallel development doing chain-of-thought type prompt engineering on GPT-4o? Is there still a big new foundational model coming, or is this it?
Learning to Reason with LLMs
891–900 of 1001 posts
Re: Learning to Reason with LLMs
#892Maybe this is improvement in some areas, still I got spurious reasoning and inability to connect three simple facts: Yes, "el presente acta de nacimiento" is correct in Spanish. Explanation: "Acta" is a feminine noun that begins with a stressed "a" sound. In Spanish, when a feminine singular noun starts with a stressed "a" or "ha", the definite article "la" is replaced with "el" to facilitate pronunciation. However,…
Proof: https://www.elcastellano.org/francisco-jos%C3%A9-d%C3%ADaz-%...
Re: Learning to Reason with LLMs
#893Re: Learning to Reason with LLMs
#8942018 - gpt1 2019 - gpt2 2020 - gpt3 2022 - gpt3.5 2023 - gpt4 2023 - gpt4-turbo 2024 - gpt-4o 2024 - o1 Did OpenAI hire Google's product marketing team in recent years?
They partnered with Microsoft, remember? 1985 – Windows 1.0 1987 – Windows 2.0 1990 – Windows 3.0 1992 – Windows 3.1 1995 – Windows 95 1998 – Windows 98 2000 – Windows ME (Millennium Edition) 2001 – Windows XP 2006 – Windows Vista 2009 – Windows 7 2012 – Windows 8 2013 – Windows 8.1 2015 – Windows 10 2021 – Windows 11
Re: Learning to Reason with LLMs
#895First shot, I gave it a medium-difficulty math problem, something I actually wanted the answer to (derive the KL divergence between two Laplace distributions). It thought for a long time, and still got it wrong, producing a plausible but wrong answer. After some prodding, it revised itself and then got it wrong again. I still feel that I can't rely on these systems.
Look where you were 3 years ago, and where you are now. And then imagine where you will be in 5 more years. If it can almost get a complex problem right now, I'm dead sure it will get it correct within 5 years
Re: Learning to Reason with LLMs
#896Earlier quoted context omitted.
LLMs perform well on small tasks that are well defined. This definition matches almost every task that a student will work on in school leading to an overestimation of LLM capabiity. LLMs cannot decide what to work on, or manage large bodies of work/code easily. They do not understand the risk of making a change and deploying it to production, or play nicely in autonomous settings. There is going to be a massive amou…
Truth is, LLMs are going to make the coding part super easy, and the ceiling for shit coders like me has just gotten a lot lower because I can just ask it to deliver clean code to me. I feel like the software developer version of an investment banking Managing Director asking my analyst to build me a pitch deck an hour before the meeting.
What is hard is gather requirements, dealing with unexpected production issues, scaling, security, fixing obscure bugs and integration with other systems.
The coding part is about 10% of my job and the easiest part by far.
Re: Learning to Reason with LLMs
#897I think what it comes down to is accuracy vs speed. OpenAI clearly took steps here to improve the accuracy of the output which is critical in a lot of cases for application. Even if it will take longer, I think this is a good direction. I am a bit skeptical when it comes to the benchmarks - because they can be gamed and they don't always reflect real world scenarios. Let's see how it works when people get to apply it…
There is probably also a practical limit at which it does truly flatten, it's probably just well past either of those points so it might as well not exist.
Re: Learning to Reason with LLMs
#898Earlier quoted context omitted.
So now it’s a question of how fast the AGI will run? :)
It's fine, it will only need to be powered by a black hole to run.
The company Oracle just announced that it is designing data centers with small modular nuclear reactors:
https://news.ycombinator.com/item?id=41505514
There are already 440 nuclear reactors operating in 32 countries today.
Sam Altman owns a stake in Oklo, a small modular reactor company. Bill Gates has a huge stake in his TerraPower reactor company. In China, 5 reactors are being built every year. You just don't hear about it... yet.
No amount of batteries can protect a solar/wind grid from an arbitrarily extended period of "bad" weather. It's like range anxiety in an electric car. If you have N days of battery storage and the sun doesn't shine for N+1 days, you're in trouble.
Nuclear fission is safe, clean, secure, and reliable.
An investor might consider buying physical uranium (via ticker SRUUF in America) or buying Cameco (via ticker CCJ).
Cameco is the dominant Canadian uranium mining company that also owns Westinghouse. Westinghouse licenses the AP1000 pressurized water reactor used at Vogtle in the U.S. as well as in China.
Re: Learning to Reason with LLMs
#899Some practical notes from digging around in their documentation: In order to get access to this, you need to be on their tier 5 level, which requires $1,000 total paid and 30+ days since first successful payment. Pricing is $15.00 / 1M input tokens and $60.00 / 1M output tokens. Context window is 128k token, max output is 32,768 tokens. There is also a mini version with double the maximum output tokens (65,536 tokens…
you need to be on their tier 5 level, which requires $1,000 total paid and [...] Good opening for OpenAI's competitors to run a 'we're not snobs' promotion.
Re: Learning to Reason with LLMs
#900Earlier quoted context omitted.
The calculator didn’t eliminate math majors. Excel and accounting software didn’t eliminate accountants and CPAs. These are all just tools. I spend very little of my overall time at work actually coding. It’s a nice treat when I get a day where that’s all I do. From my limited work with Copilot so far, the user still needs to know what they’re doing. I have 0 faith a product owner, without a coding background, can us…
Accounting mechanization is a good example of how unpredictable it can be. Initially there were armies of "accountants" (what we now call bookkeepers), mostly doing basic tasks of collecting data and making it fit something useful. When mechanization appeared, the profession split into bookkeeping and accounting. Bookkeeping became a job for women as it was more boring and could be paid lower salaries (we're in the 1…
Alternatively, I’ve thought a bit about this previously and have a slight different hypothesis. Businesses are ran by “PM types”.the only reason that developers have jobs is because pm types need technical devs to build their vision. (Obviously I’m making broad strokes here as there are also plenty of founders that ARE the dev). Now, if ai makes technical building more open to the masses, I could foresee a scenario where devs and pms actually converge into a single job title that eats up the technical-leaning PMs and the “PM-y” devs. Devs will shift to be more PM-y or else be cut out of the job market because there is less need for non-ambitious code monkeys. The easier it becomes for the masses to build because of AI, the less opportunity there is for technical grunt work. If before it took a PM 30 minutes to get together the requirements for a small task that took the entry level dev 8 hours to do, then it made sense. Now if AI makes it so a technical PM could build the feature in an hour, maybe it just makes sense to have the PM do the implementation and cut out the code monkey. And if the PM is doing the implementation, even if using some mythical AI superpower, that’s still going to have companies selecting for more technical PM’s. In this scenario I think non-technical PMs and non-pm-y devs would find themselves either without jobs or at greatly reduced wages.