Live data from Hacker News

Three Years from GPT-3 to Gemini 3

oneusefulthing.org

121–130 of 336 posts

Re: Three Years from GPT-3 to Gemini 3

#121
post #102
post #96

Earlier quoted context omitted.

But holy shit is it also a funding issue when teachers make nothing.

As far as I understand it, the problem isn’t that teachers are shit. Giving more money would bring in better teachers, but I don’t know that they’d be able to overcome the other obstacles

> Giving more money would bring in better teachers, but I don’t know that they’d be able to overcome the other obstacles

Start with the easiest thing to control? Of giving more money and see what it does?

We seem to believe in every other industry that to get the best talent pay a high salary salary, but for some reason we expect teachers to do it out of compassion for the children while they struggle to pay bills. It's absurd.

Probably one of the single most important responsibilities of a society is to prepare the next generation, and it pays enormous return. But because we can't measure it with quarterly profits we just ignore it.

The rate of return on providing society with as good education is insane.

Re: Three Years from GPT-3 to Gemini 3

#122
post #96

Earlier quoted context omitted.

Education is not just a funding issues. Policy choices, like making it impossible for students to fail which means they have no incentive to learn anything, can be more impactful.

But holy shit is it also a funding issue when teachers make nothing.

I date a lot of teachers. My last one was in the San Ramon (CA) Valley School district, she makes about $90k a year at 34 years old. Talking to her basically makes me want to homeschool my kids to make sure someone like her isn't their teacher. Paying teachers more won't do ANYTHING until we become a lot more selective about who gets to become and stay a teacher. It can't be like most government jobs where getting it is like winning the lottery and knowing you can make above market money for below market performance.

Re: Three Years from GPT-3 to Gemini 3

#124
post #50

Every time I see an article like this, it's always missing --- but is it any good, is it correct? They always show you the part that is impressive - "it walked the tricky tightrope of figuring out what might be an interesting topic and how to execute it with the data it had - one of the hardest things to teach." Then it goes on, "After a couple of vague commands (“build it out more, make it better”) I got a 14 page p…

It's gotten more and more shippable, especially with the latest generation (Codex 5.1, Sonnet 4.5, now Opus 4.5). My metric is "wtfs per line", and it's been decreasing rapidly. My current preference is Codex 5.1 (Sonnet 4.5 as a close second, though it got really dumb today for "some reason"). It's been good to the point where I shipped multiple projects with it without a problem (with eg https://pine.town being one…

Maybe the wtfs per line are decreasing because these models aren't saying anything interesting or original.

Re: Three Years from GPT-3 to Gemini 3

#125
post #111

Earlier quoted context omitted.

That's how working with junior team members or open source project contributors goes too. Perhaps that's the big disconnect. Reviewing and integrating LLM contributions slotted right into my existing workflow on my open source projects. Not all of them work. They often need fixing, stylistic adjustments, or tweaking to fit a larger architectural goal. That is the norm for all contributions in my experience. So the LL…

Junior's grow into mids, and eventually into seniors. OSS contributor's eventually learn the codebase, you talk to them, you all get invested in the shared success of the project and sometimes you even become friends. For me, personally, I just don't see the point of putting that same effort into a machine. It won't learn or grow from the corrections I make in that PR, so why bother? I might as well have written it m…

> It won't learn or grow from the corrections I make in that PR, so why bother?

That does not match my experience. As the codebases I've worked with LLMs on become more opinionated and stylized, it seems to to a better job of following the existing work. And over time the models have absolutely improved in terms of their ability to understand issues and offer solutions. Each new release has solved problems for me that the previous ones have struggled with.

Re: interpersonal interactions, I don't find that the LLM has pushed them out or away. My projects still have groups of interested folk who talk and joke and learn and have fun. What the LLMs have addressed for me in part is the relative scarcity of labor for such work. I'm not hacking on the Linux Kernel with 10,000 contributors. Even with a dozen contributors, the amount of contributed code is relatively low and only in areas they are interested in. The LLM doesn't mind if I ask it to do something super boring. And it's been surprisingly helpful in chasing down bugs.

> Maybe one day it'll reach perfect parity of what I could've written myself, but today isn't that day.

Regardless of whether or not that happens, they've already been useful for me for at least 9 months. Since O3, which is the first one that really started to understand Rust's borrow checker in my experience. My measure isn't whether or not it writes code as well as I do, but how productive I am when working with it compared to not. In my measurements with SLOCCount over the last 9 months, I'm about 8x more productive than the previous 15 years without (as long as I've been measuring). And that's allowed me to get to projects which have been on the shelf for years.

This article by an AI researcher I happen to have worked with neatly sums up feelings I've had about comments like yours: https://medium.com/@ahintze_23208/ai-or-you-who-is-the-one-w...

Re: Three Years from GPT-3 to Gemini 3

#126
post #50

Earlier quoted context omitted.

It's gotten more and more shippable, especially with the latest generation (Codex 5.1, Sonnet 4.5, now Opus 4.5). My metric is "wtfs per line", and it's been decreasing rapidly. My current preference is Codex 5.1 (Sonnet 4.5 as a close second, though it got really dumb today for "some reason"). It's been good to the point where I shipped multiple projects with it without a problem (with eg https://pine.town being one…

Maybe the wtfs per line are decreasing because these models aren't saying anything interesting or original.

No, it's because they write correct code. Why would I want interesting code?

Re: Three Years from GPT-3 to Gemini 3

#127
post #19

I find Gemini 3 to be really good. I'm impressed. However, the responses still seem to be bounded by the existing literature and data. If asked to come up with new ideas to improve on existing results for some math problems, it tends to recite known results only. Maybe I didn't challenge it enough or present problems that have scope for new ideas?

That's the inherent limit on the models, that makes humans still relevant. With the current state of architectures and training methods - they are very unlikely to be the source of new ideas. They are effectively huge librarians for accumulated knowledge, rather than true AI.

Then again, an unintelligent human librarian would be nowhere near as useful as a good LLM.

Current LLMs exist somewhere between "unintelligent/unthinking" and "true AI," but lack of agreement on what any of these terms mean is keeping us from classifying them properly.

Re: Three Years from GPT-3 to Gemini 3

#128
post #82
post #53

Earlier quoted context omitted.

> Remember 54% of US adults read at or below the equivalent of a sixth-grade level. The sane conclusion would be to invest in education, not to dump hundreds of billions of llms, but ok

Unfortunately, people are born with a certain intellectual capacity and can't be improved beyond that with any amount of training or education. We're largely hitting peoples' capacities already. We can't educate someone with 80 IQ to be you; we can't educate you (or I) into being Einstein. The same way we can't just train anyone to be an amazing basketball player.

Other countries have better outcomes. I doubt it's just because of the genetics.

Re: Three Years from GPT-3 to Gemini 3

#129
post #62
post #22

Earlier quoted context omitted.

Text is very information-dense. I'd much rather skim a transcript in a few seconds than watch a video. There's a reason keyboards haven't changed much since the 1860s when typewriters were invented. We keep coming up with other fun UI like touchscreens and VR, but pretty much all real work happens on boring old keyboards.

Here's an old blog post that explores that topic at least with one specific example: https://www.loper-os.org/?p=861 The gist is that keyboards are optimized for ease of use but that there could be other designs which would be harder to learn but might be more efficient.

>> There's a reason keyboards haven't changed much since the 1860s when typewriters were invented.

> The gist is that keyboards are optimized for ease of use but that there could be other designs which would be harder to learn but might be more efficient.

Here's a relevant trivia question; assuming a person has two hands with five digits each, what is the largest number they can count to using only same?

Answer: (2 ** 10) - 1 = 1023

Ignoring keyboard layout options (such as QWERTY vs DVORAK), IMHO keyboards have the potential for capturing thought faster and with a higher degree of accuracy than other forms of input. For example, it is common for touch-typists to be able to produce 60 - 70 words per minute, for any definition of word.

Modern keyboard input efficiency can be correlated to the ability to choose between dozens of glyphs with one or two finger combinations, typically requiring less than 2cm of movement to produce each.

Re: Three Years from GPT-3 to Gemini 3

#130

Earlier quoted context omitted.

Unix CLI utilities have been all text for 50 years. Arguably that is why they are still relevant. Attempts to impose structured data on the paradigm like those in PowerShell have their adherents and can be powerful, but fail when the data doesn't fit the structure. We see similar tendency toward the most general interfaces in "operator mode" and similar the-AI-uses-the-mouse-and-keyboard schemes. It's entirely possib…

PowerShell is completely suitable. People are just used to bash and don’t feel the incentive to switch, especially with Windows becoming less relevant outside of desktop development.

Powershell feels like it's not built to be used in a practical way, unlike Unix tools that have been built and used by and for developers, which then feels nice because they are actually used a lot, and feel good to use.

Like, to set an env variable permanently, you either have to go through 5 GUI interfaces, or use this PS command:

[Environment]::SetEnvironmentVariable ("INCLUDE", $env:INCLUDE, [System.EnvironmentVariableTarget]::User)

Which is honeslty horrendous. Why the brackets ? Why the double columns ? Why the uppercases everywhere ? I get that it's trying to look more "OOP-ish" and look like C#, but nobody wants to work with that kind of shell script tbh. It's just one example, but all the powershell commands look like this, unless they have been aliased to trick you to think windows go more unixish

Post reply on HN