Live data from Hacker News

My 2.5 year old laptop can write Space Invaders in JavaScript now (GLM-4.5 Air)

simonwillison.net

311–320 of 415 posts

Re: My 2.5 year old laptop can write Space Invaders in JavaScript now (GLM-4.5 Air)

#311

Earlier quoted context omitted.

I have a friend who has been doing just that... usually with his company he manages a handful of projects where a bulk of the development is outsourced overseas. This past year, he's outpaced the 6 devs he's had working on misc projects just with his own efforts and AI. Most of this being a relatively unique combination of UX with features that are less common. He's using AI with note taking apps for meetings to enha…

Is this the same person who posted about launching 17 "products" in one year a few days ago on HN? :)

No, he's been working on building a larger eLearning solution with some interesting workflow analytics around courseware evaluation and grading. He's been involved in some of the newer LRS specifications and some implementation details to bridge training as well as real world exposure scenarios. Working a lot with first responders, incident response training etc.

I've worked with him off and on for years from simulating aircraft diagnostics hardware to incident command simulation and setting up core infrastructure for F100 learning management backends.

Re: My 2.5 year old laptop can write Space Invaders in JavaScript now (GLM-4.5 Air)

#312

Did you understand the implementation or just that it produced a result? I would hope an LLM could spit out a cobbled form of answer to a common interview question. Today a colleague presented data changes and used an LLM to build a display app for the JSON for presentation. Why did they not just pipe the JSON into our already working app that displays this data? People around me for the most part are using LLMs to e…

The LLM is the solution.

Re: My 2.5 year old laptop can write Space Invaders in JavaScript now (GLM-4.5 Air)

#313
post #306

Earlier quoted context omitted.

Sure. If you ask ChatGPT to play chess, it will put up an amateur-level effort at best. Stockfish will indeed wipe the floor with it. But what happens when you ask Stockfish to write a Space Invaders game? ChatGPT will get better at chess over time. Stockfish will not get better at anything except chess. That's kind of a big difference.

> ChatGPT will get better at chess over time Oddly, LLMs got worse at specifically chess: https://dynomight.net/chess/ But even to the general point, there's absolutely no agreement how much better the current architectures can ultimately get, nor how quickly they can get there. Do they have potential for unbounded improvements, albeit at exponential cost for each linear incremental improvement? Or will they asymptom…

If I had to bet, I'd say current models have an asymptomatic growth converging to a merely "ok" performance

(Shrug) People with actual money to spend are betting twelve figures that you're wrong.

Should be fun to watch it shake out from up here in the cheap seats.

Re: My 2.5 year old laptop can write Space Invaders in JavaScript now (GLM-4.5 Air)

#314
I tried with Claude Sonnet 4 and it does *not* work. So looks like GLM-4.5 Air in 3bit quant is ahead.

Chat is here: https://claude.ai/share/dc9eccbf-b34a-4e2b-af86-ec2dd83687ea

Claude Opus 4 does work but is far behind of Simon's GLM-4.5: https://claude.ai/share/5ddc0e94-3429-4c35-ad3f-2c9a2499fb5d

Re: My 2.5 year old laptop can write Space Invaders in JavaScript now (GLM-4.5 Air)

#315

Earlier quoted context omitted.

To emphasize this point further, at least with my efforts, it is not even possible to buy a 64GB M4 Pro right now. 32GB, 64GB, and 128GB are all sold out. We can say that 64GB addressable by a GPU is not exceptional when compared to 128GB and it still costs less than a month's pay for a FAANG engineer, but the fact that they aren't actually purchasable right now shows that it's not as easy as driving to Best Buy and…

They're not sold out—Apple's configurator (and chip naming) is just confusing. The MacBook Pro with M4 Pro is only available in 24 or 48 GB configurations. To get 64 or 128 GB, you need to upgrade to the M4 Max. If you're looking for the cheapest way into 64 of unified memory, the Mac mini is available with an M4 Pro and 64GB at $1999. So, truly, not "exceptional" unless you consider the price to be exorbitant (it's…

thank you for providing that extra info! i agree that $2000-4000 is not an absolutely earth shattering price, but i still wonder what the benefit one receives is when they say "2.5 year old laptop" instead of "64GB M2 laptop"

Re: My 2.5 year old laptop can write Space Invaders in JavaScript now (GLM-4.5 Air)

#316
post #306

Earlier quoted context omitted.

> ChatGPT will get better at chess over time Oddly, LLMs got worse at specifically chess: https://dynomight.net/chess/ But even to the general point, there's absolutely no agreement how much better the current architectures can ultimately get, nor how quickly they can get there. Do they have potential for unbounded improvements, albeit at exponential cost for each linear incremental improvement? Or will they asymptom…

If I had to bet, I'd say current models have an asymptomatic growth converging to a merely "ok" performance (Shrug) People with actual money to spend are betting twelve figures that you're wrong. Should be fun to watch it shake out from up here in the cheap seats.

Nah, trillion dollars is about right for "ok". Percentage point of the global economy in cost, automate 2 percent and get a huge margin. We literally set more than that on actual fire each year.

For "pretty good", it would be worth 14 figures, over two years. The global GDP is 14 figures. Even if this only automated 10% of the economy, it pays for itself after a decade.

For "Ahura Mazda", it would easily be worth 16 figures, what with that being the principal God and god of the sky in Zoroastrianism, and the only reason it stops at 16 is the implausibility of people staying organised for longer to get it done.

Re: My 2.5 year old laptop can write Space Invaders in JavaScript now (GLM-4.5 Air)

#317
post #297

Earlier quoted context omitted.

Serious question: if you have to read every line of code in order to validate it in production, why not just write every line of code instead?

Because it's much, much faster to review a hundred lines of code than it is to write a hundred lines of code. (I'm experienced at reading and reviewing code.)

Simon, don't you fear "atrophy" in your writing ability?

Re: My 2.5 year old laptop can write Space Invaders in JavaScript now (GLM-4.5 Air)

#318
post #261

Earlier quoted context omitted.

I don't care if it can "understand" anything, as long as I can use it to achieve useful things.

“useful things“ like poorly drawing birds on bikes? ;) (I have much respect for what you have done and are currently doing, but you did walk right into that one)

The pelican on a bicycle is a very useful test.

Re: My 2.5 year old laptop can write Space Invaders in JavaScript now (GLM-4.5 Air)

#319
post #224

Earlier quoted context omitted.

When do you think fine tuning is worth it over prompt engineering a base model? I imagine with the finetunes you have to worry about self-hosting, model utilization, and then also retraining the model as new base models come out. I'm curious under what circumstances you've found that the benefits outweigh the downsides.

For self-hosting, there are a few companies that offer per-token pricing for LoRA finetunes (LoRAs are basically efficient-to-train, efficient-to-host finetunes) of certain base models: - (shameless plug) My company, Synthetic, supports LoRAs for Llama 3.1 8b and 70b: https://synthetic.new All you need to do is give us the Hugging Face repo and we take care of the rest. If you want other people to try your model, we…

Do you maybe know if there is a company in the EU that hosts models (DeepSeek, Qwen3, Kimi)?

Re: My 2.5 year old laptop can write Space Invaders in JavaScript now (GLM-4.5 Air)

#320

Most likely its training data included countless Space Invaders in various programming languages.

This comment is ~3 years late. Every model since gpt3 has had the entirety of available code in their training data. That's not a gotcha anymore. We went from chatgpt's "oh, look, it looks like python code but everything is wrong" to "here's a full stack boilerplate app that does what you asked and works in 0-shot" inside 2 years. That's the kicker. And the sauce isn't just in the training set, models now do post-tra…

It's amazing that none of you even try to falsify you claims anymore. You can literally just put some of the code in a search engine and find the prior art example:

https://www.web-leb.com/en/code/2108

Your "AI tools" are just "copyright whitewashing machines."

These kinds of comments are really ignoring reality.

Post reply on HN