Earlier quoted context omitted.
Who exactly said that? Can you give sources to high profile figures that said it?
Since we're in an Anthropic topic https://www.businessinsider.com/anthropic-ceo-ai-90-percent-...
Claude’s C Compiler vs. GCC
351–360 of 377 posts
Re: Claude’s C Compiler vs. GCC
#352Earlier quoted context omitted.
Maybe if AI evangelists would stop lying about what AI can do then people would hate it less. But lying and hype is baked into the DNA of AI booster culture. At this point it can be safely assumed anything short of right-here-right-now proof is pure unfettered horseshit when coming from anyone and everyone promoting the value of AI.
Your comment is a perfect example of the biases I'm talking about. "AI evangelists" are not a singular group of people.
It's not common for present-capabilities to be lied about too. But it does happen!
And the smallest presence are the users who don't work in the AI industry but rave about AI. I know of...two....people who fit that bill. A lead developer at a cybersecurity firm and someone who works heavily in statistics and data analytics. Both of which are very senior people in their fields who can articulate exactly what they're looking for without much left to interpretation.
Re: Claude’s C Compiler vs. GCC
#353Earlier quoted context omitted.
however it's something entirely different for that output to be good and maintainable People aren't prompting LLMs to write good, maintainable code though. They're assuming that because we've made a collective assumption that good, maintainable code is the goal then it must also be the goal of an LLM too. That isn't true. LLMs don't care about our goals. They are solving problems in a probabilistic way based on the c…
> People aren't prompting LLMs to write good, maintainable code though. Then they're not using the tools correctly. LLMs are capable of producing good clean code, but they need to be carefully instructed as to how. I recently used Gemini to build my first Android app, and I have zero experience with Kotlin or most of the libraries (but I have done many years of enterprise Java in my career). When I started I first ha…
The danger is, it doesn't quite scale up. The more complex the project, the more likely the AI is to get confused and start writing spaghetti code. It may even work for a while, but eventually the spaghetti piles up to the point that not even more spaghetti will fix it
I'll get that's going to get better over the next few years, with better tooling and better ways to get the AI to figure out/remember relevant parts of the code base, but that's just my guess
Re: Claude’s C Compiler vs. GCC
#354It seems like if Anthropic released a super cool and useful _free_ utility (like a compiler, for example) that was better than existing counterparts or solved a problem that hadn’t been solved before[0] and just casually said “Here is this awesome thing that you should use every day. By the way our language model made this.” it would be incredible advertising for them. But they instead made a blog post about how it w…
They made a blog post about it because it's an amazing test of the abilities of the models to deliver a working C-compiler, even with lots of bugs and serious caveats, for $20k of tokens, without a human babysitting it. I'd challenge anyone who are negative to this to try to achieve what they did by hand, with the same restrictions (e.g. generating full SSA form instead of just directly emitting code, capable of comp…
From the original blog post:
>Every agent would hit the same bug, fix that bug, and then overwrite each other's changes. Having 16 agents running didn't help because each was stuck solving the same task.
>The fix was to use GCC as an online known-good compiler oracle to compare against. I wrote a new test harness that randomly compiled most of the kernel using GCC
The blog post used the word autonomous a lot, which I suppose is true if Nicholas Carlini is not a human being but in fact a Claude agent.
>I'd challenge anyone who are negative to this to try to achieve what they did by hand, with the same restrictions (e.g. generating full SSA form instead of just directly emitting code, capable of compiling Linux), and log their time doing it.
Why would anyone do that? My point was that why does the company _not_ make a useful tool? I feel like that is a much more interesting topic of discussion than “why aren’t people that aren’t impressed by this spending their time trying to make this company look good?”
>This is the new floor.
Aside from the notion that they maybe intentionally set out to create the least useful or valuable output from their tooling (eg ‘the floor’) when they did not say that they did that, my question was “Why do they not make something genuinely useful?”. Marketing speak and imaginary engineers failing at made up challenges does not answer that question.
Re: Claude’s C Compiler vs. GCC
#355Earlier quoted context omitted.
This to me sounds a lot like the SpaceX conversation: - Ohh look it can [write small function / do a small rocket hop] but it can't [ write a compiler / get to orbit]! - Ohh look it can [write a toy compiler / get to orbit] but it can't [compile linux / be reusable] - Ohh look it can [compile linux / get reusable orbital rocket] but it can't [build a compiler that rivals GCC / turn the rockets around fast enough] - T…
All right, but perhaps they should also list the grand promises they made and failed to deliver on. They said they would have fully self-driving cars by 2016. They said they would land on Mars in 2018, yet almost a decade has passed since then. They said they would have Tesla's fully self-driving robo-taxis by 2020 and human-to-human telepathy via Neuralink brain implants by 2025–2027. > - Sure, but not by what was a…
Most of the time when you're writing a compiler for a new language, you'll be doing things that have been done before.
Because most of the concepts in your language are brought along from somewhere else.
That said: I'd always want a compiler and language designs to be well considered. Ideally, the authors have some proofs of soundness in their heads.
Perhaps LLM will make formal verification more feasible (from a cost perspective) and then our mind about what reliable software is might change.
Re: Claude’s C Compiler vs. GCC
#356I think this is a great example of both points of view in the ongoing debate. Pro-LLM coding agents: look! a working compiler built in a few hours by an agent! this is amazing! Anti-LLM coding agents: it's not a working compiler, though. And it doesn't matter how few hours it took, because it doesn't work. It's useless. Pro: Sure, but we can get the agent to fix that. Anti: Can you, though? We've seen that the more c…
Re: Claude’s C Compiler vs. GCC
#357Earlier quoted context omitted.
They made a blog post about it because it's an amazing test of the abilities of the models to deliver a working C-compiler, even with lots of bugs and serious caveats, for $20k of tokens, without a human babysitting it. I'd challenge anyone who are negative to this to try to achieve what they did by hand, with the same restrictions (e.g. generating full SSA form instead of just directly emitting code, capable of comp…
>without a human babysitting it. From the original blog post: >Every agent would hit the same bug, fix that bug, and then overwrite each other's changes. Having 16 agents running didn't help because each was stuck solving the same task. >The fix was to use GCC as an online known-good compiler oracle to compare against. I wrote a new test harness that randomly compiled most of the kernel using GCC The blog post used t…
Nothing in the article suggests it did not autonomously do the work.
> Why would anyone do that?
Because a lot of naysayers here pretend as if this is somehow trivial.
> My point was that why does the company _not_ make a useful tool?
Useful to whom? This is a researcher testing the limits of the models. Knowing those limits is highly useful to Anthropic. And it's highly useful to lots of others too, like me, as a means of understanding the capabilities of these models.
What, exactly would such a tool that'd somehow make the people dismissing this change their minds look like? Because I don't think anything would. They could produce lots of useful tools, if they aimed lower than testing the limits of the model. But it would not achieve what they set out to do, and it would not tell us anything useful.
I produce "useful tools" with Claude every day. That's not interesting. Anyone who actually uses these tools properly will develop a good understanding of the many things that can be achieved with them.
Most of us can't spend $20k figuring out where the limits are, however.
> I feel like that is a much more interesting topic of discussion than “why aren’t people that aren’t impressed by this spending their time trying to make this company look good?”
This is a ridiculous misrepresentation of the point. The point is that the people who aren't impressed by this very clearly and obviously do not have an understanding of the complexity of what they achieved, and are making ignorant statements about it.
> Aside from the notion that they maybe intentionally set out to create the least useful or valuable output from their tooling (eg ‘the floor’)
Again, you're either entirely failing to understand, or wilfully misrepresenting what I said. No, their goal was not to "set out the create the least useful or valuable output". Their goal was to test the limits of what the model can achieve. They did that.
That has far higher value than not testing the limits. Lots, and lots of people are building tools with Claude without testing the limits. We would not learn anything from that.
> my question was “Why do they not make something genuinely useful?”
Because that wasn't the purpose. The purpose was to test the limits of what the model can achieve. That you struggle to understand why what they achieved was massively impressive, does not change that.
Re: Claude’s C Compiler vs. GCC
#358Did Anthropic release the scaffolding, harnesses, prompts, etc. they used to build their compiler? That would be an even cooler flex to be able to go and say "Here, if you still doubt, run this and build your own! And show us what else you can build using these techniques."
That would still require someone else to burn 20000$ to try it themselves.
Re: Claude’s C Compiler vs. GCC
#359Earlier quoted context omitted.
Internet services have been centralised into a few ISPs and a few websites everyone visits
a few? all sorts of websites and services are thriving on the Internet even after significant consolidation of attention social media caused. Not even close to a dystopian picture parent comment paints.
Re: Claude’s C Compiler vs. GCC
#360My 2 cents: just like Cursor's browser, it seems the AI attempted a really ambitious technical design, generally matching the bells and whistles of a true industrial strength compiler, with SSA optimization passes etc. However looking at the assembly, it's clear to me the opt passes do not work, an I suspect it contains large amounts of 'dead code' - where the AI decided to bypass non-functioning modules. If a human…
> My 2 cents: just like Cursor's browser, it seems the AI attempted a really ambitious technical design, generally matching the bells and whistles of a true industrial strength compiler, with SSA optimization passes etc. Per the article from the person who directed this, the user directed the AI to use SSA form. > However looking at the assembly, it's clear to me the opt passes do not work, an I suspect it contains l…
The generated code's quality is more inline with 'undergrad course compiler backend', that is, basically doing as little work on the backend as possible, and always doing all the work conservatively.
Basic SSA optimizations such as constant propagation, copy propagation or common subexpression propagation are clearly missing from the assembly, the register allocator is also pretty bad, even though there are simple algorithms for that sort of thing that perform decently.
So even though the generated code works, I feel like something's gone majorly wrong inside the compiler.
The 300k LoC things isnt encouraging either, its way too much for what the code actually does.
I just want to point out, that I think a competent-ish dev (me?) could build something like this (a reasonably accurate C compiler), by a more human-in-the-loop workflow. The result would be much more reasonable code and design, much shorter, and the codebase wouldn't be full of surprises like it is now, and would conform to sane engineering practices.
Honestly I would certainly prefer to do things like this as opposed to having AI build it, then clean it up manually.
And it would be possible without these fancy agent orchestration frameworks and spending tens of thousands of dollars on API.
This is basically what went down with Cursor's agentic browser, vs an implementation that was recreated by just one guy in a week, with AI dev tools and a premium subscription.
There's no doubt that this is impressive, but I wouldn't say that agentic sofware engineering is here just yet.