Live data from Hacker News

AI: Accelerated Incompetence

slater.dev

211–220 of 287 posts

Re: AI: Accelerated Incompetence

#211
post #168

Earlier quoted context omitted.

As always with definitive assertions regarding LLMs incapacities, i would be more convinced if one could demonstrate those assertions with an illustrative example, on a real LLM. So far, the abilities of LLM to manipulate concepts, in practice, has been indistinguishable in practice from "true" human-level concept manipulation. And not just for scientific, "narrow" fields.

The problem with mental capacities, is that they are not measured by tests. We have no valid and reliable way of determining them. Hence why "metrics psychology" is a pseudoscience. If I give a child a physics exam, and they score 100% it could either be because they're genuinely a genius (possessing all relevant capabilities and knowledge), or because they cheated. Suppose we dont know how they're cheating, but they…

"…we would never look to linguistic competence".

On the contrary, i strongly believe that what LLM proved is the fact linguists have always told us about : that the language provides a structure on top of which we're building our experience of concepts (Sapir whorf hypothesis).

I don't think one can conceptualize much without the use of a language.

Re: AI: Accelerated Incompetence

#212

> I don't think anyone believes that a computer program is literally their companion Quoth the makers of Claude: > AI systems are no longer just specialized research tools: they’re everyday academic companions. > https://www.anthropic.com/news/anthropic-education-report-ho... To call Anthropic's opener brazen, obnoxious, or euphemistic would be an understatement. I hope it ages like milk, as it deserves, and so embar…

> companion

replica ai

Re: AI: Accelerated Incompetence

#213

The author seems to have an inflated notion of what developing software is about. Most software doesn't require "perspicacity". Think about Civil Engineering. In a very few cases you are designing a Golden Gate Bridge. The vast majority of the time you are designing yet another rural bridge over a culvert. The first case requires deep investigation into all of the factors involving materials strength, soil dynamics,…

I don't disagree except that software is different to building bridges because once you build the bridge, it's a bridge and you're done. Software engineering is like deciding after you built the bridge that it needs to now be able to open. Oh and the bridge is now busy and used heavily so we can't allow for any downtime while you rebuild the bridge to open. And as much as we hope for standards so everything cookie cu…

Great points. The only reason that Civil Engineering isn't like software engineering is because of cost and legal regulations. It costs only labor to change software. If there were robotic workers and plentiful energy and materials you'd better believe that bridges would be getting refactored all of the time.

Also bridges are never done. They are continually inspected and refurbished. Every bridge you build has an on-going cost. Just like software.

Re: AI: Accelerated Incompetence

#214

Earlier quoted context omitted.

I'm very well-familiar with the literature in this area. I understand that computer scientists, with obscene and wild abandon, will just pick whatever word suits their agenda and define it opportunistically, without regards to the confusion it will cause -- indeed, seemingly with this aim -- to "impress" the reader and make their research seem extraordinary. "Concept" is not a term from computer science, its use here…

You have biased view on the definition of "concept" based on English language and logic. In Chinese language, concept is 概念. In Chinese language, happy dog is 快乐的狗. Notice it has an extra "的" that is missing in English language. This tells you that you can't just treat English grammar and structure as the formal definition of "concept". Some languages do not have words for happiness, or dog. But that doesn't mean the…

> But that doesn't mean the concept of happiness or dog does not exist.

That would be a consequence of your position.

The person who wrote the article is english. The claim being evaluated here is from the article. The term "concept" is english. The *meaning* of that term isn't english, any more than the meaning of "two" is english.

My analysis of "concept" has nothing to do with the english language. "Happy" here stands in for any property-concept and 'dog' any term which can be a object-concept, or a property-concept, or others. If some other language has terms which are translated into terms that do not function in the same way, then that would be a bad translation for the purpose of discussing the structure of concepts.

It is you who are hijacking the meaning of "concept", ignoring the meaning the author intended, substituting one made up 5 minutes ago by self-aggrandising poorly read people in XAI -- and the going off about irrelevant translations into Chinese.

The claim the author made has nothing to do with XAI, nor chinese, nor english. It has to do with mental capacities to "bring objects under a concept", partition experience into its conceptual structure ("conceptualise"), simulate scenarios based on compositions of concepts ("the imagination") and so on. These are mental capabilities a wide class of animals possess, who know no language; that LLMs do not possess.

Re: AI: Accelerated Incompetence

#216
post #65

You know, sometimes I feel that all this discourse about AI for coding reflects the difference between software engineers and data scientists / machine learning engineers. Both often work with unclear requirements, and sometimes may face floating bugs which are hard to fix, but in most cases, SWE create software that is expected to always behave in a certain way. It is reproducible, can pass tests, and the tooling is…

> So, for MLE, working with AI that isn't always reliable, is a norm. They are accustomed to thinking in terms of probabilities, distributions, and acceptable levels of error. Applying this mindset to a coding assistant that might produce incorrect or unexpected code feels more natural. They might evaluate it like a model: "It gets the code right 80% of the time, saving me effort, and I can catch the 20%." And given…

This is where I find having a disconnect between an ML team and product team is so broken. Same for SE to be fair.

Accuracy rates, F1, anything, they're all just rough guides. The company cares about making money and some errors are much bigger than others.

We'd manually review changes for updates to our algos and models. Even with a golden set, breaking one case to fix five could be awesome or terrible.

I've given talks about this, my classic example is this somewhat imagined scenario (because it's unfair of me to accuse people of not checking at all):

It's 2015. You get an update to your classification model. Accuracy rates go up on a classic dataset, hooray! Let's deploy.

Your boss's, boss's, boss gets a call at 2am because you're in the news.

https://www.bbc.co.uk/news/technology-33347866

Ah. Turns out improving classifications of types of dogs improved but... that wasn't as important as this.

Issues and errors must be understood in context of the business. If your ML team is chucking models over the fence you're going to at best move slowly. At worst you're leaving yourself open to this kind of problem.

Re: AI: Accelerated Incompetence

#217
post #211

Earlier quoted context omitted.

The problem with mental capacities, is that they are not measured by tests. We have no valid and reliable way of determining them. Hence why "metrics psychology" is a pseudoscience. If I give a child a physics exam, and they score 100% it could either be because they're genuinely a genius (possessing all relevant capabilities and knowledge), or because they cheated. Suppose we dont know how they're cheating, but they…

"…we would never look to linguistic competence". On the contrary, i strongly believe that what LLM proved is the fact linguists have always told us about : that the language provides a structure on top of which we're building our experience of concepts (Sapir whorf hypothesis). I don't think one can conceptualize much without the use of a language.

> I don't think one can conceptualize much without the use of a language.

Well a great swath of the animal kingdom stands against you.

LLMs have invited yet more of this pseudoscience. It's a nonesense position in an empirical study of mental capabilites across the animal kingdom. Something previuosly only believed by idealist philosophers of the early 20th century and prior. Now brought back so people can maintain their image in the face of their apparent self-deception: better we opt for gross pseudoscience than admit we're fooled by a text generation machine.

Re: AI: Accelerated Incompetence

#218

> Input Risk. An LLM does not challenge a prompt which is leading ... (Emphasis mine) This has been the biggest pain point for me, and the frustrating part is that you might not even realize you're leading it a particular way at all. I mean it makes sense with how LLMs work, but a single word used in a vague enough way is enough to skew the results in a bad direction, sometimes contrary to what you actually wanted to…

What I've been doing when I want to avoid this "unexpected leading", is to tell the LLM to "Ask me 3 rounds of 5 clarifying questions each, first.". The first round usually exposes the main assumptions it's making, and from there we narrow down and clarify things.

I've read you comment about all the things you tried, and it seems you have much broader experience with LLMs than I do. But I didn't see this technique mentioned, so leaving this here in case it helps someone else :).

Re: AI: Accelerated Incompetence

#219

Earlier quoted context omitted.

I'm not sure why "The contents of `I` is nowhere modelled by "projection" because this does not model composition, and is not relevantly discrete and bounded by logical connectives." In practical terms, what do you think the LLM output cannot contain right now? Because the way I read it now is "LLM can't speculate". But that's trivial to disprove by asking for that happy dog on Mars speculation you have as an example…

LLMs are just a token->token mapping. They can output any set of tokens for any set of input tokens. So there is no output which isn't in the domain or codomain. The issue is why one (prompt, answer) pair is given. If the answer is given as a "reasoning process" over salient parts of the prompt, that, e.g., involves imagining/simulation as expected, then for {(prompt', answer')} of similar imaginings we will get reli…

> LLMs are just a token->token mapping. They can output any set of tokens for any set of input tokens. So there is no output which isn't in the domain or codomain.

This applies the same to humans hearing a question and responding. Tokens in, tokens out (whether words or sound). It's not unique to LLMs, so not useful for explaining differences.

> then for {(prompt', answer')} of similar imaginings we will get reliable mappings. If its cheating, then we wont.

You're not really showing that this is/isn't the case already. Also this would put people with quirky ideas and wild imagination in the "cheating" group if I understand your claim correctly. There's even a whole game around a similar concept - Dixit - describe an image in a way that as few people as possible will get it.

> we can give a series of prompts (p1, p2, p3...) which require increasing complexity of the imagined scenario, and we do not find O(answering) to follow O(p-complexity-increase). Rather the search strategy is always the same

You're describing most current implementations, not a property of LLMs. Gemini scales the thinking phase for example. Future models are likely to do the same. Another recent post implemented this too https://news.ycombinator.com/item?id=44112326

Re: AI: Accelerated Incompetence

#220
LLMs are amazing at writing code and terrible at owning it.

Every line you accept without understanding is borrowed comprehension, which you’ll repay during maintenance with high interest. It feels like free velocity. But it's probably more like tech debt at ~40 % annual interest. As a tribe, we have to figure out how to use AI to automate typing and NOT thinking.

Post reply on HN