Live data from Hacker News

Do AI companies work?

benn.substack.com

231–240 of 457 posts

Re: Do AI companies work?

#231
I like the article and agree with many of the arguments. A few comments though.

1. It's not that easy to switch between providers. There's no lock-in of course, but once you build a bunch of code that is provider specific (structured outputs, prompt caching, json mode, function calls, prompts designed for a specific provider, specific tools used by the openai Assistant, etc) then you need a good reason to switch (like a much better or cheaper model)

2. All of these companies do try to build some echo system around them, esp in the enterprise. The problem is that Google and Microsoft have a huge advantage here cause they have all the integrations

3. The consumer side. It's not just LLM. It's image, video, voice, and many more. You cannot ignore that ChatGPT can rival Google in a few years in terms of usage. As long as they can deliver good models, users are not going to switch so quickly. It's a huge market, just like Google. Pretty much, everyone in the world is going to use ChatGPT or some alternative in the next few years. My 9 year old and her friends already use it. No reason why they cannot monetize their huge user base like Google did.

Re: Do AI companies work?

#232

Earlier quoted context omitted.

Just the fact that I can have something proficient in language trivially accessible to me is really useful. I'm working on something that uses LLMs (language translation), but besides that I think it's brilliant that I can just ask an LLM to summarise my prompt in a way that gets the point across in far fewer tokens. When I forget a word, I can give it a vague description and it'll find it. I'm terrible at writing em…

> I can benchmark the quality of one LLM's translation by asking another to critique it Do you speak two or more languages? Anyone that does is wary of automated translations, especially across estranged cultures. > It's a new tool in the toolbox, one that we haven't had in our seventy years of working on computers, It's data analysis at scale, and reliant on scrapping what humans produced. A word processor does not…

> Do you speak two or more languages? Anyone that does is wary of automated translations, especially across estranged cultures.

I'm aware that it's imperfect. It's still pretty cool that they're multilingual as an emergent property - and rather than merely translating, you can discuss aspects of another language. Of course it hallucinates - that's the big problem with LLMs - but that doesn't make it useless. Besides, while automated translators aren't perfect, they're the only option in a lot of situations.

> It's data analysis at scale, and reliant on scrapping what humans produced. A word processor does not need TB of eBooks to do it's job.

And I'm not replacing word processors with LLMs, nor did I claim that they were trained in a vacuum.

> Because there's no wrong or right about poetry. Would you be comfortable having LLMs managing your bank account?

No, which is why I don't intend to... they're probabilistic and immensely fallible. I wasn't claiming they were gods. My point was that they're far outside of what we usually use computers for (e.g. managing your bank account), and that opens a lot of possibilities.

> That would be hand-holding, not learning.

Correct. That's why we don't do it... and in normal circumstances, it'd probably be better to deeply consider the problem in order to work out why you were wrong. But a week before the exam it's excellent. I got a healthy A*, so it doesn't seem to have hurt.

What exactly are you arguing against? I'm not convinced they're a route to AGI either, and I'm not about to replace my graphics driver with an LLM, nor code written by one - but you seem to have a vendetta against them that's lead to you sidestepping every single point I made with a snide remark against claims you seem to have imagined me making.

Re: Do AI companies work?

#233

This is like when VCs were funding all kinds of ride share, bike share, food delivery, cannabis delivery, and burning money so everyone gets subsidized stuff while the market figures out wtf is going on. I love it. More goodies for us

It invariably end in enshitifcation of the startups product AND the destruction of the original option.

Leaves you in a worse position than you started with.

Re: Do AI companies work?

#234

This is like when VCs were funding all kinds of ride share, bike share, food delivery, cannabis delivery, and burning money so everyone gets subsidized stuff while the market figures out wtf is going on. I love it. More goodies for us

Where I live the ridesharing/delivering startups didn't bring goodies, they just made everything worse. They destroyed the Taxi industry, I used to be able to just walk out to the taxi rank and get in the first taxi, but not anymore. Now I have to organize it on an app or with a phone call to a robot, then wait for the car to arrive, and finally I have to find the car among all the others that other people called. Fo…

taxis were the greatest example of regulatory capture. the post-event/airport Uber pickup situation is stupid and has obvious fixes, but, again, that's the taxicab regulatory capture where Uber has to thread a needle in order to not be a taxi, for them to legally operate. if we could clean slate, and make a working system, that would be great but we can't, because of the taxicab regulatory commission.

I can now summon a cab from the comfort of the phone I'm holding, and know that they'll accept my credit card. I know the price before I get in and I know the route they should take. I'm not going to get taken for an unnecessary scenic tourist surcharge detour.

people don't like feeling they got cheated, and pre-uber, taxis did that all the time.

Re: Do AI companies work?

#235
post #216

Earlier quoted context omitted.

* We're seeing much less of "it's making mistakes" these days.* Perhaps less than before, but still making very fundamental errors. Anything involving number I'm automatically suspicious. Pretty frequently I'd get different answers for the same question (to a human). e.g. ChatGPT will give an effective tax rate of n for some income amount. Then when asked to break down the calculation will come up with an effective t…

> Perhaps less than before, but still making very fundamental errors. Yes. Suppose someone developed a way to get a reliable confidence metric out of an LLM. Given that, much more useful systems can be built. Only high-confidence outputs can be used to initiate action. For low-confidence outputs, chain of reasoning tactics can be tried. Ask for a simpler question. Ask the LLM to divide the question into sub-questions…

I remember watching IBM's Watson soundly beat Ken Jennings on Jeapordy. One of the things that sticks out most to me about the memory is that Watson had a confidence rating score for each answer it gave (and there were a few questions for which it had very low confidence). I didn't realize it at the time, but that was actually pretty impressive given the overconfidence issues LLMs have nowadays.

Of course, citing sources would always trump any kind of confidence rating in my mind; sources provide provenance, while confidence rating can be fudged just like a bogus answer can.

Re: Do AI companies work?

#236
post #147

I lead an applied AI research team where I work - which is a mid-sized public enterprise products company. I've been saying this in my professional circles quite often. We talk about scaling laws, superintelligence, AGI etc. But there is another threshold - the ability for humans to leverage super-intelligence. It's just incredibly hard to innovate on products that fully leverage superintelligence. At some point, AI…

We're seeing much less of "it's making mistakes" these days.

Is this because it's actually making less mistakes, or is it just because most people have used it enough now to know not to bother with anything complex?

Re: Do AI companies work?

#237
post #179

Earlier quoted context omitted.

Arguably form factor, i.e. smartphones and tablets. Whether you think that's incremental or not is subjective, I guess. I would say incremental, but it does solve a different set of problems (connectedness, mobility) than the PC.

I would say smartphones and tablets are new devices, not, per se, PCs. Smartphones did bring something huge to the table, and exploded accordingly. But here we are, less than 2 decades from their introduction and... the pace of improvement has gone back to being incremental at best.

They're, at a minimum, PC replacements in the sense that I'm currently on the Internet and posting on a message board from my bathtub instead of needing to go over to my desk to do that.

Re: Do AI companies work?

#238
post #31

I've found that everything that works stops being called AI. Logic programming? AI until SQL came out. Now it's not AI. OCR, computer algebra systems, voice recognition, checkers, machine translation, go, natural language search. All solved, all not AI any more yet all were AI before they got solved by AI researchers. There's even a name for it: https://en.m.wikipedia.org/wiki/AI_effect?utm_source=perplex...

I think the best definition is

AI = machine learning

Yes, that makes a convolutional net trained on recognizing digits an AI.

Re: Do AI companies work?

#239
post #216
post #147

I lead an applied AI research team where I work - which is a mid-sized public enterprise products company. I've been saying this in my professional circles quite often. We talk about scaling laws, superintelligence, AGI etc. But there is another threshold - the ability for humans to leverage super-intelligence. It's just incredibly hard to innovate on products that fully leverage superintelligence. At some point, AI…

* We're seeing much less of "it's making mistakes" these days.* Perhaps less than before, but still making very fundamental errors. Anything involving number I'm automatically suspicious. Pretty frequently I'd get different answers for the same question (to a human). e.g. ChatGPT will give an effective tax rate of n for some income amount. Then when asked to break down the calculation will come up with an effective t…

Yes. Numbers / math is pretty much instant hallucination.

But. Try this approach instead: have it generate python code, with print statements before every bit of math it performs. It will write pretty good code, which you then execute to generate the actual answer.

Simpler example: paste in a paragraph of text, ask it to count the number of words. The answer will be incorrect most of the time.

Instead, ask it to out each word in the text in a numbered list and then output the word count. It will be correct almost always.

My anecdotal learning from this:

LLMs are pretty human-like in their mental abilities. I wouldn't be able to simply look at some text and give you an accurate word count. I would point my finger / cursor to every word and count up.

The solutions above are basically giving LLMs some additional techniques or tools, very similar to how a human may use a calculator, or count words.

In the products we've built, there is an AI feature that generates aggregations of spreadsheet data. We have a dual unittest & aggregator loop to generate correct values.

The first step is to generate some unittests. And in order to generate correct numerical data for unittests, we ask it to write some code with math expressions first. We interpret the expressions, and paste it back into the unittest generator - which then writes the unittests with the correct inputs / outputs.

Then the aggregation generator then generates code until the generated unittests pass completely. Then we have the code for the aggregator function that we can run against the spreadsheet.

Takes a couple of minutes, but pretty bulletproof and also generalizable to other complex math calculations.

Post reply on HN