Live data from Hacker News

Modern-Day Oracles or Bullshit Machines? How to thrive in a ChatGPT world

thebullshitmachines.com

551–560 of 652 posts

Re: Modern-Day Oracles or Bullshit Machines? How to thrive in a ChatGPT world

#551
post #527

Earlier quoted context omitted.

And thus we have the AI problems in a nutshell. You think it can reason because it can describe the process in well written language. Anyone who can state the below reasoning clearly "understands" the problem: > For example, in the top‐left 3×3 block (rows 1–3, columns 1–3) the givens are 7, 5, 9, 3, and 4 so the missing digits {1,2,6,8} must appear in the three blank cells. (Later, other intersections force, say, on…

Thanks for spotting this. The solution is indeed wrong. And I agree that the machine can regurgitate plausible reasoning in principle. If it run in a loop, I would bet that it could probably figure this particular problem out eventually, but not sure it matters much in the end. The only plausible way for some of these Sudoku puzzles is a SAT solver and I'm sure that if given the right environment an LLM could just co…

[deleted]

Re: Modern-Day Oracles or Bullshit Machines? How to thrive in a ChatGPT world

#552
post #522

Earlier quoted context omitted.

I feel it's impossible for me to trust LLMs can reason when I don't know enough about LLMs to know how much of it is LLM and how much of it is sugarcoating. For example, I've always felt that having the whole thing being a single textbox is reductive and must create all sorts of problems. This thing must parse natural language and output natural language. This doesn't feel necessary. I think it should have some check…

> For example, I've always felt that having the whole thing being a single textbox is reductive and must create all sorts of problems. You observation is correct, but it's not some accident of minimalistic GUI design: The underlying algorithm is itself reductive in a way that can create problems. In essence (e.g. ignoring tokenization), the LLM is doing this: next_word = predict_next(document_word_list, chaos_percent…

This is a great explanation of a point I've been trying to make for a while, when talking to friends about LLMs, but haven't been able to put quite so succinctly. LLMs are text generators, no more, no less. That has all sorts of useful applications! But (OAI and friends) marketing departments are so eager to push the Intelligence part of AI that it's become straight-up snakeoil.. there is no intelligence to be found, and there never will be as long as we stay the course on transformers-based models (and, as far as I know, nobody has tried to go back to the drawing board yet). Actual, real AI will probably come one day, but nobody is working on it yet, and it probably won't even be called "AI" at that point because the term has been poisoned by the current trends. IMO there's no way to correct the course on the current set of AI/LLM products.

I find the current products incredibly helpful in a variety of domains: creating writing in particular, editing my written work, as an interface to web searches (Gemini, in particular, is a rockstar assistant for helping with research), etc etc. But I know perfectly well there's no intelligence behind the curtain, it's really just a text generator.

Re: Modern-Day Oracles or Bullshit Machines? How to thrive in a ChatGPT world

#553

Earlier quoted context omitted.

You know, these days I think the abstracts are generated by LLMs too. And the paper. Or at least it uses something like Grammarly. If things keep going this ways typos are going to be a sign of academic integrity.

A proper LLM will include realistic rates of typos eventually. ;)

Darn.

Re: Modern-Day Oracles or Bullshit Machines? How to thrive in a ChatGPT world

#554

What I find frightening is how many are willing to take LLM output at face value. An argument is won or lost not on its merits, but by whether the LLM say so. It was bad enough when people took whatever was written on Wikipedia at face value, trusting an LLM that may have hardcoded biases and is munging whatever data it comes across is so much worse.

> frightening

Don't be scared of "the many," they're just people, not unlike you.

Re: Modern-Day Oracles or Bullshit Machines? How to thrive in a ChatGPT world

#555
post #527

Earlier quoted context omitted.

And thus we have the AI problems in a nutshell. You think it can reason because it can describe the process in well written language. Anyone who can state the below reasoning clearly "understands" the problem: > For example, in the top‐left 3×3 block (rows 1–3, columns 1–3) the givens are 7, 5, 9, 3, and 4 so the missing digits {1,2,6,8} must appear in the three blank cells. (Later, other intersections force, say, on…

Thanks for spotting this. The solution is indeed wrong. And I agree that the machine can regurgitate plausible reasoning in principle. If it run in a loop, I would bet that it could probably figure this particular problem out eventually, but not sure it matters much in the end. The only plausible way for some of these Sudoku puzzles is a SAT solver and I'm sure that if given the right environment an LLM could just co…

Yeah I think this was a wrong puzzle to try according to:

https://sudoku.com/sudoku-solver

A bummer.

Re: Modern-Day Oracles or Bullshit Machines? How to thrive in a ChatGPT world

#556
post #9

Kudos; feels very timely! I feel that one underappreciated nuance is why we cannot use human examinations to judge AI. I haven't seen this satisfactorily spelt out anywhere, so I recently wrote a Twitter thread [1], including an example with running -vs- biking. It might be worth making sure your students understand this. Happy to expand on any aspects if you seek. [1] : https://x.com/ergodicthought/status/1887774722…

Perhaps it's no longer being spelled out because it's getting outdated ? In your thread you argue we can't assume AI models generalize the same way we do (which is technically true except maybe not in the limit), but you seem to be worried about the extent of generalization ability (like learning to run vs. bike example, in terms of generalizing from either to climbing stairs). Thing is, people made these objections…

> But we went past that very rapidly, and for the past half a year or so, we've already seen models excelling at every single task listed above simultaneously. Same architecture, same basic training approach, few extra modalities, ever growing capabilities.

With due deference to the title of the top-level post, I'm tempted to call bullshit unless your claim can be justified.

Just because a single model can do a handful of things you've listed doesn't mean that its capabilities are not "jagged"; you've just cherry-picked a few things it can do among the countless things it cannot yet. If AI really were so good at every single task, then (for example) it wouldn't matter much how you prompt it.

PS: I really do want to debate this further and understand your perspective, so I will reach out for continuing discussion.

Re: Modern-Day Oracles or Bullshit Machines? How to thrive in a ChatGPT world

#557

This website is so important! Now ask yourself why AI companies don't want to be regulated or scrutinized. So many companies (users and providers) jump on the AI hype train because of FOMO. The end result might be just as destructive as this mythical "AGI". Edit: I am not saying to not use the technology. I am just on the side of caution and constant validation. The technology has to serve society. But I fear this hy…

It absolutely is destructive. I read an opinion the other day about Microsoft shoving Copilot into every product, and it kinda makes sense. Paraphrasing but: In MS's ideal world, worker 1 drafts a few bullet points and asks Copilot to expand it into a multi-paragraph email. Worker 2 asks Copilot to summarize the email back into bullet points, then acts on it. What's the point? Well, both workers are paying for Copilot licenses, so MS has already won. And management at the firm is happy because "we're using AI, we're so modern." But did it actually help, with anything, at all? Never mind the amount of wasted energy and resources blasting LLM-generated content (that no human will ever read) back and forth.

Re: Modern-Day Oracles or Bullshit Machines? How to thrive in a ChatGPT world

#558

The sys prompt given to run the turing test (from https://arxiv.org/pdf/2405.08007 ) actually works well. I'm honestly not sure I'd be able to tell (unless I test it adversarially e.g ignore the prompt, write a poem etc)

Curiously, this is the same point we were at in the 60s: https://en.wikipedia.org/wiki/ELIZA

Re: Modern-Day Oracles or Bullshit Machines? How to thrive in a ChatGPT world

#559
post #522

Earlier quoted context omitted.

I feel it's impossible for me to trust LLMs can reason when I don't know enough about LLMs to know how much of it is LLM and how much of it is sugarcoating. For example, I've always felt that having the whole thing being a single textbox is reductive and must create all sorts of problems. This thing must parse natural language and output natural language. This doesn't feel necessary. I think it should have some check…

> For example, I've always felt that having the whole thing being a single textbox is reductive and must create all sorts of problems. You observation is correct, but it's not some accident of minimalistic GUI design: The underlying algorithm is itself reductive in a way that can create problems. In essence (e.g. ignoring tokenization), the LLM is doing this: next_word = predict_next(document_word_list, chaos_percent…

>one could potentially train an LLM on musical notation of millions of songs, as long as you can find a way to express each one as a linear sequence of tokens.

That sounds like an interesting application of the technology! So you could for example train an LLM on piano songs, and if someone played a few notes it would autocomplete with the probable next notes, for example?

>The underlying algorithm is itself reductive in a way that can create problems

I wonder if in the future we'll see some refinement of this. The only experience I have with AI is limited to trying Stable Diffusion, but SD does have many options you can try to configure like number of steps, samplers, CFG, etc. I don't know exactly what each of these settings do, and I bet most people who use it don't either, but at least the setting is there.

If hallucinations are intrinsic of LLMs perhaps the way forward isn't trying to get rid of them to create the perfect answer machine/"oracle" but just figure out a way to make use of them. It feels to me that the randomness of AI could help a lot with creative processes, brainstorming, etc., and for that purpose it needs some configurability. For example, Youtube rolled out an AI-based tool for Youtubers that generates titles/thumbnails of videos for them to make. Presumably, it's biased toward successful titles. The thumbnails feel pretty unnecessary, though, since you wouldn't want to use the obvious AI thumbnails.

I hear a lot of people say AI is a new industry with a lot of potential when they mean it will become AGI eventually, but these things make me feel like its potential isn't to become the an oracle but to become something completely different instead that nobody is thinking about because they're so focused on creating the oracle.

Thanks for the reply, by the way. Very informative. :)

Post reply on HN