Earlier quoted context omitted.
And thus we have the AI problems in a nutshell. You think it can reason because it can describe the process in well written language. Anyone who can state the below reasoning clearly "understands" the problem: > For example, in the top‐left 3×3 block (rows 1–3, columns 1–3) the givens are 7, 5, 9, 3, and 4 so the missing digits {1,2,6,8} must appear in the three blank cells. (Later, other intersections force, say, on…
Thanks for spotting this. The solution is indeed wrong. And I agree that the machine can regurgitate plausible reasoning in principle. If it run in a loop, I would bet that it could probably figure this particular problem out eventually, but not sure it matters much in the end. The only plausible way for some of these Sudoku puzzles is a SAT solver and I'm sure that if given the right environment an LLM could just co…
Modern-Day Oracles or Bullshit Machines? How to thrive in a ChatGPT world
551–560 of 652 posts
Re: Modern-Day Oracles or Bullshit Machines? How to thrive in a ChatGPT world
#552Earlier quoted context omitted.
I feel it's impossible for me to trust LLMs can reason when I don't know enough about LLMs to know how much of it is LLM and how much of it is sugarcoating. For example, I've always felt that having the whole thing being a single textbox is reductive and must create all sorts of problems. This thing must parse natural language and output natural language. This doesn't feel necessary. I think it should have some check…
> For example, I've always felt that having the whole thing being a single textbox is reductive and must create all sorts of problems. You observation is correct, but it's not some accident of minimalistic GUI design: The underlying algorithm is itself reductive in a way that can create problems. In essence (e.g. ignoring tokenization), the LLM is doing this: next_word = predict_next(document_word_list, chaos_percent…
I find the current products incredibly helpful in a variety of domains: creating writing in particular, editing my written work, as an interface to web searches (Gemini, in particular, is a rockstar assistant for helping with research), etc etc. But I know perfectly well there's no intelligence behind the curtain, it's really just a text generator.
Re: Modern-Day Oracles or Bullshit Machines? How to thrive in a ChatGPT world
#553Earlier quoted context omitted.
You know, these days I think the abstracts are generated by LLMs too. And the paper. Or at least it uses something like Grammarly. If things keep going this ways typos are going to be a sign of academic integrity.
A proper LLM will include realistic rates of typos eventually. ;)
Re: Modern-Day Oracles or Bullshit Machines? How to thrive in a ChatGPT world
#554What I find frightening is how many are willing to take LLM output at face value. An argument is won or lost not on its merits, but by whether the LLM say so. It was bad enough when people took whatever was written on Wikipedia at face value, trusting an LLM that may have hardcoded biases and is munging whatever data it comes across is so much worse.
Don't be scared of "the many," they're just people, not unlike you.
Re: Modern-Day Oracles or Bullshit Machines? How to thrive in a ChatGPT world
#555Earlier quoted context omitted.
And thus we have the AI problems in a nutshell. You think it can reason because it can describe the process in well written language. Anyone who can state the below reasoning clearly "understands" the problem: > For example, in the top‐left 3×3 block (rows 1–3, columns 1–3) the givens are 7, 5, 9, 3, and 4 so the missing digits {1,2,6,8} must appear in the three blank cells. (Later, other intersections force, say, on…
Thanks for spotting this. The solution is indeed wrong. And I agree that the machine can regurgitate plausible reasoning in principle. If it run in a loop, I would bet that it could probably figure this particular problem out eventually, but not sure it matters much in the end. The only plausible way for some of these Sudoku puzzles is a SAT solver and I'm sure that if given the right environment an LLM could just co…
https://sudoku.com/sudoku-solver
A bummer.
Re: Modern-Day Oracles or Bullshit Machines? How to thrive in a ChatGPT world
#556Kudos; feels very timely! I feel that one underappreciated nuance is why we cannot use human examinations to judge AI. I haven't seen this satisfactorily spelt out anywhere, so I recently wrote a Twitter thread [1], including an example with running -vs- biking. It might be worth making sure your students understand this. Happy to expand on any aspects if you seek. [1] : https://x.com/ergodicthought/status/1887774722…
Perhaps it's no longer being spelled out because it's getting outdated ? In your thread you argue we can't assume AI models generalize the same way we do (which is technically true except maybe not in the limit), but you seem to be worried about the extent of generalization ability (like learning to run vs. bike example, in terms of generalizing from either to climbing stairs). Thing is, people made these objections…
With due deference to the title of the top-level post, I'm tempted to call bullshit unless your claim can be justified.
Just because a single model can do a handful of things you've listed doesn't mean that its capabilities are not "jagged"; you've just cherry-picked a few things it can do among the countless things it cannot yet. If AI really were so good at every single task, then (for example) it wouldn't matter much how you prompt it.
PS: I really do want to debate this further and understand your perspective, so I will reach out for continuing discussion.
Re: Modern-Day Oracles or Bullshit Machines? How to thrive in a ChatGPT world
#557This website is so important! Now ask yourself why AI companies don't want to be regulated or scrutinized. So many companies (users and providers) jump on the AI hype train because of FOMO. The end result might be just as destructive as this mythical "AGI". Edit: I am not saying to not use the technology. I am just on the side of caution and constant validation. The technology has to serve society. But I fear this hy…
Re: Modern-Day Oracles or Bullshit Machines? How to thrive in a ChatGPT world
#558The sys prompt given to run the turing test (from https://arxiv.org/pdf/2405.08007 ) actually works well. I'm honestly not sure I'd be able to tell (unless I test it adversarially e.g ignore the prompt, write a poem etc)
Re: Modern-Day Oracles or Bullshit Machines? How to thrive in a ChatGPT world
#559Earlier quoted context omitted.
I feel it's impossible for me to trust LLMs can reason when I don't know enough about LLMs to know how much of it is LLM and how much of it is sugarcoating. For example, I've always felt that having the whole thing being a single textbox is reductive and must create all sorts of problems. This thing must parse natural language and output natural language. This doesn't feel necessary. I think it should have some check…
> For example, I've always felt that having the whole thing being a single textbox is reductive and must create all sorts of problems. You observation is correct, but it's not some accident of minimalistic GUI design: The underlying algorithm is itself reductive in a way that can create problems. In essence (e.g. ignoring tokenization), the LLM is doing this: next_word = predict_next(document_word_list, chaos_percent…
That sounds like an interesting application of the technology! So you could for example train an LLM on piano songs, and if someone played a few notes it would autocomplete with the probable next notes, for example?
>The underlying algorithm is itself reductive in a way that can create problems
I wonder if in the future we'll see some refinement of this. The only experience I have with AI is limited to trying Stable Diffusion, but SD does have many options you can try to configure like number of steps, samplers, CFG, etc. I don't know exactly what each of these settings do, and I bet most people who use it don't either, but at least the setting is there.
If hallucinations are intrinsic of LLMs perhaps the way forward isn't trying to get rid of them to create the perfect answer machine/"oracle" but just figure out a way to make use of them. It feels to me that the randomness of AI could help a lot with creative processes, brainstorming, etc., and for that purpose it needs some configurability. For example, Youtube rolled out an AI-based tool for Youtubers that generates titles/thumbnails of videos for them to make. Presumably, it's biased toward successful titles. The thumbnails feel pretty unnecessary, though, since you wouldn't want to use the obvious AI thumbnails.
I hear a lot of people say AI is a new industry with a lot of potential when they mean it will become AGI eventually, but these things make me feel like its potential isn't to become the an oracle but to become something completely different instead that nobody is thinking about because they're so focused on creating the oracle.
Thanks for the reply, by the way. Very informative. :)