Earlier quoted context omitted.
It’s important because solving a math problem requires you to actually understand something and follow deliberate steps. The fact that they can’t means they’re just a toy ultimately.
No, I disagree. It is just deliberate steps. Understanding can greatly help you do the steps and remember which ones to do. Training math is likely hard because the corpus of training data is so much less because the computers themselves do our math as it relates to computers. You can draft text on a computer in just ascii but drafting long division is something that most people wouldn’t do in some sort of digital te…
ChatGPT-4o vs. Math
151–160 of 182 posts
Re: ChatGPT-4o vs. Math
#152Earlier quoted context omitted.
At the same time the ChatGPT app has access to write and run python, which the gpt can choose to do when it thinks it needs more accuracy.
The results from playing with this are really bizarre: (sorry, formatting hacked up a bit) To calculate 7^1.83 , you can use a scientific calculator or an exponentiation function in programming or math software. Here is the step-by-step calculation using a scientific calculator: Input the base: 7 Use the exponentiation function (usually labeled as ^ or x^y). Input the exponent: 1.83 Compute the result. Using these st…
Re: ChatGPT-4o vs. Math
#153Earlier quoted context omitted.
> Does anyone know how far off we are having logical AI? 1847, wasn't it? (George Boole). Or 1950-60 (LISP) or 1989 (Coq) depending on your taste? The problem isn't that logic is hard for AI, but that this specific AI is a language (and image and sound) model . It's wild that transformer models can get enough of an understanding of free-form text and images to get close, but using it like this is akin to using a batt…
By that same logic isn't that a similar process that we humans use as well ? Kind of seems like the whole point of "AI" (replicating the human experience)
Re: ChatGPT-4o vs. Math
#154Earlier quoted context omitted.
I can't wait for the day when instead of engineering disciplines solving problems with knowledge and logic they're instead focused on AI/LLM psychology and the correct rituals and incantations that are needed to make the immensely powerful machines at our disposal actually do what we've asked for. /s
"No dude, the bribe you offered was too much so the LLM got spooked, you need to stay in a realistic range. We've fine-tuned a local model on realistic bribe amounts sourced via Mechanical Turk to get a good starting point and then used RLMF to dial in the optimal amount by measuring task performance relative to bribe."
Re: ChatGPT-4o vs. Math
#155The model's first attempt is impressive (not sure why it's labeled a choke). Unfortunately gpt4o cannot discover calculus on its own.
I think this is the biggest flaw in LLMs and what is likely going to sour a lot of businesses on their usage (at least in their current state). It is preferable to give the right answer to a query, it is acceptable to be unable to answer a query - we run into real issues, though, when a query is confidently answered incorrectly. This recently caused a major headache for AirCanada - businesses should be held to the st…
Re: ChatGPT-4o vs. Math
#156Re: ChatGPT-4o vs. Math
#157Need to run this experiment on a problem that is t already on its training set.
It’d be interesting to obfuscate it so that the concepts are all the same but the subject matter would be hard to google for. Another failure mode for LLM’s is if you take something with a well-known answer and add a twist to it.
Re: ChatGPT-4o vs. Math
#158I actually have a contrarian view: being able to do elementary math is not that important in the current stage. Yes, understanding elementary math is a cornerstone for an AI to become more intelligent, but also let's be honest: LLMs are far from being AGIs and does not have common sense nor general ability to deduce or induct. If we accept such limitation of LLM, then focusing the mathematical understanding of an LLM…
if you sampled N random people on the street and asked them to solve this problem, what would the outcome be? would it be better than asking chatgpt N times? I wonder
Also HN: "ChatGPT just needs to spell its name."
Re: ChatGPT-4o vs. Math
#159Earlier quoted context omitted.
if you sampled N random people on the street and asked them to solve this problem, what would the outcome be? would it be better than asking chatgpt N times? I wonder
HN: "Tesla needs to be 100x safer than the best human drivers!!!" Also HN: "ChatGPT just needs to spell its name."
Besides, only half of HN is all self-driving has to be 100x safer. the other half keeps bringing up the fact that Waymo is here and working, just not everywhere yet.
Re: ChatGPT-4o vs. Math
#160Earlier quoted context omitted.
HN: "Tesla needs to be 100x safer than the best human drivers!!!" Also HN: "ChatGPT just needs to spell its name."
While words have power too, I'm not driving next to ChatGPT on the freeway where it's going to immediately kill or maim me if it hallucinates. Besides, only half of HN is all self-driving has to be 100x safer. the other half keeps bringing up the fact that Waymo is here and working, just not everywhere yet.
Doctors: https://www.news-medical.net/news/20240424/Opportunities-and....
Lawyers: https://www.ft.com/content/2365b275-8b0b-4ae9-bc4a-15c6e0776...
Or my fav: https://www.axios.com/2023/10/31/guardian-microsoft-generati...