Earlier quoted context omitted.
> Try some context engineering. Try customizing your harness. Try having the harness improve itself. These are AI 101 lessons If AI is really that complicated, it sounds like it would be easier to write an ordinary computer program to aggregate job boards.
Yes, one of the obvious ways to me to use these models is to tell them to write such a program. It can then go figure out data extraction and normalization. This is "the harness improving itself". Have it write tools to do its tasks.
What sort of maths are LLMs good at?
41–50 of 189 posts
Re: What sort of maths are LLMs good at?
#42You can toss it at a task with a suitable machine for transforming that raw material into action and it'll rattle through and sample "plausible human behavior" at that endpoint.
There are more clever ways to use it, but a general tool here is to upgrade any sort of stochastic search to use this new form of random sampling. It'll be way more efficient, properly conditioned, because it just won't visit implausible things nearly as often as competing random sources.
Re: What sort of maths are LLMs good at?
#43Earlier quoted context omitted.
> Try some context engineering. Try customizing your harness. Try having the harness improve itself. These are AI 101 lessons If AI is really that complicated, it sounds like it would be easier to write an ordinary computer program to aggregate job boards.
I totally agree - designing a competent AI agent with a fully customized harness to successfully pull off this task is a much more challenging engineering effort than merely creating an ordinary computer program. Had OP made chatgpt write an ordinary program instead, they likely would have succeeded in their task.
You have converted a fuzzy task into a conventional software engineering problem, and then relying on conventional software for the reliability :-)
Re: What sort of maths are LLMs good at?
#44> A good sign that LLMs have reached human level for a much wider class of problems will be if they start proving theorems using methods that, like much of the very best human mathematics, are new and surprising but that with hindsight come to seem beautiful and natural. They should also be methods that are difficult to stumble on by accident. It is hard to say precisely what would count as such a proof, but I think…
So then you need to explain ARC-AGI-3: https://arxiv.org/abs/2603.24621
"Our testing shows humans can solve 100% of the environments, in contrast to frontier AI systems which, as of March 2026, score below 1%."
Back 1996, EQP automatically solved the Robbins conjecture. But nobody concluded EQP was generally intelligent.
Re: What sort of maths are LLMs good at?
#45A thoughtful and measured post, as usual from Gowers. The final note is neat and worth pasting out here in full: > A good sign that LLMs have reached human level for a much wider class of problems will be if they start proving theorems using methods that, like much of the very best human mathematics, are new and surprising but that with hindsight come to seem beautiful and natural. They should also be methods that ar…
>>A good sign that LLMs have reached human level for a much wider class >> of problems will be if they start proving theorems using methods that, like much of the very best human mathematics, are new and surprising but that with hindsight come to seem beautiful and natural. I must be taking crazy pills and the AGI surely will pass me by... But TODAY, middle August 2026...And in the context of testing and evaluating t…
Anything that can be verified mechanically should be code. Only use LLMs to fill in the gaps where things are fuzzy. Don't fall for the idea that those harnesses are general purpose, make your own fit to your task with the guards and verification steps you need. Make the LLM create the harness even.
There is no amount of markdown that can make a machine generating plausible text generate truthful text, it just happens to be truthful because of what it was trained on. Nothing coming out of an LLM should be taken at face value.
The propaganda about LLMs being intelligent and able to "reason" is only serving the companies selling you tokens to waste on "prompt engineering".
Re: What sort of maths are LLMs good at?
#46Re: What sort of maths are LLMs good at?
#47Earlier quoted context omitted.
I found it very difficult to parse your description, "Clear, not too long not too short prompt, for LLMs to go out and research freelance consulting gigs for one specific IT domain, and in one specific country in Europe, including maybe opportunities driven from temp agencies based in geographically close countries." ( I was trying to quote a single sentence and then realised it ran on for the whole paragraph. ) Give…
You can easily test this yourself with the SOTA models....or read the corroborating literature... "General365: Benchmarking General Reasoning in Large Language Models Across Diverse and Challenging Tasks" https://arxiv.org/abs/2604.11778 "...General365, a benchmark specifically designed to assess general reasoning in LLMs. By restricting background knowledge to a K-12 level, General365 explicitly decouples reasoning…
I particularly found the note about “local collapse” helpful (near the end of section 3). The idea is that even though benchmarks contain a wide variety of different reasoning tasks, each individual problem requires only a few skills - unlike this benchmark where they deliberately construct tasks that span many categories.
Re: What sort of maths are LLMs good at?
#48Earlier quoted context omitted.
Based on the rest of your writing I’m going to assume that the prompt was the problem.
He was perfectly clear in both cases. If a human misunderstood this, they'd be a dumb human.
Re: What sort of maths are LLMs good at?
#49Earlier quoted context omitted.
What a sloppy reply. You've hijacked a thread on mathematics first to complain that your incompetent attempt to use ChatGPT to find a job failed, but it seems now that this was a ruse to instead begin arguments unrelated to the article at all where you just spam arxiv links you've never read to "prove" that AI is a scam. This comes across, frankly, as either Dunning-Kruger (classic illusory superiority), or potential…
Why are you so upset that someone is criticising LLMs that you call them schizophrenic? (I'd recommend refreshing your memory with this https://news.ycombinator.com/newsguidelines.html )
Re: What sort of maths are LLMs good at?
#50Earlier quoted context omitted.
> What does [ 13 ] even mean? It means that it is possible for someone to be wearing a white shirt and yellow pants (say), but the person in green shoes came from the set of a Wes Anderson film.
That's cute, but it's driving me crazy, I guess I'll have to sit down and solve it to figure out how it's meant to be clued, assuming it's not just a red-herring entirely.
white, blue, yellow, and is otherwise undefined.
But we know from [1], [2], [4], [5] and [11], that the order must be:
A, B, C, D, E or A, B, E, D, C.
Which makes C either the older 12 year old or the 35 year old.
The key this is that D can't be a 12 year old without A being in slot 2, but A can't be in slot 2 because slot 2 is South Africa and A is Morocco.
Trying to reach the shoes + Vanuatu clue is a complete waste of time, paying any attention to shoes or clothing is a waste of time, it feels like there ought to be a way to narrow it down to one of those two configurations, but the clothing is too ambiguous, the shoes end up irrelevant.
What a frustrating puzzle, where half the clues are seemingly redundant.