Live data from Hacker News

Caltech Mathathon – first hackathon ever devoted to research level mathematics

mathathonchallenge.com

31–40 of 113 posts

Re: Caltech Mathathon – first hackathon ever devoted to research level mathematics

#31
post #10

Should I read this as the big labs trying to move maths forward? Or the big labs trying to use professional mathematitians as (cheap?) Labour for validating LLM outputs? On yesterdays "An Alien Mind" post from openAI they openly said that maths is not a priority for them, so I personally know what to think...

Obviously the latter - this fact is betrayed by how the page lists the S^6 complex structure result, which was released as a 100 page barely readable mess (in fact even this might be too charitable), as still "unverified". Clearly a situation labs would like to avoid for future claimed results.

Re: Caltech Mathathon – first hackathon ever devoted to research level mathematics

#32
post #15

Earlier quoted context omitted.

Prompting for some of the results was almost the "Computer, do a breakthrough. Make no mistakes." meme. Just someone telling the model to keep trying a couple of times. Unfortunately we don't actually know what kind of prompting was done for the more prominent results.

It’s definitely not how I work, I’d need to read the model responses and set a direction for the model to go

Same here, it's one thing if I'm just screwing around, but if I'm trying to do anything serious, I need to at least have a handle on what it's doing, and, when thinking traces are available, keeping track of any logical errors in the model's reasoning.

Re: Caltech Mathathon – first hackathon ever devoted to research level mathematics

#33
No it's not, there are tons of programs in mathematics where you go there, form teams, work a problem as a team that's likely to get a result, and then publish the results from all the groups in the conference proceedings. They are called Research Collaboration Workshops.

Very consistent pattern from these tech companies in their mathematics press releases that shows a conspicuous lack of experience in the research math world.

Re: Caltech Mathathon – first hackathon ever devoted to research level mathematics

#34
recent caltech grad here! and know some of the organizers well

caltech's cs department is very, very weak, and has struggled to recruit top people in the last few years, and the most recent AI faculty hires have had issues. this is very slowly changing but a lot of the motivation for htis was to create a way for students to get ml "recognition" and learn about ai since it cant be done through the school right now. really glad to see hn picked this up!

Re: Caltech Mathathon – first hackathon ever devoted to research level mathematics

#36
post #32

Earlier quoted context omitted.

It’s definitely not how I work, I’d need to read the model responses and set a direction for the model to go

Same here, it's one thing if I'm just screwing around, but if I'm trying to do anything serious, I need to at least have a handle on what it's doing, and, when thinking traces are available, keeping track of any logical errors in the model's reasoning.

you know the thinking traces are redacted and summarized using another model, right? The actual thinking trace looks something like:

7♣-removal-IS-the-prerequisite-for-10♠/9♥!!)-⟹-OVERLAP-(ii)+(iv):-{6♠ J♦ 9♥ 2♣}-=-FOUR--—-UNLESS-7♣'s-seat-8♥-...-and-2♣-drains-only-at-crack-:-⟹-2♣-celled-+-9♥-celled-simultaneously-UNAVOIDABLE-in-t8-dig--—-BREAK:-9♥-drains-to-10♠-THE-MOMENT-10♠-is-free:-t8-dig-order:-[K♣→t2]-[2♣→cell]-[7♣→8♥-:-8♥-WHERE:-post-chunk-9♠-:-chunk-⟸-K♣--done-:-ORDER:-K♣→t2,-CHUNK→K♣-(cap-4!!:-cells-then:-{6♠ J♦ 9♥}-FULL--chunk-cap-=-1--—-F-F-F-F-F.-—-chunk-BEFORE-9♥-celling?!-:-9♥-celled-at-t1-dig-⟸-needed-for-5♣-⟸-4♥3♣-⟸-t2-dig-⟸-K♣-seat-⟸-chunk-:-⟹-chunkAFTER-9♥-celling-FORCED--:-chunk-cap-with-{6♠ J♦ 9♥}:-1--—-—-J♦-THE-NEW-CANCER.-—-⟹-J♦-celling-DELAYED-till-after-chunk?! :-J♦'s-celling-was-for-J♥→Q♠-(5♦-access-for-4♣):-DELAY-4♣-resolution:-4♣→CELLearly-(as-always)-then-4♣-cell→5♦-LATER-when-5♦-frees-!!!:-cells-rotation:-4♣-celled-[t2-dig-…-5♦-freed]:-5♦-freed-⟸-J♥→Q♠-⟸-J♦-celled-:-⟹-{6♠, 4♣, J♦}-overlap-window-until-4♣→5♦-drains:-then-{6♠ J♦}+1-rotator-:-—-AND-9♥?!-9♥-celled-[t1-dig…]:-OVERLAP-{6♠ 4♣ 9♥}-before-J♦-even-:-⟹-rotator-slot-SINGLE:-timeline-:-(1)-{6♠}+2:-…-(2)-+9♥-(t1-dig):-{6♠ 9♥}+1:-(3)-+4♣-(t2-dig):-{6♠ 9♥ 4♣}-FULL-:-(4)-NEED:-t6-dig-(9♦8♠→10♣--no-cell;-8♥→CELL--FULL)--—-8♥-alternative-seat-pre-chunk:-NONE-—-.-⟹-⟹-THE-TRIANGLE-{9♥ 4♣ 8♥}-verdammt.-—-⟹-dig-t6-BEFORE-t2?!:-(3')-+8♥:-{6♠ 9♥ 8♥}-FULL:-J♥→Q♠-⟸-J♦-cell--FULL--AAAAAAAAAAAARGH.

Citation: https://www-cdn.anthropic.com/d00db56fa754a1b115b6dd7cb2e3c3..., section 6.2.2

You're not going to get a handle on what it's doing. The thinking traces are there to make you feel better about yourself.

Re: Caltech Mathathon – first hackathon ever devoted to research level mathematics

#37
post #10

Should I read this as the big labs trying to move maths forward? Or the big labs trying to use professional mathematitians as (cheap?) Labour for validating LLM outputs? On yesterdays "An Alien Mind" post from openAI they openly said that maths is not a priority for them, so I personally know what to think...

> Or the big labs trying to use professional mathematitians as (cheap?) Labour for validating LLM outputs?

I am told AGI has been achieved. If so, shouldnt these systems be out and about on their own ? Looking at 1st proof submissions in batch 2 it is clear that fully autonomous AI systems have a long way to go.

AI harnessing human labor with the incentive of 2M in free tokens is the way my skeptic eye sees it, or humans being duped as reverse-centaurs.

Re: Caltech Mathathon – first hackathon ever devoted to research level mathematics

#38

Earlier quoted context omitted.

If the goal is to accomplish something then why limit yourself with available tools? I’m not a full on AI optimist but it is absolutely the most powerful tool in a host of applications. From a Hackathon perspective, obviously in the 90s it was much more unorganized, but the same ethos existed. Use all available tools to accomplish the goal/task, it’s where a lot of incredible learning came out of. The same will hopef…

It seems like making progress on math is letting the AI run fully autonomously for a few days, occasionally asking it to keep going. I'm not sure people need to organize a mathathon to wait for a computer to give a printout. They mainly need tokens.

That's how several major AI advancements have happened. I have seen no evidence that that is the fastest way to make progress right now. I expect that, much like chess engines, it will not take too long before AI is significantly better than AI + human. But right now, my bet is that we are still safely within the window where an AI + human mathematician team is still better than AI alone (at least for the case where the human has learned how to work effectively with the partner....something that this event could possible be good for teaching).

Re: Caltech Mathathon – first hackathon ever devoted to research level mathematics

#40
post #32

Earlier quoted context omitted.

Same here, it's one thing if I'm just screwing around, but if I'm trying to do anything serious, I need to at least have a handle on what it's doing, and, when thinking traces are available, keeping track of any logical errors in the model's reasoning.

you know the thinking traces are redacted and summarized using another model, right? The actual thinking trace looks something like: 7♣-removal-IS-the-prerequisite-for-10♠/9♥!!)-⟹-OVERLAP-(ii)+(iv):-{6♠ J♦ 9♥ 2♣}-=-FOUR--—-UNLESS-7♣'s-seat-8♥-...-and-2♣-drains-only-at-crack-:-⟹-2♣-celled-+-9♥-celled-simultaneously-UNAVOIDABLE-in-t8-dig--—-BREAK:-9♥-drains-to-10♠-THE-MOMENT-10♠-is-free:-t8-dig-order:-[K♣→t2]-[2♣→cell]…

I'm referring to models where the actual traces are available. Eg. I've been using a local Qwen3.8-Next-Flash lately.
Post reply on HN