Earlier quoted context omitted.
you know the thinking traces are redacted and summarized using another model, right? The actual thinking trace looks something like: 7♣-removal-IS-the-prerequisite-for-10♠/9♥!!)-⟹-OVERLAP-(ii)+(iv):-{6♠ J♦ 9♥ 2♣}-=-FOUR--—-UNLESS-7♣'s-seat-8♥-...-and-2♣-drains-only-at-crack-:-⟹-2♣-celled-+-9♥-celled-simultaneously-UNAVOIDABLE-in-t8-dig--—-BREAK:-9♥-drains-to-10♠-THE-MOMENT-10♠-is-free:-t8-dig-order:-[K♣→t2]-[2♣→cell]…
I'm referring to models where the actual traces are available. Eg. I've been using a local Qwen3.8-Next-Flash lately.
Caltech Mathathon – first hackathon ever devoted to research level mathematics
41–50 of 109 posts
Re: Caltech Mathathon – first hackathon ever devoted to research level mathematics
#42- We are a team of undergrads at Caltech. We don't represent Caltech, any Caltech departments, or any of our sponsors.
- We don't receive monetary compensation. All the funding raised goes toward paying our judges and participants.
- Our goal is to promote responsible AI use. You can read more about our commitments here: https://mathathonchallenge.com/faq.html
Re: Caltech Mathathon – first hackathon ever devoted to research level mathematics
#43Should I read this as the big labs trying to move maths forward? Or the big labs trying to use professional mathematitians as (cheap?) Labour for validating LLM outputs? On yesterdays "An Alien Mind" post from openAI they openly said that maths is not a priority for them, so I personally know what to think...
> Or the big labs trying to use professional mathematitians as (cheap?) Labour for validating LLM outputs? I am told AGI has been achieved. If so, shouldnt these systems be out and about on their own ? Looking at 1st proof submissions in batch 2 it is clear that fully autonomous AI systems have a long way to go. AI harnessing human labor with the incentive of 2M in free tokens is the way my skeptic eye sees it, or hu…
This seems like a strawman. It’s certainly not consensus that AGI has been achieved and I don’t think the people participating in this event feel like there is no value in human input or steering the AI.
Re: Caltech Mathathon – first hackathon ever devoted to research level mathematics
#44recent caltech grad here! and know some of the organizers well caltech's cs department is very, very weak, and has struggled to recruit top people in the last few years, and the most recent AI faculty hires have had issues. this is very slowly changing but a lot of the motivation for htis was to create a way for students to get ml "recognition" and learn about ai since it cant be done through the school right now. re…
Re: Caltech Mathathon – first hackathon ever devoted to research level mathematics
#45Earlier quoted context omitted.
It seems like making progress on math is letting the AI run fully autonomously for a few days, occasionally asking it to keep going. I'm not sure people need to organize a mathathon to wait for a computer to give a printout. They mainly need tokens.
That's how several major AI advancements have happened. I have seen no evidence that that is the fastest way to make progress right now. I expect that, much like chess engines, it will not take too long before AI is significantly better than AI + human. But right now, my bet is that we are still safely within the window where an AI + human mathematician team is still better than AI alone (at least for the case where…
You can probably ask the AI to come up with a list of problems itself, and rank them by the likelihood of progress.
Re: Caltech Mathathon – first hackathon ever devoted to research level mathematics
#46Earlier quoted context omitted.
I'm referring to models where the actual traces are available. Eg. I've been using a local Qwen3.8-Next-Flash lately.
Still not meaningful -- https://arxiv.org/pdf/2504.09762 ; even for local models, the reasoning traces are often filtered and summarized to sound sensible to humans. And even if not, they don't necessarily represent what the model is thinking.
Re: Caltech Mathathon – first hackathon ever devoted to research level mathematics
#47Hey guys, I'm one of the organizers. AMA. - We are a team of undergrads at Caltech. We don't represent Caltech, any Caltech departments, or any of our sponsors. - We don't receive monetary compensation. All the funding raised goes toward paying our judges and participants. - Our goal is to promote responsible AI use. You can read more about our commitments here: https://mathathonchallenge.com/faq.html
Re: Caltech Mathathon – first hackathon ever devoted to research level mathematics
#48Won't deny that this is an interesting idea, but I feel like waiting on the output of an LLM for 40 hours feels like it is completely antithetical to what makes classic Hackathons appealing / educative. More generally, I don't think the shape of a hackathon (intensely working for a short timespan) maps at all onto the way LLM Math progress has seemingly been made so far; AFAIK it mostly involves picking out something…
2. We have talked to mathematicians and frontier lab employees. We think 40 hours is enough to produce interesting results.
Re: Caltech Mathathon – first hackathon ever devoted to research level mathematics
#49Hey guys, I'm one of the organizers. AMA. - We are a team of undergrads at Caltech. We don't represent Caltech, any Caltech departments, or any of our sponsors. - We don't receive monetary compensation. All the funding raised goes toward paying our judges and participants. - Our goal is to promote responsible AI use. You can read more about our commitments here: https://mathathonchallenge.com/faq.html
Hey! Any indication on what area of mathematics theses questions are from?
Re: Caltech Mathathon – first hackathon ever devoted to research level mathematics
#50Should I read this as the big labs trying to move maths forward? Or the big labs trying to use professional mathematitians as (cheap?) Labour for validating LLM outputs? On yesterdays "An Alien Mind" post from openAI they openly said that maths is not a priority for them, so I personally know what to think...