Live data from Hacker News

A path to O1 open source

arxiv.org

71–80 of 86 posts

Re: A path to O1 open source

#71

I don’t think anyone who upvoted this has read more than one sentence of this paper.

It’s interesting even if not true or correct. You could also choose to enrich the discussion by elaborating on why you think this is worthless instead.

I have a hard time giving worth to a paper whose first sentence fails to spell intelligence correctly.

Re: A path to O1 open source

#72
post #69

From the first few paragraphs it doesn't pass the sniff test for me. "Now AI has made everything more complex!" "AI is embedded in everything we do"... Sounds like marketing gibberish and obfuscation, combined with self promotion. That's just my read at first sniff.

This is absolutely a worthless fluff paper

flagged it. more people should flag this kinda stuff.

Re: A path to O1 open source

#74
post #69

Earlier quoted context omitted.

This is absolutely a worthless fluff paper

flagged it. more people should flag this kinda stuff.

"Doesn't pass my sniff test" is not the purpose of the flag button. Furthermore, it passes my personal sniff test: hundreds of people upvoting it while the top comment is saying it's worthless. Usually the real alpha is in the comments under such things.

Re: A path to O1 open source

#76
post #64

> has claimed that the main techinique behinds o1 is the reinforcement learining. Typos in the first sentence of the paper doesn't give confidence that I am about to read something worthwhile.

I think this is both a harmful and irrational attitude. Why focus on some trivial mechanical errors and disparage the authors for it instead of the thing that is much more important, i.e., the substance of the work? And in dismissing work for such trivial reasons, you risk ignoring things you might have otherwise found interesting. In an ideal world would second-language speakers of English proofread assiduously? Of…

Well, it not exactly a research paper, more an overview of the problem and suggested techniques, but it'd still be interesting to hear some criticism based on the content rather than the (admittedly odd) omission to run it through a spell checker. I do wonder why it was written in English, apparently targeting a western audience.

Two of the authors are from "Shanghai AI Labs" rather than students, so one might hope it had at least been proofread and passed some sort of muster.

Re: A path to O1 open source

#77
post #74

Earlier quoted context omitted.

flagged it. more people should flag this kinda stuff.

"Doesn't pass my sniff test" is not the purpose of the flag button. Furthermore, it passes my personal sniff test: hundreds of people upvoting it while the top comment is saying it's worthless. Usually the real alpha is in the comments under such things.

I didn't flag it, I flamed it.

Seems like it stunk enough for others to flag it. Lol.

Re: A path to O1 open source

#78
many people are dismissing this paper because it has errors in spelling and grammar.

this is a terrible heuristic for evaluating AI papers. If you use it, you will miss a lot of good work by very strong researchers with below-average English writing skills.

I have not read this paper carefully so claim nothing one way or the other about its quality. It superficially seems like a pleasant and timely survey although a little flag-planty.

Re: A path to O1 open source

#79

Lots of folks working on open-source reasoning models trained with reinforcement learning right now. The best one atm appears to be Alibaba's 32B-parameter QwQ: https://qwenlm.github.io/blog/qwq-32b-preview/ I also recently wrote a blog explaining how reinforcement fine-tuning works, which is likely at least part of the pipeline used to train o1: https://openpipe.ai/blog/openai-rft

I don't know if I would call it "the best one" when it has "How many r in strawberry" as one of its example questions and when tried it arrives at the answer "two".

Re: A path to O1 open source

#80
post #12
post #10

Earlier quoted context omitted.

In these times, how else does one expect to advurtise that theeir text was not geeenerated by an LLM?

Excessive profanity could be a fun way to prove human authors!

LLMs are trained, among other things, on Internet forums. Creative swearing is something they can do surprisingly well.
Post reply on HN