Live data from Hacker News

AI's top startups are barely publishing their research

science.org

201–210 of 341 posts

Re: AI's top startups are barely publishing their research

#202
post #53
post #45

Earlier quoted context omitted.

No traditional research is mostly done by post-docs and have a phd level of education rather than a tech bro that passed leetcode. That seems like a good bar to have; not too mention the whole peer review thing, hard to really understand anything if you purposely withhold it and tell people to kick rocks.

Plenty of research is performed by non PhDs. The main difference is reproducibility and peer review process.

outside physics/math a lot of published research isn't reproducible at all. There's been several studies that show this.

Re: AI's top startups are barely publishing their research

#203

I can’t speak for other startups, but I applied to the most recent YC batch with my idea for making AI proactive instead of reactive, and pre-being selected I’ve published a paper on recursive self-improvement mapped to the Epoch AI data. I contacted a professor from a university in the UK and he responded since he was working on similar work, then asked me if I wanted to meet with him. We talked for about an hour si…

LinkedIn can and should be a platform like this, for professionals and catered to professionals. It is a shame that it's now a cesspool of engagement bait and larp entrepreneurs

Re: AI's top startups are barely publishing their research

#204

Earlier quoted context omitted.

That's also a reason the big labs stopped. Publishing is most valuable to people who have no other way to get the attention of smart strangers. Once you can hire nearly anyone and everyone already returns your calls, the main remaining effect of publishing is to tell your competitors which things worked. This is what happens to every field as it turns from a science into an industry. Chemists published freely until d…

> Chemists published freely until dyes started being worth money Notably this is exactly what patents are intended to combat. And while US IP law is clearly very broken it does at least largely accomplish this stated goal. Much (but certainly not all) industrial chemistry has made it into the academic literature. Not that the same logic necessarily applies to AI research (ie algorithms aka math and their implementati…

>> To your dye example, yttrium indium manganese blue was the first commercially viable inorganic blue pigment discovered in ~200 years, is the only known environmentally safe one, and was openly published in the literature. It's also under an exclusive license.

Gee, I wonder what was wrong with the previous blue pigments and why it was so important to have this one under an exclusive license.

Cobalt blue is a blue pigment made by sintering cobalt(II) oxide with aluminium(III) oxide (alumina) at 1200 °C. Chemically, cobalt blue pigment is cobalt(II) oxide-aluminium oxide, or cobalt(II) aluminate, CoAl2O4. Cobalt blue is lighter and less intense than the (iron-cyanide based) pigment Prussian blue.

https://en.wikipedia.org/wiki/Cobalt_blue

Oh right.

P.S. Don't lick your brushes.

Re: AI's top startups are barely publishing their research

#205

Earlier quoted context omitted.

Curious what you mean by proactive? Could you share a bit more?

Happily, current AI is interacted with in a reactive loop. I open the Claude app, CLI, whatever, say my prompt, get an output. I personally wanted an AI that was able to reach out to me about my life before I had to reach out to it. An example, a friend just emailed me asking to meet for at 1pm but I have class at 1:30, so a proactive AI would see that conflict and send me a notification about it, asking if the propo…

>> Curious what you mean by proactive? Could you share a bit more?

> Happily, current AI is interacted with in a reactive loop. I open the Claude app, CLI, whatever, say my prompt, get an output.

"Current AI" is not limited to LLM offerings. There are many AI algorithms which can assist in what you specify thusly:

> I personally wanted an AI that was able to reach out to me about my life before I had to reach out to it.

Consider a forward chaining inference engine ("expert system") provided with relevant asynchronous percepts from the deployed environment to reason about. This could serve as an initiator of a "proactive AI".

Re: AI's top startups are barely publishing their research

#206

Earlier quoted context omitted.

> Meaning what? Claw back the ideas ... If something is so obviously wrong then perhaps take a minute to consider that your interpretation isn't what the other party intended? If I pay you not to do something and then you breach the contract I can terminate the agreement and seek damages. Ditto if I pay you to repeatedly do something and then at some point you fail to do it. So if I pay you a recurring fee to publish…

>So if I pay you a recurring fee to publish all your research on a given topic and then you fail to make good on that I can seek damages, right? Now what if I paid you a lump sum up front? Now what if I licensed a patent to you in place of that lump sum? What if instead of a patent it was the right to make use of a piece of software? Uh-huh... This doesn't answer my question of what terminating the agreement of acces…

> I've breached the conditions, therefore you terminate the agreement, therefore you revoke access. Am I missing anything?

You're missing the part where I seek punitive and actual damages under the terms of the contract. No different than violating an NDA - I paid you a lump sum up front, after a while you breached the contract, the agreement is null and void, what's the consequence?

> Well, the idea of viral abstract ideas is stupid, so it forces me to give contrived examples.

On the contrary, presumably it was because you lacked the ability to roundly refute anything I had put forward. Otherwise I assume you would have done so.

> Therefore if you would rather walk away than gamble everything you own on the normal coin, the toss actually has a 100% chance of you losing?

But in this analogy it is you baselessly making that claim. There's every expectation that it's a fair coin, many experts have carefully inspected it and authored opinions on it, and some have put forward theories that it slightly deviates in one direction or another. Then you show up and confidently assert without any evidence that there's some wild deviation from fair, hand waving that you would have proof if only someone wanted to bother testing it.

Out of curiosity, what is it that has you so bothered about the idea of viral licenses? What do you find so objectionable about attaching arbitrary terms to contracts?

Re: AI's top startups are barely publishing their research

#208
post #66

I've been at two startups that have done genuine world first fundamental research. The first tried to publish novel results for 3 years in tier 1 journals before finally doing a preprint and telling the tier one publishers to jump in a fire. The second, and ongoing, isn't publishing anything because of my experience with the first. That and avoiding openAI and Anthropic copying our results and leaving us with nothing…

Tier 1 journals have high standards (that's why they're tier 1) and they also have hundreds of submissions every cycle. It's common to have a paper rejected. On the upside you get your work seen by some of the most experienced researchers in AI, who are also specialists in your paper's subject. For example, here's the editorial board of the Journal of Machine Learning Research:

https://jmlr.org/editorial-board.html

Scroll down where it says "JMLR Editorial board of reviewers" and where you can find the names of reviewers and their areas of expetise.

The advice you get from such reviewers, even if your paper is rejected, is an invaluable tool to help you improve your work and its presentation. You eventually learn to not be sore about it and iterate until your work gets accepted.

Of course nobody's forcing you to publish, especially in a journal (in AI and CS conference proceedings tend to have much quicker turn around times and you naturally get more feedback if you get to present at a conference) but the alternative is to sit alone in your room hacking away at your code until you discover the perpetual motion machine.

Re: AI's top startups are barely publishing their research

#210

Earlier quoted context omitted.

Everyone including myself has attempted this relentlessly, it just doesn't work out beyond some arbitrary improvement in specific tested categories instead of broader capability increase.

Did you publish on that topic? Just curious, would love reading more

There's a pretty large paper on the topic: https://arxiv.org/pdf/2606.15497, but results are somewhat mixed.
Post reply on HN