Live data from Hacker News

AI's top startups are barely publishing their research

science.org

171–180 of 341 posts

Re: AI's top startups are barely publishing their research

#171
post #66

I've been at two startups that have done genuine world first fundamental research. The first tried to publish novel results for 3 years in tier 1 journals before finally doing a preprint and telling the tier one publishers to jump in a fire. The second, and ongoing, isn't publishing anything because of my experience with the first. That and avoiding openAI and Anthropic copying our results and leaving us with nothing…

Is this the "AI startups blog about discoveries rather than submitting to journals" phenomenon that TFA noted?

Are journals just too slow, hidebound, and predatory to be useful to this industry any more?

Re: AI's top startups are barely publishing their research

#172
post #97
post #81

Earlier quoted context omitted.

It took 6 years to get from Attention Is All You Need to a consumer product. Up until ChatGPT’s release, LLMs were obscure and really were just fancier autocomplete. Deepmind was largely a speculative research division until Google decided to play catch up, and even that took a bit of time as they figured out how to proceed in a way that wouldn’t cannabalize search. So in short they didn’t realize the potential. That…

Google itself was shipping transformer-based features to consumers within 2 years, notably BERT for websearch.

and Google Translate: https://arxiv.org/abs/1804.09849

Re: AI's top startups are barely publishing their research

#173

Earlier quoted context omitted.

Recursive self improvement is the beginning of the end.

it's actually the end of the beginning, there's walls that self-improving models hit that are never overcome even when given vast amounts of time.

I don't think anyone has ever tried having a model fully autonomously train a model that is better than it.

Re: AI's top startups are barely publishing their research

#174

Earlier quoted context omitted.

There's already an RPL, incidentally: https://en.wikipedia.org/wiki/Reciprocal_Public_License Your RPL wouldn't be enforceable. Copyright doesn't deal with abstract ideas passing through people's minds. Even the GPL is kind of in a gray area because the virality feature and its definition of "derivative work" have never been tested in court, to my knowledge. Maybe under contract law, no idea. If nothing else, I'd lov…

Well it was a joke and is obviously quite silly but I believe it would be enforceable to the extent that the licensor could terminate the agreement and sue for damages. If I can agree to pay you not to talk about something (ie an NDA) or not to work in a field (ie a non-compete clause) then why can't I pay you to be required to publish all future work you do in a given area? ("All future work" might well be overly br…

>terminate the agreement

Meaning what? Claw back the ideas from people's minds? You can terminate the agreement in the sense that you revoke access to the paper, but presumably the person you find in breach has already used the research for something that you find them in breach for. You're kind of closing the gate after the horse has bolted.

>sue for damages

I honestly have no idea what damages you could claim from not publishing research. I think you would need to set a value ahead of time on the agreement.

>If I can agree to pay you not to talk about something (ie an NDA) or not to work in a field (ie a non-compete clause) then why can't I pay you to be required to publish all future work you do in a given area?

Not sure why you added the word "pay" to your clauses, but anyway. The reason is that the existing contracts have well-defined boundaries. An NDA stops you from divulging a very specific set of information. A non-compete clause stops you from working in a very specific field. Your proposal has an undefined reach. What counts as research? What counts as "related"? It would seem that if I agree to such a contract, my entire life, both private and professional, after reading the paper is covered by the contract, and anything I do might come under scrutiny. There's never a point when I can go off-duty. "What's that? You read my paper on compression and were working on a side-project that uses compression? Gonna have to see some publication on it."

>Rather IIUC no one has gone out of the way to test the GPL largely because it is clearly within bounds

No, it's because the status quo is convenient and no one wants to be the first guinea pig. It's definitely not obvious that the terms are legally valid, but it's ambiguous enough that people don't want to test it.

Re: AI's top startups are barely publishing their research

#175

Earlier quoted context omitted.

Arguably, the point of a startup isn’t to make money, it’s to attract attention and then investors. Once the business is making money it’s not really a startup any more, it’s a going concern. OP is an undergrad student. Attracting attention is what will make them stand out. And sharing research at this stage will probably expose them to research others don’t want to publish but do want to share with other researchers…

> Arguably, the point of a startup isn’t to make money, it’s to attract attention and then investors. Investors with money.

What do you mean by this?

What are investors who don’t have money?

Perhaps I should have said investors and speculators, but this doesn’t help you here.

Re: AI's top startups are barely publishing their research

#176

Earlier quoted context omitted.

That's also a reason the big labs stopped. Publishing is most valuable to people who have no other way to get the attention of smart strangers. Once you can hire nearly anyone and everyone already returns your calls, the main remaining effect of publishing is to tell your competitors which things worked. This is what happens to every field as it turns from a science into an industry. Chemists published freely until d…

Which is a very ironic and selfish situation when your business model dependent mostly on model training based on available published data, academic and non-academic.

Sure.

But they’re not dependent on my research in particular.

If I don’t publish, what, as a result of my not publishing, happens to the companies dependent on data & research?

Nothing.

Re: AI's top startups are barely publishing their research

#177
post #30

What the blogificafion of AI research has done is allowed all kinds of claims and terminology related to AI to be introduced and taken up in a manner replicating social media dynamics. And that is simply not healthy. We’re fast reaching a place where any claim can be backed up with a set of numbers from a number of experiments run in some gamified environment or the other, with little concern for if it all adds up to…

How is that different from traditional research publication? It's faster?

It's rather similar to publishing a non-peer-reviewed preprint on arXiv.

Re: AI's top startups are barely publishing their research

#178

Earlier quoted context omitted.

Well it was a joke and is obviously quite silly but I believe it would be enforceable to the extent that the licensor could terminate the agreement and sue for damages. If I can agree to pay you not to talk about something (ie an NDA) or not to work in a field (ie a non-compete clause) then why can't I pay you to be required to publish all future work you do in a given area? ("All future work" might well be overly br…

>terminate the agreement Meaning what? Claw back the ideas from people's minds? You can terminate the agreement in the sense that you revoke access to the paper, but presumably the person you find in breach has already used the research for something that you find them in breach for. You're kind of closing the gate after the horse has bolted. >sue for damages I honestly have no idea what damages you could claim from…

I tend to agree.

To paraphrase your last paragraph, an idea that hasn’t been tested is either perfect, or just bad enough that no one wants to test it, as, as you say, testing it is would be at least somewhat inconvenient.

Chances are nothing is perfect.

Re: AI's top startups are barely publishing their research

#179

Earlier quoted context omitted.

Okay interesting, so maybe openclaw does a lot less than I thought it did; really I have no excuse not to be just trying it myself regardless, given that I'm sitting on a 9070 XT. It doesn't look like you're building directly on openclaw, so is that coming from a place of different goals or philosophy, or what?

Openclaw is great from when I’d used it for a couple months, the key difference is that for Openclaw you’ll have to actually schedule the cron jobs or automated tasks yourself, so it’s a decision on your part to make the AI do a thing whereas Orb is a decision on the part of the AI (after your approval) to do the thing. So it’s a kind of shift of agency. Openclaw could definitely do many similar things, it just would…

I mostly got it to play Arc Raiders tbh, but I did kind of have in mind that it was worth stretching for 16GB so that local models would be on the table.
Post reply on HN