Earlier quoted context omitted.
You know also what was the biggest help for these companies? All the data they scrapped from the internet, the books they torrented, the open source projects whose licenses they did not respect. The least these companies can do is to publish everything they own
> The least these companies can do is to publish everything they own Well, being able to talk to superintelligent all-knowing savant 24/7 on your phone for $20 a month is also something
AI's top startups are barely publishing their research
251–260 of 341 posts
Re: AI's top startups are barely publishing their research
#252I’m not sure why this is so surprising? AI company does not automatically mean research company. The vast majority of new startups popping up over the last few years have commercial motivations, and use models built by someone else. Why are you expecting them to publish scientific papers? 50% of startups contributing to public research is actually a crazy good outcome. That’s far more than I had expected.
also, first people complain that only AI companies publish on AI, and now they complain the same companies stop those same publications :'). it is never good.
Re: AI's top startups are barely publishing their research
#253I’m not sure why this is so surprising? AI company does not automatically mean research company. The vast majority of new startups popping up over the last few years have commercial motivations, and use models built by someone else. Why are you expecting them to publish scientific papers? 50% of startups contributing to public research is actually a crazy good outcome. That’s far more than I had expected.
When the second biggest player is named 'OpenAI', I would expect some papers to be published.
Re: AI's top startups are barely publishing their research
#254Earlier quoted context omitted.
Open source doesn't mean, it's available to the general public.
What's your point?
EDIT: That's crap. I only thought about the aspect of continuing sharing, but if you also don't want your direct customers to know, then of course open source is very much relevant.
Re: AI's top startups are barely publishing their research
#255The article is vague about the companies in the paper, for some reason. In the paper, OpenAI is at the top of the chart for cumulative citations. MEGVII, Hugging Face, Waymo, Momenta, Preferred Netowkrs, Anthropic, Owkin, and Databricks, and Aibee follow (in that order). Yes, that is citations, not publications, but they explain that they're trying to use that as a proxy for significance, albeit an imperfect one. Com…
Google published a lot. I still wonder this original transformer paper, and the attention, etc. Why would they allow it to let go in the open? Perhaps because was intended for translation first before someone decided to loop it over itself? Or was so obscure to fellow researchers what do they actually publish?
Re: AI's top startups are barely publishing their research
#256Re: AI's top startups are barely publishing their research
#257Earlier quoted context omitted.
That's also a reason the big labs stopped. Publishing is most valuable to people who have no other way to get the attention of smart strangers. Once you can hire nearly anyone and everyone already returns your calls, the main remaining effect of publishing is to tell your competitors which things worked. This is what happens to every field as it turns from a science into an industry. Chemists published freely until d…
> Chemists published freely until dyes started being worth money Notably this is exactly what patents are intended to combat. And while US IP law is clearly very broken it does at least largely accomplish this stated goal. Much (but certainly not all) industrial chemistry has made it into the academic literature. Not that the same logic necessarily applies to AI research (ie algorithms aka math and their implementati…
Want the government / courts to stop your employees leaking source code? Escrow the code, and release it in 20 years.
The residuals on 20 year code is so close to zero that the costs vs benefits of longer IP protection is not in the public interest.
Re: AI's top startups are barely publishing their research
#258Earlier quoted context omitted.
Rare books seldom are under copyright and that’s not a concern then.
Is that easy to tell for each book, though? I usually hear that it's extremely hard to track down the copyright owners for many older things. That's the reason the film/video industry often gives for not digitising the backlog online, and I would expect their volume to be much easier to handle than the book industry.
It’s really not important to come with a totally bulletproof 100% accurate rule here.
Re: AI's top startups are barely publishing their research
#259Earlier quoted context omitted.
> I wish more people would publish the things that didn't work. This would be a pointless endeavour. One of the most basic mantras of science is "absence of evidence is not evidence of absence". So just because something didn't worked out for you that doesn't mean it doesn't work out for others, or even yourself in the future.
You are talking about a different thing, i.e. you have slipped in a level of abstraction that was not there before: In 'searching a path from A to B in a maze' language: The original statement was: (1) The branch to the left from A is a dead end. Your interpretation: (2) There is no path from A to B. (1) is still very useful (reducing the wasted effort) for those trying to find a path from A to B. The OP's point is t…
This is where you get things wrong at a very basic and fundamental level.
Just because you failed to explore branch A, that does not mean it is a dead end. It just means you came up empty.
That is why science is based on observations and theories: it is based on building up on ideas and what works and can be proven. Otherwise you will left with useless papers such as "Bicycles are a dead end because I tried to ride one and I fell".
Re: AI's top startups are barely publishing their research
#260I've been at two startups that have done genuine world first fundamental research. The first tried to publish novel results for 3 years in tier 1 journals before finally doing a preprint and telling the tier one publishers to jump in a fire. The second, and ongoing, isn't publishing anything because of my experience with the first. That and avoiding openAI and Anthropic copying our results and leaving us with nothing…