Live data from Hacker News

Lyft releases self-driving research dataset

medium.com

11–20 of 123 posts

Re: Lyft releases self-driving research dataset

#11
post #8

“There will be $25,000 in prizes, and we’ll be flying the top researchers to the NeurIPS Conference in December, as well as allowing the winners to interview with our team.” I guess it’s a decent opportunity if you’re trying to break into DL?

I suspect the winners will already be well established in the field.

Re: Lyft releases self-driving research dataset

#12
post #4
post #3

The post indicates there is a competition and prizes but I'm not seeing any discussion of what sort of license the data is being made available under (or the competition for that matter). Hopefully it's there and I'm just not seeing it.

The Github they link to says it's under the CC BY-NC-SA 4.0.

Hmmm. I wonder how much fun lawyers will have arguing about whether that "NC" clause means a model trained on this data cannot be used commercially by the researcher who built it?

Re: Lyft releases self-driving research dataset

#13
(I work at scale)

Hmm this blog post and the website doesn't mention that this dataset was mostly annotated by Scale (scale.ai), as part of a partnership with Lyft ... We're going to publish a blog post about this soon, but if anyone at Lyft is reading this, please figure out how to reasonably credit Scale since I doubt leaving out Scale completely from the announcement is in the spirit of the agreement. Scale should probably also be added to the bibliography and website in some form

Contrast this with the nuScenes website, which was also annotated by Scale, and whose data format set the standard for this dataset: they credit Scale pretty reasonably

Re: Lyft releases self-driving research dataset

#14

(I work at scale) Hmm this blog post and the website doesn't mention that this dataset was mostly annotated by Scale (scale.ai), as part of a partnership with Lyft ... We're going to publish a blog post about this soon, but if anyone at Lyft is reading this, please figure out how to reasonably credit Scale since I doubt leaving out Scale completely from the announcement is in the spirit of the agreement. Scale should…

I just researched scale before your comment here ! I tought : "This isn't done by scale ?"

Re: Lyft releases self-driving research dataset

#15

(I work at scale) Hmm this blog post and the website doesn't mention that this dataset was mostly annotated by Scale (scale.ai), as part of a partnership with Lyft ... We're going to publish a blog post about this soon, but if anyone at Lyft is reading this, please figure out how to reasonably credit Scale since I doubt leaving out Scale completely from the announcement is in the spirit of the agreement. Scale should…

Also, the viewer packaged with nuScenes was built by Steven Hao from Scale, and while it was packaged as part of nuScenes it should probably be called Scale's viewer instead of nuScenes' viewer. The original viewer in the nuscenes SDK has the Scale logo, but it looks like Lyft removed that in the fork. Maybe a bit of public shaming will fix that...

Dear Lyft marketing person who wrote this: we are a data labeling company, and you may think that means we have a bunch of useless bozos working here like most other data labeling companies, but that's not true - e.g, Steven is one of the smartest people in the world - https://stats.ioinformatics.org/people/3113 - he learns ridiculously quickly - e.g, gets to number one on random video games in a few weeks and learned to boulder L10 in a few months from scratch (normally takes years/decades and most climbers never get there)

Re: Lyft releases self-driving research dataset

#16
post #9
post #6

$25k in prizes seems silly given this is a multi-billion dollar market to crack.

Such competitions do not usually result in a comprehensive "solution" by themselves - pushing the state-of-the-art is more common. Also the value is not going to be derived solely from the algorithm but more from its deployment to real world applications and the surrounding infrastructure to make it possible.

But all the important IP work done in this field is currently closed...

Re: Lyft releases self-driving research dataset

#17
post #8

“There will be $25,000 in prizes, and we’ll be flying the top researchers to the NeurIPS Conference in December, as well as allowing the winners to interview with our team.” I guess it’s a decent opportunity if you’re trying to break into DL?

I suspect the winners will already be well established in the field.

Right, and 25k is their bi weekly paycheck.

Re: Lyft releases self-driving research dataset

#18
It's looking more and more like everyone is just going to have to licence Tesla's FSD when its finished.

They are the only ones with a broad real world data source and seem to have wisely taken the right path by not adopting LIDAR, focusing purely on passive vision.

Re: Lyft releases self-driving research dataset

#19
post #8

“There will be $25,000 in prizes, and we’ll be flying the top researchers to the NeurIPS Conference in December, as well as allowing the winners to interview with our team.” I guess it’s a decent opportunity if you’re trying to break into DL?

>> allowing the winners to interview with our team

This was pretentious AF. People who win such competitions _allow companies_ to interview them sometimes, not the other way around. It's not like working at Lyft is some amazing privilege.

Re: Lyft releases self-driving research dataset

#20
post #5

> Academic research accelerates innovation, but it requires costly data that is out of reach for most academic teams. This is true of pretty much any AI research. Look at Puffer[0], which was just on HN a couple of days ago. They're running a free streaming service just to get enough data to train their algorithms, and in fact mention in their FAQ that they would love to use commercial data if they could get it. Unfo…

> I wonder if there isn't some sort of governance solution to this. Like give companies big tax breaks for sharing their data with researchers, or something like that. Essentially subsidize academia indirectly.

I think that's a really great idea. Not sure how many would take advantage of it but if it could be made to work then it would be really awesome.

It would also be extremely prone to abuse, though. Patenting is already an art of pretending to explain in clear terms what you are doing, while actually describing something as broadly and vaguely as possible. It would be pretty easy for a TON of things to leave out some key things that make it impossible or unhelpful to have the information.

You could form industry-specific regulations or even an active agency to prosecute abuses like that, but it would be immediately overwhelmed. The patent office is already heavily gamed by patent trolls, who bank on long odds for small judgements. Now imagine if millions or billions of dollars of taxes were on the line, and major companies were investing significant resources to open source while protecting their IP.

Even if that were all figured out, how would you value open sourcing stuff, even something as simple as data? Do you give breaks by size, importance, proportion of profit or future profit? Cost of the research? How do you guard against overvaluations and abuse of accounting? Even if you had perfectly accurate, annually-updated solutions for all that, companies can still game the system. Lyft has decided this dataset is what they need; if they could get a bigger break by collecting more data, they'd do that. Plus- facebook and google release tons of open source stuff. Do they deserve more than say, pharmaceutical research?[1]

Similar (IIRC Nixon) tax breaks already exist for R&D, and they are a notoriously abused loophole. Simplified but illustrative example: you build your R&D lab in the shape of a factory, do your research for a while and then suddenly scale back and replace it with machinery- well, the original building was still deducted from taxes.

Pharma is actually a perfect example. It's a well known fact that R&D only accounts for 22% of pharma industry revenue (almost equal to advertising at 19%), but only ~30% of that actually goes to new drugs. The rest takes advantage of marketing and the patent system to re-release drugs that are essentially the same. Two thirds of their research is obvious changes that are only protected because they owned the original patent- those shouldn't be getting the benefit of incentives.

Post reply on HN