Live data from Hacker News

Open-Sourcing Vespa, Yahoo’s Data Processing and Serving Engine

oath.com

101–110 of 122 posts

Re: Open-Sourcing Vespa, Yahoo’s Data Processing and Serving Engine

#101
post #70
post #48

Earlier quoted context omitted.

The Register did a good job describing the dichotomy between the product business and the tech side of the business here, related to a previous open source project Yahoo published (disclaimer, I run the open source process at Yahoo) https://www.theregister.co.uk/2017/03/23/yahoo_tensorflow_on... "Over the decades Yahoo! has contributed substantially to the greater good, publishing its own code as open source. Arguabl…

What do you mean you run the Open Source process at yahoo? Hopefully Verizon's open source legacy will become even bigger and make it easier for these kinds of innovations to happen. We've had some success a few cities over in Verizon Labs... https://verizon.github.io/

Here's some more information if you're interested in running your own open source program: https://www.linuxfoundation.org/resources/open-source-guides...

Re: Open-Sourcing Vespa, Yahoo’s Data Processing and Serving Engine

#102
post #85

Earlier quoted context omitted.

is tfos production-ready right now ? because i thought it was still experimental. is it used inside Yahoo - because Vespa comes with its own tensor processing engine. I wondered who would use one over the other.

TensorFlow on Spark is for learning, Vespa is for serving. Where Vespa excels is in evaluating a learned model very quickly over lots of documents. We're working on providing support for running models learned with TensorFlow directly. For now people make the translation on their own.

I don't want to trigger any bot detection by voting all your comments up in a short amount of time. So, I will say thank you and mention that is is contributions from people like you that keep me coming back.

Re: Open-Sourcing Vespa, Yahoo’s Data Processing and Serving Engine

#103
post #87
post #60

Earlier quoted context omitted.

https://www.joelonsoftware.com/2002/06/12/strategy-letter-v/ The point of Joel's article is that "Smart companies try to commoditize their products’ complements." For what I understand, Yahoo is a media company, and as so it may try to commoditize a natural complement of today's media companies, which is data search.

That article is a classic but I don't really understand how it applies in this case. Smart companies commoditize their products' complements (something that needs to be bought with the product) so that whoever wants buy their product has a large variety of offerings to select from. For instance, MS-DOS's ability to run on any standard PC architecture machine commoditized the PC. It's not clear how commoditizing data…

Yahoo! gets code improvements back and, by being the originals and having the most experts familiar with the entire code base - and it is huge, they retain their competitive advantage. They also do this while fostering good will and, potentially, reducing the number of bugs, improving security, and increasing efficiency.

It's brilliant, in a way. There are risks, but they are small and mitigated. They may even end up selling support and customizations, or enabling that market. The upside potentials are many and the downsides are few and only risk small impacts.

Hell, you can get RedHat for the low cost of nothing, just by signing up for it. On top of that, they will give you every single last line of code you want. They'll give you all of the code, and do it for free.

Yet, they are a successful for-profit company. They don't even accept financial donations, as far as I know. They aren't the wealthiest company, but they are doing quite well and not suffering financially.

Open source doesn't mean no profit. It just means additional rights for the source code and/or user. (Different licenses prioritize different liberties and have different goals.)

Re: Open-Sourcing Vespa, Yahoo’s Data Processing and Serving Engine

#104
post #59

i wonder if this is able to squeeze something useful out of Yahoo Answers.

I am not trying to be snarky when I say this, so I'll make it longer than a single word.

No.

That is, Y! Answers was/is a great idea with terrible results. I'm not sure if it should have been moderated better, or if it should have been marketed better. Hell, maybe it should have had a basic literacy test prior to being allowed to post and answer?

I could actually come up with a few hundred ways to have made it better. We know the question and answer format works. It does on many, many sites. It failed there and, largely, that's because of the users. I'm not sure which was worse, the questions or the answers.

I'd love to be given that project and tasked with improving it. It's a great idea, but horribly implemented. I'd like to fix it because I'm fond of trying the impossible and like fixing broken things.

Improving search for it is absolutely not going to help it. No, that's not going to help in the slightest.

Re: Open-Sourcing Vespa, Yahoo’s Data Processing and Serving Engine

#106
post #89

Earlier quoted context omitted.

Even split between Java / C++, with at least 4 different build systems: 2 of which are effectively the next generation of the other 2. 2 different UNIX shells and 5 other dynamic programming languages. Not the end of the world, and there are probably good reasons for a lot of it (maybe there are bindings for various languages, maybe some of it is misindeitified, maybe they're harnessing 2 large bodies of existing wor…

No, this is run by a single, very cohesive, very remote team. There is some misidentification in your list (our .def files are nothing to do with Teamcenter). We use 2 languages because we Java and C++ have different strengths which make each suitable at different layers of the architecture. The rest is a combination of "for good reasons" and "leftover scraps" :-)

Yeah I ran cloc on a few projects I'm very familiar with, and shell scripts that are interpreted with bash but not with a shebang because they're libraries sourced by others are called Bourne too. Just explaining what OP was getting at :)

Re: Open-Sourcing Vespa, Yahoo’s Data Processing and Serving Engine

#107

Wow... I'm guessing (hoping) their lawyers made sure to go over their old agreements with a very fine comb to make sure their license for the software allowed them to open source this. Don't get me wrong, they've added a lot to it, but there's a lot of code in there that could only have come from their purchase of Overture, who had purchased AllTheWeb from FAST (which was itself purchased by Microsoft).

If they purchased it, then they would own the copyright, so they can relicense it any way they want can't they?

Re: Open-Sourcing Vespa, Yahoo’s Data Processing and Serving Engine

#108

Earlier quoted context omitted.

As an ex-employee, there could not be a better description of Yahoo! development than this.

Hey, there is no yinst/buildyblocks stuff here. That's kind of a must-have.

> yinst

[FLASHBACKS]

To be fair, yinst was the least worst part of the systems I was working on (Yahoo!Europe backend feeds stuff.)

Re: Open-Sourcing Vespa, Yahoo’s Data Processing and Serving Engine

#109
post #56
post #39

Earlier quoted context omitted.

I am not in America, but when she arrived, I was very curious to see how she could revived such a dying zombie. Clients and engineers were leaving as hell. IMHO, she has done a lot to reduce the hemorrhage and make it attractive again.

> IMHO, she has done a lot to reduce the hemorrhage and make it attractive again. What are you talking about? She's no longer employed and Yahoo has been sold off to Verizon...

I replied to the rewrite of history of yahoo, in particular that yahoo was healthy before Marissa Meyer arrived. IMHO, she had a positive effect on yahoo during her presence.

Re: Open-Sourcing Vespa, Yahoo’s Data Processing and Serving Engine

#110
post #47
post #42

Earlier quoted context omitted.

When I was there (2003-2005), the big problem already seemed to be that a lot of engineering saw Yahoo as a tech company, while a lot of the business side saw Yahoo as a media company. Especially so after Yahoo conceded Search. So one side wants to pour money into getting technical leverage. The other side largely just wanted more effective ways of publishing and monetizing content. It doesn't really matter which sid…

Yahoo's largest problem while I was there (2004-2011) was that nobody could articulate what y! is and what it does. Actually, at employee orientation in 2004, they had a clear mission: to be a part of everything you do online: either by doing it, or cobranding it, with a goal of being in the top 3 of every internet vertical. Somehow, over time, this message was lost, and it was no longer ok to be #2 or #3 in some ver…

That sounds like an incredibly unclear mission to me. It also sounds like the mission of a media company, not a tech company, but without making that distinction clear...
Post reply on HN