Live data from Hacker News

Billions of Messages a Day – Yelp's Real-Time Data Pipeline

engineeringblog.yelp.com

11–20 of 49 posts

Re: Billions of Messages a Day – Yelp's Real-Time Data Pipeline

#11
I don't understand the following:

"Yelp passed 100 million reviews in March 2016. Imagine asking two questions. First, “Can I pull the review information from your service every day?” Now rephrase it, “I want to make more than 1,000 requests per second to your service, every second, forever. Can I do that?” At scale, with more than 86 million objects, these are the same thing."

Who is making 1000 requests per second to retrieve all of the 100 million reviews? Other services within Yelp? Why would they pull all reviews every time, and not just the reviews they haven't already processed?

How is 1000 requests per second the same thing as pulling all review information every day?

This section is really confusing and needs some clearer explanation.

Re: Billions of Messages a Day – Yelp's Real-Time Data Pipeline

#12
post #4

"In 2011, Yelp had more than a million lines of code in a single monolithic repo" Did I read that correctly, the Yelp product is a MILLION lines of code? For comparison, the Doom 3 source code has 601,000.

What do Doom 3 and Yelp have to do with each other?

Re: Billions of Messages a Day – Yelp's Real-Time Data Pipeline

#13
post #8

100 million reviews/year is only 3 reviews per second on average. Sure, they seem to do more then just that, like voting, comments, etc. But it still seems like something an old school stack could handle on a single large instance. Reading between the lines it seems the problem wasn't scaling, but programmer productivity. Smaller code bases is often easier to work with so I guess they solved that by dividing it up in…

100 million is cumulative, not annual.

Re: Billions of Messages a Day – Yelp's Real-Time Data Pipeline

#14

I don't understand the following: "Yelp passed 100 million reviews in March 2016. Imagine asking two questions. First, “Can I pull the review information from your service every day?” Now rephrase it, “I want to make more than 1,000 requests per second to your service, every second, forever. Can I do that?” At scale, with more than 86 million objects, these are the same thing." Who is making 1000 requests per second…

I think the unstated assumption is that there's some sort of processing that occurs on each review every day. So with 100 million reviews every 24 hours, that's just over 1157 requests per second.

Re: Billions of Messages a Day – Yelp's Real-Time Data Pipeline

#16

Is Yelp still using pyleus and Apache Storm? Or have they migrated to Spark and Kafka Streams?

Most of the stream processing in the Data Pipeline happens inside of an internal project called PaaStorm, which is storm-like. It was built to take advantage of our platform as a service ( http://engineeringblog.yelp.com/2015/11/introducing-paasta-a... ), which handles process scheduling really well. Architecturally, it's pretty similar to Samza, with distributed processes communicating using Kafka. We do use Spark s…

Hi Justin! Thanks for sharing, very interesting stuff.

How do you scale Kafka to handle the massive amount of traffic (and storage) that you seem to generate daily?

With services talking among themselves via HTTP there is a lot of resilience built-in. Do you have anything in place to avoid this becoming a single point of failure? It must have become the most critical piece of your infra.

Re: Billions of Messages a Day – Yelp's Real-Time Data Pipeline

#17

I don't understand the following: "Yelp passed 100 million reviews in March 2016. Imagine asking two questions. First, “Can I pull the review information from your service every day?” Now rephrase it, “I want to make more than 1,000 requests per second to your service, every second, forever. Can I do that?” At scale, with more than 86 million objects, these are the same thing." Who is making 1000 requests per second…

I think it's saying that a single iteration over the entire set, would translate to 1000 requests per second for a day (if done naively as one request per object). It's really talking about the N+1 problem.

Re: Billions of Messages a Day – Yelp's Real-Time Data Pipeline

#18
post #4

"In 2011, Yelp had more than a million lines of code in a single monolithic repo" Did I read that correctly, the Yelp product is a MILLION lines of code? For comparison, the Doom 3 source code has 601,000.

Maybe counting all the dependencies? You know, you need millions of lines of code to run an npm powered js app :)

Re: Billions of Messages a Day – Yelp's Real-Time Data Pipeline

#19
post #12
post #4

"In 2011, Yelp had more than a million lines of code in a single monolithic repo" Did I read that correctly, the Yelp product is a MILLION lines of code? For comparison, the Doom 3 source code has 601,000.

What do Doom 3 and Yelp have to do with each other?

You'd hope a CRUD app used in "hello world" tutorials (https://www.fullstackreact.com/articles/react-tutorial-cloni...) would be simpler than a 3d physics engine.

Re: Billions of Messages a Day – Yelp's Real-Time Data Pipeline

#20

I don't understand the following: "Yelp passed 100 million reviews in March 2016. Imagine asking two questions. First, “Can I pull the review information from your service every day?” Now rephrase it, “I want to make more than 1,000 requests per second to your service, every second, forever. Can I do that?” At scale, with more than 86 million objects, these are the same thing." Who is making 1000 requests per second…

It has to be internal because they are incredibly protective of their API (hell they sold their firehouse to just 1 darn company). My guess it's NLP type processing. Things like review highlights, recommendations, etc. that's probably 99% of their load.

The user-centric stuff like submitting reviews, comments is trivial - as another user said , a large single instance is enough.

Post reply on HN