Live data from Hacker News

Scraping Recipe Websites

benawad.com

141–150 of 204 posts

Re: Scraping Recipe Websites

#141

I highly useful tool in my household for dealing with the SEO/tracking scourge that recipe blogs have become is https://www.paprikaapp.com/ . Hoping someday to have some spare time to integrate this with https://grocy.info/ and have a pipeline for recipe -> preparation automation.

I love Paprika, but what keeps frustrating me is the inability to share recipes with my partner. I can Airdrop a single recipe to her, but doing this for all of them, one byone, super tedious, and then she makes some changes which I'd like to have too, and there seems to be no reliable way to get her changes back on my devices. It's all the more frustrating as Paprika does sync really well, but apparently just within…

Instead of using iCloud sync create an account with their service and log in on both devices with the same account. My wife and I have been using it this way for years and it works really well.

Re: Scraping Recipe Websites

#142
post #125

Earlier quoted context omitted.

Yes, search engines monitor result clicks, bounce rates, dwell times, etc. They do affect ranking. And if you have Google analytics on your recipe site, Google has even more data about it.

Are you implying that having Google Analytics on your page increases your Google ranking?

I don't know. I've always thought it should, and as a search engine CTO, I'd use the data that way. But I have no evidence either way.

Re: Scraping Recipe Websites

#143

I've been working on something similar for the past couple of days, but the trouble comes with wanting static types. There are a few projects out there that offer either a microdata parser, or types derived from schema.org but nothing that combines the two as yet

I’ve been working on this and will have a recipe-specific solution up in a couple weeks. See https://rcpe.io

Re: Scraping Recipe Websites

#144

I highly useful tool in my household for dealing with the SEO/tracking scourge that recipe blogs have become is https://www.paprikaapp.com/ . Hoping someday to have some spare time to integrate this with https://grocy.info/ and have a pipeline for recipe -> preparation automation.

I love Paprika, but what keeps frustrating me is the inability to share recipes with my partner. I can Airdrop a single recipe to her, but doing this for all of them, one byone, super tedious, and then she makes some changes which I'd like to have too, and there seems to be no reliable way to get her changes back on my devices. It's all the more frustrating as Paprika does sync really well, but apparently just within…

You should give AnyList a try. It also features a shopping list which is handy. Plus you can both use your individual accounts and sync recipes. https://www.anylist.com/

Re: Scraping Recipe Websites

#145
post #115

Earlier quoted context omitted.

I think it's also about different types of recipe collections. There are cookbooks that are all recipes. This seems to be what the HN crowd is looking for when they search the internet. There are cookbooks where each recipe is accompanied by a little story. These seem to sell well, judging by the number of them that appear on bookstore shelves. And then there are cookbooks where all of the anecdotes are in the front…

I wouldn't mind the "recipe at the end" format if that were the actual case. But it's not at the end. The actual structure of a modern recipe article is 1) Header 2) Blogpost 3) Click to read button 4) More blogpost, structured in a format that looks like an informal recipe - bulleted ingredients without quantities, discussion of steps. 5) Ad. 6) More blogpost 7) More ad that looks confusingly like a kind of recipe.…

And that's not the worst of it. Here's what I really hate:

- Search page and eventually find the recipe.

- Start making recipe. Up to my elbows in ingredients.

- Glance over to see the next ingredient I need, but there's now a pop-over I need to dismiss before I can see the recipe.

- Look for next ingredient, and it's scrolled off the screen because one or more adverts in the page have reloaded and have different sizes.

Most recipe sites have a nearly unusable UX.

Re: Scraping Recipe Websites

#146

Paprika 3 (I use the iOS version, but I believe the Mac version has the same function) has a fantastic web scraper for recipes. I've had to correct maybe 1-2 errors across 100 recipes I've brought in from a bunch of different sites. It's super helpful to look through them in a standardized way (and you can sort by ingredient/category) to figure out what to make.

Thank you all for the Paprika recommendation. I just grabbed it and imported the recipe I did a for Mothers Day Eve, and it looks great! That recipe wasn't one of the worst offenders, but it's off to a good start for Paprika.

Re: Scraping Recipe Websites

#147
A surprisingly good UX for recipes is Google Home. Ask it for a recipe, and it will ask if you want directions or ingredients. If you ask for ingredients, it will say them one by one, and pause between them until you ask it for the next one. My son has used it to great effect to make pancakes.

Re: Scraping Recipe Websites

#148
post #33
post #20

Earlier quoted context omitted.

I have come to the conclusion that the "fluff" is what most of the recipe-reading public want. I've often heard the prevailing reason why this happens on the web is because of SEO and 'bounce rates'. More time spent on the site improves ranking, so the actual recipe is pushed below the fold so users have to scroll down thereby adding more time on the site. Have often wondered if any SEO wonks with the inside baseball…

I've also read that it's to do with copyright - the recipes themselves can't be copyrighted, but the text around them can. Scraper republishes your ingredients list: not a lot you can do. Scraper republishes your fluffy anecdote and pictures: BLAMMO!

Except the whole reason this article on scraping recipe sites was written was to ignore the fluff. Forcing copyright this way doesn't sound particularly helpful when the fluffy anecdote has so little value in comparison to the recipe. Or is it really the case that the average reader wants the anecode secondary to the recipe?

It feels like people are trying to make money around information that is fundamentally impractical to make money off of, so they're forced into doing whatever it takes to make money off of it anyway. "Whatever it takes" is defined by Google yet ruins the user experience, and so that is why recipe sites are this way.

Re: Scraping Recipe Websites

#149

Earlier quoted context omitted.

Does it pose an actual problem? When I search for recipes, I type in the food I want + recipe and then open the top 5 or so links. I quick scan for a list of ingredients. If I don't easily spot on in a few seconds I move on. I'll do this until I have a couple different lists of ingredients for making the item. This ends up taking less than a minute or two. That just isn't a significant portion of time compared to how…

You dismiss the challenges of finding a single recipe by mentioning you find several and combine them? That completely defeats the purpose of a recipe. You're literally creating your own recipe at that stage.

Finding several to combine should be harder than finding any single one of those represented among the several.

As far as the purpose of a recipe, it still informs me of the ingredients and amounts so I can make a reasonable approximation. Say I want to cook chili, something that there seems to be innumerable recipes for. And say I want to add beans to mine, though my personal recipe doesn't normally include beans. So how many beans should I add? And what kind of beans should I add? Well if I check the top 5 recipes for chili with beans and see that 4 of them use kidney beans and that they tend to use 1 cup per pound of meat average, I can now modify my own recipe in a more informed fashion. This also works for cooking something when I don't know where to begin.

Re: Scraping Recipe Websites

#150
post #56

Earlier quoted context omitted.

Does it pose an actual problem? When I search for recipes, I type in the food I want + recipe and then open the top 5 or so links. I quick scan for a list of ingredients. If I don't easily spot on in a few seconds I move on. I'll do this until I have a couple different lists of ingredients for making the item. This ends up taking less than a minute or two. That just isn't a significant portion of time compared to how…

This ends up taking less than a minute or two. For you to open 5 websites, dismiss the cookie permission request on each one, dismiss the notifications request, scroll down to find the ingredients list, dismiss the scrolling activated newsletter signup, read the ingredients list, and click the 'next page' button to see the instructions to work out what you can tweak, all in an average of 24s per site is very impressi…

I just tried it. Went to the top seven sites for chicken tikka masala. Exited one for not loading, had to mute another tab with an annoying video, but got to the recipe in 6 of them in under two minutes. No popups (though I do have an ad blocker that may have prevented them).

>read the ingredients list, and click the 'next page' button to see the instructions to work out what you can tweak

That wasn't included. I was talking about scanning to verify there was an ingredient list. I pointed that out when I said:

>That just isn't a significant portion of time compared to how long I'll spend comparing different recipes to find a common theme to follow.

Post reply on HN