Live data from Hacker News

Updates to our web search products and Programmable Search Engine capabilities

programmablesearchengine.googleblog.com

131–140 of 207 posts

Re: Updates to our web search products and Programmable Search Engine capabilities

#131

Earlier quoted context omitted.

what stops Kagi from indexing internet and makes them pay some guys to scrape search results from Google? one guy at Marginalia can do it and entire dev team at a PAID search engine can't?

As we've seen here on HN on the AI boom, it's not wonderful when a bunch of companies all use bots to scrape the entire web. Many sites only allow Google scrapers in robots.txt and the public will fight you hard if you scrape them without permission. It's just one of those things where it would be better for everyone if search engines could pay for access to the work that's done only once.

> Many sites only allow Google scrapers in robots.txt and the public will fight you hard if you scrape them without permission.

This just lets a monopoly replace the website instead of distributing power and fostering open source. The same monopoly that was already bleeding off the web's utility and taxing it.

Re: Updates to our web search products and Programmable Search Engine capabilities

#132

Earlier quoted context omitted.

Though I'd think that you'd want to weight unaffiliated sites' anchor text to a given URL much higher than an affiliated site. "Affiliation" is a tricky term itself. Content farms were popular in the aughts (though they seem to have largely subsided), firms such as Claria and Gator. There are chumboxes (Outbrain, Taboola), and of course affiliate links (e.g., to Amazon or other shopping sites). SEO manipulation is it…

Oh yeah, there's definitely room for improvement in that general direction. Indexing anchor texts is much better than page rank, but in isolation, it's not sufficient. I've also seen some benefit fingerpinting the network traffic the websites make using a headless browser, to identify which ad networks they load. Very few spam sites have no ads, since there wouldn't be any economy in that. e.g. https://marginalia-sea…

Oh, that is clever!

I'd also suspect that there are networks / links which are more likely signs of low-value content than others. Off the top of my head, crypto, MLM, known scam/fraud sites, and perhaps share links to certain social networks might be negative indicators.

Re: Updates to our web search products and Programmable Search Engine capabilities

#133
I built many products on Google PSE (Custom Search). Results were nowhere near as good as regular Google, but still useful. I usually needed to use another library to get the DOM content anyway. But it still was solid for grounding/checking data.

RIP, another one to the Google Graveyard.

Re: Updates to our web search products and Programmable Search Engine capabilities

#134

Earlier quoted context omitted.

The input on the results page doesn't work, you always need to return to the start page on which the browser history is disabled. That's just confusing behaviour.

I guess you used the return key instead of clicking on the search icon? Seems to be a bug with the return key, I'll fix that this weekend sorry.

True, didn't occur to me, that I should click on the icon instead. Once I have clicked on the search icon once, enter also works. When I input a short query (single letter) it sometimes just shows a blank page, but maybe that is just HNs hug of death. Consider putting the query term more prominently in the front of the URL, so users can edit it. Also from the startpage, the URL in the URLbar isn't updated. As I already wrote, the browser shows completion for the searchbar on the result page, but does not for the one one the startpage. For my taste I would prefer less JS trickery, which would maybe already get rid of some of these issues.

Re: Updates to our web search products and Programmable Search Engine capabilities

#135

Earlier quoted context omitted.

You should consider filtering by input language. Showing the same Wikipedia article in different languages is not helpful when I am searching in English. Also you may unify by entries by URL, it shows the same URL, just with different publish dates, which is interesting and might be useful, but should maybe be behind a toggle, as it is confusing at first.

Great feedback, agree I need to filter here. Some website localization is very hard to work around, because they will try to geo-locate the IP address of your bot and redirect it accordingly to a given language...

The issue I was having was with the query "term+wikipedia" it then shows the wikipedia article in Czech, Hungarian, Russian, some kind of Arab and other before finally showing the English version. Then also a lot of that occur 2,3,4+ times with the same URL, just differing in crawltime by a few minutes.

Re: Updates to our web search products and Programmable Search Engine capabilities

#136
post #94

Earlier quoted context omitted.

They then go on to say that they pay a 3rd party company to scrape Google results (and serve those scraped results to their users). So their search engine is indeed based on unauthorized and uncompensated use of Google's index. But since they're not using/paying for a supported API but just taking what they want, they indeed are unlikely to be impacted by this API turndown.

Congrats on saying that in the most one-sided way possible. Google makes it literally impossible for them to pay for access to search results to make the product they want (customizable subscription search with no ads), and Google also is the de-facto globally sanctioned crawler because they are the only search engine anyone gives a shit about, and also sites need to be indexed by them to survive. In short, Google ow…

>In short, Google owns the river and sells the boats, and the public built a wall around it.

That would be a monopoly if there was only 1 river in the whole world.

Re: Updates to our web search products and Programmable Search Engine capabilities

#137

Earlier quoted context omitted.

Oh yeah, there's definitely room for improvement in that general direction. Indexing anchor texts is much better than page rank, but in isolation, it's not sufficient. I've also seen some benefit fingerpinting the network traffic the websites make using a headless browser, to identify which ad networks they load. Very few spam sites have no ads, since there wouldn't be any economy in that. e.g. https://marginalia-sea…

Oh, that is clever! I'd also suspect that there are networks / links which are more likely signs of low-value content than others. Off the top of my head, crypto, MLM, known scam/fraud sites, and perhaps share links to certain social networks might be negative indicators.

You can actually identify clusters of websites based on the cosine similarity of their outbound links. Pretty useful for identifying content farms spanning multiple websites.

Have a lil' data explorer for this: https://explore2.marginalia.nu/

Quite a lot of dead links in the dataset, but it's still useful.

Re: Updates to our web search products and Programmable Search Engine capabilities

#138
post #28

Earlier quoted context omitted.

No wonder Kagi is angry. Google is a monopoly across several broad categories. They're also a taxation enterprise. Google Search took over as the URL bar for 91% of all web users across all devices. Since this intercepts trademarks and brand names, Google gets to tax all businesses unfairly. Tell your legislators in the US and the EU that Google shouldn't be able to sell ads against registered trademarks (+/- some ed…

what stops Kagi from indexing internet and makes them pay some guys to scrape search results from Google? one guy at Marginalia can do it and entire dev team at a PAID search engine can't?

I don't know about others, but we have special rules for Google, Bing, and a few others, rate-limiting them less than some random bot.

The problem is scrapers (mostly AI scrapers from what we can tell). They will pound a site into the ground and not care and they are becoming increasingly good at hiding their tracks. The only reasonable way to deal with them is to rate-limit every IP by default and then lifting some of those restrictions on known, well behaving bots. Now we will lift those restrictions if asked, and frequently look at statistics to lift the restrictions from search engines we might have missed, but it's an up hill battle if you're new and unknown.

Re: Updates to our web search products and Programmable Search Engine capabilities

#139

Google quietly announced that Programmable Search (ex-Custom Search) won’t allow new engines to “search the entire web” anymore. New engines are capped at searching up to 50 domains, and existing full-web engines have until Jan 1, 2027 to transition. If you actually need whole-web search, Google now points you to an “interest form” for enterprise solutions (Vertex AI Search etc.), with no public pricing and no guaran…

I know that duckduckgo uses Microsoft Bing Custom search and honestly it is a much more robust system since you don't have to worry about Google axing it. https://www.customsearch.ai
Post reply on HN