Live data from Hacker News

Does Google crawl dynamic content?

centrical.com

61–62 of 62 posts

Re: Does Google crawl dynamic content?

#61
post #34

Earlier quoted context omitted.

Chromium is Open Source and embeddable, so my guess is a big part of why their headless variant isn't open source is simply that there isn't really much to it.

From a naive point of view it's easy. But it has to scale and launching a new process for every request works only for a pet project. So a crawler based on a headless browser that consumes little memory and runs for weeks is a major achivement.

Why doesn't that scale? You might be underestimating Google's resources.

Re: Does Google crawl dynamic content?

#62
post #12

I have modified wikipedia pages, then googled it, to see search result "instantly" updated. Also, sneaky web sites often give different results to the googlebot user agent than to a non-google firefox user agent https://en.wikipedia.org/wiki/User_agent https://addons.mozilla.org/en-GB/firefox/search/?q=user+agen...

Google deliberately crawls in a non-Google looking way to try and detect "masking".

[deleted]
Post reply on HN