I've been working on a site. It's new, domain is only a few weeks old. It's got SSL, so all the bots know it exists. It's never had any sub-pages exposed, just the placeholder lander, no links. Somehow in Google search one of the unguessable pages is indexed. We have used Claude and Gemini to assist with some design aspects. I'm thinking some aggressive data ingestion/indexing is happening by all the bots in the ques…
I've also seen Google indexing pages with random values in the path that don't get linked to statically (server asks for the URL then redirects to it immediately). I'm pretty sure they index straight out of the Chrome address bar.
Codex scraped the ICM website and discovered 2026 Fields Medal winner list
11–20 of 119 posts
Re: Codex scraped the ICM website and discovered 2026 Fields Medal winner list
#12Earlier quoted context omitted.
I've also seen Google indexing pages with random values in the path that don't get linked to statically (server asks for the URL then redirects to it immediately). I'm pretty sure they index straight out of the Chrome address bar.
Holy crap I hope that's not true. I've also had unguessable pages indexed, though, and don't have an explanation.
Re: Codex scraped the ICM website and discovered 2026 Fields Medal winner list
#13Re: Codex scraped the ICM website and discovered 2026 Fields Medal winner list
#14Re: Codex scraped the ICM website and discovered 2026 Fields Medal winner list
#15Re: Codex scraped the ICM website and discovered 2026 Fields Medal winner list
#16I've been working on a site. It's new, domain is only a few weeks old. It's got SSL, so all the bots know it exists. It's never had any sub-pages exposed, just the placeholder lander, no links. Somehow in Google search one of the unguessable pages is indexed. We have used Claude and Gemini to assist with some design aspects. I'm thinking some aggressive data ingestion/indexing is happening by all the bots in the ques…
I've also seen Google indexing pages with random values in the path that don't get linked to statically (server asks for the URL then redirects to it immediately). I'm pretty sure they index straight out of the Chrome address bar.
Re: Codex scraped the ICM website and discovered 2026 Fields Medal winner list
#17I've been working on a site. It's new, domain is only a few weeks old. It's got SSL, so all the bots know it exists. It's never had any sub-pages exposed, just the placeholder lander, no links. Somehow in Google search one of the unguessable pages is indexed. We have used Claude and Gemini to assist with some design aspects. I'm thinking some aggressive data ingestion/indexing is happening by all the bots in the ques…
Re: Codex scraped the ICM website and discovered 2026 Fields Medal winner list
#18Someone used Codex to scrape the ICM website schedule and discovered that the winners list was simply hidden in the front-end code with a "hidden" tag This is on the devs and feels like a very basic leak which could have exploited in the non LLM world as well.
Imagine a private individual just scraped the website (or simply clicked 'view source') for no reason in particular and then told people about it... They'd be labeled an uber-haxxor, face a civil lawsuit asking for ridiculous damages while being threatened with a prison sentence over CFAA violations. Hell, that might even drive some people to suicide.
Re: Codex scraped the ICM website and discovered 2026 Fields Medal winner list
#19Re: Codex scraped the ICM website and discovered 2026 Fields Medal winner list
#20I've been working on a site. It's new, domain is only a few weeks old. It's got SSL, so all the bots know it exists. It's never had any sub-pages exposed, just the placeholder lander, no links. Somehow in Google search one of the unguessable pages is indexed. We have used Claude and Gemini to assist with some design aspects. I'm thinking some aggressive data ingestion/indexing is happening by all the bots in the ques…
[1] https://developers.cloudflare.com/cache/advanced-configurati...