I've been working on a site. It's new, domain is only a few weeks old. It's got SSL, so all the bots know it exists. It's never had any sub-pages exposed, just the placeholder lander, no links. Somehow in Google search one of the unguessable pages is indexed. We have used Claude and Gemini to assist with some design aspects. I'm thinking some aggressive data ingestion/indexing is happening by all the bots in the ques…
Codex scraped the ICM website and discovered 2026 Fields Medal winner list
31–40 of 119 posts
Re: Codex scraped the ICM website and discovered 2026 Fields Medal winner list
#32I've been working on a site. It's new, domain is only a few weeks old. It's got SSL, so all the bots know it exists. It's never had any sub-pages exposed, just the placeholder lander, no links. Somehow in Google search one of the unguessable pages is indexed. We have used Claude and Gemini to assist with some design aspects. I'm thinking some aggressive data ingestion/indexing is happening by all the bots in the ques…
Re: Codex scraped the ICM website and discovered 2026 Fields Medal winner list
#33I've been working on a site. It's new, domain is only a few weeks old. It's got SSL, so all the bots know it exists. It's never had any sub-pages exposed, just the placeholder lander, no links. Somehow in Google search one of the unguessable pages is indexed. We have used Claude and Gemini to assist with some design aspects. I'm thinking some aggressive data ingestion/indexing is happening by all the bots in the ques…
You ISP also collects and sells data to companies like Moz, and possibly to Google too.
Re: Codex scraped the ICM website and discovered 2026 Fields Medal winner list
#34Re: Codex scraped the ICM website and discovered 2026 Fields Medal winner list
#35Someone used Codex to scrape the ICM website schedule and discovered that the winners list was simply hidden in the front-end code with a "hidden" tag This is on the devs and feels like a very basic leak which could have exploited in the non LLM world as well.
Re: Codex scraped the ICM website and discovered 2026 Fields Medal winner list
#36Earlier quoted context omitted.
You ISP also collects and sells data to companies like Moz, and possibly to Google too.
URL paths over https wouldn't be transparent to the ISP though, would they?
Re: Codex scraped the ICM website and discovered 2026 Fields Medal winner list
#37Earlier quoted context omitted.
I've also seen Google indexing pages with random values in the path that don't get linked to statically (server asks for the URL then redirects to it immediately). I'm pretty sure they index straight out of the Chrome address bar.
Holy crap I hope that's not true. I've also had unguessable pages indexed, though, and don't have an explanation.
Re: Codex scraped the ICM website and discovered 2026 Fields Medal winner list
#38I've been working on a site. It's new, domain is only a few weeks old. It's got SSL, so all the bots know it exists. It's never had any sub-pages exposed, just the placeholder lander, no links. Somehow in Google search one of the unguessable pages is indexed. We have used Claude and Gemini to assist with some design aspects. I'm thinking some aggressive data ingestion/indexing is happening by all the bots in the ques…
Re: Codex scraped the ICM website and discovered 2026 Fields Medal winner list
#39Earlier quoted context omitted.
Wikipedia says Hong Wang while acknowledging that the native form is Wang Hong and that they are using the Western name order.
Nobody says Jinping Xi or Zedong Mao.
Some Indian restaurants near me sell Aloo Saag, others sell Alu Sag.