Live data from Hacker News

Codex scraped the ICM website and discovered 2026 Fields Medal winner list

phemex.com

31–40 of 119 posts

Re: Codex scraped the ICM website and discovered 2026 Fields Medal winner list

#31
post #8

I've been working on a site. It's new, domain is only a few weeks old. It's got SSL, so all the bots know it exists. It's never had any sub-pages exposed, just the placeholder lander, no links. Somehow in Google search one of the unguessable pages is indexed. We have used Claude and Gemini to assist with some design aspects. I'm thinking some aggressive data ingestion/indexing is happening by all the bots in the ques…

You ISP also collects and sells data to companies like Moz, and possibly to Google too.

Re: Codex scraped the ICM website and discovered 2026 Fields Medal winner list

#32
post #8

I've been working on a site. It's new, domain is only a few weeks old. It's got SSL, so all the bots know it exists. It's never had any sub-pages exposed, just the placeholder lander, no links. Somehow in Google search one of the unguessable pages is indexed. We have used Claude and Gemini to assist with some design aspects. I'm thinking some aggressive data ingestion/indexing is happening by all the bots in the ques…

Isn't leaking browser extension used by one of people on the team (doesn't need to be developer, could be qa or anybody with whom the access was shared) more plausible?

Re: Codex scraped the ICM website and discovered 2026 Fields Medal winner list

#33
post #31
post #8

I've been working on a site. It's new, domain is only a few weeks old. It's got SSL, so all the bots know it exists. It's never had any sub-pages exposed, just the placeholder lander, no links. Somehow in Google search one of the unguessable pages is indexed. We have used Claude and Gemini to assist with some design aspects. I'm thinking some aggressive data ingestion/indexing is happening by all the bots in the ques…

You ISP also collects and sells data to companies like Moz, and possibly to Google too.

URL paths over https wouldn't be transparent to the ISP though, would they?

Re: Codex scraped the ICM website and discovered 2026 Fields Medal winner list

#34
post #15

Earlier quoted context omitted.

Wikipedia says Hong Wang while acknowledging that the native form is Wang Hong and that they are using the Western name order.

Nobody says Jinping Xi or Zedong Mao.

Is Elton John or Jhon Elton?

Re: Codex scraped the ICM website and discovered 2026 Fields Medal winner list

#35
post #2

Someone used Codex to scrape the ICM website schedule and discovered that the winners list was simply hidden in the front-end code with a "hidden" tag This is on the devs and feels like a very basic leak which could have exploited in the non LLM world as well.

Most of what an LLM does "could have" been done by a human if you throw enough human hours at it. But the reality in this circumstance is that a new tool helped find this leak. Saying this could have happened in a "non LLM world" is analogous to "someone else could have discovered special relativity, let's not mention Einstein"

Re: Codex scraped the ICM website and discovered 2026 Fields Medal winner list

#36
post #31

Earlier quoted context omitted.

You ISP also collects and sells data to companies like Moz, and possibly to Google too.

URL paths over https wouldn't be transparent to the ISP though, would they?

They would not - GP was probably bringing up something not directly relevant, but still related. (they should have clarified though)

Re: Codex scraped the ICM website and discovered 2026 Fields Medal winner list

#37

Earlier quoted context omitted.

I've also seen Google indexing pages with random values in the path that don't get linked to statically (server asks for the URL then redirects to it immediately). I'm pretty sure they index straight out of the Chrome address bar.

Holy crap I hope that's not true. I've also had unguessable pages indexed, though, and don't have an explanation.

There's no reason to think it isn't true. It matches every pattern of behavior observed from every tech company.

Re: Codex scraped the ICM website and discovered 2026 Fields Medal winner list

#38
post #8

I've been working on a site. It's new, domain is only a few weeks old. It's got SSL, so all the bots know it exists. It's never had any sub-pages exposed, just the placeholder lander, no links. Somehow in Google search one of the unguessable pages is indexed. We have used Claude and Gemini to assist with some design aspects. I'm thinking some aggressive data ingestion/indexing is happening by all the bots in the ques…

Nothing you enter into an LLM not hosted by you, or put onto the web is safe from being collected and exploited by these "AI" companies and their LLM's voracious appetite.

Re: Codex scraped the ICM website and discovered 2026 Fields Medal winner list

#39
post #15

Earlier quoted context omitted.

Wikipedia says Hong Wang while acknowledging that the native form is Wang Hong and that they are using the Western name order.

Nobody says Jinping Xi or Zedong Mao.

Can we not just agree that transliteration is tricky business with no single canon?

Some Indian restaurants near me sell Aloo Saag, others sell Alu Sag.

Re: Codex scraped the ICM website and discovered 2026 Fields Medal winner list

#40
post #15

Earlier quoted context omitted.

Wikipedia says Hong Wang while acknowledging that the native form is Wang Hong and that they are using the Western name order.

Nobody says Jinping Xi or Zedong Mao.

Well, some do say Jinping the Eleventh …
Post reply on HN