Live data from Hacker News

Smaller, faster, safer: running Kimi and GLM at scale

blog.cloudflare.com

61–70 of 74 posts

Re: Smaller, faster, safer: running Kimi and GLM at scale

#61
post #56

Yeah, I've noticed them recently on OR and whitelisted. Then I backed-off pretty quickly after seeing the cache hit rates. It was also rather revealing to see how some provider hit rates differ when you are using them directly vs via OR.

Worth knowing a cache hit rate can be structurally zero and look identical to a broken one.

Re: Smaller, faster, safer: running Kimi and GLM at scale

#62
post #42

I think Cloudflare not providing ZDR on their inference is the biggest public indicator that Cloudlare glows. We let all traffic get MITM'd, now we're letting our AI conversation get tracked. Cloudflare reeks like a US Honeypot.

Seems like the relevant page is: https://developers.cloudflare.com/workers-ai/platform/data-u... and their privacy policy also applies: https://www.cloudflare.com/privacypolicy/

"Your inputs (e.g., text prompts, image submissions, audio files, etc.), outputs (e.g., generated text/images, translations, etc.), embeddings, and training data constitute Customer Content.

For Workers AI:

    * You own, and are responsible for, all of your Customer Content.
    * Cloudflare does not make your Customer Content available to any other Cloudflare customer.
    * Cloudflare does not use your Customer Content to (1) train any AI models made available on Workers AI or (2) improve any Cloudflare or third-party services, and would not do so unless we received your explicit consent.
    * Your Customer Content for Workers AI may be stored by Cloudflare if you specifically use a storage service (e.g., R2, KV, DO, Vectorize, etc.) in conjunction with Workers AI."
OpenRouter has a page of different providers and what OpenRouter understands their position on this to be: https://openrouter.ai/docs/guides/privacy/provider-logging#d...

For CF they write: "Cloudflare • Prompts are retained for unknown period • Does not train"

So in the sense of training models on your company's codebase or maybe running analytics on prompt content for the purposes of improving their AI product suite they won't use your data. Though presumably for purposes of security/abuse etc. there will be some level of retention as per their privacy policy.

I guess it comes down to how much you trust providers on openrouter who claim to offer absolute ZDR versus Cloudflare and how they would use retained data. Obviously if someone thinks CF is a honeypot designed to sidestep the rise of LetsEncrypt/widespread HTTPS then they wouldn't trust a ZDR claim by them in any case. Would you then trust some of these frontier labs respective claims of ZDR when they are pushing unbelievably hard to win? I don't have strong opinions for this - I have a pretty conservative approach by default and exclusively use local inference on in-office hardware for anything close to or related to customer data. I do some coding on 3rd party services.

I imagine if Mullvad offered a ZDR set of open model endpoints with similar efforts at building trust like their VPN it might be popular.

Re: Smaller, faster, safer: running Kimi and GLM at scale

#63
The Cloud killed HW trust, SW is next

One vCPU means nothing, which chip? which SIMD? what RAM speed? what storage? what latency?

Quantization is the next layer of lies

Society is moving towards offloading intelligence to clouid overlords, since nothing runs on your computer anymore, you can't trust anything

And since datacenters occupy a physical space, they are building a monopoly defacto

Unless transparency becomes mandatory, we are headed towards the biggest self sabotage mankind has ever witnessed

Re: Smaller, faster, safer: running Kimi and GLM at scale

#64
post #56

Yeah, I've noticed them recently on OR and whitelisted. Then I backed-off pretty quickly after seeing the cache hit rates. It was also rather revealing to see how some provider hit rates differ when you are using them directly vs via OR.

Worth knowing a cache hit rate can be structurally zero and look identical to a broken one.

Of course. On the other hand if you see 8 % gap for similar workloads, averaged across tens of sessions, with the same underlying model, it becomes a pretty clear signal. And I do exclude first request per provider per session from the statistic.

Re: Smaller, faster, safer: running Kimi and GLM at scale

#65

The Cloud killed HW trust, SW is next One vCPU means nothing, which chip? which SIMD? what RAM speed? what storage? what latency? Quantization is the next layer of lies Society is moving towards offloading intelligence to clouid overlords, since nothing runs on your computer anymore, you can't trust anything And since datacenters occupy a physical space, they are building a monopoly defacto Unless transparency become…

> One vCPU means nothing, which chip? which SIMD? what RAM speed? what storage? what latency?

Don't forget: When will all of this change?

Re: Smaller, faster, safer: running Kimi and GLM at scale

#66
post #32
post #27

Earlier quoted context omitted.

I'm not sure how accurate this is, but there is pricing here: https://openrouter.ai/provider/cloudflare

So they don't even support K3? What's the point. K2.7 Code is practically free already

I don't see it via openrouter, but it is in the dashboard if you log into CF.

https://imgur.com/a/Sb7jQ0A

$3/$0.30/$15 ( input / cached input / output, all per 1M )

Re: Smaller, faster, safer: running Kimi and GLM at scale

#67
post #62
post #42

I think Cloudflare not providing ZDR on their inference is the biggest public indicator that Cloudlare glows. We let all traffic get MITM'd, now we're letting our AI conversation get tracked. Cloudflare reeks like a US Honeypot.

Seems like the relevant page is: https://developers.cloudflare.com/workers-ai/platform/data-u... and their privacy policy also applies: https://www.cloudflare.com/privacypolicy/ "Your inputs (e.g., text prompts, image submissions, audio files, etc.), outputs (e.g., generated text/images, translations, etc.), embeddings, and training data constitute Customer Content. For Workers AI: * You own, and are responsible for,…

The wording of this phrase is pretty specific: "Cloudflare does not use your Customer Content to (1) train any AI models made available on Workers AI," it would seem they could use your data to train models they don't make available on Workers AI.

Re: Smaller, faster, safer: running Kimi and GLM at scale

#69
post #42

I think Cloudflare not providing ZDR on their inference is the biggest public indicator that Cloudlare glows. We let all traffic get MITM'd, now we're letting our AI conversation get tracked. Cloudflare reeks like a US Honeypot.

Cloudflare could be the end of the open web and it seems crazy that more people aren’t worried about this.

Turnstile everywhere + device attestation required is the direction this all seems to be headed. Because of all the AI bots of course (it is an excellent scapegoat).

Re: Smaller, faster, safer: running Kimi and GLM at scale

#70
post #69
post #42

I think Cloudflare not providing ZDR on their inference is the biggest public indicator that Cloudlare glows. We let all traffic get MITM'd, now we're letting our AI conversation get tracked. Cloudflare reeks like a US Honeypot.

Cloudflare could be the end of the open web and it seems crazy that more people aren’t worried about this. Turnstile everywhere + device attestation required is the direction this all seems to be headed. Because of all the AI bots of course (it is an excellent scapegoat).

It's pretty safe to say that most people don't know what cloudflare does, if they even know that it exists.
Post reply on HN