Live data from Hacker News

Viewing profile — byefruit

byefruit

HN member
Joined
Mon, Aug 27, 2012, 9:55 AM UTC
HN karma
936
Public activity
134 items

About byefruit

No profile information was provided.

Recent public activity

  1. comment
    Comment #45983989

    I've just found myself using OpenRouter if we need Google models for a project, it's worth the extra 5% just not to have to deal with the utter disaster that is their product offer…

  2. comment
    Comment #45848523

    This is the wrong interpretation of the oxcaml project. If you look at the features and work on it, it's primarily performance or parallelism safety features. The latter going much…

  3. comment
    Comment #45837580

    7.3% return, not bad. As battery prices drop it will get even better.

  4. comment
    Comment #45454055

    And even when it does copy other products, it seems to be doing a terrible job of them. Google's AI offering is a complete nightmare to use. Three different APIs, at least two diff…

  5. comment
    Comment #44815378

    How is this different from https://github.com/google-gemini/gemini-cli ? Edit: it seems this is a hosted version. Would be nice if they actually joined up some of their products.

  6. comment
    Comment #44815024

    The openrouter rankings can be biased. For example, Google's inexplicable design decisions around libraries and APIs means it's often worth the 5% premium to just use OpenRouter to…

  7. comment
    Comment #44620170

    Before just accepting this at face value, New Statesman claim this is not the case: https://www.newstatesman.com/politics/2025/07/the-british-we...

  8. comment
    Comment #43919425

    You are probably getting downvoted because you don't give any model generations or versions ('ChatGPT') which makes this not very credible.

  9. comment
    Comment #43888448

    100% this. We actually use OpenRouter (and pay their surcharge) with Gemini 2.5 Pro just because we can actually control spend via spent limit on keys (A++ feature) and prepaid cre…

  10. comment
    Comment #43751226

    Indeed, average in CA is $260/month so $5k pays off very fast in some places.

  11. comment
    Comment #43721089

    It's interesting that there's a price nearly 6x price difference between reasoning and no reasoning. This implies it's not a hybrid model that can just skip reasoning steps if requ…

  12. comment
    Comment #43682759

    "In addition, S3 Express One Zone has reduced the per-GB charges for data uploads and retrievals by 60 percent, and these charges now apply to all bytes transferred rather than jus…

  13. comment
    Comment #43674014

    > Both of our models are trained on top of DeepSeek-R1-Distill-Qwen-7B and DeepSeek-R1-Distill-Qwen-32B. Not to take away from their work but this shouldn't be buried at the bottom…

  14. comment
    Comment #43523779

    https://aws.amazon.com/snowball/pricing/ snowball seems to support getting data out of S3 though you still end up paying extortionate egress charges.

  15. comment
    Comment #43285226

    I think this is where (in the EU) a Subject Access Request could work.

  16. comment
    Comment #43251958

    This has already been proposed by the current government for wind farms: https://inews.co.uk/news/environment/new-energy-bill-discoun... And the switch to zonal energy pricing will…

  17. comment
    Comment #43126040

    I'm waiting for https://github.com/huggingface/trl/pull/2810 to land. I think this should work with the existing unsloth setup without changes.

  18. comment
    Comment #42937089

    > This generally requires thousands of examples created by an expert in the field. Or an AI model pretending to be an expert in the field... (works well in a few niche domains I ha…

  19. comment
    Comment #42890190

    I look forward to people applying the same standards to the OpenAI's O3 as they did Deepseek's R1 release and paper in the discussions last week.

  20. comment
    Comment #42824594

    That's not what I'm saying, they may be hiding their true compute. I'm pointing out that nearly every thread covering Deepseek R1 so far has been like this. Compare to the O1 syste…

  21. comment
    Comment #42824373

    It's amazing how different the standards are here. Deepseek's released their weights under a real open source license and published a paper with their work which now has independen…

  22. comment
    Comment #42813626

    What I don't understand is how you don't end up with a totally impractical number of vectors if you have one per token? Surely nobody is storing that many in any real system?

  23. comment
    Comment #42778855

    Do you have any evidence for this accusation? O1's reasoning traces aren't even shown, are you suggesting they've somehow exfiltrated them?

  24. comment
    Comment #42768824

    This is pretty harsh on DeepSeek. There are some significant innovations behind behind v2 and v3 like multi-headed latent attention, their many MoE improvements and multi-token pre…

  25. comment
    Comment #42635275

    What would you recommend for building a strong linear algebra foundation?