Live data from Hacker News

Viewing profile — Bolwin

Bolwin

HN member
Joined
Sat, Dec 21, 2024, 3:26 AM UTC
HN karma
332
Public activity
207 items

About Bolwin

No profile information was provided.

Recent public activity

  1. comment
    Comment #49228348

    Yeah and it's degraded significantly since then. Older llms were still mostly language focused and had a lot of latent knowledge about things like writing styles. Now it's crowded …

  2. comment
    Comment #49187934

    > Muse Spark 1.2 is available today in Muse Code and in Meta Model API with expanded global access Wasn't the previous one us only? This is probably the biggest part of the post An…

  3. comment
    Comment #49173001

    Two things 1. If we're using native harnesses, I'd have preferred you use kimi code, not opencode 2. The variation in the two kimi providers just shows how you can't trust n = 1 tr…

  4. comment
    Comment #49157682

    What about your examples has an llm tell? I don't trust pangram 4 much. There was a post here earlier confirming it fails for many others.

  5. comment
    Comment #49138166

    It feels like that should be a golden opportunity for competitors, but every time a competitor makes a decent replacement, the big tech company either buys it or briefly invests in…

  6. comment
    Comment #49115979

    Why is the title "Africa" and not Morocco

  7. comment
    Comment #49088306

    Screams it in fact

  8. comment
    Comment #49054536

    Jeez way to ruin of the few remaining joys of flying. Lock everyone in a tin can. You will experience reality through screens only and you will enjoy it

  9. comment
    Comment #49044383

    Llms use them a lot more than humans, including this blog post. Like all slop. There's a reason it's called slop and it's not because of restraint

  10. comment
    Comment #49044344

    I like this. It feels like more a person talking and less like a prepared speech.

  11. comment
    Comment #49043744

    For a fair comparison, you should compare to K3 (which AA has not tested yet unfortunately) and GPT 5.6 Sol also on medium or the closest equivalent

  12. comment
    Comment #48986359

    This is a lot of words to say "switch models instead of writing a plan file and starting a new session" which I do anyway. That said, it takes me a while to reads plans and often I…

  13. comment
    Comment #48972158

    X.com is not publicly accessible. I wish people would stop using it as a source

  14. comment
    Comment #48931140

    Every time a competitor for acquired, sold out, or changed business models, Kilo made a snarky post. All companies are the same

  15. comment
    Comment #48928293

    Where are you getting supposedly. It does worse in most benchmarks

  16. comment
    Comment #48884933

    Claude is very cache friendly, however there have been some inconsistencies with non anthropic endpoints that led to cache breakages

  17. comment
    Comment #48867912

    Glaciers have never been accessible to most people. The night sky has, until recently.

  18. comment
    Comment #48842656

    Yes Americans can do both, unless their boss dislikes it, but that applies the world over.

  19. comment
    Comment #48838156

    Has he made anything else interesting?

  20. comment
    Comment #48814308

    When I use it for fiction, I generally switch models 2-3 times per response. It's basically normal

  21. comment
    Comment #48792442

    Tools come with a tool description in json schema format, but yes your point stands, it is not enough for opus 4.8 which I've also noticed having tool call issues.

  22. comment
    Comment #48693854

    > Yeah, but the biggest plus for open models is that they can never be taken away. In other words, whatever capabilities they reach (even if there will never be another model), tho…

  23. comment
    Comment #48633391

    > Interleaved reasoning and function calling makes this even more dangerous. A model can call functions during the hidden reasoning phase. The reasoning may be hidden but the tool …

  24. comment
    Comment #48612328

    Hah, I noticed the same thing writing fiction with fable. Most models seem to go into a sort of "storytelling mode" where they forget their PhD level smarts. I had a character who …

  25. comment
    Comment #48575296

    For Claude models at least, you can tell to just manually think in the output and it works fine. I do it reguralrly because for creative writing and summarization, they seem to bel…