Live data from Hacker News

Viewing profile — timabdulla

timabdulla

HN member
Joined
Tue, Sep 27, 2016, 10:45 PM UTC
HN karma
204
Public activity
39 items

About timabdulla

No profile information was provided.

Recent public activity

  1. comment
    Comment #47587751

    It's a neat idea, but the internet is not very tolerant to things that don't appear to be human traffic. That's why browsers used by web automation infrastructure are often Chromiu…

  2. comment
    Comment #47076500

    Google tends to trumpet preview models that aren't actually production-grade. For instance, both 3 Pro and Flash suffer from looping and tool-calling issues. I would love for them …

  3. comment
    Comment #46633490

    I'd be curious to see screenshots or a video! I only have a Mac at my disposal, unfortunately.

  4. comment
    Comment #46573553

    This seems cool, but beware that Fly's other products are not exactly models of stability and polish. API downtime is a semi-frequent occurrence, as are transient API errors and sl…

  5. comment
    Comment #44645401

    > We conducted three runs per experiment and selected the run with the highest final accuracy for inclusion in the chart (though illustrative examples and anecdotes may be drawn fr…

  6. comment
  7. comment
    Comment #43904624

    I mean, the fact that OpenAI, at the bleeding edge of it all, has decided to buy an IDE is a rather strong hint that the future of agents handling entire engineering tickets might …

  8. comment
    Comment #43560808

    What were the human PhDs able to do after more than 48 hours of effort? Presumably given that these are top-level PhDs, the replication success rate would be close to 100%?

  9. comment
    Comment #43560792

    How does it perform on e.g. WebVoyager, WebArena, or OSWorld? These seem to be the oft-cited benchmarks when comparing computer-use agents.

  10. comment
    Comment #43493561

    This is the most interesting aspect to me. I had Claude generate a guide to all the gyms in Pokemon Red and instructions for how to quickly execute a play through [0]. It obviously…

  11. comment
    Comment #43299868

    There is no 3.6. There is 3.5 and 3.5 (New), both of which remain available.

  12. comment
    Comment #43299862

    There was never a Sonnet 3.6. They released what is commonly known as 3.6 as "Sonnet 3.5 (New)". Then, because so many folks ended up referring to it as 3.6, they decided to call t…

  13. comment
    Comment #43299848

    My feeling (totally unproven) is that in the drive to make Sonnet 3.7 more "agentic", they've lost some of its ability to actually just stick to what you asked it to do. It seems t…

  14. comment
    Comment #43173666

    It's copy on write.

  15. comment
    Comment #42919363

    I didn't actually give it a goal of writing any particular length, but I do think that perhaps given my not-so-large online footprint, it may have felt "pressured" to generate cont…

  16. comment
    Comment #42917014

    I tried it on a few things I was familiar with just to assess its reliability. The first was on a topic with which I am deeply familiar -- myself -- and it made three factual error…

  17. comment
    Comment #42916899

    I just gave it a whirl. Pretty neat, but definitely watch out for hallucinations. For instance, I asked it to compile a report on myself (vain, I know.) In this 500-word report (ok…

  18. comment
    Comment #42807591

    I think I hit all those points in my previous post, except for the fact that it's two different models, as you've noted. That said, neither of them seem to report scores for the ot…

  19. comment
    Comment #42807083

    Those numbers are not the full story. Note that GP specifically says: "Big jumps in benchmarks from _Claude's Computer Use_ though." Claude Computer Use was not SOTA for browser ta…

  20. comment
  21. comment
    Comment #42806730

    OpenAI is merely matching SOTA in browser tasks as compared to existing browser-use agents. It is a big improvement over Claude Computer Use, but it is more of the same in the spec…

  22. comment
    Comment #42791421

    I'm not too sure about HTMX in particular, but my Rails app's FE is just HTML and Stimulus/Turbo. I'm not sure why you think it simply "doesn't work". To me, it's a lot simpler. I …

  23. comment
    Comment #42781625

    That's certainly possible. I'm not convinced AGI is just around the corner either, but I can't say with a high degree of certainty that it definitely won't arrive in the next few y…

  24. comment
    Comment #42780913

    I don't think the conclusion of this article is controversial if you accept the premise: If a horizontal AI model is able to serve as a "drop-in remote worker" and all you need to …

  25. comment
    Comment #42736269

    Your take seems much more positive than theirs. What do you think the key differences are between your experience and the one here?