Viewing profile — timabdulla
timabdulla
HN member- Joined
- Tue, Sep 27, 2016, 10:45 PM UTC
- HN karma
- 204
- Public activity
- 39 items
- HN profile
- View on Hacker News ↗
About timabdulla
No profile information was provided.
Recent public activity
-
comment
Comment #47587751
It's a neat idea, but the internet is not very tolerant to things that don't appear to be human traffic. That's why browsers used by web automation infrastructure are often Chromiu…
-
comment
Comment #47076500
Google tends to trumpet preview models that aren't actually production-grade. For instance, both 3 Pro and Flash suffer from looping and tool-calling issues. I would love for them …
-
comment
Comment #46633490
I'd be curious to see screenshots or a video! I only have a Mac at my disposal, unfortunately.
-
comment
Comment #46573553
This seems cool, but beware that Fly's other products are not exactly models of stability and polish. API downtime is a semi-frequent occurrence, as are transient API errors and sl…
-
comment
Comment #44645401
> We conducted three runs per experiment and selected the run with the highest final accuracy for inclusion in the chart (though illustrative examples and anecdotes may be drawn fr…
- comment
-
comment
Comment #43904624
I mean, the fact that OpenAI, at the bleeding edge of it all, has decided to buy an IDE is a rather strong hint that the future of agents handling entire engineering tickets might …
-
comment
Comment #43560808
What were the human PhDs able to do after more than 48 hours of effort? Presumably given that these are top-level PhDs, the replication success rate would be close to 100%?
-
comment
Comment #43560792
How does it perform on e.g. WebVoyager, WebArena, or OSWorld? These seem to be the oft-cited benchmarks when comparing computer-use agents.
-
comment
Comment #43493561
This is the most interesting aspect to me. I had Claude generate a guide to all the gyms in Pokemon Red and instructions for how to quickly execute a play through [0]. It obviously…
-
comment
Comment #43299868
There is no 3.6. There is 3.5 and 3.5 (New), both of which remain available.
-
comment
Comment #43299862
There was never a Sonnet 3.6. They released what is commonly known as 3.6 as "Sonnet 3.5 (New)". Then, because so many folks ended up referring to it as 3.6, they decided to call t…
-
comment
Comment #43299848
My feeling (totally unproven) is that in the drive to make Sonnet 3.7 more "agentic", they've lost some of its ability to actually just stick to what you asked it to do. It seems t…
-
comment
Comment #43173666
It's copy on write.
-
comment
Comment #42919363
I didn't actually give it a goal of writing any particular length, but I do think that perhaps given my not-so-large online footprint, it may have felt "pressured" to generate cont…
-
comment
Comment #42917014
I tried it on a few things I was familiar with just to assess its reliability. The first was on a topic with which I am deeply familiar -- myself -- and it made three factual error…
-
comment
Comment #42916899
I just gave it a whirl. Pretty neat, but definitely watch out for hallucinations. For instance, I asked it to compile a report on myself (vain, I know.) In this 500-word report (ok…
-
comment
Comment #42807591
I think I hit all those points in my previous post, except for the fact that it's two different models, as you've noted. That said, neither of them seem to report scores for the ot…
-
comment
Comment #42807083
Those numbers are not the full story. Note that GP specifically says: "Big jumps in benchmarks from _Claude's Computer Use_ though." Claude Computer Use was not SOTA for browser ta…
- comment
-
comment
Comment #42806730
OpenAI is merely matching SOTA in browser tasks as compared to existing browser-use agents. It is a big improvement over Claude Computer Use, but it is more of the same in the spec…
-
comment
Comment #42791421
I'm not too sure about HTMX in particular, but my Rails app's FE is just HTML and Stimulus/Turbo. I'm not sure why you think it simply "doesn't work". To me, it's a lot simpler. I …
-
comment
Comment #42781625
That's certainly possible. I'm not convinced AGI is just around the corner either, but I can't say with a high degree of certainty that it definitely won't arrive in the next few y…
-
comment
Comment #42780913
I don't think the conclusion of this article is controversial if you accept the premise: If a horizontal AI model is able to serve as a "drop-in remote worker" and all you need to …
-
comment
Comment #42736269
Your take seems much more positive than theirs. What do you think the key differences are between your experience and the one here?