Viewing profile — osti
osti
HN member- Joined
- Tue, Jan 05, 2016, 8:42 PM UTC
- HN karma
- 574
- Public activity
- 235 items
- HN profile
- View on Hacker News ↗
About osti
No profile information was provided.
Recent public activity
-
comment
Comment #49158271
It's an axiom that the modern Western mind is built upon.
-
comment
Comment #49127389
Do you have any numbers on the solve quality? Exploitability numbers etc.
-
comment
Comment #49026386
You can use geekbench 5 in that case. But given that they deprecated that, it might be harder to compare to others.
-
comment
Comment #49026366
I agree. For me personally I mostly only care about single thread geekbench variant, I believe it's an excellent proxy for general performance of a CPU. Multi thread geekbench (or …
-
comment
Comment #48994037
In computer science, one definition of algorithm is basically any program that runs on a turing machine. By that definition, any LLM is an algorithm.
-
comment
Comment #48980681
Lol yet I've used Apple and Android phones extensively and would choose Android every single time.
-
comment
Comment #48963492
He's talking about the plans, you are talking about API prices.
-
comment
Comment #48963272
No idea lol, didn't even know those exist..
-
comment
Comment #48963250
It is complicated, but paying for the cheaper usd plans really don't get you much usage.
-
comment
Comment #48963242
Nah that won't work. I don't know tbh, I just used someone else's number.
-
comment
Comment #48961996
Absolutely do not pay for the kimi plans thinking they will be cheaper. If you sign up with a Chinese phone number, you can get the same plan for 200 yuan instead of 200 usd, it al…
-
comment
Comment #48961930
GPT should be better at these optimization problems given that they won the recent atcoder heuristics competition against top humans. And Anthropic is less focused on these types o…
-
comment
Comment #48851663
This was extremely impressive to me. AtCoder has the hardest problems these days, usually the human onsite final round contestants can't solve more than 2 or 3 problems. This year …
- comment
-
comment
Comment #48849903
Mythos probably wouldn't, otherwise they'd have included it in their release. Next version of Mythos probably will though. And yeah.. Reality has not been kind to LeCun.
-
comment
Comment #48849777
SWE-bench series just aren't that great by today's standard, even Anthropic previously stated Claude had memorized solutions for the non Pro version of the benchmark, I suspect the…
-
comment
Comment #48849553
Huh so that's why it's hard to find. They probably haven't properly optimized their caching, or they are just trying to make more money from there.
-
comment
Comment #48849386
And they'd be right, it's an almost saturated benchmark where even some subpar open source models score very well on. And most models are clustered within a small range so it reall…
-
comment
Comment #48849363
SWE-Bench pro is pretty much useless now even though many ppl still look at it. OpenAI published a report yesterday saying so as well. Only look at DeepSWE and FrontierCode right n…
-
comment
Comment #48849251
GPT usually performs better on DeepSWE while Claude does better on FrontierCode. These two coding benchmarks are pretty much the only ones right now that's still worth taking a loo…
-
comment
Comment #48813643
Most programming is that, but most music and literature are probably uninspired junks as well. But there are many beautiful algorithms (such as the ones in Knuths books) that are m…
-
comment
Comment #48752700
I think at the current stage of LLM, it just doesn't make sense to ever have an annual sub. Things change too quickly that you really don't want to be stuck with one model.
-
comment
Comment #48741187
If anything that'll be more obnoxious because they have to show the government that it's safe.
-
comment
Comment #48721988
This announcement also mentioned that they will release the next version (official non preview version) of v4 in mid July.
-
comment
Comment #48714062
Is it only the Indian government? I don't think that's in any way unique to India, I've seen many poor government websites.