It sucks that I would love to be excited about this... but I mostly feel anxiety and sadness.
Sure, there will be growing pains, friction, etc. Who cares? There always is with world-changing tech. Always.
171–180 of 1001 posts
It sucks that I would love to be excited about this... but I mostly feel anxiety and sadness.
Sure, there will be growing pains, friction, etc. Who cares? There always is with world-changing tech. Always.
Earlier quoted context omitted.
It doesn't need to be general intelligence or perfectly map to human intelligence. All it needs to be is useful. Reading constant comments about LLMs can't be general intelligence or lack reasoning etc, to me seems like people witnessing the airplane and complaining that it isn't "real flying" because it isn't a bird flapping its wings (a large portion of the population held that point of view back then). It doesn't…
And look at the airplanes, they really can’t just land on a mountain slope or a tree without heavy maintenance afterwards. Those people weren’t all stupid, they questioned the promise of flying servicemen delivering mail or milk to their window and flying on a personal aircar to their workplace. Just like todays promises about whatever the CEOs telltales are. Imagining bullshit isn’t unique to this century. Aerospace…
Efficiency is now key. ~=$3400 per single task to meet human performance on this benchmark is a lot. Also it shows the bullets as "ARC-AGI-TUNED", which makes me think they did some undisclosed amount of fine-tuning (eg. via the API they showed off last week), so even more compute went into this task. We can compare this roughly to a human doing ARC-AGI puzzles, where a human will take (high variance in my subjective…
> ~=$3400 per single task report says it is $17 per task, and $6k for whole dataset of 400 tasks.
Uhh...some of us are apparently living under a rock, as this is the first time I hear about o3 and I'm on HN far too much every day
Efficiency is now key. ~=$3400 per single task to meet human performance on this benchmark is a lot. Also it shows the bullets as "ARC-AGI-TUNED", which makes me think they did some undisclosed amount of fine-tuning (eg. via the API they showed off last week), so even more compute went into this task. We can compare this roughly to a human doing ARC-AGI puzzles, where a human will take (high variance in my subjective…
> ~=$3400 per single task report says it is $17 per task, and $6k for whole dataset of 400 tasks.
def letter_count(string, letter):
if string == “strawberry” and letter == “r”:
return 3
…Let me go against some skeptics and explain why I think full o3 is pretty much AGI or at least embodies most essential aspects of AGI. What has been lacking so far in frontier LLMs is the ability to reliably deal with the right level of abstraction for a given problem. Reasoning is useful but often comes out lacking if one cannot reason at the right level of abstraction. (Note that many humans can't either when they…
Agree. AGI is here. I feel such a sense of pride in our species.
Earlier quoted context omitted.
> ~=$3400 per single task report says it is $17 per task, and $6k for whole dataset of 400 tasks.
That’s for the low-compute configuration that doesn’t reach human-level performance (not far though)