Live data from Hacker News

Gemini 2.5 Computer Use model

blog.google

131–140 of 339 posts

Re: Gemini 2.5 Computer Use model

#131

Earlier quoted context omitted.

That's ridiculous, sorry. If that were so, we wouldn't have positional encodings in vision transformers.

It's not ridiculous if you understand how neural networks actually work. Your perception of the numbers has nothing to do w/ the logic of the arithmetic in the network.

Do you know what "positional encoding" means?

Re: Gemini 2.5 Computer Use model

#132

Earlier quoted context omitted.

It's not ridiculous if you understand how neural networks actually work. Your perception of the numbers has nothing to do w/ the logic of the arithmetic in the network.

Do you know what "positional encoding" means?

Completely irrelevant to the point being made.

Re: Gemini 2.5 Computer Use model

#133
post #126
post #98

Earlier quoted context omitted.

I cycle a lot. Outdoors I listen to podcasts and the fact that I can say "Hey Google, go back 30sec" to relisten to something (or forward to skip ads) is very valuable to me. Indoors I tend to cast some show or youtube video. Often enough I want to change the Youtube video or show using voice commands - I can do this for Youtube, but results are horrible unless I know exactly which video I want to watch. For other se…

Do you have a lot of dedicated cycle ways? I'm not sure I'd want to have headphones impeding my hearing anywhere I'd have to interact with cars or pedestrians while on my bike.

Lots of noise cancelling headphones have a pass-through mode that lets you hear the outside world. Alternatively, I use bone conducting headphones that leave my ears uncovered.

Re: Gemini 2.5 Computer Use model

#134
post #69

Earlier quoted context omitted.

Yeah I think it would be better to just have the model write out playwright scripts than the way it's doing it right now (or at least first navigate manually and then based on that, write a playwright typescript script for future tests). Cuz right now it's way too slow... perform an action, then read the results, then wait for the next tool call, etc.

This is basically our approach with Herd[0]. We operate agents that develop, test and heal trails[1, 2], which are packaged browser automations that do not require browser use LLMs to run and therefore are much cheaper and reliable. Trail automations are then abstracted as a REST API and MCP[3] which can be used either as simple functions called from your code, or by your own agent, or any combination of such. You ca…

Whoa that’s cool. I’ll check it out, thanks!

Re: Gemini 2.5 Computer Use model

#135

> It is not yet optimized for desktop OS-level control Alas, AGI is not yet here. But I feel like if this OS-level of control was good enough, and the cost of the LLM in the loop wasn't bad, maybe that would be enough to kick start something akin to AGI.

Funny thing is, most humans cannot properly control a computer. Intelligence seems to be impossible to define.

Intelligence is whatever an LLM can’t do yet. Fluid intelligence is the capacity to quickly move goal posts.

Re: Gemini 2.5 Computer Use model

#136
post #71

Many years ago I was sitting at a red light on a secondary road, where the primary cross road was idle. It seemed like you could solve this using a computer vision camera system that watched the primary road and when it was idle, would expedite the secondary road's green light. This was long before computer vision was mature enough to do anything like that and I found out that instead, there are magnetic systems that…

Ironically now that computer vision is commonplace, the cameras you talk about have become increasingly popular over the years because the magnetic systems do not do a very good job of detecting cyclists and the cameras double as a congestion monitoring tool for city staff.

and soon/now triple as surveillance.

Re: Gemini 2.5 Computer Use model

#137

doesn't seem like it makes sense to train AI around human user interfaces which aren't really efficient. It is like building a mechanical horse.

Why do you think we have fully self driving cars instead of just more simplistic beacon systems? Why doesn't McDonald's have a fully automated kitchen? New technology is slow due to risk aversion, it's very rare for people to just tear up what they already have to re-implement new technology from the ground up. We always have to shoe-horn new technology into old systems to prove it first. There are just so many facto…

> Why do you think we have fully self driving cars instead of just more simplistic beacon systems?

While the self-driving car industry aims to replace all humans with machines, I don't think this is the case with browser automation.

I see this technology as more similar to a crash dummy than a self-driving system. It's designed to simulate a human in very niche scenarios.

Re: Gemini 2.5 Computer Use model

#138
I think it’s related that I got an email from google, titled “ Simplifying your Gemini Apps experience”. It reads no privacy maximize AI. They are going to automatically collect data from all google apps, and users no longer have options to control access to individual apps.

Re: Gemini 2.5 Computer Use model

#139

Earlier quoted context omitted.

The original comment I replied to said "You can navigate a website without visually decoding the image of a website." I replied that decoding is necessary to know where the elements will end up in a visual arrangement, because often that carries semantics. A label that is rendered next to another element can be crucial for understanding the functioning of the program. It's nontrivial just from the HTML or whatever tr…

2D rendering is not necessary for processing information by neural networks. In fact, the image is flattened into 1D array & loses the topological structure almost entirely b/c the topology is not relevant to the arithmetic performed by the network.

I'm talking about HTML (or other markup, in the form of text) vs image. That simply getting the markup as text tokens will be much harder to interpret since it's not clear where the elements will end up. I guess I can't make this any more clear.

Re: Gemini 2.5 Computer Use model

#140
post #71

Many years ago I was sitting at a red light on a secondary road, where the primary cross road was idle. It seemed like you could solve this using a computer vision camera system that watched the primary road and when it was idle, would expedite the secondary road's green light. This was long before computer vision was mature enough to do anything like that and I found out that instead, there are magnetic systems that…

The best thing about being nerds like we are is we can just ignore this product since it's not for us.
Post reply on HN