Live data from Hacker News

Run interactive commands in Gemini CLI

developers.googleblog.com

51–60 of 80 posts

Re: Run interactive commands in Gemini CLI

#51
post #27

Earlier quoted context omitted.

Gemini doesn't seem to be trained on tool use (which Claude is) so it quiet often thinks it can't do something it certainly can and does a lot of nonsense. For me it fails nearly everytime while it's trying to read project files because it uses relative paths instead of absolute so I've put "For your "ReadFile" and "WriteFile" tool, you MUST use absolute paths to files" in my system instructions. Speaking of system i…

Gemini seems to have a poor model of both what it can and what it is allowed to do. I’ve noticed the latter with several image generation refusals I could eventually easily talk them out of (usually by mentioning fair use in a copyright/trademark context).

> Gemini seems to have a poor model of both what it can and what it is allowed to do.

Starting to feel like LLMs models are more of a representation of the culture of the company training them, than a fair representation of the world at large.

Re: Run interactive commands in Gemini CLI

#52
post #43

Earlier quoted context omitted.

Gemini CLI is definitely a much worse client than some of the other agent clients like opencode, cursor etc. But from my experience, that isn't because of the model quality. I get better quality responses from the gemini web chat interface than chatgpt, claude etc. Of course my experience is anecdotal, but we hardly have any decent benchmarks to compare these models. I suspect most benchmarks have leaked into trainin…

Also people don't talk enough about (or are bad at separating themselves) the model vs. the client tool - e.g. from your comment maybe using codex/Claude Code/aider with Gemini API would be better, best even, but people rarely make that comparison or separation, it's always 'Claude Code with Claude vs. codex with GPT-x' etc.

To be fair, most of the times, the tools works best with the models trained with those tools in mind, and vice-versa.

Not to mention not all models/inference works the same way so you can't really replicate the same experience. For example, new Harmony format means you can now inject messages while GPT-OSS is running inference, but obviously Claude Code don't support that because their models don't support that.

Re: Run interactive commands in Gemini CLI

#53

I've had a pretty poor experience with Gemini. I've had to convince it to do things it should just be able to do but thinks it can't for some reason. Like reading from a file outside of the project directory- it can do it fine, but refuses to unless you convince it that no it actually can. Also has inserted "\n" instead of newlines on a number of occasions. I'd argue these behaviors are much more important than being…

Well, those are problems with the underlying Gemini models. It's not like the team responsible for CLI could have trained a better model instead of making this feature.

Gemini 3.0 is likely to be released soon, and likely they would improve agentic coding experience.

Re: Run interactive commands in Gemini CLI

#54
post #47

Building an interactive shell inside their CLI seems like a very odd technical solution. I can’t think of any use case where the same context gathering couldn’t be gleaned by examining the file/system state after the session ended, but maybe I’m missing something. On the other hand, now that I’ve read this, I can see how having some hooks between the code agent CLIs and ghostty/etc could be extremely powerful.

LLMs in general struggles with numbers, it's easy to tell with the medium sized models that struggle with line replacement commands where it has to count, it usually takes a couple of tries to get right. I always imagined they'd have an easier time if they could start a vim instance and send search/movement/insert commands instead, not having to keep track of numbers and do calculations, but instead visually inspect…

Gotta be better than codex literally writing a python script to edit a file multiple times in a single prompt response.

Re: Run interactive commands in Gemini CLI

#55
I have used Claude Code heavily, and I've been forced to use Gemini CLI heavily (for a particular client project).

Of all my issues with Gemini CLI (and there are many), this addresses none of them. This is a fascinating product management prioritization decision. It makes me wonder if the people who build Gemini CLI actually use Gemini CLI for real work. Because I would think that if they did, they would surely have prioritized other things.

My personal biggest issue with Gemini CLI, which is a deal breaker if I have a say in the tooling I'm using, is that if you hit a per-minute rate limit (meaning it will be resolved in a few seconds) your session is forcefully and permanently switched over to using Flash and there is nothing you can do other than manually quit and restart to get back to using Pro 2.5. The status footer line will even continue to lie to you about what model you are using. I would genuinely like to understand the use cases for which this is desirable behavior. But even IF those use cases do exist, what is the harm or difficulty in giving an option to override this behavior? These models are not interchangeable. GitHub issues have been opened for months, some even with PRs attached, with no action from Google.

For comparison, Claude Code handles this situation with a simple exponential back off until the request succeeds. That's what I want, ESPECIALLY in a CLI agent that may be running headlessly in a pipeline.

Re: Run interactive commands in Gemini CLI

#56
post #54

Earlier quoted context omitted.

LLMs in general struggles with numbers, it's easy to tell with the medium sized models that struggle with line replacement commands where it has to count, it usually takes a couple of tries to get right. I always imagined they'd have an easier time if they could start a vim instance and send search/movement/insert commands instead, not having to keep track of numbers and do calculations, but instead visually inspect…

Gotta be better than codex literally writing a python script to edit a file multiple times in a single prompt response.

Personally haven't had that happen to me, been using Codex (and lots of other agents) for months now. Anecdote, but still. I wrote up a summary of how I see the current difference between the agents right now: https://news.ycombinator.com/item?id=45680796

Re: Run interactive commands in Gemini CLI

#58
post #55

I have used Claude Code heavily, and I've been forced to use Gemini CLI heavily (for a particular client project). Of all my issues with Gemini CLI (and there are many), this addresses none of them. This is a fascinating product management prioritization decision. It makes me wonder if the people who build Gemini CLI actually use Gemini CLI for real work. Because I would think that if they did, they would surely have…

These tools aren’t made to be used, they’re made to make the CEO look less bad by showing that “Google has these things too! We’re not falling behind! Don’t tell the board to vote to fire him!”

That’s it.

It’s not a “product”, it’s a keeping-up-with-the-Joneses checklist item.

Re: Run interactive commands in Gemini CLI

#59
Does anyone know / care to speculate how they actually make this work, in terms of the LLM call loop? Specifically: does it call back to the LLM after each keystroke sending it the new state of the interactive tool, or does it batch keystrokes up? If the former, isn’t that very slow? If the latter, won’t that cause it to make mistakes with a tool it hasn’t used before?

Re: Run interactive commands in Gemini CLI

#60
post #55

I have used Claude Code heavily, and I've been forced to use Gemini CLI heavily (for a particular client project). Of all my issues with Gemini CLI (and there are many), this addresses none of them. This is a fascinating product management prioritization decision. It makes me wonder if the people who build Gemini CLI actually use Gemini CLI for real work. Because I would think that if they did, they would surely have…

Google is a proving ground for building "wow" factor products to line your resume with.

There isn't a drive to actually cater to users, it's a selfish endeavor which sometimes aligns with what users want. So the game is feature pack so you can leverage it for jumping ship or spring boarding internally.

It's the absolutely worst aspect of google, and I think its something worth dumping Sundar over, in order to get in a leader that will unify goals and get people who want to make great products, not great window dressings for themselves.

Post reply on HN