Show HN: First Claude Code client for Ollama local models
11–20 of 29 posts
Re: Show HN: First Claude Code client for Ollama local models
#12Does this UI work with Open Code?
Re: Show HN: First Claude Code client for Ollama local models
#13Re: Show HN: First Claude Code client for Ollama local models
#14I was trying to get Claude code to work with llama.cpp but could never figure out anything functional. It always insisted on a phone home login for first time setup. In cline I’m getting better results with glm-4.7-flash than with qwen3-coder:30b
Re: Show HN: First Claude Code client for Ollama local models
#15There are already various proxies to translate between OpenAI-style models (local or otherwise) and an Anthropic endpoint that Claude Code can talk to. Is the advantage here just one less piece of infrastructure to worry about?
siderailing here - but got one that _actually_ works? in particular i´d like to call claude-models - in openai-schema hosted by a reseller - with some proxy that offers anthropic format to my claude --- but it seems like nothing gets to fully line things up (double-translated tool names for example) reseller is abacus.ai - tried BerriAI/litellm, musistudio/claude-code-router, ziozzang/claude2openai-proxy, 1rgs/claude…
The invocation would be like this
llsed --host 0.0.0.0 --port 8080 --map_file claude_to_openai.json --server https://openrouter.ai/api
Where the json has something like { tag: ... from: ..., to: ..., params: ..., pre: ..., post: ...}
So if one call is two, you can call multiple in the pre or post or rearrange things accordingly.This sounds like the proper separation of concerns here... probably
The pre/post should probably be json-rpc that get lazy loaded.
Writing that now. Let's do this: https://github.com/day50-dev/llsed
Re: Show HN: First Claude Code client for Ollama local models
#16What hardware are you running the 30b model on? I guess it needs at least 24GB VRAM for decent inference speeds.
Re: Show HN: First Claude Code client for Ollama local models
#17What hardware are you running the 30b model on? I guess it needs at least 24GB VRAM for decent inference speeds.
Re: Show HN: First Claude Code client for Ollama local models
#18There are already various proxies to translate between OpenAI-style models (local or otherwise) and an Anthropic endpoint that Claude Code can talk to. Is the advantage here just one less piece of infrastructure to worry about?
siderailing here - but got one that _actually_ works? in particular i´d like to call claude-models - in openai-schema hosted by a reseller - with some proxy that offers anthropic format to my claude --- but it seems like nothing gets to fully line things up (double-translated tool names for example) reseller is abacus.ai - tried BerriAI/litellm, musistudio/claude-code-router, ziozzang/claude2openai-proxy, 1rgs/claude…
But I'm surprised litellm (and its wrappers) don't work for you and I wonder if there's something wrong with your provider or model. Which model were you using?
Re: Show HN: First Claude Code client for Ollama local models
#19Earlier quoted context omitted.
siderailing here - but got one that _actually_ works? in particular i´d like to call claude-models - in openai-schema hosted by a reseller - with some proxy that offers anthropic format to my claude --- but it seems like nothing gets to fully line things up (double-translated tool names for example) reseller is abacus.ai - tried BerriAI/litellm, musistudio/claude-code-router, ziozzang/claude2openai-proxy, 1rgs/claude…
What probably needs to exist is something like `llsed`. The invocation would be like this llsed --host 0.0.0.0 --port 8080 --map_file claude_to_openai.json --server https://openrouter.ai/api Where the json has something like { tag: ... from: ..., to: ..., params: ..., pre: ..., post: ...} So if one call is two, you can call multiple in the pre or post or rearrange things accordingly. This sounds like the proper separ…
Re: Show HN: First Claude Code client for Ollama local models
#20What hardware are you running the 30b model on? I guess it needs at least 24GB VRAM for decent inference speeds.
I'd like to know this, too. I'm just getting started getting my feet wet with ollama and local models using just CPU, and it's obviously terribly slow (even 24 cores, 128GB DRAM. It's hard to gauge how much GPU money I'd need to plonk down to get acceptable performance for coding workflows.