Viewing profile — parthsareen
parthsareen
HN member- Joined
- Thu, May 27, 2021, 8:32 PM UTC
- HN karma
- 42
- Public activity
- 25 items
- HN profile
- View on Hacker News ↗
About parthsareen
No profile information was provided.
Recent public activity
-
comment
Comment #46894514
Also recently added ollama launch claude if you want to connect to cloud models from there :)
-
comment
Comment #46773970
Hey! One of the maintainers of Ollama. 8GB of VRAM is a bit tight for coding agents since their prompts are quite large. You could try playing with qwen3 and at least 16k context l…
- story
-
comment
Comment #46351544
How much ram are you running with? Qwen3 and gpt-oss:20b punch a good bit above their weight. Personally use it for small agents.
-
comment
Comment #46351516
You're welcome to go through the source: https://github.com/ollama/ollama/
-
comment
Comment #46351508
Desktop app is open-source now.
- story
-
comment
Comment #45378978
Since we shipped web search with gpt-oss in the Ollama app I've personally been using that a lot more especially for research heavy tasks that I can shoot off. Plus with a 5090 or …
-
comment
Comment #45378793
Hi - author of the post. Yes it does! The "build a search agent" example can be used with a local model. I'd recommend trying qwen3 or gpt-oss
-
comment
Comment #45378778
Hey! Author of the blogpost and I also work on Ollama's tool calling. There has been a big push on tool calling over the last year to improve the parsing. What's the issues you're …
-
comment
Comment #45369806
That's a great idea. Going to try this next :)
-
comment
Comment #45369805
Hey! I'm the author of the post. We haven't optimized sampling yet so it's running linearly on the CPU. A lot of SOTA work either does this while the model is running the forward p…
-
comment
Comment #45351322
Thank you! Maybe not "perfect" but near-perfect is something we can expect. Models like the Osmosis structure which just structure data inspired some of that thinking ( https://oll…
-
comment
Comment #45350622
Thanks for posting! Didn't expect this to get picked up – it was a bit of a draft haha. Happy to answer questions around structured outputs :)
-
comment
Comment #42371223
Yes! I have checked guidance out, as well as a few others. Planning to refactor sampling in the near future which would include improving using grammars for sampling as well. Thank…
-
comment
Comment #42351440
The constraints will always be met. It’s the data inside that might be inaccurate. YMMV with smaller models in that sense.
-
comment
Comment #42351243
Hey! Author of the blog here. The current implementation uses llama.cpp GBNF which has allowed for a quick implementation. The biggest value-add at this time was getting the featur…
-
comment
Comment #42351214
Hey! Author of the post and one of the maintainers here. I agree - we (maintainers) got to this late and in general want to encourage more contributions. Hoping to be more on top o…
-
comment
Comment #42347160
This looks really useful. Thank you!
-
comment
Comment #42346713
I authored the blog with some other contributors and worked on the feature (PR: https://github.com/ollama/ollama/pull/7900 ). The current implementation uses llama.cpp GBNF grammar…
-
comment
Comment #42346535
We’ve been keeping a close eye on this as well as research is coming out. We’re looking into improving sampling as a whole on both speed and accuracy. Hopefully with those changes …
-
comment
Comment #42346526
Hey! Author of the blog post here. Yes you should be able to use any model. Your mileage may vary with the smaller models but asking them to “return x in json” tends to help with a…
- comment
- story
-
comment
Comment #27308074
The first few of these are my fav: https://dive.sh/thread/81vD2RhjxF