Live data from Hacker News

Show HN: Semble – Code search for agents that uses 98% fewer tokens than grep

github.com

51–60 of 187 posts

Re: Show HN: Semble – Code search for agents that uses 98% fewer tokens than grep

#55
post #46

Nice, this sounds great. I want to mention a related issue here, which is that on small codebases, Claude spends a lot of time looking for stuff when it could have just dumped the whole codebase into the context in one go and used very little tokens. I found a nice workaround which is that you can just dump the whole directory into context, as a startup hook. So then Claude skips the "fumble around blindly in the dar…

Maybe aider? https://aider.chat/2023/10/22/repomap.html

Re: Show HN: Semble – Code search for agents that uses 98% fewer tokens than grep

#56
Interesting. I too have been working in this space, though I took a different approach. Rather than building an index, I worked on making a "smarter grep" by offering search over codebases (and any text content really) with ranking and some structural awareness of the code. Most of my time was spend dealing with performance, and as a result it runs extremely quickly.

I will have to add this as a comparison to https://github.com/boyter/cs and see what my LLMs prefer for the sort of questions I ask. It too ships with MCP, but does NOT build an index for its search. I am very curious to see how it would rank seeing as it does not do basic BM25 but a code semantic variant of it.

This seems to work better for the "how does auth work" style of queries, while cs does "authenticate --only-declarations" and then weighs results based on content of the files, IE where matches are, in code, comments and the overall complexity of the file.

Have starred and will be watching.

Re: Show HN: Semble – Code search for agents that uses 98% fewer tokens than grep

#58
post #18

very curious to give it a spin but why write a cli in python? would surely be faster and more portable with go or rust?

Perhaps Python is their main language (they seem to be ML peeps, which would make that most likely), which means it's easier for them to do manual reviews even if they're using AI for implementing, etc.

Yes, this is the main reason. We've released some rust stuff in the past, but Python is our main language

Re: Show HN: Semble – Code search for agents that uses 98% fewer tokens than grep

#60
What I have personally observed with such tools is that they make the AI's dumb, similar to how it makes coders dumb when relying more on AI tools.

These agentic AI's are already smart enough to figure out a highly optimized path to code exploration or search. But, with these tools, they just go very aggressive, partly because the search results from these tools almost in 100% of the cases do not furnish full details, but, just the pointers.

To confirm this behaviour, I did a small test run. This is in no way conclusive, but, the results do align with what I been observing:

---

Task: trace full ingestion and search paths in some okayish complex project. Harness is Pi.

1. With "codebase-memory-mcp": 85k/4.4k (input/output tokens).

2. With my own regular setup: 67k/3.2k.

3. Without any of these: 80k/3.2k.

As we see, such a tool made it worse (not by much, but, still). The outputs were same in quality and informational content.

---

Now, what my "regular setup" mentioned above is?:

Just one line in AGENTS.md and CLAUDE.md: "Start by reading PROJECT.md" .

And PROJECT.md contains just following: 2-3 line description of the project, all relevant files and their one-line description, any nuiances, and finally, ends with this line:

    ## To LLM
    Update this file if the changes you have done are worth updating here. The intent of this file is to give you a rough idea of the project, from where you can explore further, if needed.
Post reply on HN