Live data from Hacker News

Indexing a billion pages

blog.mwmbl.org

11–20 of 38 posts

Re: Indexing a billion pages

#11
post #4
post #3

How does the homepage of https://mwmbl.org/ not have a single sentence explaining what it is or even an "About" link? From Github: "Mwmbl is a non-profit, ad-free, free-libre and free-lunch search engine with a focus on useability and speed."

Not everything is a product that needs to be sold.

Everything is something. It is helpful to know what this particular something is - regardless if it is sold or not.

Re: Indexing a billion pages

#14
post #2

Wuite curious. What indexing and retrieval software is this using? I couldn’t find reference to it. Does it index phrases ?

"... who crawl the web using the Firefox extension and command line script"

https://addons.mozilla.org/en-GB/firefox/addon/mwmbl-web-cra...

https://github.com/mwmbl/crawler-script

Re: Indexing a billion pages

#15
> We’ve indexed over 100 million pages

> [W]e’re crawling up to a million pages a day, as you can see on our stats page.

> Given that Mwmbl is still relatively unknown, it seems plausible that we can reach our target of crawling three billion pages a day, to refresh the entire index in one month.

I think this is supposed to read "it seems plausible that we can reach our target of crawling three million pages a day."

Re: Indexing a billion pages

#17
post #9

I thought I recalled seeing this before due to its Welsh name and (as is often the case) some are from their domain and some are from the GitHub repo; the ones with over 100 comments are https://news.ycombinator.com/item?id=37561155 https://news.ycombinator.com/item?id=29690877

Thanks! Macroexpanded:

We are entering a new era of web search - https://news.ycombinator.com/item?id=38465864 - Nov 2023 (2 comments)

Mwmbl: Free, open-source and non-profit search engine - https://news.ycombinator.com/item?id=37561155 - Sept 2023 (122 comments)

The Book of Mwmbl: a free, non-profit search engine - https://news.ycombinator.com/item?id=33828087 - Dec 2022 (8 comments)

Show HN: An open source web crawler for the Mwmbl non-profit search engine - https://news.ycombinator.com/item?id=31765015 - June 2022 (4 comments)

Show HN: I'm building a non-profit search engine - https://news.ycombinator.com/item?id=29690877 - Dec 2021 (199 comments)

Re: Indexing a billion pages

#19
post #2

Wuite curious. What indexing and retrieval software is this using? I couldn’t find reference to it. Does it index phrases ?

I believe this is closer to the thing you were asking about, and the simple answer appears to be "a home grown one in python" https://github.com/mwmbl/mwmbl/blob/e544d45c374c13cdc1a5048d...

Re: Indexing a billion pages

#20
post #15

> We’ve indexed over 100 million pages > [W]e’re crawling up to a million pages a day, as you can see on our stats page. > Given that Mwmbl is still relatively unknown, it seems plausible that we can reach our target of crawling three billion pages a day, to refresh the entire index in one month. I think this is supposed to read "it seems plausible that we can reach our target of crawling three million pages a day."

> to refresh the entire index [of 1 billion] in one month.

Seems you are correct.

Post reply on HN