Live data from Hacker News

A search engine built by the crowd

blog.archify.com

31–40 of 56 posts

Re: A search engine built by the crowd

#31
Hmm, this is an interesting algorithm, but I'd challenge its major assumption for a lot of searches. I don't have metrics, so of course my own assumptions can be challenged also, feel free to.

I think that a lot of search engine enquiries are essentially questions, with an answer that can be considered correct. Absolutely not all, but I think enough that they should certainly be considered. In that case, a site which immediately and clearly answers the question should be given, I want my answer within seconds, not minutes. If you give me the site that answers my question and that users spend the most time on, that's the exact opposite of what I want in this case.

Here's an example, I search "Population of America", your site's top result is sporcle.com, a quiz site. I bet people spend ages on there guessing the population of various countries etc, but I'd prefer to just get my answer.

That said, it appears such queries are handled outside the main algorithm by your competitors. Both Google and DuckDuckGo will give a card, at the top of the result, answering my query - I don't even have to visit a website.

I guess the tl;dr is that it's awesome that this is ambitious, but I challenge the assumption your algorithm is desirable for the majority of search results. Neither is Google's really though, so maybe this is an overly harsh criticism of something Google probably did very poorly early on too.

Re: A search engine built by the crowd

#32
post #6

I was a little surprised that they haven't included anything about spam or gamification. One core advantage of pagerank is that it's (relatively) hard to get links from high-authority websites. I can't force whithouse.gov or cnn.com to link to me. If you rely on time-spent-on-page from millions of users and treat everyone equally, how to you stop spammers from faking millions of hours spent reading their content usin…

Yes that could be a potential issue, expecially if you want to keep the data of the submitters as anonymous as possible. We are monitoring this closely, but honestly we don't have a single hammer solution against spam as always it needs my small steps. I will post some thoughts about it on our blog in the days.

-Gerald Disclosure: I am the CTO of archify/Blippex.

Re: A search engine built by the crowd

#33
post #6

I was a little surprised that they haven't included anything about spam or gamification. One core advantage of pagerank is that it's (relatively) hard to get links from high-authority websites. I can't force whithouse.gov or cnn.com to link to me. If you rely on time-spent-on-page from millions of users and treat everyone equally, how to you stop spammers from faking millions of hours spent reading their content usin…

Another problem is that if you just count time browsed then sites such as Facebook, Reddit and Kongregate get super high rankings.

Re: A search engine built by the crowd

#34

Hmm, this is an interesting algorithm, but I'd challenge its major assumption for a lot of searches. I don't have metrics, so of course my own assumptions can be challenged also, feel free to. I think that a lot of search engine enquiries are essentially questions, with an answer that can be considered correct. Absolutely not all, but I think enough that they should certainly be considered. In that case, a site which…

I think you guesses are absolutely right. We never intended to compete against other engines in the field of "Answers". I think you will always get a better result for ""Population of America" if you search for ai at DuckDuckGo or Google. But in the other hand if you search for example for "NSA" on Blippex (https://www.blippex.org/?q=NSA), we are assuming that you will get those articles about the NSA which is currently the most interesting or the most read.

-Gerald Disclosure: I am the CTO of archify/Blippex.

Re: A search engine built by the crowd

#35
post #16
post #3

This algorightm might highly impact discoverability. It gives mover visibility to already popular websites making them even popular, while not very well known websites will never be discovered because very few people spend time on them. Also I don't like the idea of having to install a plugin on my browser so that the urls I visit and how much time I spent on them is tracked, even if suposedly my identity is never tr…

Well, if anyone is able for example to implement a TOR client in javascript we would love to add it to the plugin, and the sourcecode of the plugins and Android app is on github, so no cheating there.

+1 because the source code for the plugins is on github which I didn't know

Re: A search engine built by the crowd

#36
post #33
post #6

I was a little surprised that they haven't included anything about spam or gamification. One core advantage of pagerank is that it's (relatively) hard to get links from high-authority websites. I can't force whithouse.gov or cnn.com to link to me. If you rely on time-spent-on-page from millions of users and treat everyone equally, how to you stop spammers from faking millions of hours spent reading their content usin…

Another problem is that if you just count time browsed then sites such as Facebook, Reddit and Kongregate get super high rankings.

We are counting the per unique URL, which is currently not a big advantage for the popular sites, because they are hosting so many URLs. There no domain based factor in it right now.

-Gerald Disclosure: I am the CTO of archify/Blippex

Re: A search engine built by the crowd

#37
post #12
post #11

Earlier quoted context omitted.

Thank you for the find, we will fix it!

Well that was a quick response, at least that's a positive thing :). I installed the add-on. Testing the search engine, searches seem to take forever. It keeps displaying the spinning icon in the orange square next to my query. Edit: Ah the niceness of asynchronous javascript. It returns an error (I can see it in the JS console) but the page never displays that to me. Good ol' page reloads wouldn't have done that . I…

Could you please send me a link to the header modifying addon. So I can fix that.

gb@blippex.org

Re: A search engine built by the crowd

#38

The problem with this is that it's basically asking to be manipulated.

Not only is is asking to manipulated, I'm not sure if the results are manipulated or not already. I did a search for "search engine optimization" just out of curiousity: and the number one search result was "Limo service los angeles". Looks like that company's "seo firm" inserted their links on their client site: and it's ranking for "search engine optimization". Great SEO there, NOT.

Re: A search engine built by the crowd

#39
Interesting idea, but I tried simple searches :

Facebook

gmail

news ycombinator

countries in europe wiki

Did you gather enough data already ?

All of these seraches were not successful. There was no Facebook link in the first search, no Gmail link in the second one , no news.ycombinator in the 3rd one, and the only wikipedia link I got in the last search was :

http://en.wikipedia.org/wiki/National_champions

Re: A search engine built by the crowd

#40

I think this is a great idea! However, I am worried about privacy. I also feel like this algorithm may inflate the importance of certain types of content over others. For example, just because I spend more time on a news or social media website does not mean that it has higher quality content, it just means that the content takes longer to consume. Within content categories, however, I think this could do a good job…

Thanks, that is a very interesting input. We should think about running additional semantic analysis and relate them to the time spent on sites. We are very sure that our algorithm needs a lot of fine tuning and this could be a very important part of it.

-Gerald Disclosure: I am the CTO of archify/Blippex.

Post reply on HN