Live data from Hacker News

SeenBefore: A search engine for what you have seen before

seenbefore.com

51–60 of 97 posts

Re: SeenBefore: A search engine for what you have seen before

#53
post #15

Earlier quoted context omitted.

And would it be possible to configure it to use my own "cloud"?

Definitely something we are looking into. Major barrier is the cost for someone keeping a server running 24*7 in cloud(Micro instance on AWS is 175 dollars a year).

What about using your server to run the software, but flat files as storage, I could point it to dropbox for example?

Re: SeenBefore: A search engine for what you have seen before

#54
Yes, I have seen it before! I have build a personal search engine MindRetrieve back in 2005.

http://mindretrieve.net

Specifically I'm not comfortable for big web company to keep the history of my web activity. So I make it work completely locally. My project did not get much uptake, probably my lackluster marketing and other assorted issues are to blame. So good luck on this one!

Re: SeenBefore: A search engine for what you have seen before

#55
I think this is a great idea. I've been using Opera, which has a full-text search capability for history, but it's limited to the machine you're using it on.

I often find interesting articles on Hacker News while I'm at home that I want to find again when I'm at work. Being able to search by browser history across machines is fantastic for me.

Re: SeenBefore: A search engine for what you have seen before

#58

Earlier quoted context omitted.

Citation link: http://cond.org/sigir07.pdf [PDF] Information Re-Retrieval: Repeat Queries in Yahoo’s Logs Abstract: "This paper explores repeat search behavior through the analysis of a one-year Web query log of 114 anonymous users and a separate controlled survey of an additional 119 volunteers. Our study demonstrates that as many as 40% of all queries are re-finding queries. Re-finding appears to be an important be…

Wow, does 240 people even count as a sample. At Yahoo and Google log sizes its probably the error from cosmic rays in the data center.

If they selected them in a properly random way and had an effect close to 40% then yes, that probably does count as a sample.

Re: SeenBefore: A search engine for what you have seen before

#59
post #50

I don't know why I should give you my browser history. I'd like to keep it to myself.

I agree, sounds like a crazy thing to do when this could easily be achieved locally on my machine. Or am I missing something ?

But if a small piece of software was installed on your machine it wouldn't be "in the cloud". We know that makes everything better. Ok well not application performance ... or cost ... or usability ... but still, "the cloud".

Re: SeenBefore: A search engine for what you have seen before

#60
post #19

Interesting idea. Some quick questions: - How much data do you store per user? - How do I delete certain results? (preferably after the search comes back) - Another thing to consider is - After how much time does this just become as painful as finding that page through a search engine? - What version of the page gets stored? The latest or the one that I saw? I guess its one step better than Evernoting a page and addi…

Main issue with Evernoting and Bookmarking is that it requires an effort to say that today, this page is useful and I want to store it. Most pages I want to find are very things I did not think was useful at the time. Each unique page(unique as per the content) is stored per user. Our goal is to build the tools needed to find the information quickly, similiar to what hipmunk.com did for airline search. We have the added dimension of time to use.
Post reply on HN