Viewing profile — gbmatt
gbmatt
HN member- Joined
- Mon, Feb 15, 2021, 9:22 PM UTC
- HN karma
- 623
- Public activity
- 44 items
- HN profile
- View on Hacker News ↗
About gbmatt
No profile information was provided.
Recent public activity
-
comment
Comment #37862850
Only Big Tech (Microsoft,Google,Facebook) can crawl the web at scale because they own the major content companies and they severly throttle the competition's crawlers, and sometime…
-
comment
Comment #34296707
We are just robots in a human simulator, reliving our creation.
-
comment
Comment #34135154
Q: how might an AI algorithm be modified in order to return citations with its response? A: There are several ways in which an AI algorithm could be modified to return citations wi…
-
comment
Comment #29818432
I just posted this same comment on the ddg story, but I'm going to post it here as well. Google forced my search engine (gigablast) basically out of business. I had ixquick.com as …
-
comment
Comment #29818392
Yeah, Google forced my search engine basically out of business. I had ixquick.com as a big client at one time; I was providing them with search results from my custom web search en…
-
comment
Comment #29605523
everyone needs equal access to public data. right now only big tech can download the many web pages (without thottling or being ip banned) on linkedin (microsoft), youtube (google)…
- story
-
comment
Comment #29426992
the complexity of the search algorithm has also increased substantially since 2005 And, in 2005, a billion page index was pretty big. Now it's closer to 100 billion.
-
comment
Comment #29426940
thanks ben, you are too kind.
-
comment
Comment #29425980
the javascript is run by your browser, so you can fully audit it.
-
comment
Comment #29425925
hey thanks for the recognition, people. :) finally, all my problems are solved. this comment is here for hacker news karma points.
-
comment
Comment #29425870
I'd argue that a level playing field and more competition in the search space is a good thing.
-
comment
Comment #29425858
100% custom.
-
comment
Comment #29425805
that's tripped out. where did you hear about that?
-
comment
Comment #29425676
I'll admit I had not been working on the quality of single term queries as much as I should have lately. However, especially for such simple queries, having a database of link text…
-
comment
Comment #29425587
Cloudflare is not the only gatekeeper, too. Keep that in mind. There's many others and, as an upstart search engine operator, it's quite overwhelming to have to deal with them all.…
-
comment
Comment #29425535
It's not quite that easy. Have you ever tried it? See my post below. Basically, yes, I've done it, but i had to go through a lot and was lucky enough to even get them to listen to …
-
comment
Comment #29419068
So Brave is still dependent on Google and Bing it seems. Also is this Brave's CEO: https://www.bbc.com/news/technology-26868536 https://www.nytimes.com/2020/12/22/business/brave-br…
-
comment
Comment #29418728
there's some stuff here : https://github.com/gigablast/open-source-search-engine
-
comment
Comment #29418591
brave 'falls back' to bing. which in my experience is most of the time. in fact, out of all the queries i did a while back, they all seemed to come directly from bing. is there a w…
-
comment
Comment #29418205
yes, large proxy networks are potential solutions. but they cost money, and you are playing a cat and mouse game with turing tests, and some sites require a login. furthermore, peo…
-
comment
Comment #29418024
it's both storage and computational. they go hand in hand.
- comment
-
comment
Comment #29417835
both ddg and brave are bing (microsoft) in disguise.
-
comment
Comment #29417777
it's continually spidering. just not at a high rate. actually, back in the day i had real time updates while google was doing the 'google dance'. that caused quite a stir in the we…