Live data from Hacker News

Goodreads Is Broken

onezero.medium.com

261–270 of 292 posts

Re: Goodreads Is Broken

#261
post #258
post #257

Earlier quoted context omitted.

You're right, I misspelled the title! Still, an errant s should not foil the search!

Agreed. Its fixed now [1] as I added in another fuzzy search. (I was using fzf which missed some small things and I've now added Levenshtein as well to compensate). [1] https://nowwhatdoiread.com/?q=The+Glass+Beads+Game

Great! Looking through the suggestions I actually really like them, good job!

Re: Goodreads Is Broken

#262
post #94

Earlier quoted context omitted.

I face this problem with Netflix, Spotify and Youtube too. These algorithmic recommendations just need one improper dataset to throw everything out of the window. For weeks, my spotify is overloaded with instrumental songs. Youtube keeps on repeating the same stuff. Netflix believes that the only thing I watch is science fiction. I'm dying for human curation.

I call it "collapsing into the mean" - all recommendation engines (as the currently exist) will eventually corral you into the most vanilla, mass marketed set of recommendations and then fail miserably when any conflicting data is presented. We've effectively turned the web into cable TV circa 1990 - a finite set of junk food level entertainment sources that we voted for because they were "eh, good enough" and easy t…

> And the rules are unclear.

There are no rules. They just record your preferences and might retrain their black box algo in the future.

The black box decides those rules

Re: Goodreads Is Broken

#263
post #77

> The recommendations suck, the lists suck — it’s like, 100 lists telling me to read The Handmaid’s Tale and Harry Potter. I had the same experience with GR and also Amazon.com which constantly peddles the vampire romance books when I am looking for recommendations for horror/fantasy. Both Amazon and GR strategy make sense because best-selling books sell the best, so they should recommend them to increase profits. Ho…

How can you recommend books based on only one observation ? Wouldn't it make more sense to ask many books the user liked in order to make a better "user hyperplane" ?

Re: Goodreads Is Broken

#264
post #199

Earlier quoted context omitted.

My preference is to rely on the judgement of reviewers whose taste kind of matches mine. I'm into SF and Fantasy. In the past, I would read the reviews featured in Locus magazine by the various reviewers. Nowadays, I occasionally read Locus but also reviews from other places like tor com, the Magazine of Fantasy and Science Fiction and Interzone magazine.

> My preference is to rely on the judgement of reviewers whose taste kind of matches mine. I wish there were a website where I could put in all of my favorite video games, books, movies, anime, etc. And it would recommend me things based on what people with similar tastes liked. Then, I could try out the recommendation and then either like or dislike it. I would slowly acquire a recommendation network of people like…

RateYourMusic can do exactly this; find some users with similar tastes (you can see who rates an album as what score), and add them as friends (privately, i.e they won't be notified when you add them). Then you can make lists of how your friends collectively rated music. You can also find site-wide charts for specific genres.

Re: Goodreads Is Broken

#265
post #216

Earlier quoted context omitted.

I read Neuromancer in the early 90’s and thought it was very cliched at the time. Granted that’s ~10 years after publication, but it barrows heavily from earlier works. I suspect people like it for the same reason they liked their first Anime, it’s an unusual style that seems very original unless you have been reading other stuff written in the same vein by say Philip K. Dick.

I'm curious as to what aspects seemed cliched. What earlier works are you talking about? I thought I'd read them all...

It’s been a while but ...

The focus on cyborgs a year after The Six Million Dollar Man TV show kind if shows how much a product of the times it was. But, that’s the surface.

The way it portrayed both hacking and brain machine interfaces was wildly off base and basically copied from other science fiction. Virtual reality for example goes back to 1933. Main character being a druggy is fairly common in that time period, again not a big deal. As is copping tone from other works etc.

All the big stuff is forgivable, but he also copies little things like replacing liver and kidneys to better filter the blood and thus prevent someone from getting high / poisoned etc. Sounds good, but blood takes around a minute to circulate and most of it does not hit either on the way. It might reduce how long someone stays high or improve their chances when poisoned, but it’s really not enough to prevent it.

Granted I prefer hard sci-fi, but the novel’s focus is really on style over science fiction. It’s IMO somewhere between space opera and fantasy.

Re: Goodreads Is Broken

#266

Earlier quoted context omitted.

But it works? If you measure your tuning based on which rank an item is when a user clicks you might get feedback that would obviously help find more examples or create some kind of self-learning system, but if you only care about say, boosting common best sellers periodically, well, isn’t that good for (almost) everyone? ;-) My suggestion for GoodReads would be to rank exact title matches higher than partial or reor…

> “But it works?” This is the big red flag, when non-specialists hacking on Solr boosts are claiming something works because of a few qualitative test cases. “It works” is a statement that only applies after you’ve done qualitative and quantitative goodness of fit testing. You wouldn’t have a random IT employee make a stock-trading algorithm and then test it on a month of data and call it a success. For a search solu…

I wouldn’t do this with stock trading because I’m risking everything. But I wouldn’t call the people adjusting search engine parameters completely untrained either, simply not using a methodology that tests their changes against every query. I’ve found that for libraries, at least, the search engines folks are used to are simply SO BAD, so unoptimized, that a little hand tweaking and prioritizing of exact title matches will go a long way. And you’re confusing manual testing with no testing—they would very carefully watch for counter examples with a list of known good titles to search for and get back an expected set of results and were known to rollback changes when they had unexpected consequences. Effectively they have the risk appetite to test in production because the cost to end users is minimal, and the assumption when search doesn’t work is, “oh, they must not have that book” or “oh, they need to fix this particular search” and not “oh, they broke search completely and must be fired” (no one says that last one)

Re: Goodreads Is Broken

#267
post #189
post #77

> The recommendations suck, the lists suck — it’s like, 100 lists telling me to read The Handmaid’s Tale and Harry Potter. I had the same experience with GR and also Amazon.com which constantly peddles the vampire romance books when I am looking for recommendations for horror/fantasy. Both Amazon and GR strategy make sense because best-selling books sell the best, so they should recommend them to increase profits. Ho…

As far as I can tell, all recommendations lists fail compared to freely-available experts. Like, if you like literary fiction, go through the Pulitzer prize for fiction, and just read. (I'm not even half way through that list, but everything I've read on it has been really, really good.) - there's all sorts of awards for smaller niches... the nebula, the Hugo, etc... (Actually, that's a question. What is the award fo…

> I... personally don't understand why people even try to automate making better recommendation engines

There is a long tail of long tails: niches within niches within niches. Some of these don't have a single proven trustworthy reviewer, let alone enough that the rough edges of their opinions get sanded off by aggregation. For these ultra-niche interests, it'd still be nice to have a guide. ML can do that.

Re: Goodreads Is Broken

#268
post #265

Earlier quoted context omitted.

I'm curious as to what aspects seemed cliched. What earlier works are you talking about? I thought I'd read them all...

It’s been a while but ... The focus on cyborgs a year after The Six Million Dollar Man TV show kind if shows how much a product of the times it was. But, that’s the surface. The way it portrayed both hacking and brain machine interfaces was wildly off base and basically copied from other science fiction. Virtual reality for example goes back to 1933. Main character being a druggy is fairly common in that time period,…

Any recommendations for good hard sci-fi in the last generation or so? I'm asking because your comment strongly suggests I'd like what you like. I'm one of the very few who think that "Science Fiction and Fantasy" as a genre makes about as much sense as "Math Textbooks and Romance Novels".

I promise not to blame anyone for a recommendation that's flawed. They're all flawed. Anything where the story is based on the implications of known (well, currently accepted) science without any bogus magic is as hard as trying to figure out what will really happen in a large software project that hasn't begun yet. But what have you liked despite its flaws?

Re: Goodreads Is Broken

#269

That first graphic with failed title search results reminded me of a project I did working on a library catalog search using Solr. We tuned the relevancy ranking to work for exactly those sorts of searches, and put the things the OP was looking for on top (or at least under other books with exact same title). His examples look just like some of our QA searches for our relevancy ranking (in addition to standard tf/idf…

I build search engines for a living. While I appreciate the hacker spirit of your project with Solr, I also see this as a huge problem that leads to bad search experiences. Tuning boosts in Solr is not even close to a reasonable way to solve problems like this. Arguably not even for an underfunded library, but certainly not for a high web traffic consumer website. For one, you need disciplined acceptance criteria in…

I hear you, but I firmly believe the search we did there works far better than Goodreads as outlined in the OP. And could be proven as such with the kind of formalized evaluation you suggest (I am no longer there though).

I agree that underfunded "DIY" enterprise software projects that are not properly/professionally managed/implemented with the proper expertise are a problem, in academic libraries and elsewhere, for search projects and other things.

I still don't see the problem of setting up solr indexed fields and boosts to ensure that "match as phrase" is boosted higher than non-adjacent matches (a feature built into Solr), and "match _complete_ title" is boosted highest of all. This is what the Goodreads examples failed on. It is pretty simple to set up, and I don't see much risk of this causing problems or being worse than not doing it, and would solve those horrible Goodreads results specifically.

I understand since it's what you specialize in, you see the risk of "look at what us non-specialists can set up" sending someone away from... actually I'm not sure what, hiring someone like you? (Which if I were in charge of an academic library budget, which I'm not, I'd be wiling to consider -- don't get me wrong it's not a terrible idea!). In reality, I think what it steers people away from is... a relevancy search like Goodreads has. (Goodreads is _not_ a "plucky little project", and apparently they think their horrible search is good enough! It is not! And I do think they could make it a LOT better pretty easily without having to spend millions on it).

You seem to suggest that products like Solr or ElasticSearch should not be used/configured except by people as speicalized as you backed by relatively expensive search evaluation programs. While I'm sure that would result in better searches everywhere (not being sarcastic, I fully accept that), I think it's unrealistic. If you convince people their only choices are Solr/ElasticSearch/postgres-full-text out of the box using only one indexed field with no configuration for relevance tuning; no search at all; or hiring you or an equivalently expensive search program internal or external -- you're not going to get the expensive search program you want, and you're definitely not going to get "okay, we just wont' have a search at all then", you're going to get people not touching the configuration at all, and ending up with Goodreads search.

Your search really doesn't have to be as bad as Goodreads is, without having to invest in the kind of program and expertise you are suggesting, I really believe that and stand by it. If you invest in what you are suggesting, certainly it will be even better.

(PS: If you have disciplined acceptance criteria combined with qualitative feedback from experts etc -- aren't you still gonna end up tuning your Solr configuration to achieve improvement on those evaluations? I'm confused by your suggestion that turning Solr configuration with boosts etc is not the right tool. Or are you suggesting Solr is the wrong tool for... search?)

Re: Goodreads Is Broken

#270

Earlier quoted context omitted.

I build search engines for a living. While I appreciate the hacker spirit of your project with Solr, I also see this as a huge problem that leads to bad search experiences. Tuning boosts in Solr is not even close to a reasonable way to solve problems like this. Arguably not even for an underfunded library, but certainly not for a high web traffic consumer website. For one, you need disciplined acceptance criteria in…

> When this happens, executives and managers just want to punt. They want the “nobody ever got fired for buying IBM” equivalent for search, and that’s how you end up with Confluence still only supporting exact title matching and having no ability for actual content relevancy search This whole post was a wild ride. But out of curiosity, have you considered walking up to Atlassian and saying “pay me $1MM a year and I’l…

Obviously Atlassian would have to think a high-quality search would result in $1MM of additional profits for them. Judging by what Atlassian is actually doing... they apparently don't think a search better than they've got is in fact necessary for their profits or a good investment that will be returned in profits, right?
Post reply on HN