Live data from Hacker News

Google rewrites many page titles

zyppy.com

151–160 of 273 posts

Re: Google rewrites many page titles

#151
post #95

Speaking of rewriting titles... I noticed that HN reworte a self post title a few days ago. [0] Why is HN editing self post titles? [0] https://news.ycombinator.com/item?id=30053890

Mods frequently rewrite submitted titles, either cos it's not the same title used in the article, or because there's a better wording for the HN crowd ¯\_(ツ)_/¯ Check dang's (HN mod) comments: https://news.ycombinator.com/threads?id=dang

This title edit was not for an article, which is what confused me.

A user submit a self post, a rant basically which got popular, and the title was edited hours later.

Re: Google rewrites many page titles

#152

Earlier quoted context omitted.

There’s a grain of truth here, in a bit of a tangent: librarians classify all books using a system like Dewey Decimal or Library of Congress Classification. While not adjusting titles, librarians do have some influence on how a book is classified and thus filed/organised within the library. Check out the wiki article on Dewey[1] for the various options for homosexuality, which has numbers for it including under areas…

The Dewey Decimal classification system is ridiculously flawed, and no self-respecting library uses it these days (unless it always has, and hasn't got around to re-organising). Even my school's little one-room library didn't, something I found annoying at first, but came to appreciate.

"Libraries in the United States generally use either the Library of Congress Classification System (LC) or the Dewey Decimal Classification System to organize their books. Most academic libraries use LC, and most public libraries and K-12 school libraries use Dewey." [1]

[1] https://www.usg.edu/galileo/skills/unit03/libraries03_04.pht...

Re: Google rewrites many page titles

#153

Earlier quoted context omitted.

There’s a grain of truth here, in a bit of a tangent: librarians classify all books using a system like Dewey Decimal or Library of Congress Classification. While not adjusting titles, librarians do have some influence on how a book is classified and thus filed/organised within the library. Check out the wiki article on Dewey[1] for the various options for homosexuality, which has numbers for it including under areas…

The Dewey Decimal classification system is ridiculously flawed, and no self-respecting library uses it these days (unless it always has, and hasn't got around to re-organising). Even my school's little one-room library didn't, something I found annoying at first, but came to appreciate.

Disagree that it’s ridiculously flawed. It has issues like any system, but it still works well the majority of the time.

> no self-respecting library uses it these days (unless it always has, and hasn't got around to re-organising).

The vast majority of library systems have been around long enough where Dewey was the defacto choice (or LCC). Just checked a few like the British Library, the French National Library, and all the other libraries I’ve looked up now in London, all Dewey.

Re: Google rewrites many page titles

#154

I totally get this. Back in the day when I was a kid, we went to the local library and read about the world. When the librarians weren’t serving me by “checking out books” to me, they were busily putting new and improved titles on the books in receiving. /s Seriously. Google is starting to feel less like the librarian of the net (we index the world) and more like the Truman show: we craft your reality.

It’s the ads. The way Brin and Page phrased it in their 1998 paper, they considered ad-oriented search engines to be lower quality. They were going to be more academic. They thought that there was lots of user data to mine in search…for academic purposes. Then innovation #2 at the actual startup was the ad auctions and that was the beginning of the end, all the way back at the beginning. I’ve recently read a lot abou…

> we craft your reality

As mentioned above. It's also the AI.

Ads are not the fundamental problem. The fundamental problem is tracking. More on that here and about search: https://www.mojeek.com/support/ads/

Re: Google rewrites many page titles

#156

I totally get this. Back in the day when I was a kid, we went to the local library and read about the world. When the librarians weren’t serving me by “checking out books” to me, they were busily putting new and improved titles on the books in receiving. /s Seriously. Google is starting to feel less like the librarian of the net (we index the world) and more like the Truman show: we craft your reality.

There’s a grain of truth here, in a bit of a tangent: librarians classify all books using a system like Dewey Decimal or Library of Congress Classification. While not adjusting titles, librarians do have some influence on how a book is classified and thus filed/organised within the library. Check out the wiki article on Dewey[1] for the various options for homosexuality, which has numbers for it including under areas…

Those numbers were added as a consequence of the books that needed to be classified in the 1930s, and now that there are books that don't belong in the category there are new numbers.

Re: Google rewrites many page titles

#157

I totally get this. Back in the day when I was a kid, we went to the local library and read about the world. When the librarians weren’t serving me by “checking out books” to me, they were busily putting new and improved titles on the books in receiving. /s Seriously. Google is starting to feel less like the librarian of the net (we index the world) and more like the Truman show: we craft your reality.

>...they were busily putting new and improved titles on the books in receiving.

The book titles are unchanged (when you visit the site) - this is just the Librarians adding synonyms and/or simplifying titles in their catalog so that it is "Dr. Strangelove" rather than "Dr. Strangelove or: How I Learned to Stop Worrying and Love the Bomb".

Re: Google rewrites many page titles

#158

Thinly veiled content marketing for Zyppy, complete with CTA at the bottom, and mentions of themselves throughout, including: "Fortunately, here at Zyppy, we have a large database of titles thanks to our title tag analysis tool. Armed with this data, we set out to determine how often Google rewrites titles and the scenarios which trigger this behavior." Furthermore, "HTLM" instead of "HTML"? Needs proofreading. Lol.

Your point about proofreading seems fair.

Pretty much any company producing blog content is engaging in content marketing though. I’m not sure I understand the criticism. Perhaps this particular piece was overly self-promotional?

Sure, there’s a balance to be struck, but I thought the article had some decent takeaways.

Re: Google rewrites many page titles

#159

Earlier quoted context omitted.

Should I ask Walmart to kindly start relabeling products on their shelves because what’s on the tin is rarely as good for me as what the maker purports? Maybe that’s what we need. An FDA metadata label for every website served, kinda like the fav-icon, but useful. - Readable word count (protein) - Ad count (fats) - Image count (carbs) - Embedded script size (the list of nasty sounding chemicals it contains) - Average…

This might actually be the killer app for AR. Reviews of products as you look at them on the shelf.

Rather than showing the actual reviews, just lower the color saturation for lower reviewed products. So high reviewed products would pop in a sea of gray scaled items.

Sounds like something out of Black Mirror, but could be interesting.

Re: Google rewrites many page titles

#160
post #20

Earlier quoted context omitted.

HTML5 has been around for long enough that we should be able to punish sites that use completely bonkers markup at this point right? Since Google effectively has historical archives of the internet they could pretty trivially grandfather in legitimately old content (things they tracked before some date) and just start down-ranking sites that continue to misbehave with markdown but skate by with browsers running in co…

That would be a massive loss, though. A lot of content isn't in HTML5, and a lot of that pre-HTML5 content is precious and valuable. Google has sadly already tossed a lot of that by the wayside, since it often isn't served with HTTPS. I think something like 80% of the sites my crawler is aware of serve pages over plain HTTP. In general, attempts at shaping the web through search engine indexing requirements seems to…

I think it'd be a pretty good to let in historical stuff on grace - and just start penalizing new content. Google absolutely has the tools to do this the right way and the internet archive could allow most other folks to accomplish the same thing.

Enabling HTTPS is easy on most platforms. Folks that have rolled their own platform or got unlucky and are using a CMS that fell out of favor do tend to get screwed over by this - but I think its fair to de-prioritize content that fails to adhere to good practices. The HTTP vs HTTPS debate in particular can be a real security concern - with tags its more about paying down the tech debt in our browser technology.

Post reply on HN