Live data from Hacker News

Harper – an open-source alternative to Grammarly

writewithharper.com

151–160 of 167 posts

Re: Harper – an open-source alternative to Grammarly

#151
post #134

Fantastic work, I was so fed up with Grammarly and instantly installed this. I'm just a bit skeptical about this quote: > Harper takes advantage of decades of natural language research to analyze exactly how your words come together. But it's just a rather small collection of hard-coded rules: https://docs.rs/harper-core/latest/harper_core/linting/trait... Where did the decades of classical NLP go? No gold-standard r…

to someone who would like to study/learn that evolution, any good recs?

This skips over Bag of Words / N-Gram / TF-IDF and many other things, but paints a reasonable picture of the progression.

1. https://jalammar.github.io/illustrated-word2vec/

2. https://jalammar.github.io/visualizing-neural-machine-transl...

3. https://jalammar.github.io/illustrated-transformer/

4. https://jalammar.github.io/illustrated-bert/

5. https://jalammar.github.io/illustrated-gpt2/

And from there it's mostly work on improving optimization (both at training and inference time), training techniques (many stages), data (quality and modality), and scale.

---

There's also state space models, but don't believe they've gone mainstream yet.

https://newsletter.maartengrootendorst.com/p/a-visual-guide-...

And diffusion models - but I'm struggling to find a good resource so https://ml-gsai.github.io/LLaDA-demo/

---

All this being said- many tasks are solved very well using a linear model and tfidf. And are actually interpretable.

Re: Harper – an open-source alternative to Grammarly

#152

Fantastic work, I was so fed up with Grammarly and instantly installed this. I'm just a bit skeptical about this quote: > Harper takes advantage of decades of natural language research to analyze exactly how your words come together. But it's just a rather small collection of hard-coded rules: https://docs.rs/harper-core/latest/harper_core/linting/trait... Where did the decades of classical NLP go? No gold-standard r…

I'll admit it's something of a bold label, but there is truth in it. Before our rule engine has a chance to touch the document, we run several pre-processing steps that imbue semantic meaning to the words it reads. > LLMs have completely overshadowed ML NLP methods from 10 years ago, and they themselves replaced decades statistical NLP work, which also replaced another few decades of symbolic grammar-based NLP work.…

Here's an article you might find interesting: https://www.quantamagazine.org/when-chatgpt-broke-an-entire-...

Re: Harper – an open-source alternative to Grammarly

#153
post #134

Earlier quoted context omitted.

to someone who would like to study/learn that evolution, any good recs?

This skips over Bag of Words / N-Gram / TF-IDF and many other things, but paints a reasonable picture of the progression. 1. https://jalammar.github.io/illustrated-word2vec/ 2. https://jalammar.github.io/visualizing-neural-machine-transl... 3. https://jalammar.github.io/illustrated-transformer/ 4. https://jalammar.github.io/illustrated-bert/ 5. https://jalammar.github.io/illustrated-gpt2/ And from there it's mostly w…

This is indeed the previous generation, but it's not even that old. When I was coming out of undergrad word2vec was the brand-new thing that was eating-up the whole field.

Indeed, before that there was a lot of work on applying classical ML classifiers (Naive Bayes, Decision Trees, SVM, Logistic Regression...) and clustering algorithms (fancily referred to as unsupervised ML) to bag-of-words vectors. This was a big field, with some overlap with Information Retrieval, lending to fancier weightings and normalizations of bag-of-words vectors (TF-IDF, BM25). There was also the whole field of Topic Modeling.

Before that there was a ton of statistical NLP modeling (Markov chains and such), primarily focused around machine translation before neural-networks got good enough (like the early version of Google Translate).

And before that there were a few decades of research on grammars (starting with Chomsky), with a lot of overlap with compilers, theoretical CS (state-machines and such) and symbolic AI (lisps, logic programming, expert systems...).

I myself don't have a very clear picture of all of this. I learned some in undergrad and read a few ancient NLP books (60s - 90s) out of curiosity. I started around the time where NLP, and AI in general, had been rather stagnant for a decade or two, it was rather boring and niche, believe it or not, but was starting to be revitalized by the new wave of ML and then word2vec with DNNs.

Re: Harper – an open-source alternative to Grammarly

#155
post #94

Earlier quoted context omitted.

Fair enough, thanks for replying. I don't see the task of specifying a grammar as straightforward as you do, perhaps. I guess I just didn't understand the chain of comments. I find that clear-cut, rigid rules tend to be the least helpful ones in writing. Obviously this class of rule is also easy/easier to represent in software, so it also tends to be the source of false positives and frustration that lead me to disab…

When you do writing as a form of art, rules are meant to be bent or broken; it's useful to have the ability to explicitly write new ones and make new forms of the language legal, rather than wrestle with hallucinating LLMs. When writing for utility and communication, though, English grammar is simple and standard enough. Browsing Harper sources, https://github.com/Automattic/harper/blob/0c04291bfec25d0e93... seems to…

You get it!!

Re: Harper – an open-source alternative to Grammarly

#156
post #150

Earlier quoted context omitted.

When you do writing as a form of art, rules are meant to be bent or broken; it's useful to have the ability to explicitly write new ones and make new forms of the language legal, rather than wrestle with hallucinating LLMs. When writing for utility and communication, though, English grammar is simple and standard enough. Browsing Harper sources, https://github.com/Automattic/harper/blob/0c04291bfec25d0e93... seems to…

I'm certainly not disputing the existence of grammar nor do I think an LLM is a good way to implement/check/enforce one. And now I realise how my first comment landed. Thanks again!

Your first point would be more fitting if a language checker would need a complete, computable grammar that can be parsed and understood. That would be problematic for natural languages.

Re: Harper – an open-source alternative to Grammarly

#157
Interesting, curious to try this;

I wonder whether it will impact the performance (Firefox) and things will become noticeably slower...

Recently i noticed highlighting extensions in Firefox were slowing things down significantly, not just loading but also while scrolling up and down web pages.

Re: Harper – an open-source alternative to Grammarly

#158

Harper is decent. I've relied on Grammarly to spellcheck all my writing for a few years (dyslexia prevents me from seeing the errors even when reading it 10 times). However, I find its increasing focus on LLMs and its insistence on rewriting sentences in more verbose ways bothers me a lot. (It removes personality and makes human-written text read like AI text.) So I've tried out alternatives, and Harper is the closes…

What's wild is that OpenAI's earlier models were trained to guess the next word in a sentence. I wonder if GPT-2 would get "though" correct more often than the latest AI-assisted writing tools like Grammerly.

There are some areas where it seems like LLMs (or even SLMs) should be way more capable. For example, when I touch a word on my Kindle, I'd think Amazon would know how to pick the most relevant definition. Yet it just grabs the most common definition. For example, consider the proper definition of "toilet" in this passage: "He passed ten hours out of the twenty-four in Saville Row, either in sleeping or making his toilet."

Re: Harper – an open-source alternative to Grammarly

#159
post #26

Earlier quoted context omitted.

uh. yes? it's far from uncommon, and sometimes it's ludicrously wrong. Grammarly has been getting quite a lot of meme-content lately showing stuff like that. it is of course mostly very good at it, but it's very far from "trustworthy", and it tends to mirror mistakes you make.

Do you have any examples? The only time I noticed an LLM make a language mistake was when using a quantized model (gemma) with my native language (so much smaller training data pool).

Not GP, but I've definitely seen cutting edge LLMs make language mistakes. The most head scratching one I've seen in the past few weeks is when Gemini Pro decided to use and tags to emphasize something that was not code.
Post reply on HN