Viewing profile — capnrefsmmat
capnrefsmmat
HN member- Joined
- Mon, Apr 04, 2011, 2:46 PM UTC
- HN karma
- 2,130
- Public activity
- 402 items
- HN profile
- View on Hacker News ↗
About capnrefsmmat
Associate teaching professor of Statistics & Data Science, Carnegie Mellon University
Author of Statistics Done Wrong (https://www.statisticsdonewrong.com), the woefully complete guide to statistical errors
email: alex at refsmmat dot com
Recent public activity
-
comment
Comment #49250905
> Well, what do you expect? LLMs are trained on blithering, mostly from web sites. So you get blithering out. I realize this isn't entirely serious, but I can't resist pointing out…
-
comment
Comment #48759708
Being invited to conference talks around the world is a completely normal part of being an active researcher in almost any academic field, so it doesn't register as pompous to othe…
-
comment
Comment #48587841
From your link, > TerraPower must still complete construction, submit an operating license application, and satisfy all applicable safety and regulatory requirements before loading…
-
comment
Comment #47556404
Discussion is included in the Dangerous Goods Panel report, agenda item 4.3 (pages 39-41) and Appendix E (beginning page 89). https://www.icao.int/sites/default/files/DangerousGood…
-
comment
Comment #47311237
Several reasons: 1. The post mainly reiterates a single idea (Capsicum enumerates what the process can do, seccomp provides a configurable filter) in many different ways. There is …
-
comment
Comment #47297135
Probably. One common feature of LLM output is grammatical features that indicate information density, like nominalizations, longer words, participial clauses, and so on. Perhaps tr…
-
comment
Comment #47297121
Thanks for the links. You may be interested in the other LLM writing style studies I've been collecting: https://www.refsmmat.com/notebooks/llm-style.html
-
comment
Comment #47297115
I've heard the Kenya and Nigeria story, but has anyone backed it up with quantitative evidence that the vocabulary LLMs overuse coincides with the vocabulary that is more common in…
-
comment
Comment #47292658
I work on research studying LLM writing styles, so I am going to have to steal this. I've seen plenty of lists of LLM style features, but this is the first one I noticed that menti…
-
comment
Comment #46911977
No, that doesn't really work so well. A lot of the LLM style hallmarks are still present when you ask them to write in another style, so a good quantitative linguist can find them:…
-
comment
Comment #46823787
The first sentence is a reference to prior research work that has found those productivity gains, not a summary of the experiment conducted in this paper.
-
comment
Comment #46325787
Most of the tedious formatting requirements do not match what the final typeset article looks like. The requirements are instead theoretically to benefit peer reviewers, e.g., by h…
-
comment
Comment #46320571
Outside of disciplines that use LaTeX, the ability of authors to do typesetting is pretty limited. And there are other typesetting requirements that no consumer tool makes particul…
-
comment
Comment #45349195
It didn't "survey" devs. It paid them to complete real tasks while they were randomly assigned to use AI or not, and measured the actual time taken to complete the tasks vs. just t…
-
comment
Comment #45046228
Sure, if you're learning to write and want lots of examples of a particular style, LLMs can generate that for you. Just don't assume that is a normal writing style, or that it matc…
-
comment
Comment #44599601
I don't think the AI companies are systematically working to make their models sound more human. They're working to make them better at specific tasks, but the writing styles are, …
-
comment
Comment #44499251
Probably because the article uses the Unicode right single quotation mark instead of apostrophes, due to some automated smart-quote machinery. I'll have to adjust the tagger to han…
-
comment
Comment #44495386
In our studies of ChatGPT's grammatical style ( https://arxiv.org/abs/2410.16107 ), it really loves past and present participial phrases (2-5x more usage than humans). I didn't see…
-
comment
Comment #44190922
OpenAI is the custodian of the user data, so they are responsible. If you wanted the court (i.e., the plaintiffs) to find specific infringing chatters, first they'd have to get the…
-
comment
Comment #44190892
If the output is interpreting sources rather than just regurgitating quotes from them, you need to exert judgment to verify they support its claims. When the LLM output is about so…
-
comment
Comment #44190863
Courts have always had the power to compel parties to a current case to preserve evidence. (For example, this was an issue in the Google monopoly case, since Google employees were …
-
comment
Comment #44164696
For introductory problems, the kind we use to get students to understand a concept for the first time, the AI would likely (nearly) nail it on the first try. They wouldn't have to …
-
comment
Comment #44163981
> Well, if you’re a novice, don’t do that. I agree, and it sounds like you're getting great results, but they're all going to do it. Ask anyone who grades their homework. Heck, it'…
-
comment
Comment #44163920
I agree with all of this. But it's already very difficult to do even in a college setting -- to force students to get deliberate practice, without outsourcing their thinking to an …
-
comment
Comment #44163527
Sure, you could learn about grammar, plot structure, narrative style, etc. and become a reasonable novel critic. But imagine a novice who wants to learn to do this and has access t…