Live data from Hacker News

Viewing profile — capnrefsmmat

capnrefsmmat

HN member
Joined
Mon, Apr 04, 2011, 2:46 PM UTC
HN karma
2,130
Public activity
402 items

About capnrefsmmat

https://www.refsmmat.com

Associate teaching professor of Statistics & Data Science, Carnegie Mellon University

Author of Statistics Done Wrong (https://www.statisticsdonewrong.com), the woefully complete guide to statistical errors

email: alex at refsmmat dot com

Recent public activity

  1. comment
    Comment #49250905

    > Well, what do you expect? LLMs are trained on blithering, mostly from web sites. So you get blithering out. I realize this isn't entirely serious, but I can't resist pointing out…

  2. comment
    Comment #48759708

    Being invited to conference talks around the world is a completely normal part of being an active researcher in almost any academic field, so it doesn't register as pompous to othe…

  3. comment
    Comment #48587841

    From your link, > TerraPower must still complete construction, submit an operating license application, and satisfy all applicable safety and regulatory requirements before loading…

  4. comment
    Comment #47556404

    Discussion is included in the Dangerous Goods Panel report, agenda item 4.3 (pages 39-41) and Appendix E (beginning page 89). https://www.icao.int/sites/default/files/DangerousGood…

  5. comment
    Comment #47311237

    Several reasons: 1. The post mainly reiterates a single idea (Capsicum enumerates what the process can do, seccomp provides a configurable filter) in many different ways. There is …

  6. comment
    Comment #47297135

    Probably. One common feature of LLM output is grammatical features that indicate information density, like nominalizations, longer words, participial clauses, and so on. Perhaps tr…

  7. comment
    Comment #47297121

    Thanks for the links. You may be interested in the other LLM writing style studies I've been collecting: https://www.refsmmat.com/notebooks/llm-style.html

  8. comment
    Comment #47297115

    I've heard the Kenya and Nigeria story, but has anyone backed it up with quantitative evidence that the vocabulary LLMs overuse coincides with the vocabulary that is more common in…

  9. comment
    Comment #47292658

    I work on research studying LLM writing styles, so I am going to have to steal this. I've seen plenty of lists of LLM style features, but this is the first one I noticed that menti…

  10. comment
    Comment #46911977

    No, that doesn't really work so well. A lot of the LLM style hallmarks are still present when you ask them to write in another style, so a good quantitative linguist can find them:…

  11. comment
    Comment #46823787

    The first sentence is a reference to prior research work that has found those productivity gains, not a summary of the experiment conducted in this paper.

  12. comment
    Comment #46325787

    Most of the tedious formatting requirements do not match what the final typeset article looks like. The requirements are instead theoretically to benefit peer reviewers, e.g., by h…

  13. comment
    Comment #46320571

    Outside of disciplines that use LaTeX, the ability of authors to do typesetting is pretty limited. And there are other typesetting requirements that no consumer tool makes particul…

  14. comment
    Comment #45349195

    It didn't "survey" devs. It paid them to complete real tasks while they were randomly assigned to use AI or not, and measured the actual time taken to complete the tasks vs. just t…

  15. comment
    Comment #45046228

    Sure, if you're learning to write and want lots of examples of a particular style, LLMs can generate that for you. Just don't assume that is a normal writing style, or that it matc…

  16. comment
    Comment #44599601

    I don't think the AI companies are systematically working to make their models sound more human. They're working to make them better at specific tasks, but the writing styles are, …

  17. comment
    Comment #44499251

    Probably because the article uses the Unicode right single quotation mark instead of apostrophes, due to some automated smart-quote machinery. I'll have to adjust the tagger to han…

  18. comment
    Comment #44495386

    In our studies of ChatGPT's grammatical style ( https://arxiv.org/abs/2410.16107 ), it really loves past and present participial phrases (2-5x more usage than humans). I didn't see…

  19. comment
    Comment #44190922

    OpenAI is the custodian of the user data, so they are responsible. If you wanted the court (i.e., the plaintiffs) to find specific infringing chatters, first they'd have to get the…

  20. comment
    Comment #44190892

    If the output is interpreting sources rather than just regurgitating quotes from them, you need to exert judgment to verify they support its claims. When the LLM output is about so…

  21. comment
    Comment #44190863

    Courts have always had the power to compel parties to a current case to preserve evidence. (For example, this was an issue in the Google monopoly case, since Google employees were …

  22. comment
    Comment #44164696

    For introductory problems, the kind we use to get students to understand a concept for the first time, the AI would likely (nearly) nail it on the first try. They wouldn't have to …

  23. comment
    Comment #44163981

    > Well, if you’re a novice, don’t do that. I agree, and it sounds like you're getting great results, but they're all going to do it. Ask anyone who grades their homework. Heck, it'…

  24. comment
    Comment #44163920

    I agree with all of this. But it's already very difficult to do even in a college setting -- to force students to get deliberate practice, without outsourcing their thinking to an …

  25. comment
    Comment #44163527

    Sure, you could learn about grammar, plot structure, narrative style, etc. and become a reasonable novel critic. But imagine a novice who wants to learn to do this and has access t…