Earlier quoted context omitted.
How do you decide who is listed first? And does the ampersand symbolize something that the word and doesn't, or is that just citation style?
In biomedical research or tangential fields, author order generally follows these guidelines: First author(s): the individual(s) who organized and conducted the study. Typically there is only a single first author, but nowadays there are often two first authors. This is because the amount of research required to generate “high impact” publications simply can’t be done by a single person. Typically, the first author i…
Classifying all of the pdfs on the internet
111–117 of 117 posts
Re: Classifying all of the pdfs on the internet
#112Did some similar work with similar visualizations ~2009, on ~5.7M research articles (PDFs, private corpus) from scientific publishers Elsevier, Springer: Newton, G., A. Callahan & M. Dumontier. 2009. Semantic Journal Mapping for Search Visualization in a Large Scale Article Digital Library. Second Workshop on Very Large Digital Libraries at the European Conference on Digital Libraries (ECDL) 2009. https://lekythos.li…
Nice article, thanks for sharing. I can imagine mining all of these articles was a ton of work. I’d be curious to know how quickly the computation could be done today vs. the 13 hour 2009 benchmark :) Nowadays people would be slamming those data through UMAP!
Re: Classifying all of the pdfs on the internet
#113Earlier quoted context omitted.
How do you decide who is listed first? And does the ampersand symbolize something that the word and doesn't, or is that just citation style?
In biomedical research or tangential fields, author order generally follows these guidelines: First author(s): the individual(s) who organized and conducted the study. Typically there is only a single first author, but nowadays there are often two first authors. This is because the amount of research required to generate “high impact” publications simply can’t be done by a single person. Typically, the first author i…
I just realized that the pre-print of the paper is available at the NRC's Publications Archive: https://nrc-publications.canada.ca/eng/view/object/?id=63e86...
Re: Classifying all of the pdfs on the internet
#114Earlier quoted context omitted.
In biomedical research or tangential fields, author order generally follows these guidelines: First author(s): the individual(s) who organized and conducted the study. Typically there is only a single first author, but nowadays there are often two first authors. This is because the amount of research required to generate “high impact” publications simply can’t be done by a single person. Typically, the first author i…
Agree with this but it does not apply to all fields. Economists have a norm of alphabetizing author names unless the contributions were very unevenly distributed. That way authors can avoid squabbling over contributions.
Re: Classifying all of the pdfs on the internet
#115Hi! Author here, I wasn't expecting this to be at the top of HN, AMA
I dug through the code and it seemed like a ton of things I'm not familiar with, probably a lot of techniques I don't know rather than python ofc.
Re: Classifying all of the pdfs on the internet
#116Earlier quoted context omitted.
Agree with this but it does not apply to all fields. Economists have a norm of alphabetizing author names unless the contributions were very unevenly distributed. That way authors can avoid squabbling over contributions.
That's how education is too. That's why I asked; it wasn't in alphabetical order.
I always have wondered with these approaches, is there anything in the paper that indicates who was the “lead” author?
Also, to me, the alphabetical order approach reinforces issues with lower rank last names having various advantages. E.g, lower rank alphabetical names doing better in school [0]. Do you have any counterpoint to this that I’m missing?
[0] https://news.umich.edu/keeping-up-with-the-joneses-when-it-c...
Re: Classifying all of the pdfs on the internet
#117Earlier quoted context omitted.
It doesn't take away the torrents, no?
One of my pet peeves is the way people use words like slurp, hoover, take, vaccuum, suck up, or steal, when in reality they mean copy. I mean if Chegg manages to sell something you can get for free, then all the more power to them lol. Though we could probably do more to educate the younger generation on the magic of torrents. Ignoring angry textbook publishers, of course.
You get your copy. Then send out fake dmca noticed or buy out places. Then sell your copy after everyone else's copies aren't available anymore.
It's a very standard practice in all walks of life. You gain access to a device then improve security and fix issues so that others can't get in anymore. That's a bad person approach. The legal way is to pull up the proverbial ladder or legal loopholes behind you. Good, bad, or whatever. Tons of people and places do it.
It's far more nefarious than just so innocently copying