Live data from Hacker News

Unix Admin Horror Story Summary (1992)

www-uxsup.csx.cam.ac.uk

41–50 of 94 posts

Re: Unix Admin Horror Story Summary (1992)

#41
I found my favorite story buried in the middle of this, from 1986. It's a classic on par with The Story of Mel, a Real Programmer. Reproduced here for your reading pleasure:

  Have you ever left your terminal logged in, only to find when you came
  back to it that a (supposed) friend had typed "rm -rf ~/*" and was
  hovering over the keyboard with threats along the lines of "lend me a
  fiver 'til Thursday, or I hit return"?  Undoubtedly the person in
  question would not have had the nerve to inflict such a trauma upon
  you, and was doing it in jest.  So you've probably never experienced the
  worst of such disasters....
  
  It was a quiet Wednesday afternoon.  Wednesday, 1st October, 15:15
  BST, to be precise, when Peter, an office-mate of mine, leaned away
  from his terminal and said to me, "Mario, I'm having a little trouble
  sending mail."  Knowing that msg was capable of confusing even the
  most capable of people, I sauntered over to his terminal to see what
  was wrong.  A strange error message of the form (I forget the exact
  details) "cannot access /foo/bar for userid 147" had been issued by
  msg.  My first thought was "Who's userid 147?; the sender of the
  message, the destination, or what?"  So I leant over to another
  terminal, already logged in, and typed
          grep 147 /etc/passwd
  only to receive the response
          /etc/passwd: No such file or directory.
  
  Instantly, I guessed that something was amiss.  This was confirmed
  when in response to
          ls /etc
  I got
          ls: not found.
  
  I suggested to Peter that it would be a good idea not to try anything
  for a while, and went off to find our system manager.
  
  When I arrived at his office, his door was ajar, and within ten
  seconds I realised what the problem was.  James, our manager, was
  sat down, head in hands, hands between knees, as one whose world has
  just come to an end.  Our newly-appointed system programmer, Neil, was
  beside him, gazing listlessly at the screen of his terminal.  And at
  the top of the screen I spied the following lines:
          # cd
          # rm -rf *
  
  Oh, shit, I thought.  That would just about explain it.
  
  I can't remember what happened in the succeeding minutes; my memory is
  just a blur.  I do remember trying ls (again), ps, who and maybe a few
  other commands beside, all to no avail.  The next thing I remember was
  being at my terminal again (a multi-window graphics terminal), and
  typing
          cd /
          echo \*
  I owe a debt of thanks to David Korn for making echo a built-in of his
  shell; needless to say, /bin, together with /bin/echo, had been
  deleted.  What transpired in the next few minutes was that /dev, /etc
  and /lib had also gone in their entirety; fortunately Neil had
  interrupted rm while it was somewhere down below /news, and /tmp, /usr
  and /users were all untouched.
  
  Meanwhile James had made for our tape cupboard and had retrieved what
  claimed to be a dump tape of the root filesystem, taken four weeks
  earlier.  The pressing question was, "How do we recover the contents
  of the tape?".  Not only had we lost /etc/restore, but all of the
  device entries for the tape deck had vanished.  And where does mknod
  live?  You guessed it, /etc.  How about recovery across Ethernet of
  any of this from another VAX?  Well, /bin/tar had gone, and
  thoughtfully the Berkeley people had put rcp in /bin in the 4.3
  distribution.  What's more, none of the Ether stuff wanted to know
  without /etc/hosts at least.  We found a version of cpio in
  /usr/local, but that was unlikely to do us any good without a tape
  deck.
  
  Alternatively, we could get the boot tape out and rebuild the root
  filesystem, but neither James nor Neil had done that before, and we
  weren't sure that the first thing to happen would be that the whole
  disk would be re-formatted, losing all our user files.  (We take dumps
  of the user files every Thursday; by Murphy's Law this had to happen
  on a Wednesday).  Another solution might be to borrow a disk from
  another VAX, boot off that, and tidy up later, but that would have
  entailed calling the DEC engineer out, at the very least.  We had a
  number of users in the final throes of writing up PhD theses and the
  loss of a maybe a weeks' work (not to mention the machine down time)
  was unthinkable.
  
  So, what to do?  The next idea was to write a program to make a device
  descriptor for the tape deck, but we all know where cc, as and ld
  live.  Or maybe make skeletal entries for /etc/passwd, /etc/hosts and
  so on, so that /usr/bin/ftp would work.  By sheer luck, I had a
  gnuemacs still running in one of my windows, which we could use to
  create passwd, etc., but the first step was to create a directory to
  put them in.  Of course /bin/mkdir had gone, and so had /bin/mv, so we
  couldn't rename /tmp to /etc.  However, this looked like a reasonable
  line of attack.
  
  By now we had been joined by Alasdair, our resident UNIX guru, and as
  luck would have it, someone who knows VAX assembler.  So our plan
  became this: write a program in assembler which would either rename
  /tmp to /etc, or make /etc, assemble it on another VAX, uuencode it,
  type in the uuencoded file using my gnu, uudecode it (some bright
  spark had thought to put uudecode in /usr/bin), run it, and hey
  presto, it would all be plain sailing from there.  By yet another
  miracle of good fortune, the terminal from which the damage had been
  done was still su'd to root (su is in /bin, remember?), so at least we
  stood a chance of all this working.
  
  Off we set on our merry way, and within only an hour we had managed to
  concoct the dozen or so lines of assembler to create /etc.  The
  stripped binary was only 76 bytes long, so we converted it to hex
  (slightly more readable than the output of uuencode), and typed it in
  using my editor.  If any of you ever have the same problem, here's the
  hex for future reference:
          070100002c000000000000000000000000000000000000000000000000000000
          0000dd8fff010000dd8f27000000fb02ef07000000fb01ef070000000000bc8f
          8800040000bc012f65746300
  
  I had a handy program around (doesn't everybody?) for converting ASCII
  hex to binary, and the output of /usr/bin/sum tallied with our
  original binary.  But hang on---how do you set execute permission
  without /bin/chmod?  A few seconds thought (which as usual, lasted a
  couple of minutes) suggested that we write the binary on top of an
  already existing binary, owned by me...problem solved.
  
  So along we trotted to the terminal with the root login, carefully
  remembered to set the umask to 0 (so that I could create files in it
  using my gnu), and ran the binary.  So now we had a /etc, writable by
  all.  From there it was but a few easy steps to creating passwd,
  hosts, services, protocols, (etc), and then ftp was willing to play
  ball.  Then we recovered the contents of /bin across the ether (it's
  amazing how much you come to miss ls after just a few, short hours),
  and selected files from /etc.  The key file was /etc/rrestore, with
  which we recovered /dev from the dump tape, and the rest is history.
  
  Now, you're asking yourself (as I am), what's the moral of this story?
  Well, for one thing, you must always remember the immortal words,
  DON'T PANIC.  Our initial reaction was to reboot the machine and try
  everything as single user, but it's unlikely it would have come up
  without /etc/init and /bin/sh.  Rational thought saved us from this
  one.
  
  The next thing to remember is that UNIX tools really can be put to
  unusual purposes.  Even without my gnuemacs, we could have survived by
  using, say, /usr/bin/grep as a substitute for /bin/cat.
  
  And the final thing is, it's amazing how much of the system you can
  delete without it falling apart completely.  Apart from the fact that
  nobody could login (/bin/login?), and most of the useful commands
  had gone, everything else seemed normal.  Of course, some things can't
  stand life without say /etc/termcap, or /dev/kmem, or /etc/utmp, but
  by and large it all hangs together.
  
  I shall leave you with this question: if you were placed in the same
  situation, and had the presence of mind that always comes with
  hindsight, could you have got out of it in a simpler or easier way?
  Answers on a postage stamp to:
  
  Mario Wolczko

Re: Unix Admin Horror Story Summary (1992)

#42
post #5

This takes me back to roughly 1993. I was in a department running on a mix of Wyse green-screen terminals and, later, X terminals, when we got a budget upgrade that would roll out actual individual PCs -- 486s running SCO Open Desktop -- to everyone. (This was not cheap, it cost about £4000 for the hardware per seat, although the software was free because, er, this was back in the day when SCO was a respectable UNIX…

heh I have an almost identical war story.

Re: Unix Admin Horror Story Summary (1992)

#43

2003 or 2004. Customer called in and said that his dedicated server was hacked. I restored from backup. An hour later, he calls back. Hacked again . Restored again. An hour later, he calls back. He realizes that the hacker is him ! He's doing a thing, but doesn't know what he's doing wrong. So I have him email me the last thing he typed on his server, as root: rm -rf /home/user/path/to/thing /home/otheruser/path/to/s…

This is precisely why I always `-v` when `rm`ing recursively. It might be closing the barn door after the proverbial horse has bolted; but at least the fuck up is visible and in some circumstances you have a fighting chance to kill `rm` before too much damage has been done.

Re: Unix Admin Horror Story Summary (1992)

#44
> But the most important thing that can be learned from this is not that you have to make backups (we all know that, right? ;-) ). More important than making backups is to make sure your backups are complete and verified

C’est plus ça change…

Re: Unix Admin Horror Story Summary (1992)

#45

My very favorite is more of a "recovery legend", telling the heroic tale of recovering a a Unix system after an errant "rm -rf" deleted most of the system's critical files: https://www.ee.ryerson.ca/~elf/hack/recovery.html

Nice! When I saw the thread title I was hoping this story would get posted somewhere. I read this a long time ago and hadn't been able to find it for years!

Thanks for sharing!

Re: Unix Admin Horror Story Summary (1992)

#47
post #16

Earlier quoted context omitted.

I remember being assigned to look into Solaris when working as a volunteer sysadmin in grad school, where we were a SunOS shop. I took a sparcstation, wiped it, and installed Solaris. This was 1992 or so, so it must have been 5.0 or 5.1. I hated it, but I don't remember very many specifics about why I didn't like it. I think it was partially the unbundled compilers, combined with everything just being "different", co…

That was around the time that GCC finally started to get some wind, due to the unbundling of UNIX SDK.

The unbundling of the free C compiler and the high price of the unbundled C compiler and AT&T's shitty bloated C++ compiler was emblematic of what was so bad about Sun abandoning their Berkeley BSD roots and getting into bed with AT&T System V with Solaris. And that provided an opportunity for Cygnus Solutions.

Not coincidentally, after he founded Cygnus Solutions (which Red Hat later bought), Michael Tiemann worked closely with Sun to support GCC on their platform.

https://web.archive.org/web/20160310075610/http://www.toad.c...

>We had the grandiose idea that major computer companies like Sun, SGI, and DEC would fire their compiler departments and use our free compilers and debuggers instead, paying us a million dollars a year for support and development. That wasn't quite right, but before we starved, we stumbled into the embedded systems market, doing jobs for Intel (the i960, a now-forgotten RISC chip), AMD (their now-forgotten but nice 29000 RISC), and various companies like 3Com and Adobe who had to port major pieces of code to these chips. In that market, once we fixed the tools to support cross-compiling, we had major advantages over the existing competitors, and we swarmed right through the market for 32-bit embedded system programming tools. And ultimately, we did get million-dollar contracts, such as one from Sony for building Playstation compilers and emulators. This allowed game developers to start working a year before the Playstation hardware was available. This enabled the Playstation to come to market sooner, with more and better games.

https://web.archive.org/web/20150701032848/http://www.toad.c...

>Michael Tiemann, President, has been writing free software since 1987. He wrote the code for GNU C's function inlining. He wrote a portable instruction scheduler which boosted GNU C's performance by 30\% on the SPARC. He is the author of GNU C++, the first available native code C++ compiler. Mr. Tiemann has ported the GNU compiler to the SPARC, Motorola 88000, and National 32032 architectures, as well as adding support for Sun's FPA board on Sun 3s. He ported the GNU debugger to the SPARC and Intel 80386 architectures, extended the debugger and linker to handle C++ features, and ported the linker to SPARC.

https://www.oreilly.com/openbook/opensources/book/tiemans.ht...

>The real bombshell came in June of 1987, when Stallman released the GNU C Compiler (GCC) Version 1.0. I downloaded it immediately, and I used all the tricks I'd read about in the Emacs and GDB manuals to quickly learn its 110,000 lines of code. Stallman's compiler supported two platforms in its first release: the venerable VAX and the new Sun3 workstation. It handily generated better code on these platforms than the respective vendors' compilers could muster. In two weeks, I had ported GCC to a new microprocessor (the 32032 from National Semiconductor), and the resulting port was 20% faster than the proprietary compiler supplied by National. With another two weeks of hacking, I had raised the delta to 40%. (It was often said that the reason the National chip faded from existence was because it was supposed to be a 1 MIPS chip, to compete with Motorola's 68020, but when it was released, it only clocked .75 MIPS on application benchmarks. Note that 140% * 0.75 MIPS = 1.05 MIPS. How much did poor compiler technology cost National?) Compilers, Debuggers, and Editors are the Big 3 tools that programmers use on a day-to-day basis. GCC, GDB, and Emacs were so profoundly better than the proprietary alternatives, I could not help but think about how much money (not to mention economic benefit) there would be in replacing proprietary technology with technology that was not only better, but also getting better faster.

Re: Unix Admin Horror Story Summary (1992)

#48
This one made me LOL:

   My mistake on SunOS (with OpenWindows) was to try and clean up all the
   '.*' directories in /tmp. Obviously "rm -rf /tmp/*" missed these, so I
   was very careful and made sure I was in /tmp and then executed
   "rm -rf ./.*".

   I will never do this again. If I am in any doubt as to how a wildcard
   will expand I will echo it first.
I read this, and just had to go try it because I couldn't picture it in my brain. Here it is:

   $ echo ./.*
   ./. ./..
So if you're in /tmp/ and do 'rm -rf ./.*', it's

  rm -rf ./. ./..
and ./.. is .. which from tmp is /. Thankfully we have protections against this now. Back then, not so much.

Re: Unix Admin Horror Story Summary (1992)

#49
post #32

Ironically, while I love Unix, I have spent most of my career shepherding Windows boxes. The only real horror story I got was a new coworker (turning me from "the IT guy" into "half the IT department") who looked through the Active Directory tree and found the GPO management part had replicated the organizational structure. Since there were no GPOs at the time, he considered this wasteful and confusing, so he went an…

This story worked out surprisingly well, usually, not so much (:

Re: Unix Admin Horror Story Summary (1992)

#50

My very favorite is more of a "recovery legend", telling the heroic tale of recovering a a Unix system after an errant "rm -rf" deleted most of the system's critical files: https://www.ee.ryerson.ca/~elf/hack/recovery.html

Nice! When I saw the thread title I was hoping this story would get posted somewhere. I read this a long time ago and hadn't been able to find it for years! Thanks for sharing!

I once spent an hour or more finding it again, so I bookmarked it ;)
Post reply on HN