Live data from Hacker News

Viewing profile — mota7

mota7

HN member
Joined
Sun, Jul 08, 2018, 12:29 AM UTC
HN karma
141
Public activity
27 items

About mota7

No profile information was provided.

Recent public activity

  1. comment
    Comment #48649029

    Heat is mostly driven by leakage current and gate capacitance. The big issue today is leakage currents. They typical account for around 30%-50% of total chip thermal budget, and th…

  2. comment
    Comment #48350231

    If you have two threads on different cores that write to the same cacheline, the CPU has to enforce write ordering. The way it does this is for one of the cores to acquire a write …

  3. comment
    Comment #44998668

    Is there really that big a different in TFLOPS between the GB100 and GB202 chips? The GB100 has fewer SMs than the GB202, so I'm confused about where the 10x performance would be c…

  4. comment
    Comment #44791218

    Land the western australia wheat belt sells for less than $1000/acre. Is that very expensive?

  5. comment
    Comment #43785321

    Yes that bugged me too. If you replace 'precisely' with 'approximately' everywhere in the article it becomes much improved ;)

  6. comment
    Comment #43668478

    There's basically a difference in philosophy. GPU chips have a bunch of cores, each of which is semi-capable, whereas TPU chips have (effectively) one enormous core. So GPUs have ~…

  7. comment
    Comment #42391794

    The hand-waving explanation: The slower you're going, the easier (cheaper) it is to change direction. And for eliptical orbits, the outer-most part of the orbit is where you're goi…

  8. comment
    Comment #41890236

    Not quite: It's taking advantage of (1+a)(1+b) = 1 + a + b + ab. And where a and b are both small-ish, ab is really small and can just be ignored. So it turns the (1+a)(1+b) into 1…

  9. comment
    Comment #41781277

    I had the same thought: Just eye-balling the graphs, the result of the subtraction looks very close to just reducing the temperature. They're effectively doing softmax with a fixed…

  10. comment
    Comment #35099766

    The paper says "... optimized on next-word prediction only". Which is absolutely correct in 2023. ChatGPT (and indeed all recent LLMs) using much more complex training methods than…

  11. comment
    Comment #35048768

    Like most things, it's more complex than that, and as a result it can be either faster or slower than 'median(RTT to each DC in quorum)'. It's a delicate balance based on the locat…

  12. comment
  13. comment
    Comment #34432985

    > Gradient accumulation doesn't work with batch norms so you really need that memory. Last I looked, very few SOTA models are trained with batch normalization. Most of the LLMs use…

  14. comment
    Comment #34188620

    It's hard to convey just how ridiculously complex and expensive 5nm is compared with 90nm. 5nm is a multi-billion dollar fab, absurdly high running costs, and very expensive wafers…

  15. comment
    Comment #33219198

    Excellent example: "reflective cistern tank with a reflection of the back of a truck transporting stop signs" People tend to forget just how hard the edge cases in vision are!

  16. comment
    Comment #33217648

    Do you know why the blood thinners helped? (I'm assuming there was some underlying condition that this treated?)

  17. comment
    Comment #33088291

    The problem is that predicting a pixel requires knowing what the pixels around it looks like. But if we start with lots of noise, then the neighboring pixels are all just noise and…

  18. comment
    Comment #32600809

    Minimum is around 1 kWh/m^3 for sea-water levels of salt concentration. (It varies a fair bit depending on salinity, the actual salts involved, the temperature etc etc).

  19. comment
    Comment #32600773

    This is surprisingly poor production? Peak insolation varies widely, but 1000W/m^2 is a typical value. 5.8L/hr/m^2 means that it's using something like 180kWh/m^3 on raw solar inso…

  20. comment
    Comment #31454119

    Nuts, you are correct that it isn't optimal. It's not quite as simple as just 2 bits per digit, because the last digit can only be large if the earlier digits are small. This makes…

  21. comment
    Comment #31452071

    For the first case, you can't pick a number lower than any existing number. So if 4 is already picked, then only 5 or 6 can be picked, but if we pick 6, then there's no third numbe…

  22. comment
    Comment #31443308

    Yes. The formulation in the blog post is a bit messy, but possibly more general way to think about it is something like: 0. set y = 1. set x = 0. 1. How many possible choices are t…

  23. comment
    Comment #24820868

    It's still interesting as it reflects that despite 60,000 images, there's a very small amount of data that the network actually learns. The total entropy in 10 images (even careful…

  24. comment
    Comment #24820514

    The difficultly here is that there's an implicit assumption: 'Noise' is implicitly defined to be anything that isn't learned by the network. Now in same cases that may indeed be ac…

  25. comment
    Comment #24772070

    Can't up-vote this hard enough!