Live data from Hacker News

Viewing profile — rshemet

rshemet

HN member
Joined
Sat, Aug 06, 2022, 11:32 PM UTC
HN karma
41
Public activity
40 items

About rshemet

No profile information was provided.

Recent public activity

  1. comment
    Comment #49263229

    Did you give it a tool to increase temperature, or only one that sets temperature to an absolute value? Either way, setting temperature to 5° is obviously wrong - even if it knew t…

  2. comment
    Comment #49253383

    hey Kenny, Roman from Cactus here - could you say more? What kind of home assistant / what stack

  3. comment
    Comment #49252767

    it stands for Lets-not-be-sarcastic :)

  4. comment
    Comment #49252731

    Roman from Cactus here - yes you're right, there's only so much a 14MB model can do. Needle excels at in-conext inference, with tightly defined environments. In our experience: acc…

  5. comment
    Comment #49252657

    there are android binaries you can ship in your own app - https://huggingface.co/Cactus-Compute/needle2/tree/main but if you're just looking for somewhere to try the model, use our…

  6. comment
    Comment #49250960

    Hey! Roman here from Cactus - yes, we're putting putting together a detailed guide for ESP32. In the meantime, if you have enough RAM for the current model (≈28MB), our repo will g…

  7. story
    Show HN: Cactus v2 – On-device AI with cloud fallback

    Hi HN, Roman and Henry here from Cactus ( https://github.com/cactus-compute/cactus ). We just shipped the biggest upgrade to our on-device inference platform: - Built-in model conf…

  8. comment
    Comment #45300133

    Yes! Cactus is optimized for mobile CPU inference and we're finishing internal testing of hybrid kernels that use the NPU, as well other chips. We don't advise using GPUs on smartp…

  9. comment
    Comment #45300081

    indeed, this is exactly the goal! The license grants rights to commercial use, unlocks additional hardware acceleration, includes cloud telemetry, and offers significant savings ov…

  10. comment
    Comment #44907399

    you can run it in Cactus Chat (download from the Play Store)

  11. comment
    Comment #44907392

    you can also run it on Cactus - either in Cactus Chat from the App/Play Store or by using the Cactus framework to integrate it into your own app

  12. comment
    Comment #44907329

    THIS IS THE BOMB!!! So excited for this one. Thanks for putting cool tech out there.

  13. comment
    Comment #44840769

    if you ever end up trying to take this in the mobile direction, consider running on-device AI with Cactus – https://cactuscompute.com/ Blazing-fast, cross-platform, and supports ne…

  14. comment
    Comment #44534214

    https://play.google.com/store/apps/details?id=com.rshemetsub...

  15. comment
    Comment #44534199

    thank you! Very kind feedback, and we'll add your feedback to our to-dos. re: "question would get stuck on the last phrase and keep repeating it without end." - that's a limitation…

  16. comment
    Comment #44534187

    say more about "community tools"?

  17. comment
    Comment #44534184

    in the app you mean? Adding shortly!

  18. comment
    Comment #44534177

    that's our mission! if you are passionate about the space, we look forward to your contributions!

  19. comment
    Comment #44534168

    no, good observation - not hidden; we don't have a "clear conversation" button. to your previous point - Cactus fully supports tool calling (for models that have been instruction-t…

  20. comment
    Comment #44534130

    looking forward to your feedback!

  21. comment
    Comment #44527328

    hot off the press in our latest feature release :) we support cloud fallback as an add-on feature. This lets us support vision and audio in addition to text.

  22. comment
    Comment #44527107

    great observation - this data is not from a controlled environment; these are metrics from our Cactus Chat use (we only collect tok/sec telemetry). S25 is an outlier that surprised…

  23. comment
    Comment #44526739

    thank you! We're continue to add performance metrics as more data comes in. A Qwen 2.5 500M will get you to ≈45tok/sec on an iPhone 13. Inference speeds are somewhat linearly inver…

  24. comment
    Comment #44526336

    Great question. Currently, each app is sandboxed - so each model file is downloaded inside each app's sandbox. We're working on enabling file sharing across multiple apps so you do…

  25. comment
    Comment #44526266

    reminds me of - "You are, undoubtedly, the worst pirate i have ever heard of" - "Ah, but you have heard of me" Yes, we are indeed a young project. Not two weeks, but a couple of mo…