Live data from Hacker News

Viewing profile — sid-the-kid

sid-the-kid

HN member
Joined
Sat, Oct 14, 2023, 1:30 AM UTC
HN karma
40
Public activity
40 items

About sid-the-kid

No profile information was provided.

Recent public activity

  1. comment
    Comment #46792019

    never head of InWorld. Pretty impressive.

  2. comment
    Comment #46791985

    ooof. You saw the Chinese text. Yup, that's super annoying. We are trying to squash that hallucination. Thanks for the feedback! That's helpful!

  3. comment
    Comment #46789061

    Thank you! We are considering to release an open-source version of the model. Somebody will do it soon. Might as well be us. We are mostly concerned with the additional overhead of…

  4. comment
    Comment #46787771

    One thing that is interesting: LLMs pipelines have been highly optimize for speed (since speed is directly related to cost for companies). That is just not true for real-time DiTs.…

  5. comment
  6. comment
    Comment #46787732

    Good question! Software gets democratized so fast that I am sure others will implement similar approaches soon. And, to be clear, some of our "speed upgrades" are pieced together f…

  7. comment
    Comment #46787110

    it's a fair concern. but, we don't know r0fl. and we are not astroturfing. even I am surprised with how many opnely positive comments we are getting. it's not been our experience i…

  8. comment
    Comment #46785916

    Makes sense. The init should be about 10s. But, after that, it should be real time. TBH, this is probably a common confusion. So thanks for calling it out.

  9. comment
    Comment #46785623

    Thank you! Yes, right now we are using Qwen for the LLM. They also released a super fast TTS model that we have not tried yet, which is supposed to be very fast.

  10. comment
    Comment #46785361

    thanks for the feedback. that's helpful. Ya, some avatars have worse lip synch than others. It depends a little on how zoomed in you are. I am double checking now to make 100% sure…

  11. comment
    Comment #46785214

    And, just like that, Max Headroom is back: https://lemonslice.com/try/agent_ccb102bdfc1fcb30

  12. comment
    Comment #46785117

    1) yes on Max Headroom. we are on it. 2) it already is real time...?

  13. comment
    Comment #46785072

    Fix deployed! This is why it's good to launch on hacker news. thanks for the tip.

  14. comment
    Comment #46785025

    glad we found somebody who likes it as much as us! BTW, biggest thing we are working to improve is speed of the response. I think we can make that much faster.

  15. comment
    Comment #46785000

    curious what avatar you think is poor quality? Or, what you think is poor quality. i want to know :)

  16. comment
    Comment #46784967

    thanks! it just barley worked last year, but not much else. this year it's actually good. we got lucky: it's both new tech and turned out to be good quality.

  17. comment
    Comment #46784803

    Good catch! Working on a fix now.

  18. comment
    Comment #46784489

    Our text control is good, especially for emotions. For example, you can add the text prompt: "a person talking. they are angry", and agent will have an angry expression. You can al…

  19. comment
    Comment #46784434

    Good question. When using the API, you can bring any voice agent (or LLM). Our API takes in what the agent will say, and then streams back the video of the agent saying it. For the…

  20. comment
    Comment #46784301

    thank you! it's by far the thing I have worked on that I am most proud of.

  21. comment
    Comment #46784052

    Very cool! Thanks for sharing. I love your use-case of turning an AI coding agent into more of an AI employee. Will be interesting to see if users can connect better with the produ…

  22. comment
    Comment #46783963

    hey HN! one of the founders here. as of today, we are seeing informational avatars + roleplaying for training as the most common use cases. The roleplaying use-case was surprising …

  23. comment
    Comment #43787933

    We use modal ( https://modal.com/ ). They give us GPUs on-demand, which is critical for us so we are only paying for what we are using. Pricing is about $2/hr per GPU (as a baselin…

  24. comment
    Comment #43787788

    that's a good idea! Would be especially cool if the human is charismatic and does a good job driving the convo. Maybe we can try it out with a streamer.

  25. comment
    Comment #43787727

    Looked it up. Cool reference.