Live data from Hacker News

Viewing profile — wujerry2000

wujerry2000

HN member
Joined
Tue, Mar 01, 2022, 10:15 PM UTC
HN karma
285
Public activity
21 items

About wujerry2000

No profile information was provided.

Recent public activity

  1. comment
    Comment #45982657

    Hi all! Sharing some of our recent work around building RL envs and sims for agent training. There are a lot more technical details on building the benchmark in the post. If you ar…

  2. story
  3. comment
    Comment #44876391

    This is a really important question. I definitely think as companies begin optimizing for an "Agent first" economy, they will start figuring out how to optimize their sites for age…

  4. comment
    Comment #44876364

    Lets do it! There's a cal link on our website if you wanna chat more

  5. comment
    Comment #44872382

    Self driving cars are a really good place to derive intuitions. Robotics as well! Both those spaces are still optimizing on the last mile performance gains that get exponentially h…

  6. comment
    Comment #44870548

    OpenAI agent is very impressive! That being said, there are still a lot of use cases its not good at, and also looking at long trajectory tasks, enterprise work tasks, etc. I imagi…

  7. comment
    Comment #44870540

    I think this is totally going to be the case! AI vibe coding tools already prefer some solutions over others, probably because of training data distribution/post training preferenc…

  8. comment
    Comment #44867362

    Theses are really good questions! we share the public/consumer simulators, but we also build bespoke environments on a per customer basis (think enterprise sites or even full VMs l…

  9. comment
    Comment #44866426

    Computer use agents are starting to perform well on websites/apps that are in their training distribution, but still struggle a lot when dealing with tasks outside their distributi…

  10. comment
    Comment #44866363

    A few common ones we've heard Engineering: QA automation is huge, closes the loop on "fully automated" software engineering if another computer use system is able to click around a…

  11. comment
    Comment #44866295

    UI refreshes knocking down simulator realism is a real issue that we're still trying to solve. I think this will probably be a mixture of automated QA/engineering and scale. Anothe…

  12. comment
    Comment #44866266

    We agree that as a demo flight booking is probably overused. However, in talking with my AI Labs, their perspective on flight booking is a little different. "Solving" flight bookin…

  13. comment
    Comment #44865741

    Yea haha ... early idea was illuminate + hallucinations. Naming isn't our strength :)

  14. story
    Launch HN: Halluminate (YC S25) – Simulating the internet to train computer use

    Hi everyone, Jerry and Wyatt here from Halluminate ( https://halluminate.ai/ ). We help AI labs train computer use agents with high quality data and RL environments. Training AI ag…

  15. comment
    Comment #44188435

    Running a test with browser agents and asked it to leave a comment on a Medium article. Came back 1 month later and found it had risen to be the most liked comment on the thread. I…

  16. story
  17. comment
    Comment #42788658

    For fun, I calculated how this stacks up against other humanity-scale mega projects. Mega Project Rankings (USD Inflation Adjusted) The New Deal: $1T, Interstate Highway System: $6…

  18. comment
    Comment #42763790

    My takeaways (1) Companies will probably increasingly invest in building their own evals for their use cases because its becoming clear public/allegedly private benchmarks have mis…

  19. comment
  20. story
  21. story
    How are generative AI companies monitoring their systems in production?

    Companies with LLM-based products (specifically Retrieval-Augmented Generation)deployed in production - how are you monitoring outputs for hallucinations? What's your process?