Viewing profile — wujerry2000
wujerry2000
HN member- Joined
- Tue, Mar 01, 2022, 10:15 PM UTC
- HN karma
- 285
- Public activity
- 21 items
- HN profile
- View on Hacker News ↗
About wujerry2000
No profile information was provided.
Recent public activity
-
comment
Comment #45982657
Hi all! Sharing some of our recent work around building RL envs and sims for agent training. There are a lot more technical details on building the benchmark in the post. If you ar…
- story
-
comment
Comment #44876391
This is a really important question. I definitely think as companies begin optimizing for an "Agent first" economy, they will start figuring out how to optimize their sites for age…
-
comment
Comment #44876364
Lets do it! There's a cal link on our website if you wanna chat more
-
comment
Comment #44872382
Self driving cars are a really good place to derive intuitions. Robotics as well! Both those spaces are still optimizing on the last mile performance gains that get exponentially h…
-
comment
Comment #44870548
OpenAI agent is very impressive! That being said, there are still a lot of use cases its not good at, and also looking at long trajectory tasks, enterprise work tasks, etc. I imagi…
-
comment
Comment #44870540
I think this is totally going to be the case! AI vibe coding tools already prefer some solutions over others, probably because of training data distribution/post training preferenc…
-
comment
Comment #44867362
Theses are really good questions! we share the public/consumer simulators, but we also build bespoke environments on a per customer basis (think enterprise sites or even full VMs l…
-
comment
Comment #44866426
Computer use agents are starting to perform well on websites/apps that are in their training distribution, but still struggle a lot when dealing with tasks outside their distributi…
-
comment
Comment #44866363
A few common ones we've heard Engineering: QA automation is huge, closes the loop on "fully automated" software engineering if another computer use system is able to click around a…
-
comment
Comment #44866295
UI refreshes knocking down simulator realism is a real issue that we're still trying to solve. I think this will probably be a mixture of automated QA/engineering and scale. Anothe…
-
comment
Comment #44866266
We agree that as a demo flight booking is probably overused. However, in talking with my AI Labs, their perspective on flight booking is a little different. "Solving" flight bookin…
-
comment
Comment #44865741
Yea haha ... early idea was illuminate + hallucinations. Naming isn't our strength :)
-
story
Launch HN: Halluminate (YC S25) – Simulating the internet to train computer use
Hi everyone, Jerry and Wyatt here from Halluminate ( https://halluminate.ai/ ). We help AI labs train computer use agents with high quality data and RL environments. Training AI ag…
-
comment
Comment #44188435
Running a test with browser agents and asked it to leave a comment on a Medium article. Came back 1 month later and found it had risen to be the most liked comment on the thread. I…
- story
-
comment
Comment #42788658
For fun, I calculated how this stacks up against other humanity-scale mega projects. Mega Project Rankings (USD Inflation Adjusted) The New Deal: $1T, Interstate Highway System: $6…
-
comment
Comment #42763790
My takeaways (1) Companies will probably increasingly invest in building their own evals for their use cases because its becoming clear public/allegedly private benchmarks have mis…
- comment
- story
-
story
How are generative AI companies monitoring their systems in production?
Companies with LLM-based products (specifically Retrieval-Augmented Generation)deployed in production - how are you monitoring outputs for hallucinations? What's your process?