Live data from Hacker News

TREX: An AI code reviewer that runs your code

greptile.com

1–10 of 14 posts

Re: TREX: An AI code reviewer that runs your code

#3
We've found that methods like this substantially increase the quality and reliability of coding agent output. The ability to run code in a sandbox, drive an interactive session using a browser or API calls or other apps, and visually confirm output via vision models all adds up to plugging a big hole in the feedback loop for agent modifying a complex codebase.

We've had agents go as far as interactively testing how our product responds in video calls by launching our full stack in a set of docker containers (app, api, db, queues, etc.), all inside a larger sandbox, populating test data, connecting the mock system to a real video call solution like Google meet, and injecting audio and video to test the response. End-to-end, like a real user flow.

It's not perfect yet, but if you are a skeptic on the ability for AI agents to productively modify a complex product, I'd highly encourage you to play with a setup like this before ossifying your conclusions.

Re: TREX: An AI code reviewer that runs your code

#5
post #4

How does this work when a projet have many external dependencies, like an S3 bucket, a secret manager, a third party API, etc?

In my experience, you still are left with these annoying parts. (Ie, figuring out how to give appropriate access to your agents)

Re: TREX: An AI code reviewer that runs your code

#6
post #4

How does this work when a projet have many external dependencies, like an S3 bucket, a secret manager, a third party API, etc?

we're working on a way for you to expose creds safely into our sandbox. But for now, it's limited to mocks API calls, clicks around the UI, and unit tests.

Re: TREX: An AI code reviewer that runs your code

#8
post #6
post #4

How does this work when a projet have many external dependencies, like an S3 bucket, a secret manager, a third party API, etc?

we're working on a way for you to expose creds safely into our sandbox. But for now, it's limited to mocks API calls, clicks around the UI, and unit tests.

Are the "clicks around the UI" converted into end-to-end tests eventually? e.g. via playwright.

Re: TREX: An AI code reviewer that runs your code

#10
post #8
post #6

Earlier quoted context omitted.

we're working on a way for you to expose creds safely into our sandbox. But for now, it's limited to mocks API calls, clicks around the UI, and unit tests.

Are the "clicks around the UI" converted into end-to-end tests eventually? e.g. via playwright.

Not yet - but we want to do this. Similarly true for the ephemeral unit tests that greptile writes.
Post reply on HN