WatchMachineGo is basically a visual simulator that simulates how a local LLM model gets loaded, prefilled and then used for inference, while showing the effects of different hardware parameters like memory bandwitch or setups like no GPU, two GPUs and so on.
All free and without ads, forever.
I plan to open source it too, but want to think about how first, still.
Note: It is still under construction!
If you check it out: Thank you very much and I hope that it will be time well spent! :)
Show HN: WatchMachineGo – A visualizer to show hardware performing LLM inference
watchmachinego.com