Live data from Hacker News

Viewing profile — zacksiri

zacksiri

HN member
Joined
Tue, Jan 11, 2022, 7:49 AM UTC
HN karma
244
Public activity
132 items

About zacksiri

agentic movie database - https://memovee.com blog - https://zacksiri.dev

Recent public activity

  1. comment
    Comment #49219782

    Yes, I think a proper comprehensive test would be a better judge of the outcome. I may do round 2 given my first batch of models is already outdated.

  2. comment
    Comment #49218829

    I’m comparing same / similar settings between models. I can’t use high on one and low on others it’s not a fair test. Not sure why I was downvoted. But seems the downvoter is quick…

  3. comment
    Comment #49218541

    I'm not sure about all these benchmarks, I did some very simple tests (I have my own benchmarks https://upmaru.com/llm-tests ) and these models fail, not sure if it's the inference…

  4. comment
    Comment #48998781

    Gemini 3.5 flash-lite is more expensive than Gemini 3.1 flash-lite. Every upgrade is getting more expensive.

  5. story
    Show HN: Graph Context Engine for Reliable AI

    Been working on this for the last 2 years. There is a video on the site explaining how it works. There is also a production app powered by this engine. Give it a try, I'm happy to …

  6. comment
    Comment #48848206

    They want you to buy their hosted service, that's where the convenience is sold. If they give you a one liner script you can paste in or a docker compose that does everything from …

  7. story
  8. story
    Show HN: I built custom movie dashboards for Apple TV

    At home we have Netflix, Disney+, Amazon Prime Video, HBO Max, and Apple TV+, but somehow we still end up just scrolling around, giving up, and watching YouTube. There must be over…

  9. comment
    Comment #48707319

    I watched this movie recently, and the same thought crossed my mind. I was going to write a blog post about it, but was too afraid to read the responses. I'm glad you did. If I mee…

  10. comment
    Comment #48694852

    They cherry pick and choose the parts of her work that fits their agenda, while forgetting the other parts.

  11. comment
    Comment #48694788

    This reminds me of the following quote "When you see that in order to produce, you need to obtain permission from men who produce nothing - When you see that money is flowing to th…

  12. comment
    Comment #48528876

    Thank you for publishing this. I've been following Paul Graham and his works for a long time. It's refreshing to see everything written down in a document like this. This is the bi…

  13. comment
  14. comment
    Comment #48202416

    Do you have similar math for the flash-lite variant of the models? I'd be curious. Based on my testing / benchmark i think it's around the 100-120B mark. With the Pro variant being…

  15. story
  16. comment
    Comment #47836903

    It's a lot more than an LLM, it's backed by database + LLM. Gemini / ChatGPT on it's own believe it or not has hallucinated movies and imdb links. While google can show you results…

  17. comment
    Comment #47836719

    Well with JustWatch and other service aggregators you cannot search for movies by themes / use natural language. For example: - "Zombie or post apocalyptic movies released in the l…

  18. story
    Show HN: My First iOS App

    Honestly I didn't think it would be possible when I started, my background is Elixir / Phoenix / Backend engineering. But thx to Codex the app turned out to be something I'm quite …

  19. comment
    Comment #47687065

    I hope that one day humanity learns that in war there are no winners. We're all just brothers and sisters born on different corners of the planet. We share the same home. I hope th…

  20. story
  21. comment
    Comment #47409730

    It's ok, it's not the best. There are models that do better, I'd use it for some basic tasks but not actual complex tasks like query generation and retrieval.

  22. comment
    Comment #47408944

    I tested the model in an agentic workflow. Here is the report: https://upmaru.com/llm-tests/simple-tama-agentic-workflow-q1...

  23. comment
    Comment #47396260

    I tested this model in an agentic workflow, it failed at some very basic tasks: https://upmaru.com/llm-tests/simple-tama-agentic-workflow-q1...

  24. comment
    Comment #47235782

    Yes, my workflows use caching intensively. It's the only way to keep things fast / economical.

  25. comment
    Comment #47235401

    This is going to be a fun one to play with. I've been conducting tests on various models for my agentic workflow. I was just wishing they would make a new flash-lite model, these t…