Live data from Hacker News

Viewing profile — xfalcox

xfalcox

HN member
Joined
Thu, Nov 07, 2013, 3:57 PM UTC
HN karma
1,035
Public activity
313 items

About xfalcox

[ my public key: https://keybase.io/falcofantastic; my proof: https://keybase.io/falcofantastic/sigs/_PKYsKf2wmCyt834lEh6N4POje9RoICd3Ta7qezTzJE ]

Recent public activity

  1. comment
    Comment #48950142

    I'm wondering the same! How that article has no links is beyond me.

  2. comment
    Comment #48852402

    Question to the OP, have you tested this on a machine where the entire model and context fit in RAM ?

  3. comment
    Comment #48852393

    README covers that https://github.com/JustVugg/colibri#ssd-wear-warning

  4. comment
    Comment #48377823

    Given my dev machine has 32GB of RAM and 32GB of VRAM that sits mostly idle when I'm not running AI models, this is not that bad of an idea.

  5. comment
    Comment #47616947

    Comparing a model you can downloads weights for with an API-only model doesn't make much sense.

  6. comment
    Comment #47606119

    Our CEO did that at our company and found 33 CVEs. Rails also did that and found 7 or 8.

  7. comment
    Comment #46714929

    I just made a new installer for Discourse on CharmRuby, now I gotta check this out and see if porting is feasible. Hopefully this reduces the app size, that is quite large with Cha…

  8. comment
    Comment #46619383

    That is a great fit for the GIF integration in Discourse. I was able to quickly add support for it at https://github.com/discourse/discourse-gifs/pull/107 Love to see WEBP support.…

  9. comment
    Comment #46529703

    First time I was in San Francisco and someone introduced themselves like that, going even beyond, was indeed a super weird experience being a brazilian.

  10. comment
    Comment #46175375

    We have vLLM for running text LLMs in production. What is the equivalent for this model?

  11. comment
    Comment #46081938

    I am partial to https://huggingface.co/Qwen/Qwen3-Embedding-0.6B nowadays. Open weights, multilingual, 32k context.

  12. comment
    Comment #45926515

    It's the Amazon own model. I'm baffled someone would pick it, even more that someone would test Llama 4 for a task in an age where Sonnet 4.5 is already out, so in the last 45 days…

  13. comment
    Comment #45807277

    > what does the rag for uploaded files do in discourse? You can upload files that will act as RAG files for an AI bot. The bot can also have access to forum content, plus the abili…

  14. comment
    Comment #45805394

    We host thousands of forums but each one has its own database, which means we get a sort of free sharding of the data where each instance has less than a million topics on average.…

  15. comment
    Comment #45805239

    I was taken back when I saw what was basically zero recall loss in the real world task of finding related topics, by doing the same thing you described where we over capture with b…

  16. comment
    Comment #45801650

    In Discourse embeddings power: - Related Topics, a list of topics to read next, which uses embeddings of the current topic as the key to search for similar ones - Suggesting tags a…

  17. comment
    Comment #45800076

    Also worth mentioning that we use quantization extensively: - halfvec (16bit float) for storage - bit (binary vectors) for indexes Which makes the storage cost and on-going perform…

  18. comment
    Comment #45799162

    > Nobody’s actually run this in production We do at Discourse, in thousands of databases, and it's leveraged in most of the billions of page views we serve. > Pre- vs. Post-Filteri…

  19. comment
    Comment #44994870

    Depends on your needs. You surely don't want 32k long chunks for doing the standard RAG pipeline, that's for sure. My use case is basically a recommendation engine, where retrieve …

  20. comment
    Comment #44994043

    Just migrated all embeddings to this same model a few weeks ago in my company, and it's a game changer. Having 32k context is a 64x increase when compared with our previous used mo…

  21. comment
    Comment #44894359

    Having a public tokenizer is quite useful, specially for embeddings. It allows you to do the chunking locally without going to the internet.

  22. comment
    Comment #44859645

    Qwen 3 is not slow by any metrics. Which model, inference software and hardware are you running it on? The 30BA3B variant flies on any GPU.

  23. comment
    Comment #44206871

    You'd be surprised how often people in enterprise can be left waiting months to get an API key approved for an LLM provider.

  24. comment
    Comment #44056099

    This looks like a great fit for allowing people to monetize their Discourse forums, by having partners stores and plugging those instead of ads. Will build a quick poc integration.…

  25. comment
    Comment #43939287

    This looks super cool, exactly what I've been wanting to create some useful widgets! Thanks for sharing!