Live data from Hacker News

Show HN: NCompass Technologies – yet another AI Inference API, but hear us out

ncompass.tech

31–36 of 36 posts

Re: Show HN: NCompass Technologies – yet another AI Inference API, but hear us out

#31
post #17

Earlier quoted context omitted.

Hey, great that you mentioned this. We actually had BAAI/bge-m3 on our list of models to put up in the near future to see if people had use for it over an API. It's great to hear that this is something you're looking for. If you could let us know if there was a specific model you wanted to run, we can look into getting that put up soon.

Colbert, colqwen are underserved would benefit from a latency optimized inference service

Awesome, we really appreciate the suggestions! We'll look into getting these up and running shortly!

Re: Show HN: NCompass Technologies – yet another AI Inference API, but hear us out

#32
Interesting approach to model serving - the 2-4x lower TTFT compared to vLLM is impressive, but I'd be curious to see detailed benchmarks across different batch sizes and model architectures to validate those performance claims. The no rate limits policy is bold but could get expensive fast if you're not doing some clever GPU utilization under the hood.

Re: Show HN: NCompass Technologies – yet another AI Inference API, but hear us out

#33
post #19
post #18

Unrelated: During the dot-com boom, there was a company called nCompass Labs that developed one of the first content management systems ( https://en.wikipedia.org/wiki/NCompass_Labs_Inc ). Microsoft bought them in 2001. Their product was, "a plug-in for hosting ActiveX controls in Netscape Navigator named ScriptActive." ActiveX itself was a novelty, using C++ templates to define reusable and _downloadable_ web compon…

It now makes sense that when we tested the domain ncompass.com it took us to a Microsoft home page, which is why we're ncompass.tech :)

That’s hilarious. I bet if you reach out to Microsoft, they will give you that domain. There’s no way they’re using the trademark anymore.

Re: Show HN: NCompass Technologies – yet another AI Inference API, but hear us out

#34

Interesting approach to model serving - the 2-4x lower TTFT compared to vLLM is impressive, but I'd be curious to see detailed benchmarks across different batch sizes and model architectures to validate those performance claims. The no rate limits policy is bold but could get expensive fast if you're not doing some clever GPU utilization under the hood.

Thanks for your comments. Absolutely, as we were mentioning in one of the other threads, we are really keen on building towards having a reproducible dashboard of efficiency and other metrics.

Also regarding the no rate limits, we agree this is a real challenge and it's part of why we're interested in building this as well. I think the clever GPU utilization tricks are exactly what we're building out and also looking forward to see what the various issues we're going to run into at such scale.

Re: Show HN: NCompass Technologies – yet another AI Inference API, but hear us out

#36
post #35

this sounds like black magic, kudos to you. i'd love to chat, dm me on https://twitter.com/swyx if you'd find it useful to chat with someone like me.

Thank you! Absolutely, I'll send over a DM and we can take it from there!
Post reply on HN