Live data from Hacker News

Launch HN: Traceloop (YC W23) – Detecting LLM Hallucinations with OpenTelemetry

news.ycombinator.com

21–30 of 76 posts

Re: Launch HN: Traceloop (YC W23) – Detecting LLM Hallucinations with OpenTelemetry

#21
post #19

Earlier quoted context omitted.

Are there any approaches today that you've found are at least mostly reliable? Bonus points if it is somewhat clear/easy/predictable to know when it isn't or won't be. We use human evaluation but that is naturally far from scalable, which has especially been a problem when working on more complicated workflows/chains where changes can have a cascading effect. I've been encouraging a lot of dev experimentation on my t…

I tend to find classic NLP metric more predictable and stable than "LLM as a judge" metrics so I'd try to see if you rely on them more. We've written a couple of blog posts about some of them: https://www.traceloop.com/blog

for your blog can i offer a big downvote for the massive ai generated cover image thing? its a trend for normies but for developers its absolutely meaningless. give us info density pls

Re: Launch HN: Traceloop (YC W23) – Detecting LLM Hallucinations with OpenTelemetry

#22
post #21
post #19

Earlier quoted context omitted.

I tend to find classic NLP metric more predictable and stable than "LLM as a judge" metrics so I'd try to see if you rely on them more. We've written a couple of blog posts about some of them: https://www.traceloop.com/blog

for your blog can i offer a big downvote for the massive ai generated cover image thing? its a trend for normies but for developers its absolutely meaningless. give us info density pls

roger that! I like them though (am I a normie then?)

Re: Launch HN: Traceloop (YC W23) – Detecting LLM Hallucinations with OpenTelemetry

#23
post #15

Earlier quoted context omitted.

I have it internally, I can share it if you want! But to the point of comparison between these and tools like Traceloop - it's interesting to see this space and how each platform takes it's own path and finds its own use cases. LangSmith works well within the LangChain ecosystem together with LangGraph, LangServe. But if you're using LlamaIndex, or even just vanilla OpenAI you'll be spending hours to set up your obse…

I'd love to see that list!

Ping me over slack (traceloop.com/slack) or email nir at traceloop dot com

Re: Launch HN: Traceloop (YC W23) – Detecting LLM Hallucinations with OpenTelemetry

#24
post #8
post #2

Congratulations on launch. This is a crowded market, and there are many tools doing the same thing. How are you differentiating yourself from other tools like: Langfuse Portkey Keywords ai Promptfoo

not to mention langsmith? braintrust? humanloop? does that count? not sure what else - lets crowdsource a list here so that people can find them in future

Im not sure which ones are Otel compliant. Im only aware of 3 that are Otel compliant:

1. Traceloop Otel 2. Langtrace.ai Otel 3. OpenLIT Otel 4. Portkey 5. Langfuse 6. Arize LLM 7. Phoniex SDK 8. Truera LLM 9. Truelens 10. Context 11. Braintrust 12. Parea 13. Context AI 14. openlayer.com 15. Deepchecks 16. langsmith 17. Confident AI 18. Helicone 19. Langwatch.ai 20. Arthur 21. Aporia 22. scale.com 23. Whylabs 24. gentrace.ai 25. humanloop.com 26. fixpoint.co 27. W n B Traces 28. Langtail 29. Fiddler 30. Evidently Ai 31. Superwise 32. Exxa 33. Honeyhive 34. Flowstack 35. Log10 36. Giskard 37. Raga AI 38. AgentOps 39. Patronus AI 40. Mona 41. Bricks Ai 42. Sentify 43. LogSpend 44. Nebuly 45. Autoblocks 46. Radar / Langcheck 47. Dokulabs 48. Missing studio 49. Lunary.ai 50. Censius.ai 51. ML flow 52. Galileo 53. trubrics 54. Prompt Layer 55. Athina 56. getnomos.com 57. c3.ai 58. baselime.io 59. Honeycomb llm

Re: Launch HN: Traceloop (YC W23) – Detecting LLM Hallucinations with OpenTelemetry

#25
post #8

Earlier quoted context omitted.

not to mention langsmith? braintrust? humanloop? does that count? not sure what else - lets crowdsource a list here so that people can find them in future

Im not sure which ones are Otel compliant. Im only aware of 3 that are Otel compliant: 1. Traceloop Otel 2. Langtrace.ai Otel 3. OpenLIT Otel 4. Portkey 5. Langfuse 6. Arize LLM 7. Phoniex SDK 8. Truera LLM 9. Truelens 10. Context 11. Braintrust 12. Parea 13. Context AI 14. openlayer.com 15. Deepchecks 16. langsmith 17. Confident AI 18. Helicone 19. Langwatch.ai 20. Arthur 21. Aporia 22. scale.com 23. Whylabs 24. gen…

dear god, where is this list from? surely not hand curated?

Re: Launch HN: Traceloop (YC W23) – Detecting LLM Hallucinations with OpenTelemetry

#26
Check out these Wikipedia articles:

Confabulation https://en.m.wikipedia.org/wiki/Confabulation

Hallucination https://en.m.wikipedia.org/wiki/Hallucination

What drove the AI industry to blow off accepted naming from psychopathology and use the word for PERCEPTUAL errors to refer to LANGUAGE OUTPUT errors?

When AI hallucinates, and AI people already use the preferred term “hallucination” to label confabulations, then what’s the new word for “hallucinations?”

How will we avoid serious errors in understanding if hallucination in AI means confabulation in humans and $NEW_TERM in AI means hallucination in humans?

Just seems harmful to gloss over this humongous vocabulary error.

How can we claim to respect the difficulty of naming things if we all select the wrong answer to a basic undergrad psychology multiple choice question with only two options?

It feels like painting ourselves into a corner which will inevitably make computer scientists look dumb. Who here wants to look dumb for no reason?

I don’t want to be negative, but is using the blatantly wrong word for confabulation a good idea in the long term?

Re: Launch HN: Traceloop (YC W23) – Detecting LLM Hallucinations with OpenTelemetry

#27
post #16

there's a well known artist named traceloops who has a prolific/longstanding body of work. why did you choose this name?

I know! When we started every time I was googling "traceloop" this was the first result. 2 reasons why we chose it (in this order): 1. traceloop.com was available 2. we work with traces

an available .com is basically the only reason you should use https://paulgraham.com/name.html

Re: Launch HN: Traceloop (YC W23) – Detecting LLM Hallucinations with OpenTelemetry

#28
post #11
post #4

Just wanted to say great work on standardizing otel for LLM applications ( https://github.com/open-telemetry/semantic-conventions/tree/... ] and opensourcing OpenLLMetry. We're also building in this space, focusing more on eval (agenta). I think using otel would make the whole space move much faster.

Thanks so much! I always say that I'm a strong believer in open protocols so I'd love to assist you if you want to use OpenLLMetry as your SDK. We onboarded other startups / competitors like Helicone and Honeyhive and it's been tremendously successful (hopefully that's what they'll tell you as well)

HoneyHive founder here.

Nir and team have built an amazing OSS package and have been fantastic to collaborate with (despite being competitors)! As an industry, I think more of us need to work together to standardize telemetry protocols, schemas, naming conventions, etc. since it’s currently all over the place and leads to a ton of confusion and headache for developers (which ultimately goes against the whole point of using devtools in the first place).

We recently integrated OpenLLMetry into our SDKs with the sole purpose of offering standardization and interoperability with customers’ existing DevSecOps stacks. Customers have been loving it so far!

Re: Launch HN: Traceloop (YC W23) – Detecting LLM Hallucinations with OpenTelemetry

#29

there's a well known artist named traceloops who has a prolific/longstanding body of work. why did you choose this name?

I doubt anyone would be confused with Traceloops the artist vs Traceloop the LLM Observability Platform

Re: Launch HN: Traceloop (YC W23) – Detecting LLM Hallucinations with OpenTelemetry

#30
Big congrats on the official launch!

Slightly tooting my own horn here, but at OpenPipe we've got a collaboration set up with Traceloop. That means you can record your production traces in Traceloop then export them to OpenPipe where you can filter/enrich them and use them to fine-tune a super strong model. :)

Post reply on HN