Viewing profile — rocauc
rocauc
HN member- Joined
- Sat, Jan 07, 2017, 6:06 PM UTC
- HN karma
- 1,024
- Public activity
- 129 items
- HN profile
- View on Hacker News ↗
About rocauc
prev built NLP products, farms
Recent public activity
- story
- story
- story
-
comment
Comment #45986139
As someone that works on a platform users have used for labeling 1B images, I'm bullish SAM 3 can automate at least 90% of the work. Data prep is flipped to models being human-assi…
-
comment
Comment #45985334
A brief history. SAM 1 - Visual prompt to create pixel-perfect masks in an image. No video. No class names. No open vocabulary. SAM 2 - Visual prompting for tracking on images and …
- comment
-
comment
Comment #45985218
The model supports batch inference, so all prompts are sent to the model, and we parse the results.
-
comment
Comment #45985159
I tried it on transparent glass mugs, and it does pretty well. At least better than other available models: https://i.imgur.com/OBfx9JY.png Curious if you find interesting results …
-
comment
Comment #45985087
Yes. But also note that redistribution of SAM 3 requires using the same SAM 3 license downstream. So libraries that attempt to, e.g., relicense the model as AGPL are non-compliant.…
-
comment
Comment #45985024
Yes. It's a custom license with an Acceptable Use Policy preventing military use and export restrictions. The custom license permits commercial use.
-
comment
Comment #45977005
yes, downdetectorsdowndetectorsdowndetectorsdowndetector is available.
-
comment
Comment #45678263
The bike lane compliant vehicle category is exciting. Infinite Machine (infinitemachine.com) made me aware of this category with their Olto model, which is at a (surprisingly) supe…
- comment
-
comment
Comment #45624614
Not nearly enough gradient for a vibe coded site :)
-
comment
Comment #45562919
One of the most common uses for edge AI not listed in this course is computer vision. You similarly want real-time inference for processing video. Another open source project that …
-
comment
Comment #44884764
Reminds me of NY Cerebro, semantic search across New York City's hundreds of public street cameras: https://nycerebro.vercel.app/ (e.g. search for "scaffolding")
-
comment
Comment #44555946
both the endeavor and the site are super cool - congrats on 10 years. interaction on the graphics would be a nice touch to select into a specific run. went looking for the code on …
-
comment
Comment #43382928
In the 2019 fatal Tesla Autopilot crash, the Tesla failed to identify a white tractor trailer crossing the highway: https://www.washingtonpost.com/technology/interactive/2023/t...
-
comment
Comment #43382764
I wonder how long until techniques like Depth Anything ( https://depth-anything-v2.github.io/ ) provide parity with human depth perception. In Mark Rober's tests, I'm not sure even…
-
comment
Comment #42994461
Meta deeply comprehends the impact of GPT-3 vs ChatGPT. The model is a starting point, and the UX of what you do with the model showcases intelligence. This is especially pronounce…
-
comment
Comment #41774036
A suggestion: I'd swap llava for Florence-2 for your open set text description. Florence-2 seems uniformly more descriptive in its outputs.
-
comment
Comment #41219761
SAM 2's key contribution is adding time-based segmentation to apply to videos. Even on images alone, the authors note [0] the image-based segmentation benchmark does exceed SAM 1 p…
-
comment
Comment #41104928
One thing its enabled is automated annotations for segmentation, even on out-of-distribution examples. e.g. in the first 7 months of SAM, users on Roboflow used SAM-powered labelin…
-
comment
Comment #41097434
"Load Example" was very helpful to get a sense of what this does. Awesome build. +1 to the other comment wanting a breakdown of what the colors mean. Also, combining this with real…
-
comment
Comment #40768521
i work on roboflow. seeing all the creative ways people use computer vision is motivating for us. let me know (email in bio) if there's things you'd like to be better.