Live data from Hacker News

Meta Segment Anything Model Audio

ai.meta.com

21–30 of 40 posts

Re: Meta Segment Anything Model Audio

#22

From my brief testing in the playground, it is not very good. Maybe it needs better prompting than the 1 word examples.

For me it either worked great or not at all. Extracting footsteps, the air conditioner noise, voices, one particular persons voice (identified by gender), all worked great (across multiple clips for most of those).

A few prompts failed almost entirely though, "train noises", "background noise" and "clatter"... so definitely sensitive to either prompting or the kind of noise being extracted.

Re: Meta Segment Anything Model Audio

#23
post #19

Can this be used to nuke the laugh tracks?!?

It’s really a shame how popular it was to mar shows with this… I saw a DVD set of a show once with a no-laugh-track version. It sucked because the actors pause for the laughs after each line. This is bad enough with the laugh track in place, but if it’s just dead air it makes every scene feel awkward.

Re: Meta Segment Anything Model Audio

#24
Would be interesting to leverage the non spoken/environment noises to guide what level of detail and style of speech a chatbot replied with, for instance being more casual, gentle, with a touch more detail if in a quiet home/office environment, but more curt and concise with emphasized diction if the person is traveling, such as in a noisy train concourse. People tend to do that subconsciously but bots ignorantly wittering on can be annoying and hard to use because they miss the cues.

Re: Meta Segment Anything Model Audio

#26
post #8

Earlier quoted context omitted.

Basically the same thing musicians said about the synth and music made by computers back in the day

100%. The music world has gone through the "but what will we do now?" at least 6-7 times. Music videos ("video killed the radio star"), sampling, the DAW (and time aligning), home studios, auto tune, plugins and amp simulators, napster/piracy, etc, etc.

> napster/piracy

This one rankles me because of a) the benefits piracy has (third world consumers can now discover you, for starters) and b) the absolute bad faith way in which the industry acts, screwing over artists, unethically going after Pirate Bay by making it into a trade war with Sweden (I think)

Re: Meta Segment Anything Model Audio

#27
post #13

Playing with the background I tried to Isolate just the espresso machine and the train sounds in one of their demos and it seemed to fail. Maybe not the desired use case, but I thought it was odd that I could break it so easily on the sample material.

Footsteps worked pretty well when I tried that on the other hand. I wonder if lot of it has to do with how well the model understands what the english description of the sound should sound like...

i do think that’s the case. i tried a few different ways to write x and got meaningfully varied results

Re: Meta Segment Anything Model Audio

#29
post #19

Can this be used to nuke the laugh tracks?!?

It’s really a shame how popular it was to mar shows with this… I saw a DVD set of a show once with a no-laugh-track version. It sucked because the actors pause for the laughs after each line. This is bad enough with the laugh track in place, but if it’s just dead air it makes every scene feel awkward.

AI can remove those pauses by the actors too so maybe that would work.

Re: Meta Segment Anything Model Audio

#30
post #19

Can this be used to nuke the laugh tracks?!?

It’s really a shame how popular it was to mar shows with this… I saw a DVD set of a show once with a no-laugh-track version. It sucked because the actors pause for the laughs after each line. This is bad enough with the laugh track in place, but if it’s just dead air it makes every scene feel awkward.

I don't even mind awkward pauses. I tried using the laugh track silencer on an episode of Black Adder, and it worked out OK.
Post reply on HN