Viewing profile — JonathanFly
JonathanFly
HN member- Joined
- Tue, Jul 09, 2019, 8:01 PM UTC
- HN karma
- 822
- Public activity
- 163 items
- HN profile
- View on Hacker News ↗
About JonathanFly
No profile information was provided.
Recent public activity
-
comment
Comment #47239305
> This is correct and also increasingly affecting me as my eyes age. I had to give my Studio Display to my wife because my eyes can't focus at a reasonable distance anymore, and if…
-
comment
Comment #46794444
> While I do agree with the content, this tone of writing feels awfully similar to LLM generated posts > Commenter's history is full of 'red flags': - "The real cost of this comple…
-
comment
Comment #46762994
Is there no way to play Doom with just the earbuds? There's a mod that adds audio cues to make Doom playable for the blind: https://www.youtube.com/watch?v=vtoAo__2kYo Adding high …
-
comment
Comment #46363066
> Every time the LLM is slightly off target, ask yourself, "What could've been clarified? Better than that, ask the LLM. Better than that, have the LLM ask itself. You do still hav…
-
comment
Comment #45784110
For me it's the motion clarity that I notice the most. Higher FPS is just one way to get more clarity though, with other methods like black frame insertion then even 60 fps feels l…
-
comment
Comment #45578332
Set nproc_per_node-1 instead of 8 (or run the training script directly instead of using torchrun) and set device_batch_size=4 instead of 32. You may be able to use 8 with a 5090, b…
-
comment
Comment #45576140
> Batch size of 8 would imply 20gb mem, no? I'm running it now and I had to go down to 4 instead of 8, and that 4 is using around 22-23GB of GPU memory. Not sure if something is wr…
-
comment
Comment #43759095
Yes, see: https://github.com/nari-labs/dia/blob/main/example/voice_clo...
-
comment
Comment #43758498
> first time I've seen such expressiveness in TTS for laughs, coughs, yelling about a fire, etc! The old Bark TTS is noisy and often unreliable, but pretty great at coughs, throat …
-
comment
Comment #42600171
So this a new method that simulates a CRT and genuinely reduces motion blur on any type of higher framerate displays, starting a 120hz. But it doesn't dim the image like black fram…
-
comment
Comment #42151940
>I love the way “take a break” is presented as an available option. I guarantee that for many caregivers it’s absolutely not. I had the same first reaction - why didn't I think of …
-
comment
Comment #41694331
It's almost certainly Google SoundStorm, a traditional TTS trained on dialogs from last year: https://x.com/jonathanfly/status/1675987073893904386
-
comment
Comment #41694315
>But why would I buy those books or listen to those podcasts that are synthetic affectations of no substance? A randomly selected NotebookLM podcast is probably not substantial eno…
-
comment
Comment #41694109
"Like" is a filler word I barely notice, along with lower key words like "right" or "uh uh". But the NotebookLM constantly exclaiming "Exactly" and "Precisely" stand out and are dr…
-
comment
Comment #41694058
Apparently people are already spamming podcast sites with NotebookLM: https://x.com/ListenNotes/status/1840470094708899992 >do you have tools to detect if audio is generated by not…
-
comment
Comment #41694005
You aren't doing anything wrong - Bark out the box uses a randomly generated voice and I like to think it's modeling the world of random voices which includes bad microphones/audio…
-
comment
Comment #41693614
Bark can sound as good, but Google is using SoundStorm which was specifically trained on dialogs. Surprisingly Bark can even sort of match it without being trained to do so, but no…
-
comment
Comment #41693591
Lawncareguy85, the creator of the viral "Podcasters discover they are AI" podcast has some other fun creations in this thread: https://www.reddit.com/r/notebooklm/comments/1fs7ka3/…
-
comment
Comment #40437121
> What is a technique/library that can take an image of a 3d environment/drawing of a room and detect a rough mesh highlighting ground, walls, barriers ? Well just in case it wasn'…
-
comment
Comment #40390239
Creating 3D spaces from inconsistent source images! Super fun idea. I tried a crude and terrible version of something like this a few years ago, but not just inconsistent spaces wi…
-
comment
Comment #39467098
From: https://twitter.com/EMostaque/status/1760660709308846135 Some notes: - This uses a new type of diffusion transformer (similar to Sora) combined with flow matching and other i…
-
comment
Comment #38755108
> I love this product, the discord community sort of seems like a dumpster fire though Do you mind saying more? (via dm if you like, I'm in that Discord same name)
-
comment
Comment #38754066
I don't work for Suno, I just Chirp and Bark a lot. You'll have to contact Suno via email, or reply to one of the Suno team that popped up in the comments here.
-
comment
Comment #38754026
It's really awkward and hard to see in IOS, even on iPad. I think the Suno devs said they're working on it.
-
comment
Comment #38753973
Are you using lyrics? They tend not to be chiptuney with sung lyrics, though they can sound cool as a hybrid: https://app.suno.ai/song/c13fab80-7f07-4d76-bad4-b79a28bb245... https:…