It takes a lot of computing power to reconstruct an audio sample well. Once reconstructed it takes massive massive amounts of storage.
It's all about the Shannon Nyquist sampling theorum.
"Strictly speaking, the theorem only applies to a class of mathematical functions having a Fourier transform that is zero outside of a finite region of frequencies. Intuitively we expect that when one reduces a continuous function to a discrete sequence and interpolates back to a continuous function, the fidelity of the result depends on the density (or sample rate) of the original samples. The sampling theorem introduces the concept of a sample rate that is sufficient for perfect fidelity for the class of functions that are band-limited to a given bandwidth, such that no actual information is lost in the sampling process. It expresses the sufficient sample rate in terms of the bandwidth for the class of functions. The theorem also leads to a formula for perfectly reconstructing the original continuous-time function from the samples." [1]
But, HOLD ON: "Practical digital-to-analog converters produce neither scaled and delayed sinc functions, nor ideal Dirac pulses. Instead they produce a piecewise-constant sequence of scaled and delayed rectangular pulses (the zero-order hold), usually followed by a lowpass filter (called an "anti-imaging filter") to remove spurious high-frequency replicas (images) of the original baseband signal." [1]
In case that was too verbose.....
"sample rate that is sufficient for perfect fidelity".... "such that no actual information is lost in the sampling process"... "theorem also leads to a formula for perfectly reconstructing the original continuous-time function from the samples"
PERFECTLY!!! Zero information loss between sampling and reconstruction. Still WTFs me to this day.
So yeah, you can feed bits to a PWM. But that not music.
[1] https://en.m.wikipedia.org/wiki/Nyquist%E2%80%93Shannon_samp...