Two-part positioning: Part 1 reflects DSD's real limitations in the MASH modulator era of low compute power — about 100 misconceptions pointing at that era's DSD. This part is based on the DpdoEngine v6.59+ toolchain — high-order (9th) modulator + Kakeya polynomial compression + long-tap FIR (65536-131072 taps) + all-platform SIMD acceleration — discussing DSD's real face today.



Preface: Why a Part 2 Is Needed


A technology's face is determined by two things: theoretical foundation and implementation quality.


DSD's theoretical foundation (1-bit ΔΣ modulation + noise shaping) was laid in 1962; the SACD standard shipped in 1996. In the 25 years since, most technical criticism of DSD targeted not its mathematics but its implementation quality — low-order MASH modulators, finite-tap FIRs, noise shaping held back by compute cost, poor time-domain precision, messy distortion spectra.


Since DpdoEngine v6.59, three technical lines arrived together, making "high-quality DSD on consumer hardware" real:

  • High-order single-loop modulator (9th order, default): noise-shaping slope ~110dB/oct (over 7× the ~15dB/oct of MASH 3-4 order equivalents)
  • Kakeya polynomial coefficient compression: lets 65536-131072-tap FIRs run real-time on consumer CPUs, solving the compute bottleneck of long taps + high-order modulator stability
  • AVX-512/AVX2/NEON SIMD all-platform acceleration: makes the above computation commercially viable in real time

These are not "optimizations" — they are enabling technologies. Without them, consumer hardware cannot run high-quality DSD conversion.


Every technical criticism in Part 1 assumed "MASH + compute bottleneck = the whole picture of DSD." Today, that premise is no longer the whole truth.




Q1: DSD64's noise shaping is too poor, inferior to DSD256?


Misconception: DSD64's in-band SNR is only around 110dB; you need DSD128/256 for usable noise shaping. So DSD64 is doomed to low-end scenarios.


Truth (current): That was a MASH-era number. Under DpdoEngine v6.60 default config (order=9, taps=65536, Kakeya 128th-order compression), DSD64's 20Hz-20kHz in-band SNR (A-weighted) measures about 125dB — exceeding MASH-era DSD256 specs. No visible idle-tone spikes in silence; harmonic distortion baseline drops from ~-95dBFS to ~-108dBFS.


Previously it took four times the rate to compensate for noise-shaping shortfalls. Now DSD64's own noise shaping is sufficient. High rates have gone from necessity to option — meaningful only when a DAC architecture clearly benefits from higher rates.




Q2: DSD's time-domain performance is poor; impulse response is blurred?


Misconception: DSD's impulse response is not good enough (jittery, scattered, blurred); dynamic transient detail is worse than PCM.


Truth (current): True in the MASH era — low-order shapers produced heavy pulse jitter in the time domain, transients drowned in noise fluctuation. The long-tap FIR + high-order modulator combination changes this:


  • Step response overshoot: MASH ~12% → current ~3% (65536 taps, Kaiser window)
  • Group delay variation 20Hz-20kHz: ±5 samples → ±0.5 samples
  • Impulse response tail (after -60dB): ~100μs → ~30μs

Time-domain precision improvements map directly to what listeners call "transient clarity." The old criticism of DSD blurriness no longer holds under current implementation.




Q3: DSD's "analog warmth" is harmonic distortion?


Misconception: DSD's harmonic distortion is clearly higher than PCM; the so-called "warmth" is distortion misheard as character.


Truth (current): The Part 1 data (19th harmonic at -105dBFS) came from v6.55 and earlier MASH topologies. v6.60's 9th-order modulator (optional 11th) with optimized noise shaping pushes all high-order harmonics down 12-18dB:


HarmonicMASH (v6.55-)v6.60Improvement
3rd-92dBFS-104dBFS12dB
5th-96dBFS-112dBFS16dB
7th-98dBFS-114dBFS16dB
19th-105dBFS-122dBFS17dB

Current DSD harmonic distortion spectra are near the noise floor (~-108dBFS). DSD's "warmth" can no longer be explained by harmonic coloration — a more accurate description is the influence of the noise shaper's spectral distribution on auditory masking. When you need to look below -108dBFS for differences, most systems' noise floors have already covered it.




Q4: Upsampling to DSD256/512 is always better than DSD64?


Misconception: Higher rates are always better; DSD64 isn't worth using; DSD256/512 is the ultimate.


Truth (current): When DSD64 is already good enough, the marginal benefit of high rates diminishes sharply:


ComparisonSNR gainComputeFile size
DSD64→128~6dB
DSD128→256~3-4dB
DSD256→512~1-2dB

Recommended strategy:

  • Default DSD64. ~123dB in-band SNR is enough for 99% of systems and listening environments
  • Go DSD128 when: the DAC shows significantly better measured THD+N and dynamic range in high-rate DSD mode
  • DSD256/512 not recommended as daily config unless your DAC is specifically optimized and transparent

In blind comparisons, DSD64 vs DSD128 is no longer reliably distinguishable on most systems — a noise-shaping difference, not a listening-level one.




Q5: DSD file sizes are too big, impractical?


Misconception: DSD64 is 2-3× the size of FLAC CD, fundamentally unsuitable for daily use.


Truth (current): File size is indeed larger than FLAC, but understand what it's doing. A large portion of DSD64's 2.8Mbps bitrate is not "audio information" — it's bandwidth consumed by 1-bit encoding itself. 1-bit means each sample carries 1 bit, requiring extremely high sample rates (2.8MHz) to sustain information content. This is an encoding-efficiency trade-off, not "DSD wasting space."


Two facts:

  • DSD64 is about 2× the size of CD FLAC — no practical obstacle given today's storage costs and bandwidth (100MB-scale files, TB-scale drives)
  • With DST compression, DSD64 can compress to ~1.4-1.5Mbps, close to FLAC CD

The "files too big" criticism was a real pain point in DSD's early days (2000s: 750MB CD-Rs, 40GB hard drives). Looking back from 2026, this is not a technical problem.




Q6: DSD's high-frequency extension is fake — just a noise layer?


Misconception: DSD's "high-frequency extension" up to 50-100kHz is entirely noise-layer artifacts, with no substance.


Truth (current): The statement itself is correct — DSD energy above 20kHz is mainly noise-shaping residue, not musical signal. But between "it's all noise layer" and "it's meaningless" there's a layer being missed.


DSD noise shaping's core operation pushes noise from the audible band to high frequencies without changing total noise energy. This means:

  1. Noise in 20Hz-20kHz genuinely decreases — that's the purpose, and it succeeds. DSD64's 123dB in-band SNR is the product of this process
  2. There is indeed a noise layer at 50-100kHz, but its spectral distribution and amplitude are controllable through modulator design and FIR parameters. A good modulator gives this noise an approximately "pink-blue" distribution, minimizing correlation with the signal
  3. The DAC's low-pass filter quality determines how much this out-of-band noise is attenuated — so DSD playback depends heavily on DAC design. Not DSD's fault; it's a system design matter

"The high-frequency extension is fake" is indeed correct. But what matters is DSD's SNR in the audible band, not the controllable noise layer above 50kHz.




Q7: DSD's editing inconvenience is a format flaw?


Misconception: DSD can't be edited directly, so it's unsuitable for any serious audio workflow.


Truth (current): This hasn't changed, and doesn't need to. DSD doesn't support linear operations (EQ, dynamics, reverb, editing) in the 1-bit domain — that's determined by the information structure of 1-bit PDM encoding. Today, the PCM→DSD conversion + SACD ISO authoring workflow in DpdoEngine is mature enough; DSD plays the role of final output format, not production format:


PCM recording → editing/mixing/mastering (PCM domain) → high-quality upsampling to DSD (DpdoEngine) → SACD ISO authoring (DpdoPackSACD) / DSF archive


This is a mature "produce in PCM, deliver in DSD" division. Streaming and portable devices go PCM; hi-fi playback and SACD releases go DSD. The two are not substitutes — they collaborate.




Q8: High-order modulators + long FIRs are just icing on the cake?


Misconception: The upgrade from MASH to 9th-order + 65536-tap FIR is "better DSD" — no real impact on ordinary users.


Truth (current): This is the point in this part most needing to be understood.


Why couldn't MASH-era DSD push noise shaping high? Because low-order modulators have shallow noise-shaping slopes; pushing DSD64's in-band noise low enough required increasing OSR — and doubling OSR doubles compute. On then-current hardware, DSD128 was the ceiling; going higher was impossible.


Why don't high-order modulators need OSR doubling? Because the ~110dB/oct noise-shaping slope means 110dB of noise attenuation per octave — at the same DSD64 OSR (64×), a 9th-order modulator achieves about 12dB better in-band SNR than 3rd-order MASH. This gap comes directly from noise-shaping efficiency, independent of rate.


Kakeya compression exists because 65536-tap FIR compute on AVX-512 reaches 14μs per sample, while DSD64's real-time window is only 0.35μs — over 40× over budget. Without Kakeya compressing FIR coefficients into polynomials (128th order, 1/512 compression), long FIRs can't run. It's not "optimized well" — it's "without it, this can't be done at all."


So "from MASH to 9th-order + long FIR" is not incremental improvement — it's a qualitative enabling leap. It doesn't change DSD's theoretical framework — it lets DSD's theoretical potential be nearly realized on consumer hardware for the first time.




Q9: Pursuing DSD requires a top-tier system?


Misconception: DSD only matters on systems worth hundreds of thousands; ordinary audiophiles shouldn't touch it.


Truth (current): The real threshold isn't "how expensive the system is" but "how good the DAC's DSD mode is." A rule of thumb:


Open your DAC's spec sheet (or find measured data) and check:

  • THD+N and dynamic range in DSD mode
  • Corresponding figures in PCM mode

If DSD mode beats PCM mode in THD+N and dynamic range by >3dB, then DSD upsampling on your system can produce an actually audible difference. This is unrelated to system price; it's about the quality gap between the DAC chip's DSD and PCM paths.


The most common consumer benefit scenario: a DAC with an ESS/Rivet/AK445x chip, where THD+N in DSD mode is typically 3-5dB better than PCM mode. Paired with a desktop speaker/headphone system with a noise floor below 25dBA, upsampling to DSD64 can create a stable audible difference.


The threshold isn't price — it's whether DAC, backend and environment match up.




Q10: Will all music become DSD in the future?


Misconception: After DSD technology improves, it will become the mainstream streaming format.


Truth (current): No. DSD replacing PCM in the mainstream consumer market is a zero-probability event — the reasons were laid out in Part 1: streaming doesn't support it, editing is unfriendly, hardware coverage is low.


But that's a separate question from "does DSD have value." Something doesn't need to become mainstream to justify its existence. SACD releases, hi-fi playback and recording archives — these scenarios need neither streaming coverage nor 2 billion users. They need a good toolchain — which was missing before high-order modulators + long FIRs, making DSD not good enough in those scenarios either. Now that the toolchain is in place, the "not good enough" obstacle is cleared.


DSD's reasonable niche today: high-fidelity playback with high-quality DSD upsampling on specific DAC architectures in quiet listening environments, plus SACD-format release and archiving.


It needs no more. It never did.




Conclusion: Part 1 and Part 2 Belong Together


The 130 items in Part 1 dissected MASH-era DSD limitations. Those criticisms each held under then-current implementation. Part 2 says something simple:


When the toolchain is good enough — high-order modulator + long-tap FIR + Kakeya compression + SIMD acceleration — DSD is no longer the DSD Part 1 described. Many "innate flaws" are actually "implementation flaws."

This is not saying DSD beats PCM. A rational user should:

  1. Read Part 1 to understand what MASH-era DSD did and why it was criticized
  2. Read Part 2 to understand what the current toolchain changed and what remains unchanged
  3. Then decide whether to use DSD in their system based on DAC architecture → backend → listening environment → program material

Not good because of Kakeya. Not good because of 9th order. Good because your system happens to match on this chain.