Two bug reports, one bug — rebuilding the CW decoder's timing
Two CW decoder complaints that looked unrelated turned out to be the same defect pointing in opposite directions. Here's the root cause, the rebuild, and how I tested a decoder without a radio.
Two complaints came in about the CW decoder. The first: nothing decodes above 25 WPM. The second: single letters shatter into their own elements — M comes out as T T.
They read like unrelated problems. They’re the same bug, pointing in opposite directions, and it’s a bug I wrote on purpose and then described as a feature in the original post about building the decoder. Here’s what I said at the time:
WPM range clamping prevents the adaptive timing from drifting too far. If the estimated dit duration would imply a WPM outside the configured range, it gets clamped. This stops the decoder from locking onto noise or drifting wildly during pauses.
Every sentence there is true. The conclusion was still wrong.
What the clamp actually did
The decoder measures dit and dah durations as they arrive and keeps a rolling average. That average is the basis for every timing decision it makes — whether a gap is between elements, between letters, or between words. The clamp took that measured average and squeezed it into the range implied by the two WPM sliders:
let avg = ditDurations.reduce((a, b) => a + b, 0) / ditDurations.length;
const minDit = 1200 / wMax;
const maxDit = 1200 / wMin;
return Math.max(minDit, Math.min(maxDit, avg));
And wMax was hard-capped at 25, in the function and in the slider markup both.
So the decoder would measure the sender correctly and then throw the measurement away. Every gap threshold downstream was derived from the clamped number rather than the signal. That’s the whole defect. It’s not that the estimate drifted — it’s that the estimate was never allowed to be right.
Symptom one — the wall at 25 WPM
With wMax pinned at 25, the dit estimate can never go below 1200 / 25 = 48 ms. The letter-gap threshold is 2.5 dits, so it’s stuck at 120 ms.
At 25 WPM a real letter gap is 144 ms, comfortably over the threshold. At 30 WPM it’s 120 ms — exactly on it. At 35 WPM it’s 103 ms, and the decoder never sees a letter boundary at all. Elements pile into one ever-growing string that matches nothing in the lookup table.
The tell is that element classification still works up there. The dit/dah split lands at 84 ms, which cleanly separates a 40 ms dit from a 120 ms dah. Only the framing is broken. That’s why it looked tantalisingly close to working rather than obviously dead, and why the output was a run of □ — the decoder’s “I got elements but they spell nothing” character — instead of silence.
The live WPM readout was computed from the same clamped value, so it couldn’t display above 25 either. Someone copying a 35 WPM sender saw a confident 25 and no hint that anything was out of range. The instrument that should have exposed the bug was downstream of it.
Symptom two — M becoming T T
Same mechanism, opposite direction. If the sender is slower than the WPM minimum, the dit estimate is pinned too small.
A character shatters when a one-dit gap between its elements exceeds the letter-gap threshold. Crank wpmMin to 25 and that threshold sits at 120 ms. A sender at 8 WPM has a dit of 150 ms. Every single gap inside every single letter now reads as a letter boundary, so M decodes as T T, O as T T T, S as E E E.
The two reports were almost certainly one person in one session: chasing a fast sender by raising wpmMin, then finding slow copy had broken. One slider, both failure modes.
The fix
The structural problem is that gap thresholds derived from what the operator declared rather than what the decoder measured. So:
The WPM range now only rejects outliers. It decides whether a duration is plausible enough to enter the timing window. It has no say in the thresholds. Setting it wide costs you almost nothing now, which is why the default range widened to 5–40 WPM and the sliders go to 60.
Timing comes from clustering, not averaging. The old code kept separate dit and dah buffers, which meant a misclassified element was filed in the wrong bucket permanently and dragged the threshold with it — a guess compounding a guess, with no way out except the rolling window ageing it off. There’s now one window holding every element duration, split by a small 1-D 2-means pass. Cluster medians rather than means, so a single noise-truncated element can’t drag the estimate, and the split is validated: if the two clusters aren’t roughly a dit and a dah apart, the window is single-mode — every element in it is the same symbol — and treating it as two clusters would invent a distinction that isn’t there. That check is precisely what stops M shattering.
The dit/dah threshold is the geometric mean of the two estimates rather than the arithmetic one. Durations cluster roughly log-normally, and it gives a saner bootstrap: 1.73 dits instead of 2.
Two things I only found by testing
I couldn’t test this against a radio at a known speed on demand, so I wrote a simulator that generates ideal Morse at any WPM, quantises every edge to the worklet’s 13.3 ms block grid — that quantisation is what made the original failures feel intermittent — and runs it through both the old and new decode paths. A few hundred lines, and worth every one of them.
It immediately showed two problems the design hadn’t anticipated, both at the start of a transmission before the estimator has anything to work with.
The first character was being lost. The decoder classified each element the instant it arrived and never revisited it. But the first element arrives when the decoder knows nothing — so it gets classified against a guess, and by the time the estimate is good the character is already wrong and gone. The in-progress character is now re-derived from its stored durations every time the estimate improves. S sent at 8 WPM with the sliders wrong used to come out as D; now it’s S.
The gaps were free information I was ignoring. A gap between two elements of the same letter is exactly one dit. That’s a direct dit measurement available partway through the very first character, long before there are enough elements to cluster. The shortest completed gap now seeds the estimate during bootstrap.
Neither is exotic. Both are obvious in hindsight. Neither occurred to me until I could watch the decoder fail the same way twenty-five times in a row.
The results
Accuracy on a test transmission, scored by edit distance with 10% element jitter, averaged over 25 runs — old decoder first, new one second:
- 5 WPM — 6% → 100%
- 8 to 18 WPM — 100% → 100%
- 25 WPM — 4% → 93%
- 30 WPM — 4% → 99%
- 45 WPM — 4% → 98%
- 55 WPM — 2% → 68%
The 4% figures aren’t a typo. Above the clamp the old decoder produced essentially nothing usable, which is exactly what was reported.
The other bug — it wasn’t the filters
Separately, the decoder felt sluggish while running. The obvious suspect was the DSP: a Goertzel filter and a biquad chewing through 48,000 samples a second sounds expensive.
It isn’t, and it’s worth being precise about why. The Goertzel is one multiply and two adds per sample. With a block size of sampleRate / 75 it runs 75 blocks a second of 640 iterations each — 48,000 iterations per second, or exactly one per input sample. The biquad is native code running a single second-order section. Together they’re well under 1% of one core.
The cost was all on the main thread, competing with layout and paint:
The waterfall was blitting the canvas onto itself. To scroll down one pixel it drew the canvas into itself, then wrote the new row with putImageData. That combination forces a GPU-to-CPU readback, a CPU-side write, and a re-upload — every frame, defeating canvas acceleration entirely, plus a fresh ImageData allocation each time for the GC to deal with. It now keeps one persistent buffer, scrolls it with data.copyWithin() — a memmove, no GPU involvement — and does a single upload per frame.
The DOM was being written at 75 Hz. The magnitude handler fires off the audio worklet’s message port, so it isn’t aligned to frames and runs more often than the display refreshes. Each call did a dozen DOM reads interleaved with style writes, re-parsing slider values that only change when someone drags them. Textbook layout thrash. Slider values are now cached in module scope and updated by their input handlers; the meter writes moved into the animation frame loop. The audio callback no longer touches anything that affects layout.
There was also a trap in there for anyone profiling it. The worklet contains a scan mode that runs a full Goertzel bank — every candidate frequency, every sample, inside one render quantum. That genuinely would be expensive. It has never executed once: the auto-detect feature was reimplemented on the main thread using FFT data and the message that would trigger it is never sent. Anyone opening that file looking for the bottleneck would find the nested loop and stop there. It’s gone from the new worklet.
What’s still broken
The block size sets a hard floor on timing resolution. At 13.3 ms per block, a 60 WPM dit is 20 ms — a block and a half — so it can measure as one block or two with nothing in between. That’s why 55 WPM sits at 68% and won’t go much higher. Shortening the block would fix it and simultaneously widen the Goertzel bin, costing the selectivity that makes the thing usable on a crowded band. That’s a real trade, not an oversight, and I haven’t decided which side of it to be on.
The honest caveat: everything above was verified by build and by simulation. The rendering changes in particular have never been exercised against a real signal — no performance capture, no on-air testing. Ideal Morse with synthetic jitter is not a pileup on 40m. I expect real RF to find things the simulator can’t.
That’s why it’s a beta and why it’s on its own page.
Try it
The beta is at skipzone.co.uk/tools/cw-decoder-v2. The original is untouched at /tools/cw-decoder and stays there — separate page, separate code, separate worklet, so the beta can’t break it.
If it mis-decodes, that’s useful. What helps most: the speed you were sending at, where the WPM sliders were, and what you got versus what was sent. The old bug was invisible partly because the readout that should have shown it was computed from the broken value — so if the live WPM disagrees with reality, I especially want to know.
Thanks to Brad — a brilliant electronics and software engineer, and a great friend — for both reports, and for the rigour behind them. He didn’t stop at “the decoder feels slow”. He worked out the arithmetic, checked the gap thresholds against real sender speeds, spotted that two complaints which looked unrelated were one defect, and flagged that the dead scan-mode loop would send anyone profiling the worklet down the wrong path entirely.
His second report saved me from optimising the one part of the pipeline that was already fast. His first found a bug I had written, documented, and then defended in print. That’s the kind of report you rarely get and always want, and this rebuild is his as much as mine.
73 de MM7IUY