It’s finally time for me to wrap up this ZX Spectrum series, with a look at how to get the simple 1-bit beeper on every Spectrum model to emit high-quality digital audio. The results are quite good: it’s the last sound played on last week’s demo reel, and the difference between it and the simple PCM playback system before it is night and day.
I have also, at this point, gotten hardware verification of the demo program on both period hardware and modern clones: the vagaries of the Internet mean that identification is usually by handle, but “gmc” via Mastodon was able to run it on a 48K Spectrum, and Tom Harte got results from the Omni 128 system, which is a modern rebuild of the system. Emulation support is also quite widespread; I used Fuse to capture the demo above, but it also works fine on, at minimum, EightyOne, Zesarux, Spectaculator, and Clock Signal.
The practicality of the technique is a bit limited; the sound is very soft on hardware without an external amplifier, and the RAM required for samples of the length and quality I’m experimenting is effectively “all of it.” But it does work.
Getting it to work, though, was a journey. It took me two complete rewrites before I landed on something that worked at all, and even then it barely fit into the timing constraints I’d set for myself. I had to pull out several techniques I haven’t used before on this system. I mostly work in assembly language here, just because it’s generally more comfortable to work there on the systems I’m playing with—as long as there’s suitable compiler support like there was on DOS you usually can do just fine in C. Today, I would not have done just fine in C. Even leaving aside how I needed exactly timed instruction sequences for the audio effects, the timing constraints were so tight and the register pressure so fierce that I needed direct chip access to make it work.
This isn’t even the hardest form of the problem, either, and there are some straightforward ways to improve the technique I use here which I don’t need for my tests. I’ll be wrapping up today by looking at Dmitry Milk’s Spectrum sound engines as well as Michael J. Mahon’s RTSynth system for the Apple IIe; they’re solving a slightly different problem from the one I faced, so even though our sound output approaches are similar, they end up more similar to each other than to my playback system despite being on totally different CPUs.
Prior Work
I first ran into 1-bit digital samples by way of early-1990s DOS games. In particular, the illustrated text adventure Eric the Unready could play sound effects through the PC speaker with a technology its manual called “RealSound,” and Star Control II: The Ur-Quan Masters used Amiga-style tracker music for its soundtrack and could deliver it out your PC speaker just as cheerfully as it could your SoundBlaster.
The PC’s 1-bit speaker is markedly easier to program than the Spectrum or the Apple II; I went through both simple tone generation and PWM playback about a decade ago. Everything there is still functional and accurate as far as I know, but we don’t really need it for what we’re doing today. Its approaches do motivate our design and strategy, though.
The fundamental physical principles here are that circuits have capacitance and speakers have mass, so it takes time for a 1-bit signal to actually travel from 0 to 1 or vice versa, and then it takes more time for the speaker’s cone to actually move across its own space to actually produce the sound waves we hear. If we interrupt the signal partway through, the signal and the speaker won’t make it all the way across the space and we’ll get a more finely-varying sound wave.
This was unreasonably easy on the PC—while the speaker’s setting can be directly commanded via an I/O port the way the Spectrum’s is, it could also be driven from a highly-programmable hardware timer that ran independent of the system clock interrupt. This meant that instead of having to constantly baby-sit the beeper with cycle-counted code, you could just configure timing information only when you needed to change the signal being sent. For normal beeping that’s just when you’re changing frequencies, but you could also configure one-shot pulse widths with microsecond precision. That meant that you could set a CPU timer interrupt at 16kHz, or whatever your sample rate was, and then set the sound timer to run for a time based on that sample’s PCM value. The timer chip itself thus ended up serving as a sort of digital-analog converter.
We don’t have a timer on the Spectrum, either for enforcing a sample rate or for configuring a pulse width, but we should be able to do both of those things in software directly. The general idea will be to have a loop that lasts as long as each sample, and it will turn the speaker on then off each loop, with two short but variable-length delays on either side of “turn the speaker off again.” The first delay sets the pulse width, and the second makes sure that the total block of code takes the same amount of time no matter how wide the pulse was.
Misguided Attempt 1: Actual Delay Loops
My first implementation attempt started with some napkin math.
- My PC playback system accepted pulse widths that ranged from 1 to 127 microseconds.
- An 8kHz sample rate means that we need a pulse every 125 microseconds. This is the first point where I lift an eyebrow; my test program used a 16kHz sample rate which means I would be overrunning myself if the sample were really using the whole range. Perhaps my sample was softer than I thought. All the same, it’s broadly the same sample here so we should be in reasonable shape.
- The Spectrum CPU is about 3.5 MHz, so 125 microseconds is 438 cycles.
- 438 cycles is comfortable enough that it fits 15 iterations of our default delay loop: We need 18 cycles to turn the speaker off after turning it on, so our delay range goes up to
13*15+2+18=215cycles. That’s close enough to the 16kHz mark itself that I’m motivated to stay in 8kHz mode for my first experiments.
I already have a version of the audio clip that’s 7-bit audio at 8kHz; I used it for the NES awhile ago. There’s plenty of time in this 8kHz window to just use that data directly and convert it into 4-bit PCM on the fly. The pseudocode for the first draft turned into this:
- Read a sample byte, N, and truncate it to 4 bits.
- Turn on the speaker.
- Run N iterations of a do-nothing loop.
- Turn off the speaker.
- Run 16-N iterations of the same do-nothing loop.
- Increment sample pointer and quit if we’re at the end of the sample.
- Wait an appropriate amount of time until we hit our 125 microsecond mark and then return to step 1.
The key insight here is that while steps 3 and 5 might take variable amounts of time individually, together they would always take the same amount.
Unfortunately for me, this build didn’t work at all; all I got was a piercing high-pitched whistle. I concluded at the time that I was delaying too long—we were not succeeding in interrupting the speaker mid-transition and were instead just producing an 8kHz tone. Truncating to shorter bit widths didn’t really help, either.
Misguided Attempt 2: Dynamic Delays
My first thought was that I was having a “constant factor” problem—the sample was pretty soft which meant mostly samples in the 6-10 range, and maybe I needed to get those numbers down. I didn’t really know what I should be aiming for, though, and jumps of 13 cycles seemed a bit excessive. I decided to rework it to be finer grained. Way back at the start of the blog I looked at fine-tuned programmable delays on the 6502, borrowing a technique I first saw on the Atari 2600. There were some CPU-specific shenanigans there that we can’t replicate on the Sinclair’s Z80, but we also don’t really need to. Those tricks allowed the 1MHz 6502 to get cycle-exact precision, and that means 4-cycle precision will be good enough for us.
The general technique was to lay down a long stream of delay instructions and achieve a variable delay by doing a computed jump into the middle of the sequence. If we have that “long stream of delay instructions” just all be NOP instructions, we’ll get 1.14-microsecond precision in the changes of our delay.
The result was still mostly that 8kHz tone, but now I could just barely hear a whisper of the sound clip playing behind the tone. This was a hint that the technique would work, but still no sense that it would function acceptably or even that any emulators would produce good results, or if success or failure on emulation could mean anything about what real hardware would do. At that point I was stuck and started asking more knowledgeable enthusiasts.
This was where I learned about Dmitry Milk’s demos, and that they worked both on hardware and in the Fuse emulator (which was the one I was using). I also learned that the physical speaker on the Spectrum’s beeper was much smaller and lighter than the on on the PC. That at least lent some credibility to the idea that I was wildly overshooting the length of my puzzles. This also meant that having each value in my 4-bit signal correspond to an extra 4 cycles gave us a dynamic range of about 20 microseconds on our pulses, and that might be fine as is.
I now knew that the goal was possible. What I needed now were some actual numbers to aim at.
Finding a Path
Now that I had confirmation that Fuse was good enough to produce these effects, I dug into its source code to see if I could find information on how long it took the audio signal to go from low to high. What I found was the Blip Buffer system with some low-pass filters afterwards to smooth the signal. That’s reason to believe the impulse response would let our technique work, but I was looking for microsecond counts, not impulse response or equalizer configurations. I eventually took the easy way out, logging the 16-bit 44.1kHz audio output on its way to the system speakers and just looking at a BEEP command in an editor. That gave me a number range: a transition ran 4 to 7 samples, or 97-170 microseconds. That was a little surprising—it’s longer than our 8kHz sample rate, but it’s also in the range we’d been aiming at. This, I think, was simple foolishness on my part: I had forgotten that the system’s physical inertia applied not merely to the part where the speaker starts moving, but the part where it slows down and changes direction, and for much of that journey it will still be moving forward.
At this point, I wanted to have my pulses be as short as I could make them: turn the speaker on, run 0-15 NOP instructions, and turn the speaker off. My pulses would be between 18 and 78 cycles long.
My first attempt at this was bad. I laid down 17 no-ops, and then just before entering it, would overwrite two bytes of it with my OUT ($FE),A instruction, rewriting them back into NOP instructions on the way out. This was slow enough I had to do serious calculations to make sure I still fit in 8 kHz, and it also was very fragile and difficult to debug.
Before I actually finished debugged it, I came up with a much better idea and scrapped it. I didn’t need self-modifying code: there are only 16 possible values and the logic is very rigid, so I could write a macro that created each possibility one at a time, and then swap between them with a jump table. The dispatch would look something like this, assuming A was a value between 0 and 15:
add a ; double value
ld hl,.table
add l ; align .table so this never carries
ld l,a
ld e,(hl) ; Load .table[A*2] into DE
inc hl
ld d,(hl)
ex de,hl ; Put it in HL
jp (hl) ; And jump to it
.continue:
;; Delay, rest of loop, etc..table: defw .s0,.s1,.s2,.s3,.s4,.s5,.s6,.s7
defw .s8,.s9,.s10,.s11,.s12,.s13,.s14,.s15;; Each value is a macro
.s0: pcmstep 0
.s1: pcmstep 1
.s2: pcmstep 2
;; ... and so on...
The pcmstep macro would look something like this:
macro pcmstep i
ld a,$17 ; Turn speaker on
out ($fe),a
xor $10 ; Value for speaker off
repeat i
nop
endrepeat
out ($fe),a ; Turn speaker off
repeat 15-i
nop
endrepeat
jp .continue
endmacro
I wasn’t sure what to do about volume yet, so instead of trying to actually read my sample clip I just had it produce a sawtooth wave, incrementing a value from 0 to 15 at a rate of 1 per sample. This was a simple loop to set up, and a 8kHz it produced a plausible-sounding sawtooth-wave buzz at a reasonable frequency (500 Hz is a bit flat of High C, a completely reasonable tone), and the captured waveform looked wonderful.

I was now sure the technique was sound. Now I just had to finish the job.
From Sawtooth to Sample
The main challenge in advancing from a simple procedural sound wave to a proper sample played out of a memory buffer is that there’s too much competition for our 16-bit registers. We need a 16-bit counter for the outermost loop, a pointer into the sample buffer, a pointer into the jump table, a pointer to hold the value read out of the jump table, and we also need to keep the B register in particular free to be my count register when doing precise delays.
As long as I stayed at 8kHz, I had time to spill out DE and HL and B to RAM and restore them at appropriate times. I realized while writing this, though, that I could instead use the EXX instruction to swap between the normal and shadow registers and use the shadow registers for the last three of those values. If I did this I would have enough time to actually keep to a 16kHz cadence!
I would not, however, have enough memory to keep to it. My clip at 16kHz is about 40KB long, and I can’t fit that in the 32KB of reliable doesn’t-fight-with-the-video-chip memory. I also found I didn’t have time to store the data compressed and just unpack it at run time. I could, however, just have my pcmstep macro send each pulse twice.
At this point I was almost done: turning the sawtooth wave into a sample gave me a good result, but my sound clip was still inaudible. The solution was pretty clearly that I’d need to amplify it, but there’s no need to preprocess my wave data when I could just alter the jump table to do the amplification and clipping for me! I subtracted five from each value, multiplied by three, and then capped the extremes at 0 and 15 themselves. I now had the sample coming clearly out of the emulator, but there was one final little complication; when my playback program exited, the whole system crashed, often wildly distorting the display as it did so.
It turns out that the values in the shadow registers are actually important to the system ROM. Fortunately, that’s easy to deal with: I just added a little extra code around the main playback call to preserve the registers so that the next time an interrupt happened it didn’t explode the whole system.
The Final Playback Code
The routine takes a pointer to the sound buffer in HL and the length of that buffer in BC. The top-level function blocks interrupts for the duration and then preserves and restores the shadow registers, calling out to the main loop as a subroutine (which lets me use the cheaper RET Z to leave the loop later on).
playwav:
di ; Disable interrupts
exx ; Preserve shadow registers
push bc
push de
push hl
exx
call .lp ; Play sound
exx ; Restore shadow registers
pop hl
pop de
pop bc
exx ; Make the shadow registers shadow again
ei ; Re-enable interrupts
ret
The pcmstep macro has also gotten more sophisticated. The overall loop take 438 cycles to run through, and I want my OUT ($FE),A instructions to end on cycles 219 and 438. I set up my cycle counting so that the first thing we do in the macro is turn on the speaker; that means that we enter the macro at cycle 438-11=427.
macro pcmstep i
out ($fe),a ; + 11 (438)
xor $10 ; + 7 ( 7)
repeat i
nop ; 4*i
endrepeat
out ($fe),a ; 11
repeat 15-i ; 60-4*i
nop
endrepeat ; + 71 ( 78)
ld b,9 ; + 7 ( 85)
1 djnz 1B ; +112 (197)
nop ; + 4 (201)
xor $10 ; + 7 (208)
out ($fe),a ; + 11 (219)
xor $10 ; + 7 (226)
repeat i
nop
endrepeat
out ($fe),a
repeat 15-i
nop
endrepeat ; + 71 (297)
jp .step ; + 10 (307)
endmacro
We leave the macro at cycle 307 of the next unit. In the event, it turns out that we only have 19 cycles to spare after doing all the rest of the work of the loop, so I take a weird shortcut: I define the end of the loop first. The .step label, replacing the earlier .continue, is actually defined above .lp and falls through into it:
.step: ;; Enter .step on cycle 307
ld b,0 ; + 7 (314) Delay
ld b,0 ; + 7 (321)
exx ; + 4 (325) Restore counter/buffer ptr
dec bc ; + 6 (331) Decrement/test counter
ld a,b ; + 4 (335)
or c ; + 4 (339)
ret z ; + 5 (344) Return if zero
;; Fall through into .lp
The .lp code itself is pretty close to the earlier code, but we have to account for the shadow register swap, and we also need to convert the 7-bit PCM data into a value between 0 and 15 that we then double. We can make that operation simpler by just shifting right twice and masking out all but what were once the top four bits:
.lp: ld a,(hl) ; + 7 (351)
inc hl ; + 6 (357)
rrca ; + 4 (361)
rrca ; + 4 (365)
and $1e ; + 7 (372)
I can also combine some of the other operations together when computing the jump table address, so this code is also a bit simplified.
exx ; + 4 (376)
ld h,.table / 256 ; + 7 (383)
add .table & $ff ; + 7 (390)
ld l,a ; + 4 (394)
ld e,(hl) ; + 7 (401)
inc l ; + 4 (405)
ld d,(hl) ; + 7 (412)
ex de,hl ; + 4 (416)
ld a,$17 ; + 7 (423)
jp (hl) ; + 4 (427)
To make sure that all my additions and increments never carry, I align the jump table to a 32-byte boundary. We also see the “built-in amplification” here.
align 32
.table: dw .s0,.s0,.s0,.s0
dw .s0,.s0,.s3,.s6
dw .s9,.s12,.s15,.s15
dw .s15,.s15,.s15,.s15
The sound wave looks… pretty respectable, overall, at least once I amplify it so that it’s visible in Audacity.

Possible Future Work
This is a good place for me to wrap up, but it’s certainly not the end of all possible roads.
I’ve put this sound clip through a lot of different systems at this point, and that’s given me a number of ways of encoding the data. This is usually demanded by the hardware itself, but sometimes I can be flexible about it. On the C64, I ended up using a scheme that run-length encoded 4-bit PCM data and could get 16 kHz playback out of even the 1MHz 6502 chip. I’d like to use that scheme here as well, but I face the difficulty of only having 14 cycles to spare in my loop.
I’m pretty sure the solution here is to break up the macro system and instead put the logic into the gaps between the two pulses. Since we’ve split each possible output into its own standalone branch, we should have plenty of room to interleave our logic between the pulses. I didn’t actually do this for this project because I didn’t have to to prove out the technique.
The other main option would be to go the other direction: start dividing the signal more finely. After amplification I’m using very little of the actual space I coded up, and we could probably at least make it to 5 bit sound or even 6 without changing much of the existing machinery. This is another case I just skipped out on because 4-bit sounded fine.
More ambitiously, the playback routine could be built into something more like a full music engine. This would be similar to the Softsoniq engine I wrote for the Tandy Color Computer and its 6-bit PCM output. This ends up being a different problem than simply playing clips, as I learned back then, and I feel no need to try to expand it in part because it’s already been done twice.
Which brings us to…
Related Work
Dimitry Milk’s demos Beeperdrive and We are Vocoders were built on just such a wavetable synthesizer, and he has published the “CORE5” engine code on his GitHub. I mostly kept this code at arms length, and in particular I did not look at it until after I’d perfected my own system here—he’s publishing under a license that’s incompatible with what I use and I didn’t want any cross-contamination. But even briefly scanning some of the code and reading the explanations is enough to see the general technique.
It is, overall, pretty close to a hybrid of SoftSoniq and the “split into different process paths” technique I finally settled on, just pushed much harder. There are five voices each with their own sample counter, with instruments limited to 256 samples each. (Unless “We Are Vocoders” enhanced that, this makes the voice clips in that demo much more impressive, if it’s using this engine! It’s possible an extra long-clip voice was mixed in though; it wouldn’t seriously change the technique.) There are 256 branches in the table here instead of the 16 I used, as well; the synthesizer admits full 8-bit audio samples.
That’s far too much typing to do by hand, even with macros—the engine appears to rely on dedicated code generators written in Python. Since I was keeping the code at arm’s length I don’t know if there’s any more logic interleaved in there, but I imagine that some of that is happening. The labels make it pretty clear that these options are also themselves cycle-perfect instead of the 4-cycle jumps I programmed in or the 12-cycle jumps I found myself actually using. The final generated jump table also reveals a fun additional consequence of the jump table technique I didn’t use but which would make a synthesizer like this much friendlier: there’s no actual requirement for your jump table entries to be in order, and the difference between unsigned and signed sample inputs is just a question of how you order the entries in the table.
There is also one other generic Z80 technique that showed up when I first looked at the demo running in a debugger: he uses a much faster technique for going through the jump table. Both of us compute an address in HL that holds the address of the location in the jump table that holds the actual address we want to go to. I used a fairly traditional technique of getting that value into DE and then HL for the jump:
ld e,(hl)
inc hl
ld d,(hl)
ex de,hl
jp (hl)
Dmitry’s code, however, accomplishes the same thing like this:
ld sp,hl
ret
That’s much faster, it leaves DE free, but since I do not have liquid nitrogen for blood it did not occur to me to repurpose the stack pointer as a place from which to pop indirect call targets. It’s a good thing interrupts are disabled.
(Once again, I didn’t look too closely at this actual code, but I don’t think this is actually as destructive as it looks. You’d just need to spill it to an absolute address in RAM somewhere and restore it when you’re done, same as I did with the shadow registers.)
As fun as all this is, though, it also wasn’t what inspired me to attempt it in the first place. I had originally been inspired by the 2017 album Class Apples, by the band 8-Bit Weapon, which is a number of classical and classical-adjacent songs played through the beepers of Apple II computers. The playback here seems to be based on even earlier work by Michael J. Mahon, which in turn goes back to work in early 2000s, and ultimately a much older system simply called “SoftDAC” which seems to have been from the Apple II’s actual heyday. The physics should all be identical to what we’ve seen here, but the logic will be slightly simplified because Apple’s control of the 1-bit speaker involves writing a memory location to toggle the speaker instead of actually needing two different values to turn it on or off. That makes this easier, though it would make last week’s chord-players a bit more challenging, I think. It’s not clear from the album description, but it sounds like it might also not have a true mixer the way Softsoniq and CORE5 do, with multitracking recorded used to assemble the final output.
Mahon’s writeup of the RTSynth system indicates that he also felt it necessary to double his 8kHz pulses in the synthesizer to get good results, and, well, after hearing the whine of the 8kHz carrier wave on the Spectrum, it’s hard to disagree.
A Fair Shake, Given
Back in May I decided to “give the ZX Spectrum a fair shake” and put sime time into putting the system through its paces properly. My initial TODO list looked pretty short:
- Do a little tour in BASIC of the things the system would offer ordinary users.
- For all of the things that BASIC let you do with the hardware, show how to also do those things in machine language directly.
- Go beyond BASIC to do “sprite” systems on its bitmap, ideally by porting my shooting gallery “Rosetta Stone” program.
- Play around a little with the sound chip, to get parity with the PC Speaker’s system and maybe see if I could replicate some of the tricks folks got up to back in the day.
It has been The Sinclair Show pretty much all the time from then on, with only occasional breaks that usually still wound up tying into the main projects somehow. But every one of my initial boxes is now properly checked, and I feel comfortable putting the Spectrum back on the shelf again until I have another idea where it’s a comfortable fit.
Looking back over the tour, it’s actually quite interesting to see how it’s different from many of the other systems I look at. A lot of C64 work relies on working around surprising but predictable behaviors in its hardware that were never really formally specified as part of its operation, or on taking advantage of similarly predictable behaviors when the chips are pushed outside of the specifications of their normal operation. Advanced Spectrum work, even the audio effects here, were not particularly like that. What you see is what you get, and you get the things you want simply by asking for them directly. The challenge, and the surprising rewards, come from determining what these capabilities permit you to ask for and what constraints they incidentally impose on you. It’s quite a different story from the C64 or the Amiga, but it got more common as systems got more powerful.
The Spectrum’s overall capability is a very comfortable “sweet spot” of raw power and modest specifications, and of its potential peers, I think it also did the best in terms of presenting its capabilities to all its users in a friendly and usable way.
Good stuff. But all the same, I think it’s time for a break.