
There is something both exciting and comfortable about the introduction of the 65816 into the Atari world. Exciting because it adds more raw capability to be explored and used in interesting new ways. Comfortable because it remains simple and true to the 6502 philosophy. Indeed, anyone still programming the 6502 in the mid-1990s probably isn't going to be too keen on something that overturns their way of thinking and adds unwanted complexity. Heck, you can slide into the 65816 without even knowing it; if you ignore the 16-bit stuff, it basically is a 6502.
65C02 Pleasantries
Well, not really a 6502, but a 65C02. The 65C02 adds some tweaks to the 6502 instruction set. Tweaks that may be minor, but will strike a note of rightness with longtime 6502 programmers.
You can now push and pull the X and Y registers directly with PHX, PLX, PHY, and PLY, and transfer data between them with TXY and TYX. You can increment the accumulator with INC A (or INA as it is sometimes called). You can zero a memory location without loading a register with zero using STZ-Store Zero. There's a two-byte unconditional branch with a snickery name: BRA. And there are two nifty bit manipulation instructions that, unfortunately, are somewhat restrictive. TSB (Test and Set Bits) sets the bits of the operand that have corresponding bits set in the accumulator. TRB (Test and Reset Bits) clears the bits of the operand that have corresponding bits set in the accumulator. An example is definitely in order:
LDA #1
TSB TEMP
TSB TEMP
This sets the low order bit of the value stored at TEMP Unfortunately, these instructions are somewhat awkward to use in common situations, because indexed addressing modes are not available.
The last gleeful tweak is a variation on the indirect indexed addressing mode-good of "(ZPAGE),Y". The new addressing mode is simply "indirect." It drops the comma-Y from the indirect indexed mode and works as you'd expect. "LDA (TEMP)", where TEMP is a zero page variable, uses the values at TEMP and TEMP+1 as an address and fetches the value at that address.
The Gateway To 16-Bits
The 65C02 enhancements are old hat to anyone who ever got bored and read the obscure portions of the MAC/65 manual. Let's get to the good stuff!
I expected the back-door to 16-bit mode to be through the mysterious unused bit in the processor status register, as this bit was always mentioned as being for future expansion, but this is not the case. The truth is even more obscure.
Deep inside the 65816 somewhere is a bit that indicates whether the processor is in 65C02 mode or 65816 mode. But this bit is not addressable; not directly anyway. There is a quirky little instruction, XCE, that exchanges the carry bit and the "emulation" bit, as the hidden bit is called. When it is set, the processor acts as a tried and true 65C02. When it is clear, it becomes a full-bore 16-bit processor. Whisper the incantation:
CLC
XCE
and you're on your way.XCE
Big Deal #1:
24-Bit Addressing
There are two major features you have at your disposal when you switch to native 65816 mode: 24-bit addressing and 16-bit processing. 24bit addressing is the showier of the two. It's definitely useful, but considering the impressive software that has been written in 32K or even 16K, it's a bit flighty for many applications. In a Super Nintendo game I wrote in 1994, I reached out of the lower 64K in fewer than twenty of 12,000 lines of code. At the risk of miffing those few programmers that do need to address monstrous amounts of data, I'm going to only lightly touch the 24bit topic so I can spend more time on Big Deal #2.
You can directly read from or write to any address in the 24-bit address range with a single instruction. The greater than sign marks 24-bit addresses so "LDA #>$020032" grabs data from that 6-digit (!) address. The restriction is that only a few addressing modes can be used with 24-bit addresses: absolute, indexed X, indirect indexed, and the new indirect mode. These last two modes use brackets instead of parentheses and reference 3-byte addresses in page zero instead of the usual 2-byte:
LDY #8
LDA [TEMP],Y
LDA [TEMP],Y
In true 6502 fashion, the highest byte of the address is stored in TEMP+2.
Now here's where things get muddled a bit. The 65816, in some situations, "sees" memory as being made up of 64K "banks." This works out well in a nifty visual sorta way. Given a 24-bit address, the left most two hex digits are the bank number. "$01FFFF," refers to address $FFFF in bank $01. This bank nonsense is mostly to allow addressing shortcuts. There is a register that contains the bank number to use for 16-bit absolute addresses. So when you use a traditional absolute instruction, like LDA $A000, the default bank number is tacked on to the front, making a full 24-bit address. This register defaults to $00, which is one of the ways the 65816 emulates a 6502.
Big Deal #2:
16-Bit Processing
After executing a CLC/XCE pair, the 65816 still acts as an 8-bit processor. The A, X, and Y registers are all 8-bits, as are all accesses to memory. But be assured, things have changed. All of a sudden you have access to new instructions and addressing modes. The stack pointer is now 16-bits, letting you move the stack anywhere in the first 64K of memory. And there is a more subtle change as well: the b (break) bit of the processor status register has been removed and both it and the aforementioned unused bit now have new purposes. They control the sizes of the accumulator and index (X and Y) registers.
There are a couple of ways to directly change bits in the processor status register of the 6502. There's PLP, which pops the register from the stack, and the special purpose instructions like CLC and SED which affect specific bits. With the 65816 there's another option: the SEP and REP instructions. SEP sets bits in the P register and REP clears bits. Each instruction takes a bytesized operand representing the bits to mess with.
These instructions are most often used to change the new bits of the P register: the bit corresponding to $10 which controls the size of memory and the accumulator, and the bit corresponding to $20 which controls the size of the index registers. A set bit means 8-bit operation and clear means 16-bit. REP $30 and all your registers are 16-bits.
Now that we've executed this funky REP instruction, what happened to the processor? For starters, all load and store operations now involve 16-bits of data. LDA $80 loads the bytes at $80 and $81 into the accumulator. All other memory referencing instructions, like INC and PLA, are now 16-bit as well. LDA $1000,X adds the 16-bit value in X to $1000 then loads the word at this address into the accumulator. In other words, business as usual, except that everything affects words instead of bytes.
Immediate values pose an interesting problem. Yes, they are now 16-bits-you can say LDA #$FFFF-but this requires three bytes of object code, instead of the usual two. With an 8-bit accumulator, LDA #0 is two bytes long. With a 16-bit accumulator, it is 3-bytes. In both cases the opcode is $A9. How does the assembler know what the current size of the accumulator is so it can generate the proper object code? You have to tell it. Whenever you change the size of the accumulator or index registers, you have to accompany it with a directive to let the assembler know the new sizes. This usually looks something like this:
REP $10 ;Switch to 16-bit accumulator
LONGA ON ;Tell assembler this is so
That's ugly, I admit, but you usually don't bink back and forth between sizes, at least not frequently. If you think this size-switching stuff is awkward, look at it this way: instead of using bits within each instruction to determine the size of the operand, causing all instructions to be larger, they have been factored out into separate instructions. This keeps all opcodes the same size as they were on the 6502 and makes for dense object code.
What happens to the upper half of a register when the register size is 8-bits? For the index registers the answer is "nothing interesting," but the accumulator is a different story. The upper half of the accumulator is still there. The XBA instruction swaps the high and low halves of the accumulator giving, basically, a hidden 8-bit register for storing temporary values. XBA stands for "exchange B and A registers," the B register being someone's pet name for the upper half of the accumulator. Taking this silly naming scheme even further, sometimes the whole 16-bit accumulator is called the C register. (And to show how arbitrary and strained these naming conventions are, there is an unused-but-reserved-for-future expansion opcode with the mnemonic WDM. That's the designer of the 65816, William D. Mensch Jr.) XBA still works when the accumulator is 16bits, simply swapping the high and low bytes.
Pleasantries Revisited
There's are some nice little instruction tweaks that were quietly added to the 65816, but are often missed in the glare of the other, major additions. JMP and JSR have a new addressing mode, which looks like "JMP (MEM,X)." This is like the obscure indexed indirect mode of the 6502, except that MEM doesn't have to be in page zero. X is added to MEM and the data fetched from X+MEM is used as the jump target. This is handy for setting up jump tables. Assuming X is even, instead of this:
LDA TABLE,X ;Get 16-bit
jump target
STA TEMP
JMP (TEMP)
you can use the cleaner alternative:
JMP (TEMP,X)
You can do this with JSR too, which is great because there isn't a "JSR (MEM)" instruction on either the 6502 or 65816.
Another small but right-on-the-money change is that the decimal flag is always clear upon startup, after a reset, and after an NMI. You no longer need a "CLD" at the beginning of your code and interrupt handlers.
Oh Yeah, Those Block Move Instructions
I saved these for last because, though they are the most talked about and impressive features of the 65816, they are also rather unusual. More than instructions, they can be thought of as built-in subroutines to move blocks of data around, like memcpy in C and CMOVE in Forth. Before using the proper mnemonic, you have to set up the proper parameters.
The two block move instructions are MVN and MVP, which stand for "Block Move Next" and "Block Move Previous." The Next and Previous designations indicate the direction of the move. MVN moves data upward in memory, from low to high addresses, and is the more common of the two. MVP moves data downward, from high to low addresses.
Here's the calling sequence: put the lower 16-bits of the source address in X, the lower 16-bits of the destination address in Y, and the number of bytes to move minus one in the 16-bit accumulator. The actual block move instruction has two operands: the source bank and the destination bank. Here's an example to move the character set from $E000 to CHSET
LDX #$E000
LDY #CHSET
LDA #1023
MVN 0,0
Doesn't it look as if the MVN should be a JSR to a block move subroutine? But MVN is happening in hardware, and moves data at the rate of one byte every seven cycles. I remember being very impressed when I first looked at a 65816 book and saw this instruction and I guess I still am.
Apologies And Coming Attractions
My approach to explaining the 65816 is a bit unorthodox. Most books I have seen jump into, right up front, some of the specialized bank registers that the 65816 has and go into all sorts of quirks about register sizes and bank addressing. I opted for the simpler "you're a 6502 programmer and want to slip into the 65816" approach, at the expense of some lesser used details, instructions, and addressing modes (like stack addressing, which is handy for writing compilers but not used as often in pure assembly programs).
Stick around for part two: optimization. Things aren't quite as straightforward as they look. It's time to retire some of those 6502 programming methods....
