Live data from Hacker News

Make the most of compiled C loops on the 68000

dciabrin.net

11–20 of 20 posts

Re: Make the most of compiled C loops on the 68000

#13

SNK were the gods of the 68000. I still remember back in the day getting a bug report on my 68000 emulator: When playing King of Fighters, the time counter would go down to 0 and then wrap around to 99, effectively preventing the round from ending. Eventually I tracked it down to the behavior of SBCD (Subtract Binary Coded Decimal): Internally, the chip actually does update the overflow flag reliably (it's marked as…

You don't need division to convert to decimal, though it will still be slower than using BCD operations.

Re: Make the most of compiled C loops on the 68000

#14

SNK were the gods of the 68000. I still remember back in the day getting a bug report on my 68000 emulator: When playing King of Fighters, the time counter would go down to 0 and then wrap around to 99, effectively preventing the round from ending. Eventually I tracked it down to the behavior of SBCD (Subtract Binary Coded Decimal): Internally, the chip actually does update the overflow flag reliably (it's marked as…

You don't need division to convert to decimal, though it will still be slower than using BCD operations.

Oh? How do you do it? Some kind of lookup table?

Re: Make the most of compiled C loops on the 68000

#15

Interestingly, gcc-amigaos-gcc 6.5 uses dbra without having to jump through any of those contortions, as long as the optimisation level is set to at least -O1: _clear_screen: move.w #28672,3932160 move.w #1,3932164 move.l #3932162,a0 move.w #-13570,d1 move.w #1279,d0 .L2: move.w d1,(a0) dbra d0,.L2 rts

I was once into 68k so I may be rusty, but shouldn't it be move.w d1,(a0)+ (increment the target address after each step)?

Re: Make the most of compiled C loops on the 68000

#16

SNK were the gods of the 68000. I still remember back in the day getting a bug report on my 68000 emulator: When playing King of Fighters, the time counter would go down to 0 and then wrap around to 99, effectively preventing the round from ending. Eventually I tracked it down to the behavior of SBCD (Subtract Binary Coded Decimal): Internally, the chip actually does update the overflow flag reliably (it's marked as…

You don't need division to convert to decimal, though it will still be slower than using BCD operations.

Technically no, but they were also always fighting against the ROM size, trying to keep costs down. Every byte helped.

Re: Make the most of compiled C loops on the 68000

#17
post #15

Interestingly, gcc-amigaos-gcc 6.5 uses dbra without having to jump through any of those contortions, as long as the optimisation level is set to at least -O1: _clear_screen: move.w #28672,3932160 move.w #1,3932164 move.l #3932162,a0 move.w #-13570,d1 move.w #1279,d0 .L2: move.w d1,(a0) dbra d0,.L2 rts

I was once into 68k so I may be rusty, but shouldn't it be move.w d1,(a0)+ (increment the target address after each step)?

The hardware increments an internal pointer after each access. The view to that address is through value in a0.

Re: Make the most of compiled C loops on the 68000

#18

Earlier quoted context omitted.

You don't need division to convert to decimal, though it will still be slower than using BCD operations.

Oh? How do you do it? Some kind of lookup table?

https://en.wikipedia.org/wiki/Double_dabble

A software implementation with masks and shifts will beat traditional CISC dividers.

Re: Make the most of compiled C loops on the 68000

#19

Interestingly, gcc-amigaos-gcc 6.5 uses dbra without having to jump through any of those contortions, as long as the optimisation level is set to at least -O1: _clear_screen: move.w #28672,3932160 move.w #1,3932164 move.l #3932162,a0 move.w #-13570,d1 move.w #1279,d0 .L2: move.w d1,(a0) dbra d0,.L2 rts

One thing that I heard from folks who do development for retro Atari platforms is that the 68k support in GCC has been getting worse as time has gone on, and it's very difficult to get the maintainers to accept patches to improve it, since 68k is not exactly widely used at this point.

Specifically, I heard that the 68k backend keeps getting worse, whilst the front-end keeps getting better. So choosing a GCC version is a case of examining the tradeoffs between getting better AST-level optimisations from a newer version, or more optimised assembly language output from an earlier version.

I imagine GCC 6.5 probably has a backend that makes better use of the 68k chip than the GCC 11.4 that ngdevkit uses (such as knowing when to use dbra) but is probably worse in other ways due to an older and less capable frontend.

Re: Make the most of compiled C loops on the 68000

#20
post #7
post #3

The step with declaring hw registers in assembly reminds me how assignment of value to pointer is IIRC at best implementation defined, and at worst UB, and playing around with volatile saves you not from zealous optimizer. Arguably every hardware register should be declared that way as a symbol

That was a common feature on Borland and Microsoft compilers for MS-DOS.

I mean, at its most basic, it's a feature from even the earliest compilers for C.

Not sure if Borland or MS shipped big fat symbol tables for all hardware registers of an IBM PC though?

Post reply on HN