\

Register deprivation: spills and runtime under forced register scarcity

19 points - last Thursday at 12:20 PM

Source
  • Scene_Cast2

    today at 5:48 PM

    I wonder how much effect the CPU's register remapping has on these results. Also, they ran it on an ancient CPU, but I'm not sure how much that matters.

      • CalChris

        today at 8:44 PM

        How much effect the CPU's register remapping has? Probably a lot and the article doesn't discuss it at all.

        However, the Intel Xeon E-2236 was released in 2019; so it's hardly ancient and the current E-2400 isn't so different.

        As for hardware register renaming, Tomasulo dates to 1967 and out-of-order ROBs were common in the 90s. I don't know that there's been any major new ideas microarchitectural ideas in register renaming since the 90s.

        I still think it's a meaningful experiment but could be improved with consideration of microarchitectural features.

        • hinkley

          today at 7:11 PM

          Probably not that much because one of the legacies of x86 is that library calling conventions are based on this same registry scarcity. To make an API call there is a convention that a limited number of registers are treated as callee-preserves and the rest are caller-preserves. So your compiler spills them defensively before calling out.

          You can make a programming language that uses whatever convention you want, but every time you make a system call or a call to a C or Rust library, you're using essentially an FFI and your FFI logic will have to handle the register spill and recovery to match the convention. And if you convert your code to a library, then it'll have to be dumbed down as well.

      • derdi

        today at 7:38 PM

        This is pretty meaningless without showing any assembly code. What does the inner loop for siphash look like? How many GPRs does it use? Where are the spills placed? What does perf say about any of this?

        • today at 7:48 PM