Dual Xeon Server revisit: Quad Channel Memory and Process Lasso for more speed

In today’s episode of I’m an idiot and made a mistake, I wanted to add another 32 gigabytes of ECC RAM to this server’s existing 32 gigabytes, but somehow bought eight four gig sticks rather than the four eight gig sticks that I needed. It’s not dyslexia, it’s just stupidity.

But every cloud has a silver lining. This does let me test something on this dual Xeon Machinist motherboard that claims to have an X99 chipset, but is really has a C612 chipset, namely whether it works better with a quad channel memory setup, as opposed to the dual channel. Or quad and octa channel as it is sometimes reported, erroneously.

I’ve ran through a short suite of tests with the four old eight gig sticks, and I’ve repeated them with four of the new four gig sticks, and then with the final form of the eight four gig sticks to see if it’s making any difference.

We’re starting off with AIDA64’s Cache & Memory Benchmark. Looking at the Read speeds, here in megabytes per second, the old ram gets 65,879, with four of the new sticks getting a slightly lower by about 2% result of 64,656. Adding the rest of the four gig sticks increases that by a whopping 47% to 120,920. Impressive.

We can also use AIDA64 to measure the memory latency in nanoseconds, where the results were pretty much the same across all the tests, or within margin of error. The old RAM scored 84.6, it was 1%ish better with four of the new sticks at 83.8, then another 1% better with eight sticks at 82.8.

I also double checked this with Intel’s command line Memory Latency Checker utility, looking at the Peak Injection Memory bandwidth in MB/sec. The old ram gets 65,005, with four of the new sticks again getting a slightly lower result of 62,891. That’s about 3% down, but adding the rest of the four gig sticks increases that by 46% to 117,509.

For reference, my main system’s 6400 MT/s DDR5 only gets 60,248 MB/sec, so this is pretty solid performance no matter how you slice it.

So, at least we can confirm the motherboard is behaving as it should, but does this make any difference to anything outside of benchmarks? That’s kind of hard to answer because the only way to really confirm is by testing a repeated series of actions, in other works, benchmarking.

There are at least some benchmarks that test closer to real world activates rather than just testing the memory bandwidth, so let’s give them a try.

Another command line tool, and while I’m not normally calculating Pi to 500 digits in my day-to-day, yCruncher does act as a proxy for intense computations that the processors might be doing. Here with the old RAM, the total computation time in its benchmark was 15.2 seconds, with four of the four gig sticks, it took 15% longer at 17.8 seconds, and that was reduced by 51% to 11.8 seconds with the eight sticks installed. That’s about 30% faster than the old RAM setup.

The exceedingly useful compression utility 7-Zip also has a built in benchmark, which I wasn’t aware of until researching this. This gave a similar pattern to y-cruncher, but with less dramatic results. The old RAM gets 155.6 GIPS, four of the new sticks gets 141 GIPS, 10% lower, and loading it out with eight sticks gets us 164 GIPS, back up 14%, although that’s only about 5% faster than the original RAM results. What is a GIPS, you ask? I have no idea and I choose not to find out. Gonk Incidents Per Street, would be my guess, I think it’s a Cyberpunk thing.

Geekbench 6’s CPU benchmark runs though a wide range of activities such as image processing and HTML rendering to spit out a somewhat representative general computing use score. The old setup scores 7341, with four of the new sticks we drop 6% to 6929, and then we go up 24% to 9060 with eight of the new sticks. That’s 19% up from the original setup.

Cinebench R23 is a reasonably good proxy for 3D rendering programs, so let’s give that a prod. The old RAM setup scores 22212, surprisingly it goes up slightly, by 4%, to 23086 with four of the new sticks, and up a further 1% to 23246 with the full eight sticks, although hardly a dramatic increase in any case.

It seems the more realistic the test, the less of a change we see. From a bit of googling this is as expected. In terms of the games we’re normally concerned with on this channel, I’m told there’s the best chance of seeing a difference in something like a city builder, it being one of the more memory bashing types of games. I’ve ran through the Cities Skylines II benchmark at 1080p High settings, where the old setup gives 18.98 fps, the four stick 16GB setup gives 19.64 fps, and the eight stick setup gives 21.02 fps. Technically a few percentage points difference between them, but hardly conclusive.

So I suppose that’s pretty much proving what I expected, the right sort of task sees a monster boost from the quad channel memory speed, but in general use, it’s not exactly a game changer.

Speaking of games, in one of our previous videos we found that this servers dual CPU setup is challenging for some games, showing a big hit in performance compared to pulling one of them out and running the same tests. We should be able to use the Process Lasso utility to force games to run on one of the CPUs, so I tried that with Cities Skylines 2.

Surprisingly this saw a slight drop from the 21.02 fps to either 20.11 or 19.7 fps, depending on the NUMA node used, which just means the gestalt of the CPU and the memory channels wired to it. However, all of these results are margin of error stuff, so let’s take a quick look at some of the other problematic games tested last time that showed a significant drop off.

A lot of the biggest performance drops were in the 3DMark synthetics, so not the biggest deal but let’s see if we can’t fix it regardless.

In Fire Strike we saw a 10% difference in the favour of the single Xeon setup at 21486 vs 19351 points for the out of the box dual Xeon setup. After restricting the app with Process Lasso, we get 19811 and 19810 points for NUMA0 and NUMA1 respectively. That’s technically 2% better than not using it, but still roughly 8% down on unplugging one of the processors. Not a great start.

There was even more stark difference in Time Spy, where the single Xeon setup was 24% faster than the dual CPU machine, at 9897 vs 7541 points. In this round of testing I’m actually seeing worse results with Process Lasso, with a NUMA0 result of 7434 and NUMA1 of 7426, a couple of percent worse, and a huge 25% drop from the one Xeon result.

After double checking I’m doing this correctly, there was at least one result closer what I’d hoped for. Previously Baldur’s Gate 3 at 1080p Ultra settings with the usual run around Lower City gave a chunky 29% difference of 55 fps vs 39 fps, in favour of the single CPU setup. Here limited to NUMA0 we get 58 average FPS, so that’s marginally better than the last single run. Oddly using NUMA1 we’re only getting 45 FPS, still a 15% improvement but 18% off where it really should be.

Horizon Zero Dawn’s benchmark at 1080p Favour Quality settings previously gave a 5% win for the single Xeon setup with the 1080Ti, at 110 vs 104 fps, so not a massive difference but constraining it to NUMA0 gets us to 109 FPS, so closer to where it should be, although weirdly NUMA1 gets 97 fps, a 7% drop over the dual CPU results. Inexplicable, I cannot explic this.

The dual CPU setup in Total War: Warhammer III’s Mirrors of Madness benchmark at 1080P Ultra Settings last time gave us a single CPU result of 48 fps vs 32fps with two, a 33% drop. Now we are getting a NUMA0 result of 32.1fps, so no meaningful improvement, and a NUMA1 result of 29.9 FPS, so technically a 7% drop over the two CPU results.

‘s benchmark at 1080p High settings again saw the single CPU setup edge out a 9% win at 102 fps vs 93 fps. With Process Lasso, we get NUMA0 results of 101 FPS, pretty much where it ought to be, and NUMA1 results of 94 FPS, nowhere near where it ought to be and barely any difference at all.

I think the scientific conclusion here would be to say that further study is required, which is boffin speak for beats me, guv. I’m open to suggestions, but I’m pretty sure I’m using Process Lasso as intended and it’s only given the expected results half the time. There’s clearly something else in the dual CPU architecture that’s putting a spanner in the works, but as I don’t intend to have Windows on this for much longer I’m not going to spend too much time puzzling over it.

So, what is next for this lumbering beast? Well, I’m going to drop in my slightly Frankensteined 2080Ti in place of the 1080Ti, but first I want to run that through the same battery of testing that we subjected the 1080Ti to. After that we can get Proxmox installed and set this up as a file, media, and game streaming server, take over PiHole duties and I’m sure a bunch of other nonsense that I will document and report back on.

Until then, if you have any questions or want further details please leave a comment down below, and if you enjoyed this videotronic missive then consider subscribing. Until next time, take care of yourself, and each other.