The SC #124

Issue #124 of The SC. Weekly supercomputing news. Not AMD sponsored but you would be forgiven for thinking it is! If you strip away the AI financing news the interesting bits left behind mostly came from AMD this week. GPU dominant of course but not your typical AI related.

The SC #124

Quite an interesting week for supercomputing news, albeit a little AMD dominant though of course it’s hard to talk about anything GPU these days without Nvidia being mentioned in the same breath.

Remember Talaas? They had a fair bit of fanfare earlier this year with AI inference token throughput rates that were off the charts. AMD just bought them. So now Nvidia has Groq, AMD has Talaas. As I’ve said in the past, when a use case becomes sufficiently prevalent it ends up being baked into the silicon. LLM inference will be no different and that’s across consumer, edge and data centre. Or at least I think so.

When we released our scheduler selection tool last year, I didn’t think we’d need to make many changes. I mean, the HPC scheduler market seemed pretty mature, and companies weren’t exactly falling over themselves to create new ones now were they. Yea how wrong was I. AMD just dropped one too. ROCm Spur. Obviously, its targeting GPU scheduling as the primary use case, is Slurm compatible and of course written in Rust to keep up with the cool kids. Having said all that, I’m not sure what meaningful addition this brings to the table. Guess we need to update the scheduler selector tool again!

Want to do something with your GPUs that isn’t just 4 bit cat picture generation and needs real precision. AMD has always been supportive of this use case with Nvidia pushing emulation using the Ozaki scheme instead and still hoping to win the top trumps game. Looks like AMD isn’t messing about. Their native hardware FP64 support in MI430X GPUs weighs in faster than Nvidia’s Rubin GPUs even including Ozaki emulation. Crikey. I do wonder how much longer that CUDA moat will hold out.

The scientific community seems to agree with the new DOE supercomputers being AMD based.

Let’s talk CPUs for a minute. Nividia’s new Vera chips seem pretty great, just not quite as great as their marketing material would have you believe. I’d quite like to get my hands on one to see how it stacks up for financial risk analytics workloads rather than the nonsense of Agentic AI that people seem to think needs an insane amount of compute.


All the News in Depth

Cloud and vendor releases for Supercomputing & HPC (and AI too if you change the filters)

Vera is a pretty good CPU; shame they are marketing it as though it isn’t and resorting to fairly petty things.

AMD buys Talaas the AI inference ASIC company. Reuters give you the quick take but if you want meandering tale the ends with surprisingly little detail about Talaas but gives you lots of context then the Next Platform article might be more your speed.

AMD is also rattling Nvidia for real FP64 performance. It was never a question if you needed real non emulated FP64 but even if you can accept Ozaki emulated FP64 it looks like AMD still comes out ahead!

AMD has even dropped their own, Slurm compatible, scheduler. Honestly I’m not sure who this is targeted at or how it moves the game forward if at all.


HMx Labs Updates

We’re starting to look at what version 2 of the COREx benchmark should look like. Want to provide some input?

Benchmarks Should Represent Your Workload: COREX v2
This is your reminder that a benchmark is pointless if it does not represent your own workload. Period.

We’d also like to know which venue you prefer for HPC Club

HPC Club: Which Venue Did You Prefer?
What’s more important to you, a casual setting for the conversations or a better environment for the presentation?

Off Topic

Most open source projects you’ve heard of probably aren’t developed by lone developers tolling away under the midnight oil out of the goodness of their hear (though for sure that exists). No most of them are probably paid handsomely by big tech. Ever wondered why? This is a pretty solid take which I mostly agree with.

This was quite a good watch.. and I certainly agree that if you take away the noise then there really has been very little of note in the AI space lately.


Know someone else who might like to read this newsletter? Forward this on to them or even better, ask them to sign up here: https://cloudhpc.news