ack_complete
2 days ago
AArch64 definitely has a much more comprehensive baseline than x86-64, but there are some optional extensions that are situationally impactful, including the Crypto extension and some of the newer accumulation / dot product instructions. And unlike Intel, ARM has no portable equivalent to CPUID for querying feature flags and is terrible at documenting which intrinsics require specific FEAT_* flags.
The ARM-based CPU manufacturers make this worse by posting almost no low-level documentation for their CPUs. For basically any mainstream x86 CPU, it's trivial to find documentation listing what ISA level it supports and general execution widths and latencies for common operations. For the majority of ARM CPUs, there's absolutely nothing. ARM only has optimization guides for selected Cortex cores, and NVIDIA published info for their Olympus core. But execution details had to be reverse engineered for Apple M1, and there is nothing for Oryon. This is especially bad for in-order cores, which unfortunately is still relevant because new CPUs are still being shipped with in-order efficiency cores.
eqvinox
2 days ago
> some optional extensions that are situationally impactful, including the Crypto extension and some of the newer accumulation / dot product instructions.
Or LSE, adding a whole different set of atomic ops... which are way faster on some CPUs...
brohee
2 days ago
"The ARM-based CPU manufacturers make this worse by posting almost no low-level documentation for their CPUs."
Are those not based on standard ARM cores or ARM documentation on its cores not deep enough for your purpose?
ack_complete
2 days ago
Both. The specific documentation I'm referring to is their optimization guides, such as the Cortex-A72 optimization guide:
https://support.arm.com/documentation/uan0016/a/
This has detailed information on latencies and throughput, which are important when optimizing SIMD code. But ARM doesn't publish optimization guides for all their cores.
On top of that, the cores are often modified in significant ways. Snapdragon CPUs, for instance, have used modified Cortex cores in the past and can have performance differences from the original core.
To be fair, Intel's been slacking a lot on this too, not even bothering to update their own optimization guide for their latest cores. But that's made up for the community mining this information in great detail on sites like uops.info, and also being a lot less x86 cores to deal with. On the ARM side, there's practically not much analogous other than the Apple M1 microarchitecture analysis.
officialchicken
2 days ago
I really hope this is my last x86-64 CPU, Intel has become an incompetent steward. The "killer feature" for AVX-2(56) at the time of original release was basically lag/jitter-free video playback. IMO, 512 should have never been released for desktop CPUs and restricted to servers. One day I will to switch to a mainstream Neoverse dev box running linux. I also target Cortex-M in rust, so it's got a lot of the typical issues related to missing docs (e.g. bringup of non-heterogenous cores, meaning that M3/M4 still can't be used in a big.little chip)
Tuna-Fish
2 days ago
Why? AVX-512 is the best SIMD ISA anywhere right now. And the width is its least important feature.
nextaccountic
2 days ago
It's not available everywhere, which means you can't count on it if you don't know your deployment target
0x457
2 days ago
Neither is AVX or AVX2. AVX512 isn't avaiable on some of the latest intel CPUs is purely because intel can't say goodbye to Skylake for some reason.
ack_complete
2 days ago
AVX2 is much more prevalent than AVX-512 and thus tenable to require. RHEL is switching to x86-64-v3 baseline, for instance, which requires AVX2. AVX-512, on the other hand, has been moving backwards since Intel has not shipped any consumer-level CPUs with it for years now. The Steam Hardware Survey has AVX2 at 95.4% while AVX512F is still far behind at 23.9%.
0x457
2 days ago
So? AVX-512 shouldn't exist because it can't be retroactively enabled on older hardware? or because Intel can't say goodbye to Skylake? I don't understand the logic.
> AVX-512, on the other hand, has been moving backwards since Intel has not shipped any consumer-level CPUs with it for years now.
Good thing Intel hasn't shipped a viable consumer-level CPU in years either. What should never have happened is Intel removing AVX-512.
> Steam Hardware Survey
While it's a good source, it's not a definitive source. Four out of five amd64 machines at my place support AVX-512F, but none of them participate in that survey. And the one machine that doesn't have it... I wish it did.
At one point, Intel integrated graphics were literally the most common individual GPU in the Steam Hardware Survey - Intel HD 4000 held the #1 spot until the GTX 970 overtook it in late 2015. Does that mean developers shouldn't have targeted discrete GPUs either, just because most survey participants didn't have the hardware they wanted to target?
nextaccountic
2 hours ago
> So? AVX-512 shouldn't exist because it can't be retroactively enabled on older hardware?
If the issue were just older hardware, we could just wait out and let people upgrade a bit. But new hardware are shipping without AVX512 unfortunately
ack_complete
2 days ago
>So? AVX-512 shouldn't exist because it can't be retroactively enabled on older hardware? or because Intel can't say goodbye to Skylake? I don't understand the logic.
No one's saying it shouldn't exist, but it's a different story to ship software that requires it unconditionally, particularly if you are targeting the consumer market instead of only servers.
Not sure why you keep referring to Skylake since Intel hasn't shipped a Skylake-based CPU in years, and the most problematic CPUs currently shipping that lack AVX-512 support are several architectural generations ahead of Skylake, for both the P-cores and E-cores.
> Good thing Intel hasn't shipped a viable consumer-level CPU in years either. What should never have happened is Intel removing AVX-512.
Agreed on not removing, but saying that Intel hasn't shipped a viable consumer-level CPU in years is silly. They still ship huge volumes of CPUs, especially in laptops where AMD is still underrepresented.
> At one point, Intel integrated graphics were literally the most common individual GPU in the Steam Hardware Survey - Intel HD 4000 held the #1 spot until the GTX 970 overtook it in late 2015. Does that mean developers shouldn't have targeted discrete GPUs either, just because most survey participants didn't have the hardware they wanted to target?
Targeting a discrete GPU is different than not supporting integrated GPUs at all, which would be analogous. And many games do have to support integrated graphics even if the performance is not great, precisely because they don't aim high enough in the market to be able to ignore iGPUs entirely.
I also mention the Steam Hardware Survey because it tends to overrepresent users with higher end rigs. If you look at the non-gaming market, the hardware level tends to be considerably worse, and as a result a lot of productivity programs still ship plain SSE2 as their baseline required target.
0x457
2 days ago
The comment I originally replied to said: "IMO, 512 should have never been released for desktop CPUs and restricted to servers."
> Not sure why you keep referring to Skylake since Intel hasn't shipped a Skylake-based CPU in years, and the most problematic CPUs currently shipping that lack AVX-512 support are several architectural generations ahead of Skylake, for both the P-cores and E-cores.
I could have swore Intel's first E-cores were skylake based. My bad. Point is that its Intel issue, not AVX-512 issue.
> They still ship huge volumes of CPUs, especially in laptops where AMD is still underrepresented.
They physically shipped them, yes, but the products were lackluster. I don't recall when Intel had a good desktop CPU last time.
> Targeting a discrete GPU is different than not supporting integrated GPUs at all, which would be analogous.
The analogous case would be targeting a specific graphics API feature level, say, D3D 9.x, because that's what the majority of GPUs on the market support. Except game developers somehow figured out that they can support multiple D3D feature levels instead of being permanently stuck targeting whatever the majority happens to have.
nixon_why69
2 days ago
> IMO, 512 should have never been released for desktop CPUs and restricted to servers.
Why not? To save die space?
i80and
2 days ago
I think they meant it should have been deployed to all SKUs instead of the very silly ecosystem split Intel did for a long time.